Skip to main navigation Skip to search Skip to main content

Cross-Modal Semantic Enhancement via Co-Training With Supplementary Features and Labels for Multimodal Sentiment Analysis

  • Zhejiang Industry and Trade Vocational College
  • Hefei University of Technology
  • Wuhan Institute of Technology

Research output: Contribution to journalArticlepeer-review

8 Downloads (Pure)

Abstract

Multimodal sentiment analysis seeks to recognize and interpret emotional expressions from heterogeneous sources such as text, images, and speech. However, existing methods still face significant limitations in robustness and generalization due to strong modality heterogeneity, alignment challenges, and label sparsity. In this paper, we propose a cross-modal collaborative semantic enhancement method that leverages supplementary features and labels to improve the model's perceptual sensitivity to multimodal sentiment information and structural adaptability. Specifically, we design three composite interaction structures based on Multi-Head Cross Attention to model high-order dependencies across modalities, thereby enabling explicit cross-modal feature fusion and complementarity. To further address the challenges posed by limited supervision, we introduce image captioning and speech recognition as auxiliary tasks, establishing external semantic generation pathways. The intermediate representations produced by these tasks are leveraged as supplementary features to help improve the primary sentiment analysis task, while their predicted labels are utilized as supplementary supervision signals to alleviate the problem of label sparsity. This collaborative framework not only facilitates semantic transfer and structural sharing between the main and auxiliary tasks, but also improves the model's robustness to missing modalities and noise. Extensive comparative and ablation experiments conducted on the public CMU-MOSEI dataset demonstrate that our approach consistently outperforms existing state-of-the-art models.

Original languageEnglish
Pages (from-to)89705-89727
Number of pages23
JournalIEEE Access
Volume14
Early online date15 Jun 2026
DOIs
Publication statusE-pub ahead of print - 15 Jun 2026

Bibliographical note

2026 The Authors. This work is licensed under a Creative Commons Attribution 4.0 License.
For more information, see https://creativecommons.org/licenses/by/4.0/

Funding

General Project of Wenzhou Science and Technology Bureau S2023013, Second Batch of Teaching Reform Projects for Zhejiang Provinces Higher Vocational Education during the 14th Five-Year Plan Period jg20240277, Major Project of Wenzhou Science and Technology Bureau ZG2024046, Open Project of the State Key Laboratory of Computer Aided Design and Computer Graphics of Zhejiang University A2313.

FundersFunder number
Wenzhou Municipal Science and Technology BureauS2023013, ZG2024046
State Key Laboratory of Computer Aided Design and Computer GraphicsA2313

    Keywords

    • Modeling
    • Labeling
    • Sentiment analysis
    • Head
    • Training
    • Learning (artificial intelligence)
    • Bidirectional long short term memory
    • Decoding
    • Multitasking
    • Conferences
    • Multimodal sentiment analysis
    • , multi-task collaborative training
    • multi-head cross attention
    • supplemental features
    • supplemental labels

    Fingerprint

    Dive into the research topics of 'Cross-Modal Semantic Enhancement via Co-Training With Supplementary Features and Labels for Multimodal Sentiment Analysis'. Together they form a unique fingerprint.

    Cite this