Abstract
Multimodal sentiment analysis seeks to recognize and interpret emotional expressions from heterogeneous sources such as text, images, and speech. However, existing methods still face significant limitations in robustness and generalization due to strong modality heterogeneity, alignment challenges, and label sparsity. In this paper, we propose a cross-modal collaborative semantic enhancement method that leverages supplementary features and labels to improve the model's perceptual sensitivity to multimodal sentiment information and structural adaptability. Specifically, we design three composite interaction structures based on Multi-Head Cross Attention to model high-order dependencies across modalities, thereby enabling explicit cross-modal feature fusion and complementarity. To further address the challenges posed by limited supervision, we introduce image captioning and speech recognition as auxiliary tasks, establishing external semantic generation pathways. The intermediate representations produced by these tasks are leveraged as supplementary features to help improve the primary sentiment analysis task, while their predicted labels are utilized as supplementary supervision signals to alleviate the problem of label sparsity. This collaborative framework not only facilitates semantic transfer and structural sharing between the main and auxiliary tasks, but also improves the model's robustness to missing modalities and noise. Extensive comparative and ablation experiments conducted on the public CMU-MOSEI dataset demonstrate that our approach consistently outperforms existing state-of-the-art models.
| Original language | English |
|---|---|
| Pages (from-to) | 89705-89727 |
| Number of pages | 23 |
| Journal | IEEE Access |
| Volume | 14 |
| Early online date | 15 Jun 2026 |
| DOIs | |
| Publication status | E-pub ahead of print - 15 Jun 2026 |
Bibliographical note
2026 The Authors. This work is licensed under a Creative Commons Attribution 4.0 License.For more information, see https://creativecommons.org/licenses/by/4.0/
Funding
General Project of Wenzhou Science and Technology Bureau S2023013, Second Batch of Teaching Reform Projects for Zhejiang Provinces Higher Vocational Education during the 14th Five-Year Plan Period jg20240277, Major Project of Wenzhou Science and Technology Bureau ZG2024046, Open Project of the State Key Laboratory of Computer Aided Design and Computer Graphics of Zhejiang University A2313.
| Funders | Funder number |
|---|---|
| Wenzhou Municipal Science and Technology Bureau | S2023013, ZG2024046 |
| State Key Laboratory of Computer Aided Design and Computer Graphics | A2313 |
Keywords
- Modeling
- Labeling
- Sentiment analysis
- Head
- Training
- Learning (artificial intelligence)
- Bidirectional long short term memory
- Decoding
- Multitasking
- Conferences
- Multimodal sentiment analysis
- , multi-task collaborative training
- multi-head cross attention
- supplemental features
- supplemental labels
Fingerprint
Dive into the research topics of 'Cross-Modal Semantic Enhancement via Co-Training With Supplementary Features and Labels for Multimodal Sentiment Analysis'. Together they form a unique fingerprint.Cite this
- APA
- Standard
- Harvard
- Vancouver
- Author
- BIBTEX
- RIS