|
JCSE, vol. 20, no. 3, pp.166-176, September, 2026
DOI: http://dx.doi.org/10.5626/JCSE.2026.20.3.166
Speech Emotion Computing Model Based on Multi-modal Quantized Representation and Attention Mechanism
Yuan Lu Humanities College, Gandong College, Fuzhou, China
Abstract: Aiming at the feature redundancy, noise sensitivity and insufficient cross-modal interaction modeling in traditional speech emotion computing models, this study proposes a speech emotion computing model based on multi-modal quantized representation and attention mechanism. The model achieves feature discretization through codebook learning, and uses gated attention for cross-modal fusion. Finally, an end-to-end joint training strategy is used to collaboratively optimize the representation, fusion, and classification processes. On the IEMOCAP dataset, the weighted accuracy of the research model reached 74.8%, and the unweighted accuracy was 73.2%. Ablation experiments showed that after removing the quantization module, the weighted accuracy decreased by 3.5%, confirming its feature purification effect. In cross-dataset testing, the weighted accuracy on the MSP-IMPROV and CREMA-D datasets were 68.5% and 65.1%, respectively, showing good generalization ability. In addition, in the noise robustness analysis, the model performed robustly to white noise, but its performance dropped significantly under background music interference. In summary, the proposed model effectively improves the recognition accuracy and robustness through structured representation and dynamic fusion, and provides an effective solution for multi-modal emotion calculation in complex environments.
Keyword:
No keyword
Full Paper: 2 Downloads, 3 View
|