|
JCSE, vol. 20, no. 3, pp.119-130, September, 2026
DOI: http://dx.doi.org/10.5626/JCSE.2026.20.3.119
Music Emotion Recognition Model Based on Long Short-Term Memory Network and Gated Recurrent Unit
Xinxin Lai School of Art, Zhengzhou Technology and Business University, Zhengzhou, China
Abstract: To address the challenges of capturing long-term temporal dependencies, insufficient extraction of key emotional features, and the inherent limitations of single-modal information in music emotion recognition tasks, this study proposes a novel multi-modal framework. The core idea of the research is to break through the limitations of single mode in the expression of complex music emotions by separating and deeply mining the acoustic physical dynamics of audio and the music theory context of Musical Instrument Digital Interface (MIDI), and making adaptive fusion at the decision-making level. Specifically, the framework constructs a dual channel model based on long short-term memory (LSTM) networks and gated recurrent units (GRU). First, for symbolic music, a bidirectional LSTM with a self-attention mechanism is designed to capture music theory context and dynamically focus on crucial emotional phrases. Second, for audio music, a bidirectional GRU based on temporal pattern attention is constructed to perform deep temporal modeling of Mel spectrograms. Experimental results demonstrate that the symbolic model achieves an average accuracy of 88.68% on the MidiEmotion dataset, while the audio model attains 88.15% on the EMOPIA dataset. The multi-modal fusion model integrates the strengths of both, achieving a superior final accuracy of 91.55% and an F1-score of 91.04%. These findings prove that the proposed model effectively leverages the complementarity between symbolic semantics and acoustic features, substantially enhancing both the precision and robustness of emotion detection in music.
Keyword:
No keyword
Full Paper: 0 Downloads, 4 View
|