Music emotion classification method based on LSTM + GAT
By combining the timing modeling of LSTM and the relationship learning of GAT, a hybrid model for music emotion classification was constructed, which solved the problem that the existing technology was difficult to capture the complex characteristics of music, and achieved high accuracy and robust feature representation.
Patent Information
- Application Number
- CN202510006459.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-06
AI Technical Summary
Existing music emotion classification methods are difficult to fully capture the complex hierarchical relationships and long-distance dependence characteristics within music, resulting in limited classification performance.
The music emotion classification method based on LSTM+GAT is adopted, and timing modeling is performed through LSTM, relationship diagrams are constructed, and feature learning is used using double-layer GAT, and classification prediction is finally performed through full connection layer.
It effectively improves the accuracy of music emotion classification, reaching a classification accuracy of 99.25%, enhancing the robustness of feature representation, and having good interpretability.
Smart Images

Figure CN119939339A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of deep learning and emotion recognition, and in particular to a music emotion classification method based on LSTM+GAT. Background Art
[0002] Music sentiment classification is an important research direction in the field of music information retrieval (MIR). With the rapid expansion of digital music libraries, manual music sentiment classification has become extremely challenging, which has led to the rapid development of automated music sentiment classification technology. Existing music sentiment classification methods mainly include single-modality-based deep learning methods, multimodal fusion methods, and feature extraction and model comparison studies. However, these methods often simply regard music as a temporal feature sequence, which makes it difficult to fully capture the complex hierarchical relationships and long-distance dependency features within music.
[0003] At present, the mainstream music emotion feature extraction methods include low-level descriptor (LLD) feature extraction and Mel-frequency cepstral coefficient (MFCC) extraction. Although these methods can obtain the basic acoustic characteristics of music, they often ignore the multidimensional correlation between music features. Traditional machine learning methods such as KNN and SVM are difficult to effectively model the complex interactive relationship between features when processing these features, resulting in limited classification performance. In addition, the subjectivity and diversity of music emotions also bring great challenges to traditional classification methods.
[0004] The development of deep learning technology has brought new opportunities for music emotion classification. Long short-term memory networks (LSTM) perform well in temporal modeling, while graph attention networks (GAT) are good at capturing the relationship between features. However, most existing studies use a single deep learning model and fail to fully utilize the advantages of different models. Especially when processing data such as music with complex temporal structures and feature associations, a single model often fails to achieve satisfactory classification results.
[0005] In view of the shortcomings of existing technologies, a hybrid model that can simultaneously utilize the advantages of temporal modeling and relational learning is urgently needed to improve the accuracy of music sentiment classification. This model should be able to effectively handle the multi-dimensional correlation of music features and overcome the limitations of existing methods in the feature extraction and classification process. Summary of the invention
[0006] Purpose of the invention: The purpose of the invention is to add emotion category labels to music, provide basic data support for retrieval or recommendation, and overcome the shortcomings of the prior art. A music emotion classification method based on LSTM+GAT is proposed, which can effectively capture the temporal relationship and multi-dimensional correlation of music features and improve the accuracy of music emotion classification.
[0007] Technical solution: To achieve the purpose of the present invention, the technical solution adopted by the present invention is: a music emotion classification method based on LSTM+GAT, comprising the following steps:
[0008] S1. Music feature extraction: using a 44.1kHz sampling rate, the OpenSMILE tool is used to extract the low-level descriptor (LLD) features and Mel-frequency cepstral coefficient (MFCC) features of music; the low-level descriptor features include frequency-related features, energy-related features, spectrum features, and statistics of these features;
[0009] S2, time series feature learning, using LSTM network to perform time series modeling on the extracted features; constructing a fixed-length time series feature sequence for each time point; and filling the sequence with zeros when the sequence length is insufficient;
[0010] S3, graph structure construction, build the relationship graph between features based on the k-nearest neighbor algorithm; add self-loop connections to each node; obtain edge features by calculating the average value of connected node features;
[0011] S4, graph attention network processing, building a two-layer GAT structure: the first layer of GAT uses a multi-head attention mechanism, the second layer of GAT uses a single-head attention mechanism, and dropout is applied in each layer of GAT for regularization;
[0012] S5, classification prediction, maps the features to the target emotion category space through the fully connected layer; uses the log-softmax function to obtain the final category prediction probability.
[0013] Beneficial effects: Compared with the prior art, the technical solution of the present invention has the following beneficial technical effects:
[0014] 1) The present invention effectively improves the accuracy of music emotion classification by integrating the temporal modeling capability of LSTM and the relational learning advantage of GAT, achieving a classification accuracy of 99.25% on the test data set, which is 1.75 percentage points higher than the traditional KNN and SVM methods.
[0015] 2) The three-stage "timing-relationship-fusion" feature learning framework proposed in this invention realizes progressive learning from low-level temporal features to high-level semantic features, and can more comprehensively capture the multi-dimensional correlation between music features.
[0016] 3) The multi-head attention mechanism and double-layer GAT structure design adopted in this invention enhance the model's ability to learn different feature patterns and improve the robustness of feature representation.
[0017] 4) The feature extraction and selection strategy of the present invention comprehensively considers a variety of audio features and obtains the most discriminative feature subset through multi-level screening, thereby improving the computational efficiency and generalization ability of the model.
[0018] 5) While maintaining high classification accuracy, the present invention has good interpretability and can analyze the contribution of different music features to emotion classification through attention weights. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is an algorithm flow chart of an embodiment of the present invention;
[0020] Figure 2 is a schematic diagram of an embodiment of the present invention. DETAILED DESCRIPTION
[0021] The present invention is further explained below in conjunction with the accompanying drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, various equivalent forms of modifications to the present invention by those skilled in the art all fall within the scope defined by the claims attached to this application.
[0022] The music emotion classification method based on multimodal learning described in the present invention is Figure 1 , which shows the algorithm flow of an embodiment of the present invention, Figure 2 This is the architecture diagram of the present invention, which specifically includes the following steps:
[0023] S101. Low-level descriptor (LLD) feature extraction. For each music clip, three types of key low-level descriptor features are extracted: frequency-related features, energy-related features, and spectrum features. In order to capture the changing patterns of these LLD features in the time dimension, a series of statistical functions are calculated for each feature, including mean, standard deviation, skewness, kurtosis, range, quartiles, extreme points and other statistics.
[0024] S102, Mel Frequency Cepstral Coefficients (MFCCs) extraction: pre-emphasize the audio signal to enhance the high-frequency components. The specific formula is:
[0025] y[t]=x[t]-αx[t-1]
[0026] Among them, α is set to 0.97; use the Hamming window for frame processing, calculate the fast Fourier transform of each window, and get the spectrum:
[0027]
[0028] Convert the frequency domain signal to the Mel frequency scale and calculate the Mel frequency:
[0029]
[0030] Then use the Mel filter bank to process the FFT result and take the logarithm of the Mel filter result. Finally, perform discrete cosine transform on the Mel frequency energy after taking the logarithm to obtain MFCC:
[0031]
[0032] Where n is the dimension of MFCC and M is the number of Mel filters.
[0033] S103, time series feature learning, using LLD and MFCCs as LSTM input to obtain features. In order to obtain sufficient data size for applying deep learning for feature extraction, each music sample is divided into multiple music clips. For each music clip, the acoustic features are extracted using a short-time analysis method. The Hamming window is used for frame processing. MFCCs are extracted for each frame signal, and an MFCCs feature matrix is obtained for each clip. These features are input into LSTM to generate LSTM-based features for each music clip. Thus, multi-dimensional LSTM-based MFCCs features are obtained for each music sample.
[0034] S104, feature selection, uses a multi-level screening strategy to obtain the most discriminative feature subset: combine mutual information (MI) and F-score for preliminary feature evaluation. Mutual information measures the statistical dependence between features and category labels:
[0035]
[0036] Where P(x,y) is the joint probability distribution of feature X and label Y, and P(x) and P(y) are their marginal probability distributions respectively.
[0037] The F value evaluates the discriminative ability of features by calculating the ratio of the between-group variance to the within-group variance:
[0038]
[0039] MSB is the mean square error between groups, and MSW is the mean square error within groups.
[0040] In order to balance the two evaluation indicators, we perform min-max normalization on the MI score and F value and use a weighted combination:
[0041] Score c ombined=0.5×Score M I n orm+0.5×Scornorm
[0042] The top k features with the highest scores are selected based on the combined scores. Subsequently, a second round of feature selection is performed using a random forest classifier to further screen features by evaluating their importance scores in the decision tree.
[0043] S105, graph structure construction, based on the processed feature sequence, the topological structure of the graph is constructed using the k nearest neighbor algorithm, and each node is connected to the k most similar nodes in its feature space to form an initial edge set. In order to ensure the complete transmission of information and the retention of the node's own characteristics, a self-loop connection is added to each node, and the edge features of the graph are obtained by calculating the average value of the connected node features. The stratified sampling strategy is used to divide the data set into a training set and a test set in a ratio of 8:2 to ensure that the divided data set maintains a balanced distribution in each category, providing a reliable data basis for model training and evaluation.
[0044] S106, GAT feature learning, construct a two-layer graph attention network structure for feature learning. For node i and its neighbor node j, the attention coefficient is calculated using the following formula:
[0045]
[0046] The normalized attention weight is obtained by the softmax function:
[0047]
[0048] Feature update of node i:
[0049]
[0050] Where W is the weight matrix, is the feature vector of node i, is the neighbor set of node i, σ is the nonlinear activation function, α ij is the attention weight. In order to improve the stability and expressiveness of the model, a multi-head attention mechanism is used to concatenate or average the outputs of K independent attention mechanisms. The expression is:
[0051]
[0052] in represents the concatenation operation, and K is the number of attention heads.
[0053] S107, emotion classification decision, map the features output by GAT to the target emotion category space through the fully connected layer, use the log-softmax function to obtain the final category prediction probability distribution, and realize the classification of music emotion. Use the cross entropy loss function for model training, and use the Adam optimizer for parameter optimization.
[0054] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A music emotion classification method based on LSTM+GAT, characterized in that: The following steps are involved: S1. Music feature extraction. Using a 44.1kHz sampling rate, the OpenSMILE tool is used to extract the low-level descriptor (LLD) features and Mel-frequency cepstral coefficient (MFCC) features of music. S2. Time series feature learning: Use LSTM network to perform time series modeling on the extracted features; construct a fixed-length time series feature sequence for each time point. S3, graph structure construction, build the relationship graph between features based on the k-nearest neighbor algorithm; add self-loop connections to each node; obtain edge features by calculating the average value of the connected node features; S4. Graph attention network processing, constructing a two-layer GAT structure: the first layer of GAT adopts a multi-head attention mechanism, the second layer of GAT uses a single-head attention mechanism, and dropout is applied in each layer of GAT for regularization. S5, classification prediction, maps the features to the target emotion category space through the fully connected layer; uses the log-softmax function to obtain the final category prediction probability.
2. The music emotion classification method according to claim 1, characterized in that: The low-level descriptor features in step S1 include: frequency-related features, energy-related features, and statistics of spectrum features.
3. The music emotion classification method according to claim 1, characterized in that: The MFCC feature extraction in step S1 includes: pre-emphasis processing on the audio signal; frame processing using a Hamming window; calculating a fast Fourier transform to obtain a spectrum; converting the frequency domain signal to a Mel frequency scale; processing the FFT result using a Mel filter bank and taking the logarithm; and performing a discrete cosine transform on the Mel frequency energy after taking the logarithm.
4. The music emotion classification method according to claim 1, characterized in that: The graph attention network processing in step S4 includes: calculating the attention coefficient between nodes; obtaining normalized attention weights through the softmax function; updating node features; and using a multi-head attention mechanism to improve the stability and expressiveness of the model.
5. The music emotion classification method according to claim 1, characterized in that: The feature extraction and selection process also includes: combining mutual information and F-value for preliminary feature evaluation; performing min-max normalization on mutual information scores and F-values; calculating feature scores using a weighted combination method; and using a random forest classifier for a second round of feature selection.
6. The music emotion classification method according to claim 4, characterized in that: The calculation formula of the attention coefficient is: Where W is the weight matrix, is the node feature vector and a is the learnable attention vector.
7. The music emotion classification method according to claim 4, characterized in that: The feature update of the multi-head attention mechanism adopts the following formula: in represents the concatenation operation, and K is the number of attention heads.
8. The music emotion classification method according to claim 1, characterized in that: The method achieves progressive learning from low-level temporal features to high-level semantic features by integrating the temporal modeling capabilities of LSTM and the relational learning advantages of GAT, and can achieve a classification accuracy of 99.25% on the test dataset.