Music Emotion Recognition Method and System Based on Cross-Modal Fusion
Through the cross-modal fusion music emotion recognition method, using audio, text and image features, and using Transformer Encoder and fully connected classifiers, the problem of insufficient multimodal data fusion accuracy in the prior art is solved, and higher recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202310058201.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-16
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2043-01-16
AI Technical Summary
The existing music emotion recognition methods are not very accurate in multimodal data fusion, making it difficult to effectively utilize relevant data other than audio, and have high requirements for training data, and lack recognition accuracy and robustness.
The cross-modal fusion method is adopted to extract audio, text and image features, feature fusion is performed using a pre-trained Transformer Encoder, and emotional recognition is performed using a fully connected classifier, and feature extraction is performed by combining BERT and Resnet50 models.
It improves the accuracy and robustness of music emotion recognition, reduces the requirements for training data, is suitable for practical application scenarios, and can maintain high recognition accuracy when some inputs are missing.
Smart Images

Figure CN116010902B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of music emotion recognition, and particularly relates to a music emotion recognition method and system based on cross-modal fusion. Background Art
[0002] Music emotion recognition (MER) refers to the task of identifying the emotional information contained in a given music segment. Its input is the original audio file and other relevant information such as lyrics, song names, etc., and the output is the category of emotion or the valence / activation value. Music emotion recognition is a process of condensing and extracting high-level semantic information in music and is also an important subtask of music content understanding. Therefore, it occupies an extremely important and special position in the field of music information retrieval. Since emotion is considered the soul of music, music emotion recognition is an important part of understanding music content. The main application fields of music emotion recognition are music information retrieval based on music emotion, including tasks such as music classification and music recommendation.
[0003] Current commonly used music service providers in the industry, such as Deezer, Spotify, Apple Music, and domestic NetEase Cloud Music, Tencent Music, etc., all provide emotion-based music recommendation services in their APPs. For example, in the recommended playlists of Apple Music, there are playlists such as "Happy Mood" and "Anxious" recommended according to the emotional state of the listeners. There is also emotion classification in the classification retrieval system of music software. In these application scenarios, since the song data is often extremely large in scale, using manual emotion annotation often leads to too high system development costs and too low efficiency. The research task of the present invention, namely music emotion recognition, aims to automatically complete the emotion recognition of a large amount of song data using a computer algorithm model, so that a large-scale emotion-based music information retrieval system can be realized.
[0004] Early music emotion recognition methods mainly relied on traditional middle-level features of music or used traditional machine learning models to learn implicit features (latent features) in audio. Mion et al. explored middle-level features affecting music emotion expression using feature selection methods and principal component analysis methods, and found that information such as attack time, number of notes per second, and highest sound level was highly correlated with music emotion. Schmidt et al. fused middle-level features such as spectral centroid and MFCC (Mel Frequency Cepstrum Coefficient) and designed a representative music emotion recognition model based on these features.
[0005] In terms of traditional machine learning methods, in 2012, Wang et al. proposed a method using Gaussian Mixture Model (GMM) to learn implicit features in music, and named it Acoustic Emotion Gaussian (AEG). AEG had a great impact on the field of music emotion recognition and was still applied in emotion recognition methods by Chen et al. until 2017. In addition to AEG, traditional machine learning methods such as Support Vector Machine (SVM) and Linear Regression (LR) are also often used for emotion recognition in the audio domain.
[0006] Since deep learning methods have gradually developed and matured, most of the recent music emotion recognition research is based on deep learning. In 2016, Li et al. proposed a model based on Deep Bidirectional Long Short Term Memory (DBLSTM) and used multi-resolution feature fusion to achieve the state-of-the-art (SoTA) effect on the MediaEval dataset. Many other LSTM-based methods also emerged during the same period. After that, until now, Convolutional Neural Network (CNN) and Recurrent Neural Network (RNN) are still the mainstream in music emotion recognition. The bidirectional CRNN model proposed by Dong et al. in 2019 still has a great influence today. At present, the performance of emotion analysis in the audio domain has encountered a major bottleneck and has developed relatively slowly in recent years.
[0007] Since not all factors reflecting and influencing music emotion exist in a single modality, it is very feasible in theory to use multi-modal methods that combine text and audio to solve emotion problems. Early multi-modal methods were also based on traditional machine learning models. In 2008, both Laurier et al. and Yang et al. proposed emotion recognition models that use SVM to analyze audio content and lyric content simultaneously. In 2016, Huang et al. proposed a multi-modal emotion recognition model based on Deep Boltzmann Machine (DBM).
[0008] In recent years, a relatively representative multi-modal emotion method was proposed by Delbouys et al. in 2018. Delbouys et al. used LSTM to analyze the lyrics content and 1D CNN to analyze the audio content. They also proposed a high-quality multi-modal music dataset. In 2021, Konstantinos et al. proposed a multi-modal method that uses pre-trained BERT as the text model and 2D CNN as the audio model, and achieved SoTA performance.
[0009] Based on the above research background, it is not difficult to summarize the following common problems existing in the existing music emotion recognition methods:
[0010] 1. The emotion recognition accuracy is not high enough. In actual application scenarios, it is often necessary to classify music emotions with more labels, but the existing methods perform poorly in music datasets with more than four emotion categories and are difficult to achieve commercial accuracy.
[0011] 2. Although the methods based on traditional machine learning and middle-level features have good generalization and low requirements for the amount of data, the recognition accuracy is extremely low, and they cannot handle inputs with complex music structures and emotion information.
[0012] 3. Although the methods based on deep learning have relatively high accuracy, there is still a large gap from the accuracy required for actual applications. They also have high requirements for the quantity and annotation quality of training data and perform poorly in cross-dataset tests.
[0013] 4. On the one hand, the existing multi-modal methods are difficult to utilize all available relevant data, such as album names, song names, album covers, etc. On the other hand, the feature fusion efficiency of each modality is relatively low. The existing methods generally use methods such as direct concatenation or cross-modal attention for feature fusion between modalities. Summary of the Invention
[0014] The purpose of the present invention is to provide a music emotion recognition method and system based on cross-modal fusion to improve the accuracy of music emotion recognition.
[0015] The music emotion recognition method based on cross-modal fusion provided by the present invention specifically includes the following steps:
[0016] (1) Obtain the music information of the music to be recognized; the music information includes audio files, lyrics, song names, singer names, album names, and album covers;
[0017] (2) Extract the music emotion features of the music information; the music emotion features include audio features, text features, and image features; the audio features are the features of the audio files; the text features include the features of the lyrics, the song names, the singer names, and the album names; the image features are the features of the album covers;
[0018] (3) Preprocess the music emotion features to obtain a feature fusion sequence with positional encoding;
[0019] (4) Input the feature fusion sequence into a pre-trained Transformer Encoder[1] to obtain the fused music emotion features;
[0020] (5) Input the fused music emotion features into a pre-trained fully connected classifier to obtain the predicted music emotion categories; the music emotion categories include happy, sad, angry, and relaxed.
[0021] Furthermore, the extraction of the music emotion features of the music information in step (2) specifically includes:
[0022] (1) Calculate the Mel spectrogram of the audio file;
[0023] (2) Adjust the size of the spectrogram of the Mel spectrogram to a preset size;
[0024] (3) Input the adjusted spectrogram into a pre-trained 2D-CNN model for audio feature extraction to obtain audio features;
[0025] (4) Concatenate the lyrics, the song name, the singer name, and the album name to obtain a concatenated text;
[0026] (5) Calculate the word embedding representation of the concatenated text;
[0027] (6) Input the word embedding representation into a pre-trained BERT[2] model to obtain text features;
[0028] (7) Input the album cover into a pre-trained Resnet50[3] model to obtain image features.
[0029] Optionally, the 2D-CNN includes a convolutional layer and a pooling layer; the convolutional layer includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, and a fifth convolutional layer; the pooling layer includes a first pooling layer, a second pooling layer, a third pooling layer, a fourth pooling layer, and a fifth pooling layer;
[0030] The first convolutional layer, the first pooling layer, the second convolutional layer, the second pooling layer, the third convolutional layer, the third pooling layer, the fourth convolutional layer, the fourth pooling layer, the fifth convolutional layer, and the fifth pooling layer are connected in sequence;
[0031] The convolutional kernel size of the convolutional layer is all 3*3; the pooling window size of the pooling layer is all 2*2.
[0032] Further, the preprocessing of the music emotion features in step (3) to obtain a feature fusion sequence with positional encoding specifically includes:
[0033] (1) Convert the music emotion features into one-dimensional vectors;
[0034] (2) Segment the music emotion features converted into one-dimensional vectors to obtain a feature sequence with equal-length segments;
[0035] (3) Perform positional encoding on the segments in the feature sequence with equal-length segments to obtain a feature fusion sequence with positional encoding.
[0036] Optionally, the format of the audio file is the wav format.
[0037] Optionally, the format of the album cover is the jpg format.
[0038] Further, the Transformer Encoder in step (4) is a 6-layer TransformerEncoder. The Transformer mechanism consists of a multi-head self-attention mechanism and a feed-forward network. The input of the Transformer Encoder is the feature fusion sequence (the embedded representation sequence of multimodal fusion). Input the feature fusion sequence into the pre-trained 6-layer Transformer Encoder, and the output is the feature after multimodal fusion, that is, the fused music emotion feature.
[0039] Further, the input of the fused music emotion feature into the pre-trained fully connected classifier in step (5) to obtain the predicted music emotion category specifically includes:
[0040] Input the fused music emotion feature into the pre-trained fully connected classifier to obtain the predicted probabilities of each music emotion category; according to the predicted probabilities of each music emotion category, obtain the predicted music emotion category.
[0041] Based on the above cross-modal fusion music emotion recognition method, the present invention also mentions a corresponding cross-modal fusion music emotion recognition system, and the system includes:
[0042] An acquisition module for acquiring the music information of the music to be recognized; the music information includes an audio file, lyrics, song name, singer name, album name, and album cover;
[0043] An extraction module for extracting the music emotion features of the music information; the music emotion features include audio features, text features, and image features; the audio features are the features of the audio file; the text features include the features of the lyrics, the song name, the singer name, and the album name; the image features are the features of the album cover;
[0044] A preprocessing module for preprocessing the music emotion features to obtain a feature fusion sequence with positional encoding;
[0045] A fusion module for inputting the feature fusion sequence into a pre-trained Transformer Encoder to obtain the fused music emotion features;
[0046] A prediction module for inputting the fused music emotion features into a pre-trained fully connected classifier to obtain the predicted music emotion categories; the music emotion categories include happy, sad, angry, and relaxed.
[0047] These five modules perform the operations of the five steps of the music emotion recognition method based on cross-modal fusion.
[0048] The present invention also provides an electronic device, including a memory and a processor. The memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the above-mentioned music emotion recognition method based on cross-modal fusion.
[0049] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned music emotion recognition method based on cross-modal fusion.
[0050] According to the specific embodiments provided by the present invention, the following technical effects are disclosed by the present invention:
[0051] The music emotion recognition method based on cross-modal fusion provided by the present invention applies other music contents in addition to audio information, including lyrics, song names, singer names, album names, and album covers, and can obtain more accurate music emotion features. Moreover, cross-modal Transformer is applied to fuse the features of each modality. Compared with the existing multi-modal methods using cross-modal attention or direct feature splicing, it has better performance and robustness when facing input data with more complex emotion information. Therefore, when the pre-trained fully connected classifier is used for music emotion recognition prediction, more accurate recognition results can be obtained. Description of the Drawings
[0052] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0053] Figure 1 It is a flowchart of the music emotion recognition method based on cross-modal fusion provided by the present invention.
[0054] Figure 2 It is a structural diagram of the music emotion recognition model based on cross-modal fusion provided by the present invention.
[0055] Figure 3 It is a module diagram of the music emotion recognition system based on cross-modal fusion provided by the present invention.
[0056] Reference numerals in the figure: 1 is the acquisition module, 2 is the extraction module, 3 is the preprocessing module, 4 is the fusion module, and 5 is the prediction module. Specific Embodiments
[0057] The present invention will be further introduced below in combination with the embodiments and the accompanying drawings. The described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0058] Embodiment 1
[0059] As Figure 1 shown, the present invention provides a music emotion recognition method based on cross-modal fusion. The method includes:
[0060] Step S1: Obtain the music information of the music to be recognized; the music information includes an audio file, lyrics, song name, singer name, album name, and album cover; specifically, the format of the audio file is the wav format. The format of the album cover is the jpg format.
[0061] Step S2: Extract the music emotion features of the music information; the music emotion features include audio features, text features, and image features; the audio features are the features of the audio file; the text features include the features of the lyrics, the song name, the singer name, and the album name; the image features are the features of the album cover.
[0062] S2 specifically includes:
[0063] Step S21: Calculate the Mel spectrogram of the audio file.
[0064] Step S22: Adjust the size of the spectrogram of the Mel spectrogram to a preset size.
[0065] Step S23: Input the adjusted spectrogram into a pre-trained 2D-CNN model for audio feature extraction to obtain audio features. Specifically, the 2D-CNN includes a convolutional layer and a pooling layer; the convolutional layer includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, and a fifth convolutional layer; the pooling layer includes a first pooling layer, a second pooling layer, a third pooling layer, a fourth pooling layer, and a fifth pooling layer; the first convolutional layer, the first pooling layer, the second convolutional layer, the second pooling layer, the third convolutional layer, the third pooling layer, the fourth convolutional layer, the fourth pooling layer, the fifth convolutional layer, and the fifth pooling layer are connected in sequence; the convolutional kernel size of the convolutional layer is 3*3; the pooling window size of the pooling layer is 2*2.
[0066] In practical applications, for the input audio file, the melspectrogram() function in the librosa library of python is used to calculate its Mel spectrogram, and the resize() function is used to adjust its size to the same size. Specifically, the Mel spectrogram is a spectrogram with the horizontal axis representing time, the vertical axis representing frequency, and the energy distribution represented by the color shade. The generated spectrogram is a graph with a size of x*y, which here means adjusting this size to be consistent. After calculating the Mel spectrogram, a five-layer 2D-CNN is used for audio feature extraction. Each of these five convolutional layers has 32, 64, 128, 256, and 128 filters respectively, and a 3*3 convolutional kernel size is used, and the stride is set to 1. In order to downsample the features, a 2D max pooling layer is used after each convolutional layer, and the pooling window size is set to 2*2. Through the above calculations, the audio features of the music are obtained.
[0067] Step S24: Concatenate the lyrics, the song name, the singer name, and the album name to obtain a concatenated text.
[0068] Step S25: Calculate the word embedding representation of the concatenated text.
[0069] Step S26: Input the word embedding representation into a pre-trained BERT model to obtain text features.
[0070] In practical applications, for the input lyrics, song titles, singer names, and album names, these text information are concatenated, and the word embedding representation corresponding to the concatenated text information is calculated. Specifically, the BertEmbeddings() function in the Transformers library can be used for quick calculation to convert the input text information into a word embedding representation form that can be processed by the BERT model. After calculating the word embeddings corresponding to these text information, the pre-trained BERT model is used to extract the text features corresponding to the music lyrics, song titles, singer names, and album names. The input of the pre-trained BERT model is the word embedding representation of the text information, and the output is the extracted text features.
[0071] Step S27: Input the album cover into the pre-trained Resnet50 model to obtain image features.
[0072] In practical applications, for the input album cover, the pre-trained Resnet50 model in the torchvision library is used to extract its features, and the image features corresponding to the song album cover are obtained. Among them, the input of the pre-trained Resnet50 model is a picture, and the output is the extracted image features.
[0073] Step S3: Preprocess the music emotion features to obtain a feature fusion sequence with positional encoding.
[0074] S3 specifically includes:
[0075] Step S31: Convert the music emotion features into one-dimensional vectors.
[0076] Step S32: Segment the music emotion features converted into one-dimensional vectors to obtain a feature sequence with equal-length segments.
[0077] Step S33: Perform positional encoding on the segments in the feature sequence with equal-length segments to obtain a feature fusion sequence with positional encoding.
[0078] In practical applications, the above-obtained audio features, text features, and image features are all flattened into one-dimensional vectors, and they are segmented into a feature sequence with equal-length segments for each segment. Positional encoding is added to the above feature sequence to obtain a multi-modal fusion embedding representation sequence. Specifically, each segment is regarded as an Embedding (embedding representation), and the positional encoding corresponding to the segment is directly added or multiplied to it to obtain the multi-modal fusion embedding representation sequence. This multi-modal fusion embedding representation sequence is the feature fusion sequence with positional encoding.
[0079] Step S4: Input the feature fusion sequence into a pre-trained Transformer Encoder to obtain the fused music emotion features. Specifically, the pre-trained Transformer Encoder is a 6-layer TransformerEncoder. The Transformer mechanism consists of a multi-head self-attention mechanism and a feed-forward network. The input of the Transformer Encoder is the feature fusion sequence (the embedded representation sequence of multi-modal fusion). Inputting the feature fusion sequence into the pre-trained 6-layer Transformer Encoder, the output is the feature after multi-modal fusion, that is, the fused music emotion features.
[0080] Step S5: Input the fused music emotion features into a pre-trained fully connected classifier to obtain the predicted music emotion categories; the music emotion categories include happy, sad, angry, and relaxed.
[0081] S5 specifically includes:
[0082] Step S51: Input the fused music emotion features into a pre-trained fully connected classifier to obtain the predicted probabilities of each music emotion category. Specifically, using a fully connected classifier for the fused music emotion features can output the predicted probabilities (i.e., the probabilities of the input music corresponding to each emotion category). The music emotion categories depend on the actual application scenario and the training data set, including but not limited to happy, sad, angry, and relaxed.
[0083] Step S52: Obtain the predicted music emotion category according to the predicted probabilities of each music emotion category.
[0084] In practical applications, the music emotion recognition method based on cross-modal fusion provided by the present invention constitutes a music emotion recognition model. This model takes music information as the input and the music emotion recognition result as the output; this model includes an initially pre-trained Transformer Encoder, an initially pre-trained 2D-CNN, a pre-trained BERT model, and an initially pre-trained Resnet50; in the model, the initially pre-trained Transformer Encoder, the initially pre-trained 2D-CNN, the pre-trained BERT model, and the initially pre-trained Resnet50 can be loaded with open-source pre-trained models.
[0085] When the overall training of the music emotion recognition model is carried out, the music information can be obtained from the music database. During the training process, open-source datasets can be used for training. For example, some subsets of the Million Song dataset have emotion label annotations. During the training process of the music emotion recognition model, since the Transformer Encoder, 2D-CNN, BERT model, and Resnet50 are all pre-trained, when conducting the overall training of the music emotion recognition model, only the parameters of the initially pre-trained Transformer Encoder, initially pre-trained 2D-CNN, and initially pre-trained Resnet50 need to be fine-tuned, saving the training time of the music emotion recognition model and obtaining the pre-trained Transformer Encoder, pre-trained 2D-CNN, and pre-trained Resnet50. The pre-trained BERT model does not need to update its parameters during the overall training process of the music emotion recognition model, and when conducting the overall training of the music emotion recognition model, a pre-trained fully connected classifier is obtained.
[0086] Train the music emotion recognition model and calculate the cross-entropy loss function between the output probability of the recognition model and the label. Using this loss function for backward gradient transmission can complete the training of the model. Specifically, there are two criteria for completing the training. First, when the loss function basically converges during the training process, it is considered that the training is completed. Second, during the training, the performance of the current model is tested on the validation set every certain number of rounds, and the version with the best performance is taken as the trained version. In addition, the initially pre-trained Transformer Encoder, initially pre-trained 2D-CNN, pre-trained BERT model, and initially pre-trained Resnet50 do not require additional fine-tune (transfer learning) operations. Moreover, directly using the open-source dataset to train the overall model loaded with the initially pre-trained Transformer Encoder, initially pre-trained 2D-CNN, pre-trained BERT model, and initially pre-trained Resnet50, during the training process of the music emotion recognition model, the parameters of the initially pre-trained Transformer Encoder, initially pre-trained 2D-CNN, and initially pre-trained Resnet50 will be adaptively adjusted accordingly.
[0087] When the music emotion recognition method based on cross-modal fusion provided by the present invention is actually used, some inputs can be defaulted.
[0088] Compared with the prior art, the advantages of the music emotion recognition method based on cross-modal fusion provided by the present invention are as follows:
[0089] 1. The music emotion recognition method based on cross-modal fusion provided by the present invention can more efficiently utilize other music contents besides audio information, such as lyrics, song names, singer names, album names, and album pictures, etc. It has higher recognition accuracy compared with existing methods and is more suitable for implementation in actual application scenarios;
[0090] 2. The feature extraction models used in the music emotion recognition method based on cross-modal fusion provided by the present invention are mostly open-source pre-trained models. Therefore, only a small amount of labeled data is required for fine-tuning during actual use, and the requirement for training data is lower compared with existing methods, and the cost during actual application is lower;
[0091] 3. The music emotion recognition method based on cross-modal fusion provided by the present invention can still run when some inputs are missing, and has higher robustness in the face of special situations, and can predict emotions through partial inputs. (For example, relatively accurate predictions can also be made by only inputting the song name and lyrics);
[0092] 4. The music emotion recognition method based on cross-modal fusion provided by the present invention uses a high-performance cross-modal Transformer for fusing features of each modality. Compared with existing multi-modal methods using cross-modal attention or direct feature concatenation, it has better performance and robustness when facing input data with more complex emotion information.
[0093] Example 2
[0094] In order to execute the method corresponding to the above Example 1 to achieve the corresponding functions and technical effects, a music emotion recognition system based on cross-modal fusion is provided below, as Figure 3 shown. The system includes:
[0095] An acquisition module 1, configured to acquire music information of the music to be recognized; the music information includes an audio file, lyrics, a song name, a singer name, an album name, and an album cover.
[0096] An extraction module 2, configured to extract music emotion features of the music information; the music emotion features include audio features, text features, and image features; the audio features are the features of the audio file; the text features include the features of the lyrics, the song name, the singer name, and the album name; the image features are the features of the album cover.
[0097] A preprocessing module 3, configured to preprocess the music emotion features to obtain a feature fusion sequence with position encoding.
[0098] A fusion module 4, configured to input the feature fusion sequence into a pre-trained Transformer Encoder to obtain fused music emotion features.
[0099] A prediction module 5, configured to input the fused music emotion features into a pre-trained fully-connected classifier to obtain predicted music emotion categories; the music emotion categories include happy, sad, angry, and relaxed.
[0100] Embodiment 3
[0101] The present invention provides an electronic device, including a memory and a processor, where the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the music emotion recognition method based on cross-modal fusion in Embodiment 1.
[0102] Optionally, the above electronic device may be a server.
[0103] In addition, an embodiment of the present invention further provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it implements the music emotion recognition method based on cross-modal fusion in Embodiment 1.
[0104] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0105] Embodiment 4
[0106] The method proposed by the present invention is trained and tested on the publicly available MSDD (Million Song Dataset annotated by Deezer)[4] dataset. The total number of data is 2000 tracks. 60% of the data is used as the training set, 20% as the validation set, and 20% as the test set.
[0107] The experimental parameter settings are as follows: The number of layers of the 2D-CNN in the audio feature extraction part is set to 5 layers. A max-pooling layer with a window size of 2*2 is set after each convolutional layer. The 5 convolutional layers have 32, 64, 128, 256, and 128 convolutional kernels respectively, and the convolutional kernel size is 3*3 with a stride of 1. The text feature extraction part uses a pre-trained BERT model, the BERT-base model pre-trained on an English corpus, and this model has 12 layers of transformer encoders. The image feature extraction part loads the resnet50 model pre-trained on the ImageNet dataset. In the processing part of the multi-modal fusion features, the original 6-layer transformer encoder structure proposed in [1] is used.
[0108] The method proposed in this invention achieves an R^2 metric of 0.352 on the MSDD dataset, representing a significant performance improvement of 0.043 compared to the existing SoTA method [5]. The experimental results demonstrate that the method proposed in this invention has better recognition performance compared to the existing methods.
[0109] References
[0110] [1] Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need[J]. Advances in neural information processing systems, 2017, 30.
[0111] [2] Devlin J, Chang M W, Lee K, et al. Bert: Pre-training of deep bidirectional transformers for language understanding[J]. arXiv preprint arXiv:1810.04805, 2018.
[0112] [3] He K, Zhang X, Ren S, et al. Deep residual learning for image recognition[C] / / Proceedings of the IEEE conference on computer vision and pattern recognition. 2016: 770 - 778.
[0113] [4] Delbouys R, Hennequin R, Piccoli F, et al. Music Mood Detection Based On Audio And Lyrics With Deep Neural Net[J]. 2018.
[0114] [5] Zhao J, Ru G, Yu Y, et al. Multimodal music emotion recognition with hierarchical cross-modal attention network[C] / / 2022 IEEE International Conference on Multimedia and Expo(ICME). IEEE, 2022: 1 - 6.
Claims
1. A music emotion recognition method based on cross-modal fusion, characterized in that, The specific steps are as follows: (1) Obtain the music information of the music to be recognized; the music information includes an audio file, lyrics, song name, singer name, album name, and album cover; (2) Extract the music emotion features of the music information; the music emotion features include audio features, text features, and image features; the audio features are the features of the audio file; the text features include the features of the lyrics, the song name, the singer name, and the album name; the image features are the features of the album cover; (3) Preprocess the music emotion features to obtain a feature fusion sequence with positional encoding; (4) Input the feature fusion sequence into a pre-trained Transformer Encoder to obtain the fused music emotion features; (5) Input the fused music emotion features into a pre-trained fully connected classifier to obtain the predicted music emotion categories; the music emotion categories include happy, sad, angry, and relaxed; The extraction of the music emotion features of the music information in step (2) specifically includes: Calculate the Mel spectrogram of the audio file; Adjust the size of the spectrogram of the Mel spectrogram to a preset size; Input the adjusted spectrogram into a pre-trained 2D-CNN model for audio feature extraction to obtain audio features; The 2D-CNN (2-Dimensional Convolutional Neural Network) includes a convolutional layer and a pooling layer; the convolutional layer includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, and a fifth convolutional layer; the pooling layer includes a first pooling layer, a second pooling layer, a third pooling layer, a fourth pooling layer, and a fifth pooling layer; The first convolutional layer, the first pooling layer, the second convolutional layer, the second pooling layer, the third convolutional layer, the third pooling layer, the fourth convolutional layer, the fourth pooling layer, the fifth convolutional layer, and the fifth pooling layer are connected in sequence; The convolutional kernel size of the convolutional layer is 3*3; the pooling window size of the pooling layer is 2*2; Concatenate the lyrics, the song name, the singer name, and the album name to obtain a concatenated text; Calculate the word embedding representation of the concatenated text; Input the word embedding representation into a pre-trained BERT model to obtain text features; Input the album cover into a pre-trained Resnet50 model to obtain image features; The preprocessing of the music emotion features in step (3) to obtain a feature fusion sequence with positional encoding specifically includes: Convert the music emotion features into a one-dimensional vector; Slice the music emotion features converted into a one-dimensional vector to obtain a feature sequence with equal-length segments; Perform positional encoding on the segments in the feature sequence with equal-length segments to obtain a feature fusion sequence with positional encoding.
2. The music emotion recognition method according to claim 1, characterized in that The format of the audio file is the wav format; the format of the album cover is the jpg format.
3. The music emotion recognition method according to claim 1, wherein The Transformer Encoder described in step (4) is a 6-layer Transformer Encoder; the Transformer mechanism consists of a multi-head self-attention mechanism and a feed-forward network; the feature fusion sequence is input into the pre-trained 6-layer Transformer Encoder, and the output is the feature after multi-modal fusion, that is, the fused music emotion feature.
4. The music emotion recognition method according to claim 1, wherein Step (5) inputs the fused music emotion feature into the pre-trained fully-connected classifier to obtain the predicted music emotion category, specifically including: Inputting the fused music emotion feature into the pre-trained fully-connected classifier to obtain the predicted probabilities of each music emotion category; according to the predicted probabilities of each music emotion category, obtaining the predicted music emotion category.
5. A music emotion recognition system based on cross-modal fusion for the music emotion recognition method according to any one of claims 1-4, characterized in that, Including: An acquisition module for acquiring the music information of the music to be recognized; The music information includes an audio file, lyrics, song name, singer name, album name, and album cover; An extraction module for extracting the music emotion features of the music information; the music emotion features include audio features, text features, and image features; the audio features are the features of the audio file; the text features include the features of the lyrics, the song name, the singer name, and the album name; the image features are the features of the album cover; A preprocessing module for preprocessing the music emotion features to obtain a feature fusion sequence with position encoding; A fusion module for inputting the feature fusion sequence into the pre-trained Transformer Encoder to obtain the fused music emotion feature; A prediction module for inputting the fused music emotion feature into the pre-trained fully-connected classifier to obtain the predicted music emotion category; the music emotion categories include happy, sad, angry, and relaxed; These 5 modules perform the operations of the 5 steps of the music emotion recognition method based on cross-modal fusion.
6. An electronic device, characterized in that, Including a memory and a processor, the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the music emotion recognition method based on cross-modal fusion according to any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that, It stores a computer program, and when the computer program is executed by the processor, it implements the music emotion recognition method based on cross-modal fusion according to any one of claims 1 to 4.
Citation Information
Patent Citations
Audio and video multi-mode sentiment classification method and system
CN113408385A
Multi-modal sentiment analysis method and system based on Transformer and multi-task learning
CN114091466A