Song emotion feature modeling and recommending method based on large-scale multi-modal model
By constructing lyrics emotional summary and information fusion prompt sentences, combining large multimodal models and user historical behavior data, dynamically adjusting user preferences and emotional feature weights, solving the problem of insufficient emotional feature extraction in song recommendations, realizing personalized song recommendations, and improving user experience.
Patent Information
- Application Number
- CN202510496542.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-08-01
AI Technical Summary
The existing song recommendation method based on large multimodal models ignores the emotional attributes of the song and cannot accurately identify the user's short-term emotional characteristics and long-term music preferences, resulting in the recommendation system mistakenly recommending emotional songs when the user's mood fluctuates, affecting the user experience.
By constructing lyrics, emotional summary prompt sentences and information fusion prompt sentences, a large multimodal model is used to extract the emotional characteristics of the song, and combining user historical behavior data, an attention mechanism and emotional degree evaluation module are introduced, and the weight of user preferences and emotional characteristics is dynamically adjusted to achieve personalized song recommendations.
Accurately perceive the emotional characteristics of the song, improve the emotional matching of the recommendation system, avoid misrepresenting emotional songs, and improve the user's listening experience.
Smart Images

Figure CN120407843A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data mining and recommendation, and specifically to a method for modeling song emotional features and recommendation based on a large multi-modal model. Background Art
[0002] In today's digital age, the rapid development of the Internet and the widespread application of big data technology have greatly promoted the research and practice of recommendation systems. As an intelligent information filtering tool, the recommendation system aims to provide personalized content recommendations for users by analyzing their historical behavior data and preference information, thereby improving the efficiency and experience of users' information acquisition. With the continuous increase in the number of Internet users and the massive accumulation of user behavior data, the recommendation system faces unprecedented opportunities and challenges. On the one hand, the rich data provides a basis for more accurate modeling of user preferences for the recommendation system; on the other hand, the diversification and dynamics of user needs require the recommendation system to be able to perceive and adapt to the changes in users' interests in real time. In the field of music recommendation, traditional recommendation methods mainly rely on collaborative filtering technology, which mines users' potential preferences by analyzing the interaction behavior between users and songs. However, these methods often ignore the emotional attributes of the songs themselves and the emotional needs of users in different situations.
[0003] In recent years, with the rise of large multi-modal models, the recommendation system has been empowered with more powerful semantic understanding and content perception capabilities, which also brings new opportunities to the field of music recommendation. By analyzing the text information of songs such as lyrics, music styles, and singers, as well as the cover pictures and audio information of songs, large multi-modal models can provide rich content features of songs for the recommendation system, thereby enhancing the recommendation performance. However, in the existing methods for extracting song content features combined with large multi-modal models, the emotional understanding ability of large multi-modal models for songs is often ignored, and content feature modeling containing emotional analysis is carried out. How to utilize the powerful semantic understanding ability, image and audio information processing ability, and multi-modal information fusion ability of large multi-modal models to extract content information containing the emotional tendency of songs and perform song emotional feature modeling for the special property of songs having emotional attributes different from other recommendation scenarios is the first technical problem.
[0004] In the user-side feature modeling of music recommendation, how to accurately model the short-term emotional features and long-term music preference features of users based on their historical behavior data and recommend songs that not only meet their music style preferences but also satisfy their current emotional needs is the second technical problem.
[0005] When users are in a calm state, their music listening needs usually tend to their own music genre preferences. However, when listening to music in an emotional state, they need to pay more attention to their emotional needs. How to make the recommendation model combine traditional ID features such as user portrait information with the user's historical behavior sequence, accurately identify the current user's emotional state, distinguish whether the user listens to music with emotional needs, and prevent the misrecommendation of emotional songs from affecting the user's music listening experience is the third technical problem. Summary of the Invention
[0006] The purpose of the present invention is to overcome the deficiencies of the prior art and propose a method for modeling song emotional features and recommendation of a large multi-modal model, so as to more accurately and finely perceive song emotional features, and realize song recommendations that conform to the user's mood and music preferences according to the song sequence of the user's historical interactions.
[0007] In order to achieve the above purpose, the specific technical solutions adopted by the present invention are as follows:
[0008] A method for modeling song emotional features and recommendation based on a large multi-modal model, including the following steps:
[0009] (1) Collect user music listening records and multi-modal content data of all songs.
[0010] The user song interaction records include the song records that the user has fully played, the song records marked as liked, and the corresponding playing time of each record; the song multi-modal content data includes images (such as the cover picture of the song), audio, and text (such as artist / singer information, album information, style information, language information, and lyrics).
[0011] (2) Utilize the text, image, and audio information understanding and fusion capabilities of the large multi-modal model, and based on the music content data in step (1), model the song emotional content features, specifically including:
[0012] (2.1) Extract the lyrics of each song as text information, construct a summary prompt statement for the emotional content of the lyrics, input it into the large multi-modal model, and generate a concise summary text of the lyrics' emotions. The summary prompt statement for the emotional content of the lyrics is in the following format: "Summarize the emotions and emotional content expressed by the lyrics of this song, less than 20 words. Lyrics: {the lyrics text of the song}". Thus, obtain the {lyrics emotion summary text} output by the large multi-modal model.
[0013] (2.2) Combine the lyric emotion summary text with other text information of the song (such as song name, singer name, language, music genre type, etc.) to form an information fusion prompt statement, and input it together with the song's picture and audio data into a large multi-modal model. The format of the information fusion prompt statement is: "Based on the picture, audio, and song introduction, model the song emotion features. The song introduction is as follows: The song name is: {xxx}, the singer is: {xxx}, the region to which the song belongs is: {xx}, the language is: {xx}, the primary music genre is {xx}, the secondary music genre is {xx}, the instrument is {xx}, {lyric emotion summary text}."
[0014] (2.3) Extract the hidden vectors of each token from the last layer of the large multi-modal model, and perform average pooling operation to obtain the song's emotion content features. The calculation formula is:
[0015]
[0016] where: E m is the song's emotion content feature vector, N is the number of tokens in the hidden vectors of the last layer of the large multi-modal model, and E token is the hidden vector of each token.
[0017] (3) Based on the user's song interaction records, according to the playing time, construct the songs marked as liked by the user into a user preference sequence, and construct the songs completely played by the user into a user emotion sequence, and through the attention mechanism, model the user's preference features and emotion features, specifically including:
[0018] (3.1) Extract the song records marked as liked by the user from the user's song interaction records, sort them by the interaction time, and form a user preference sequence where k is the length of the user preference sequence, is the index number of the corresponding song.
[0019] (3.2) Extract the song records completely played by the user from the user's song interaction records, sort them by the interaction time, and form a user emotion sequence where l is the length of the user emotion sequence, is the index number of the corresponding song.
[0020] (3.3) Through the attention mechanism, calculate the emotion content features of the target song and the similarity scores between the emotion content features of each song in the user preference sequence and the user emotion sequence. The calculation formula is as follows:
[0021]
[0022] where S irepresents the similarity score between the target song and the i-th song, is the emotional content feature vector of a song in the user preference sequence or user emotion sequence. H(·) represents element-wise multiplication, D(·) represents element-wise subtraction, ReLU(·) is the ReLU activation function, W is the weight matrix, T represents the transpose operation, and concat(·) represents the vector concatenation operation.
[0023] (3.4) According to the similarity scores, weighted summations are respectively performed on the two sequences to obtain the user's preference feature (denoted as U p ) and emotion feature (denoted as U q ), and the calculation formulas are respectively:
[0024]
[0025] where: represents the emotional content feature vector of the i-th song in the sequence, K is the length of the user preference sequence, and L is the length of the user emotion sequence.
[0026] (4) Introduce an emotional degree evaluation module. Based on the user portrait information and emotion features, comprehensively evaluate the user's current emotional degree to obtain the emotional degree evaluation factor β, so as to dynamically adjust the weights of the model for the user preference features and emotion features. The specific calculation formula for the emotional degree evaluation factor is as follows:
[0027]
[0028] where σ(·) is the sigmoid activation function, W p and W q are weight matrices, and E p is the portrait feature obtained after encoding the user's gender, age and other information.
[0029] (5) Based on the user's current emotional degree evaluation factor β obtained in step (4), dynamically adjust the weights of the user's preference feature U p and emotion feature U q , and concatenate them with the user portrait feature E p and the emotional content feature of the target song , and jointly input them into a multi-layer perceptron (MLP) for feature mapping, and finally obtain the logical score of the user for the target song. The calculation formula is as follows:
[0030]
[0031] where: MLP is the multi-layer perceptron.
[0032] The final model sorts all candidate target songs according to the logical scores and extracts the top several songs to recommend to the user.
[0033] The present invention has the following characteristics and beneficial effects:
[0034] For the first time, the present invention utilizes the powerful semantic understanding ability, image processing ability and multi-modal information fusion ability of a large multi-modal model to combine text content information such as lyrics, music styles, and singers with image and audio information, and extracts the emotional features of songs. By constructing a lyric emotion summary prompt statement and an information fusion prompt statement, the large multi-modal model can accurately capture the emotional tendency of songs, solve the problem of insufficient accuracy in emotional feature extraction by traditional methods, provide richer and finer-grained song emotional features for the music recommendation system, and improve the emotional matching degree of the recommendation service; the present invention distinguishes different types of user listening records and uses the attention mechanism to realize the joint modeling of user preference features and emotion features, making up for the deficiency of traditional music recommendation methods in user emotion modeling; the present invention innovatively introduces an emotional degree evaluation module to dynamically evaluate the user's current emotional state and adjust the weights of user preference features and emotion features, so that the system can recommend songs that conform to the user's music preferences when the user is calm, and recommend songs that meet the user's emotional needs when the user's emotions fluctuate. This mechanism effectively avoids the situation of mis-recommending emotional songs and improves the user's music listening experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 It is a schematic diagram of the system architecture of the music recommendation method of the present invention.
[0036] Figure 2 It is a schematic diagram of the user music demand prediction process in the music recommendation method of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0037] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0038] A method for modeling and recommending song emotional features of a large multi-modal model, as Figure 1 shown, includes the following steps:
[0039] Step 1, collect user song interaction records and multi-modal content data of all songs, where the user song interaction records include the song records that the user has fully played, the song records marked as liked and the corresponding play times; the multi-modal content data of the songs includes image, audio and text information;
[0040] Among them, the image generally refers to the cover picture of the song, and the text information includes singer information, album information, style information, language information, and lyrics.
[0041] Step 2: Use a large multi-modal model to understand and fuse the text, image, and audio information of the song, and model the emotional content features of the song;
[0042] Step 2-1: Extract the lyrics of each song, construct a summary prompt statement of the emotional content of the lyrics, input it into the large multi-modal model, and generate a summary text of the lyrics' emotions;
[0043] In this embodiment, taking the song "Nocturne" as an example, the summary prompt statement of its lyrics' emotions is:
[0044] "Summarize the emotions and emotional content expressed in the lyrics of this song, less than 20 characters. Lyrics: {lyric text of Nocturne}".
[0045] Thus, the summary text of the lyrics' emotions output by the large multi-modal model may be:
[0046] "{The emotional content expressed by this song is melancholy, loneliness, and nocturnal contemplation}".
[0047] Step 2-2: Combine the generated summary text of the lyrics' emotions with other text information to form an information fusion prompt statement;
[0048] Furthermore, taking "Nocturne" as an example, the format of its information fusion prompt statement is:
[0049] "Based on the picture, audio, and song introduction, model the emotional features of the song. The song introduction is as follows: The song name is: Nocturne, the singer is: {×××}, the region to which the song belongs is: {Taiwan Province, China}, the language is: {Chinese}, the primary music genre is {Pop}, the secondary music genre is {R&B}, the instrument is {piano}, {The emotional content expressed by this song is melancholy, loneliness, and nocturnal contemplation}.".
[0050] Step 2-3: Input the obtained information fusion prompt statement together with the image and audio data into the large multi-modal model to obtain the emotional content features of the song.
[0051] Specifically, the extraction method of the emotional content features of the song is:
[0052] The large multi-modal model extracts the hidden layer of the last layer to obtain the hidden vector of each token, and then performs average pooling operation:
[0053]
[0054] where E m is the emotional content feature vector of the song, N is the number of tokens, Etoken is the token hidden vector.
[0055] It can be understood that for the text input into the large multi-modal model, it will first divide it into individual words, such as English words like "apple" and "like", or Chinese words like "eating", "sleeping", and "washing clothes", which are called tokens. For modalities such as images and audio, the original data will be compressed or sampled, and the compressed elements will be regarded as tokens.
[0056] It can be understood that the emotional content features E of all songs m will be stored in a feature database for subsequent user feature modeling and recommendation modules to use.
[0057] It should be noted that in this embodiment, M-CLIP is selected as the large multi-modal model mentioned above.
[0058] It should be noted that the choice of large multi-modal models is diverse, and any large multi-modal model with the ability to extract features of image, text, and audio modalities can replace the large multi-modal model referred to in this embodiment. Such as the Flava large multi-modal model, etc.
[0059] The deployment and invocation of large multi-modal models belong to conventional technical means, and will not be specifically described in this embodiment.
[0060] Step 3: Construct a user preference sequence and a user emotion sequence based on the user's historical play records, and calculate the similarity scores between the target song and the songs in the user preference sequence and the user emotion sequence through the attention mechanism, respectively generating user preference features and user emotion features;
[0061] Specifically, it includes the following sub-steps:
[0062] Step 3-1: Extract the song records marked as liked by the user from the user's song interaction records, and sort them by interaction time as the user preference sequence where k is the length of the user preference sequence, is the index number of the corresponding song.
[0063] Step 3-2: Extract the song records that the user has fully played from the user's song interaction records, and sort them by interaction time as the user emotion sequence where l is the length of the user emotion sequence, is the index number of the corresponding song.
[0064] Step 3-3: Through the attention mechanism, calculate the emotional content features of the target song and the emotional content features of each song in the user preference sequence and the user emotion sequence (denoted as ) The similarity score is calculated as follows:
[0065]
[0066] Where S i represents the similarity score between the target song and the i-th song. is the emotional content feature vector of a song in the user preference sequence or user emotion sequence. H(·) represents element-wise multiplication, D(·) represents element-wise subtraction, ReLU(·) is the ReLU activation function, W is the weight matrix, T represents the transpose operation, and concat(·) represents the vector concatenation operation.
[0067] Step 3-4: According to the similarity scores between each song in the sequence and the target song, perform weighted summation operations on the two sequences respectively to obtain the user's preference feature (denoted as U p ) and emotion feature (denoted as U q ). The calculation formulas are as follows:
[0068]
[0069] Where: represents the emotional content feature vector of the i-th song in the sequence. K is the length of the user preference sequence, and L is the length of the user emotion sequence.
[0070] Step 4: Based on the user portrait features and the user emotion features, calculate the emotional degree evaluation factor β through the emotional degree evaluation module, and dynamically adjust the weights of the user preference features and user emotion features by applying the emotional degree evaluation factor;
[0071] Specifically, the calculation formula for the emotional degree evaluation factor is as follows:
[0072]
[0073] Where σ(·) is the sigmoid activation function, W p and W q are weight matrices, and E p is the portrait feature obtained after encoding the user's gender, age, etc.
[0074] In this embodiment, the encoding method uses one-hot encoding (also known as one-hot encoding), which is then multiplied by the embedding layer to obtain the feature vector corresponding to each attribute information. For example, user gender includes three types: [male, female, unknown]. Then, for a male user, the user profile feature obtained after one-hot encoding is a vector [1,0,0]. This feature vector is input into the deep learning model to obtain the corresponding embedding. For example, multiplying a 3*64 matrix by [1,0,0] will produce a feature vector with a length of 64.
[0075] It can be understood that the dynamic adjustment method is: taking the emotionality evaluation factor β as the coefficient of the user emotional feature and (1-β) as the coefficient of the user preference feature, and then splicing and inputting them into the final MLP.
[0076] Step 5: Based on the user's current emotionality evaluation factor β obtained in step (4), dynamically adjust the user's preference feature U p and emotional characteristics U q The weight of the user portrait feature E p , emotional content characteristics of the target song The results are concatenated and fed into a multi-layer perceptron (MLP) for feature mapping, and the user's logical score for the target song is finally obtained. The calculation formula is as follows:
[0077]
[0078] Among them: MLP is a multi-layer perceptron.
[0079] The final model ranks all candidate target songs according to the logical scores and extracts the top several songs to recommend to users.
[0080] Preferably, the text information includes singer information, album information, style information, language information and lyrics.
[0081] Figure 1The architecture of the song emotion feature modeling and recommendation method based on a large multi-modal model in this embodiment is shown. This recommendation method is divided into two main modules: the song emotion content feature modeling module and the song recommendation module. In the song emotion content feature modeling module, first, the multi-modal content information of the song is extracted, including text, images, and audio, and the large multi-modal model is used to summarize the emotion of the lyrics to generate a concise lyric emotion summary text; then the emotion summary text is fused with other text information of the song (such as song name, singer name, language, etc.) to form an information fusion prompt statement, which is input into the large multi-modal model together with the image and audio data, and the emotion content features of the song are extracted. In the song recommendation module, the emotion sequence and preference sequence of the user are constructed, and the similarity scores between the songs in the sequence and each target song in the candidate target song library are calculated to model the user's preference features and emotion features. Finally, the system combines the user portrait information and the emotion level evaluation module to dynamically adjust the weights of the preference features and emotion features, calculates the logical scores of the user and each candidate target song, and is used for the selection and ranking of the recommended songs to achieve personalized music recommendations. Figure 2 The detailed steps of the joint modeling of the user's emotion features and preference features are shown. First, the system extracts the song records marked as liked by the user from the user's historical interaction records to construct the user preference sequence; at the same time, it extracts the song records that the user has fully played recently to construct the user emotion sequence. Then, through the attention mechanism, the similarity scores between the target song and each song in the preference sequence and emotion sequence are calculated respectively, and the weighted sum is used to obtain the user's preference features and emotion features.
[0082] Embodiment
[0083] The present invention is compared with several traditional recommendation models and two music models for song emotion-based recommendations based on the music dataset Music4all, where: (1) DNN model: a neural network based on a multi-layer perceptron; (2) DeepFM model: a neural network based on a factorization machine and a multi-layer perceptron; (3) DIN model: a recommendation model that models the user's historical behavior sequence based on the attention mechanism; (4) PEIA model: an emotion-based music recommendation model based on the attention mechanism and a multi-layer perceptron; (5) EPMR model: an emotion-based music recommendation model that models the user's historical behavior sequence based on a time recurrent neural network.
[0084] Table 1 shows the comparison results of the method provided in this embodiment with other models
[0085] Model Name AUC Index Logloss Index DNN 0.6019 1.147 DeepFM 0.6103 1.159 DIN 0.6091 0.615 PEIA 0.6130 1.085 EPMR 0.6167 1.163 The method of the present invention 0.6426 0.609
[0086] As shown in Table 1, the AUC metric is a commonly used evaluation metric for measuring the accuracy of predicting user behavior in the field of recommendation. Its value range is between 0 and 1, and the closer it is to 1, the better the prediction effect. The Logloss metric is used to measure the convergence of the model, and its value range is from 0 to positive infinity. The closer it is to 0, the better the convergence effect of the model. The AUC metric of the method of the present invention is improved by 4.2% compared with the sub-optimal model EPMR; the Logloss metric is improved by 0.1% compared with the sub-optimal model DIN.
[0087] The above-described embodiments are only a preferred solution of the present invention, but they are not intended to limit the present invention. Those of ordinary skill in the relevant technical field can still make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all technical solutions obtained by adopting the means of equivalent replacement or equivalent transformation fall within the protection scope of the present invention.
[0088] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not intended to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for modeling song emotion features and recommendation based on a large multi-modal model, characterized in that It includes the following steps: Step 1: Collect user song interaction records and multimodal content data of all songs. The user song interaction records include records of songs that the user has fully played, records of songs marked as liked, and the corresponding playing times. The multimodal content data of the songs includes image, audio, and text information. Step 2: Use a large multimodal model to understand and fuse the text, image, and audio information of the songs, and model the emotional content features of the songs. Step 3: Construct a user preference sequence and a user emotion sequence based on the user's historical play records, and calculate the similarity scores between the target song and the songs in the user preference sequence and user emotion sequence through an attention mechanism, and generate user preference features and user emotion features respectively. Step 4: Based on the user portrait features and the user emotion features, calculate the emotional degree evaluation factor through an emotional degree evaluation module, and dynamically adjust the weights of the user preference features and user emotion features by applying the emotional degree evaluation factor. Step 5: Concatenate the adjusted user preference features, user emotion features, user portrait features, and the emotional content features of the target song, and input them into a multi-layer perceptron to calculate the logical score of the user for the target song, and generate a recommendation list according to the score ranking.
2. The method according to claim 1, wherein The text information includes singer information, album information, style information, language information, and lyrics.
3. The method according to claim 1, characterized in that, Step 2 includes the following sub-steps: Step 2-1: Extract the lyrics of each song, construct a summary prompt statement of the emotional content of the lyrics, input it into the large multimodal model, and generate a summary text of the emotional content of the lyrics. Step 2-2: Combine the generated summary text of the emotional content of the lyrics with other text information to form an information fusion prompt statement. Step 2-3: Input the obtained information fusion prompt statement together with the image and audio data into the large multimodal model to obtain the emotional content features of the song.
4. The method according to claim 1, wherein In Step 2-3, the extraction method of the emotional content features of the song is as follows: The large multimodal model uses the last hidden layer to extract the hidden vectors of each token and performs an average pooling operation: Among them, E m is the emotional content feature vector of the song, N is the number of tokens, and E token is the token hidden vector.
5. The method according to claim 1, wherein In Step (3), extract the records of songs marked as liked by the user from the user song interaction records, sort them by interaction time, and form a user preference sequence; extract the records of songs that the user has fully played from the user song interaction records, sort them by interaction time, and form a user emotion sequence.
6. The method according to claim 1, wherein In Step 3, the calculation formula for the similarity score is: Among them, is the emotional content feature vector of the target song, is the emotional content feature vector of each song in the user preference sequence or user emotion sequence. H(·) represents element-wise multiplication, D(·) represents element-wise subtraction, ReLU(·) is the ReLU activation function, W is the weight matrix, T represents the transpose operation, and concat(·) represents the vector concatenation operation.
7. The method according to claim 6, characterized in that In Step 3, according to the similarity scores, perform weighted summation on the two sequences respectively to obtain the user's preference features and emotion features. The calculation method is as follows: Among them, U p represents the user preference feature, U q represents the user emotion feature, represents the emotional content feature vector of the i-th song in the sequence. K is the length of the user preference sequence, and L is the length of the user emotion sequence.
8. The method according to claim 1, wherein The user portrait features are obtained through one-hot encoding according to the user portrait information.
9. The method according to claim 1, wherein In Step 4, the acquisition method of the emotional degree evaluation factor is: Among them: β is the evaluation factor of the user's current emotional level, σ(·) is the sigmoid activation function, W p and W q are weight matrices, and E p is the user portrait feature.
10. The method according to claim 9, characterized in that, In the said step 5, the logical score is calculated according to the formula: Where MLP is a multi-layer perceptron.
11. The method according to claim 8, characterized in that The user portrait information includes at least one of gender and age.
12. The method according to claim 1, wherein The large multimodal model is a pre-trained deep learning model that can process text, image, and audio information and perform multimodal fusion.
Citation Information
Cited By
Intelligent customer service dynamic intention recognition system based on semantic analysis model
CN120745648A