Song recommendation method and device and storage medium

By obtaining and analyzing user's conversation content, using search engines and feature extraction models to recommend songs, the problem of low matching between song recommendations and user needs in the prior art is solved, and a more efficient song recommendation effect is achieved.

CN120216719APending Publication Date: 2025-06-27TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510278562.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing song recommendation method cannot recommend songs based on the current needs of users, and the recommended songs match the user's needs less.

Method used

By obtaining the conversation content currently entered by the user's account and determining that it belongs to the song recommendation intention, search for matching song extension information using search engines, and use the dialogue content and historical dialogue content as context information, enter a pre-trained feature extraction model, extract target text features to represent user emotions and song information, and finally recommend songs to the user based on these features.

Benefits of technology

It realizes real-time recommendation of songs that meet user needs based on the conversation content currently entered by the user, improving the matching degree and user experience of recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216719A_ABST
    Figure CN120216719A_ABST
Patent Text Reader

Abstract

The invention provides a song recommendation method and device and a storage medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining the current input dialogue content of a user account, the current input dialogue content comprising the basic information of a song; when it is determined that the currently input dialogue content belongs to the song recommendation intention, song extension information matched with the basic information of the current song is searched through a search engine; taking the currently input dialogue content and at least one piece of historical dialogue content associated with the currently input dialogue content as context information; inputting the context information and the song extension information into a pre-trained feature extraction model to obtain target text features; the target text feature is used for representing user emotion and / or song information; and determining a to-be-recommended song based on the target text feature, and displaying related information of the to-be-recommended song to the user account. By adopting the method, songs meeting the current requirements of the user can be recommended to the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of machine learning technology, and particularly to a method, device, and storage medium for song recommendation. Background Art

[0002] Song recommendation is an important means of song promotion. The common song recommendation methods are to set song recommendation sections in the interface of music applications, such as "Guess You Like", "30 Songs a Day", etc.

[0003] However, the songs recommended by the current song recommendation methods are fixed, and mostly recommend the songs that the user may want to listen to today based on the user's historical listening behavior.

[0004] The above song recommendation methods cannot recommend songs according to the user's current needs, and the matching degree between the recommended songs and the user's needs may be low. Summary of the Invention

[0005] Embodiments of this application provide a method, device, and storage medium for song recommendation, which can recommend songs that more meet the user's current needs. The technical solutions are as follows:

[0006] In a first aspect, a method for song recommendation is provided, characterized in that the method includes:

[0007] Obtain the conversation content currently input by the user account, where the currently input conversation content includes song basic information;

[0008] When it is determined that the currently input conversation content belongs to the song recommendation intention, search for song extension information that matches the song basic information included in the currently input conversation content through a search engine;

[0009] Use the currently input conversation content and at least one historical conversation content associated with the currently input conversation content as context information;

[0010] Input the context information and the song extension information into a pre-trained feature extraction model to obtain target text features; the target text features are used to characterize user emotions and / or song information;

[0011] Determine the songs to be recommended based on the target text features, and display the relevant information of the songs to be recommended to the user account.

[0012] In a possible implementation, the obtaining of the conversation content currently input by the user account includes:

[0013] Input the current conversation content entered by the user account into the intent recognition model to obtain the intent indication information output by the intent recognition model; wherein the intent indication information is used to indicate whether the current conversation content entered belongs to the song recommendation intent.

[0014] In a possible implementation, the querying the search engine for extended information matching the current conversation content entered includes:

[0015] Extract the basic song information in the current conversation content entered and input it into the Retrieval-Augmented Generation (RAG) model, so that the RAG model calls the search engine to query the song search results matching the basic song information;

[0016] Select at least one song search result that meets the preset conditions from the song search results, and use the summary information of the at least one song search result as the song extended information.

[0017] In a possible implementation, the method further includes: obtaining the song preference information corresponding to the user account; the song preference information is the text features extracted by the feature extraction model according to the historical conversation content of the user account;

[0018] The inputting the context information and the song extended information into the pre-trained feature extraction model to obtain the target text features includes: inputting the context information, the song extended information, and the song preference information into the pre-trained feature extraction model to obtain the target text features.

[0019] In a possible implementation, the feature extraction model is a large language model; the inputting the context information and the song extended information into the pre-trained feature extraction model to obtain the target text features includes:

[0020] Generate a specific type of guiding statement, which is used to indicate that the input of the large language model is the context information and the song extended information and to indicate that the output of the large language model is the target text features of the specific type and the output format of the target text features;

[0021] Input the specific type of guiding statement into the pre-trained large language model to obtain the target text features of the specific type and conforming to the output format.

[0022] In a possible implementation, the output format of the target text features is the JSON format, the key of the JSON format target text features is the identifier of the specific type, and the key of the JSON format target text features is the text features corresponding to the specific type.

[0023] In a possible implementation, searching, by a search engine, for song extension information that matches the basic song information included in the currently input conversation content includes:

[0024] Inputting the currently input conversation content into an intent recognition model, and the intent recognition model outputs a target content type;

[0025] In the case where the target content type is used to indicate that the currently input conversation content includes knowledge-based information, searching, by a search engine, for song extension information that matches the basic song information included in the currently input conversation content.

[0026] In a possible implementation, inputting the context information and the song extension information into a pre-trained feature extraction model to obtain target text features includes:

[0027] Inputting the context information and the song extension information into a pre-trained feature extraction model;

[0028] Determining, by the pre-trained feature extraction model, whether the context information contains information for characterizing the user's emotion. If it contains information for characterizing the user's emotion, extracting the information for characterizing the user's emotion from the context information;

[0029] Determining, by the pre-trained feature extraction model, whether the context information contains information for characterizing song information. If it contains information for characterizing song information, extracting the information for characterizing song information from the context information;

[0030] Using, by the pre-trained feature extraction model, the information for characterizing the user's emotion, the information for characterizing song information, and the target song information included in the song extension information as target text features.

[0031] In a possible implementation, the song information includes at least one of song type, lyric description, song language, suitable scene of the song, singer name, song name, and album name.

[0032] In a second aspect, a song recommendation device is provided. The device includes at least one module, and the at least one module is configured to perform the operations executed by the song recommendation method as described in the first aspect and any possible implementation of the first aspect above.

[0033] In a third aspect, a computing device is provided. The computing device includes a processor and a memory. At least one instruction is stored in the memory, and the instruction is loaded and executed by the processor to implement the operations executed by the song recommendation method as described in the first aspect and any possible implementation of the first aspect above.

[0034] In a fourth aspect, a computer-readable storage medium is provided, in which at least one instruction is stored, and the instruction is loaded and executed by a processor to implement the operations performed by the method for song recommendation as described in the first aspect above and any possible implementation of the first aspect.

[0035] In a fifth aspect, a computer program product is provided, in which at least one instruction is stored, and the instruction is loaded and executed by a processor to implement the operations performed by the method for song recommendation as described in the first aspect above and any possible implementation of the first aspect.

[0036] The beneficial effects brought by the technical solution provided in this application are as follows:

[0037] In the technical solution provided in the embodiments of this application, when a user inputs conversation content and the conversation content triggers a song recommendation behavior, a search engine is used to expand the conversation content, and the expanded information and the conversation content input by the user are jointly used as the input of a feature extraction model. The feature extraction model extracts text features, which not only include the features of the conversation content but also the features of the expanded information, and can perform extended recommendations as much as possible on the premise of combining the user's needs. Furthermore, relevant information of the to-be-recommended songs that match the text features is obtained from the song library, and a reply content is generated and replied to the user. It can be seen that in this solution, songs are recommended to the user in combination with the conversation content input by the user in real time, and the recommended songs meet the user's current needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0039] Figure 1 is a schematic diagram of the interface of a music application program provided by an embodiment of this application;

[0040] Figure 2 is a schematic diagram of the interface of a music application program provided by an embodiment of this application;

[0041] Figure 3 is a flowchart of the method for song recommendation provided by an embodiment of this application;

[0042] Figure 4 is a schematic diagram of the interface of a music application program provided by an embodiment of this application;

[0043] Figure 5 It is a schematic diagram of text feature extraction provided by an embodiment of the present application;

[0044] Figure 6 It is a schematic diagram of a method for text feature extraction provided by an embodiment of the present application;

[0045] Figure 7 It is a schematic diagram of the interface of a music application provided by an embodiment of the present application;

[0046] Figure 8 It is a schematic diagram of the device structure for song recommendation provided by an embodiment of the present application;

[0047] Figure 9 It is a schematic diagram of the structure of a terminal provided by an embodiment of the present application;

[0048] Figure 10 It is a schematic diagram of the structure of a server provided by an embodiment of the present application. Detailed implementation manners

[0049] To make the objectives, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.

[0050] An embodiment of the present application provides a method for song recommendation. This method can be implemented by a computing device, which can be a terminal or a server. Among them, the terminal can be a mobile phone, a tablet computer, a desktop computer, a laptop computer, etc., and the server can be a single server, a server cluster, etc.

[0051] The implementation scenario of the song recommendation method provided by the embodiment of the present application will be described below:

[0052] The user terminal can be installed with a music application. In the interface of the music application, dialogue options can be displayed, such as Figure 1 shown. After the user selects this dialogue option and enters the dialogue mode, the user terminal displays a dialogue box as shown in Figure 2 shown. The user can input dialogue content in ways such as voice and text. Furthermore, the user terminal obtains the dialogue content and, based on the dialogue content, uses the song recommendation method provided by the embodiment of the present application to recommend songs to the user, or can also send the dialogue content to the server. The server, based on the dialogue content, uses the song recommendation method provided by the embodiment of the present application to recommend songs to the user. In the embodiment of the present application, the specific execution entity of this method is not limited to the terminal or the server.

[0053] In the song recommendation method provided in the embodiments of the present application, a computing device obtains the current conversation content input in the user dialog box. When the conversation content triggers a song recommendation behavior, a search engine is used to expand the conversation content, and the expanded information and the conversation content input by the user are jointly used as the input of a feature extraction model. The feature extraction model extracts text features, which not only include the features of the conversation content but also the features of the expanded information, and can perform extended recommendations as much as possible on the premise of combining the user's needs. Furthermore, relevant information of the to-be-recommended songs that match the text features is obtained from the song library, and based on the relevant information, a reply content to the conversation content input by the user is generated and displayed in the dialog box. It can be seen that in this solution, songs are recommended in combination with the conversation content input by the user in real time, and the recommended songs meet the user's current needs.

[0054] The following describes the song recommendation method provided in the embodiments of the present application. Refer to Figure 3 , and the processing of this method may include the following steps:

[0055] Step 301: Obtain the conversation content currently input by the user account.

[0056] Among them, the currently input conversation content includes basic song information.

[0057] In implementation, the user first logs in to the user account in the music application. Then, if the user wants to have a conversation with the song recommendation assistant, the user can select the conversation option in the interface of the music application, such as Figure 1 shown. After the terminal detects the selection operation of the conversation option, a dialog box is displayed. After the dialog box is displayed, the music application can first send out the conversation content. As Figure 4 shown, the conversation content first sent by the music application can be "Hello, welcome to the song recommendation assistant! I can help you find the most suitable songs for you. First of all, I need to understand your music taste. What is your favorite music genre? Such as pop, rock, classical, country, jazz, electronic dance music, etc.". Then, the user can input the conversation content by means of voice, text input, etc. The conversation content input by the user may include basic song information, and the basic song information here refers to information that is input by the user and is related to the song. As Figure 4 shown, the user inputs the conversation content "I like pop songs, preferably in Cantonese", where pop songs and Cantonese are the basic song information.

[0058] Furthermore, the terminal can obtain the conversation content currently input by the user.

[0059] Step 302: When it is determined that the current input conversation content belongs to the song recommendation intention, search for song extension information that matches the basic song information included in the current input conversation content through a search engine.

[0060] In implementation, after the terminal obtains the current input conversation content of the user, it can perform subsequent steps locally based on the conversation content, or send the conversation content to the server, and the server performs subsequent steps based on the conversation content. The specific execution is the same, but the execution entity is different. In the embodiments of the present application, it is taken as an example that after the terminal obtains the current input conversation content of the user, it sends the conversation content to the server, and the server performs subsequent steps based on the conversation content for illustration.

[0061] After the server receives the conversation content input by the user, it inputs the conversation content into the intention recognition model, and the intention recognition model outputs intention indication information. Among them, the intention recognition model is a type of neural network model. Among them, the intention indication information is used to indicate the intention of the current input conversation content, and the intention can include chatting, continuing to ask, song recommendation, obtaining song details, obtaining recommendation reasons, etc.

[0062] If the intention indicated by the intention indication information output by the intention recognition model is to continue asking, the server generates inquiry reply content and sends it to the terminal, which is displayed in the above dialog box to ask the user for more information. For example, Figure 4 As shown, when the conversation content input by the user is "I like pop songs, preferably in Cantonese", the reply content is "Do you have any particularly favorite singers or bands? For example, singer A, singer B", etc. Then, the user continues to input the conversation content "Singer A, I'm recently lovelorn", and the server obtains this conversation content. For this conversation content, if the intention indicated by the intention indication information output by the intention recognition model is to continue asking, the server generates the inquiry reply content "I'm very sorry to hear about your recent experience. Music is a great way to heal. I will recommend some songs that can comfort you. Do you prefer happy songs to boost your mood, or sad songs to accompany your emotions?", and sends it to the terminal. Then, the user continues to input the conversation content "Come with some sad songs", and the server obtains this conversation content and continues to judge the intention of this conversation content.

[0063] If the intention indicated by the intention indication information output by the intention recognition model is the song recommendation intention, the server searches for song extension information that matches the basic song information included in the current input conversation content through a search engine.

[0064] If the intention indicated by the intention indication information output by the intention recognition model is to obtain song details, the server obtains the details of the specified song as the reply content and sends it to the terminal for display in the above dialog box. The specified song refers to the song included in the user's input conversation content.

[0065] If the intention indicated by the intention indication information output by the intention recognition model is to obtain the recommendation reason, the server can generate a recommendation reason based on the user's input conversation content and the songs recommended to the user as the reply content and send it to the terminal for display in the above dialog box.

[0066] If the intention indicated by the intention indication information output by the intention recognition model is casual chat, the server obtains the general reply content corresponding to the user's input conversation content and sends it to the terminal for display in the above dialog box.

[0067] The following is an explanation of how the server searches for song extension information that matches the basic song information included in the current input conversation content through a search engine:

[0068] The above intention recognition model can be a module in the RAG (Retrieval-Augmented Generation) model. The intention recognition model can also determine whether the current input conversation content includes knowledge-based information, which can include the performing singer, song language, release year, emotional type (such as sad songs, upbeat songs), suitable scenarios, tempo, genre, associated film and television works, etc. For example, "Cantonese songs", "songs suitable for weddings", "Japanese songs from the 1990s", "songs from XX anime", etc.

[0069] When the current input conversation content includes knowledge-based information, the RAG model calls the search engine to search for at least one song search result that meets the preset conditions with the basic song information included in the current input conversation content, and uses the summary information of the at least one song search result as the song extension information. Each song search result corresponds to a matching degree with the above song basic information. Accordingly, the preset condition can be the N song search results with the highest matching degree, and the value of N can be configured by relevant personnel according to actual needs, such as N being 1.

[0070] For example, as Figure 5 shown, the current input conversation content is "songs from XX anime", where "XX anime" is the basic song information, and the song extension information searched by the search engine is "X music playlist 'Relight the Passionate Memories: XX Anime Original Soundtrack', songs: Song 1, Song 2, Song 3...". The song extension information may not be displayed to the user.

[0071] Step 303: Use the currently input conversation content and at least one piece of historical conversation content associated with the currently input conversation content as context information.

[0072] In implementation, the server can use the currently input conversation content of the user and at least one piece of historical conversation content associated with the currently input conversation content as context information. As Figure 4 shown, where the currently input conversation content is "Play some sad songs", and the associated historical conversation contents are "I like pop songs, preferably in Cantonese", "Singer A, I'm recently lovelorn", then use "Play some sad songs", "I like pop songs, preferably in Cantonese", "Singer A, I'm recently lovelorn" as context information.

[0073] Step 304: Input the context information and song extension information into a pre-trained feature extraction model to obtain target text features.

[0074] Among them, the target text features are used to represent user emotions and / or song information. The song information includes at least one of song type, lyric description, song language, suitable scene of the song, singer name, song name, album name, and the name of the movie or TV drama it belongs to.

[0075] In implementation, input the context information and song extension information into a pre-trained feature extraction model. Use the pre-trained feature extraction model to determine whether the context information contains information for representing user emotions. If it contains information for representing user emotions, extract the information for representing user emotions from the context information. Use the pre-trained feature extraction model to determine whether the context information and song extension information contain information for representing song information. If it contains information for representing song information, extract the information for representing song information from the context information and song extension information. Use the pre-trained feature extraction model to output the information for representing user emotions and the information for representing song information as target text features.

[0076] In a possible implementation, the feature extraction model is an LLM (Large Language Model), and the output format of the target text features is in JSON (JavaScript Object Notation) format. The JSON format is a lightweight data exchange format, which is convenient for machine parsing. Outputting the text features in JSON format can make subsequent machine processing faster and more accurate, thus effectively improving the song recommendation efficiency. The keys of the target text features in JSON format are identifiers of specific types, and the keys of the target text features in JSON format are the text features corresponding to specific types. Correspondingly, the processing of the above Step 304 can be as follows:

[0077] Generate a specific type of guide sentence, which is used to indicate that the input of the large language model is context information and song extension information, and is used to indicate that the output of the large language model is a specific type of target text feature and the output format of the target text feature. Input the specific type of guide sentence into the pre-trained large language model to obtain a specific type of target text feature that conforms to the output format. Figure 5 Explanation of the processing of LLM:

[0078] like Figure 5 As shown, a guide sentence of song type is generated: "Please summarize my preferences based on the above information. Is there any mention of the type of song I like? If so, please provide it in the form of JSON {"song type": [type 1, type 2]", where the above information can be context information and song extension information. Then, the guide sentence of the song type is input into LLM, and LLM extracts and summarizes the song types in the context information and song extension information to obtain the summary result "Your preferences are summarized as follows: You like Japanese anime songs, especially the theme songs and interludes of XX anime, including "Song 1", "Song 2", "Song 3", etc. Your favorite song types are in the form of JSON: {"song type": ["Japanese anime songs"]}".

[0079] Generate a guide sentence for user emotions: "Please summarize my preferences based on the above information. Do you mention my current mood? If yes, please provide it in the form of JSON {"feelings": [type 1, type 2]}" The above information can be context information. Then, the guide sentence for user emotions is input into LLM, and LLM extracts and summarizes the user emotions in the context information to obtain the summary result "In the conversation, you did not mention your current mood or feelings, so the relevant JSON format cannot be provided."

[0080] Generate a guide sentence for the film and TV series: "Please summarize my preferences based on the above information. Are there any films and TV series / anime that I like? If so, please provide them in the form of JSON {"movie":[type 1, type 2]}", where the above information can be context information and song extension information. Then, input the guide sentence for the film and TV series into LLM, which extracts and summarizes the user emotions in the context information and obtains the summary result "Your preferences are summarized as follows: The film and TV series to which your favorite song belongs is XX anime. The film and TV series to which your favorite song belongs is in the form of JSON: {"movie":["XX anime"]}".

[0081] Generate a guiding statement for the singer's name: "Please summarize my preferences based on the above information. If there is any mention of my favorite singer / band, please provide it in JSON format as {"singer": [type 1, type 2]}." Here, the above information includes both context information and song extension information. Then, input the guiding statement for the singer's name into the LLM. The LLM extracts and summarizes the singers in the context information and song extension information, and gets the summary result: "In the conversation, you did not explicitly mention your favorite singer, so it is impossible to provide the relevant JSON format."

[0082] Generate a guiding statement for the song name / album name: "Please summarize my preferences based on the above information. If there is any mention of my favorite song / album name, please provide it in JSON format as {"song": [type 1, type 2]}." Here, the above information includes both context information and song extension information. Then, input the guiding statement for the song name / album name into the LLM. The LLM extracts and summarizes the song name / album name in the context information and song extension information, and gets the summary result: "Your preferences are summarized as follows: The songs you like include "Song 1", "Song 2", "Song 3", etc. Your favorite songs in JSON format are {"song": ["Song 1", "Song 2", "Song 3"]}."

[0083] Through the above guiding statements, the target text features can finally be obtained, including {"song type": ["Japanese anime songs"]}, {"movie": ["XX anime"]}, {"song": ["I Want to Shout Out My Love", "Gaze Only at You", "Till the End of the World"]}.

[0084] In addition, Figure 5 The guiding statements and summary results shown in can be not shown to the user. The generation order and content of the guiding statements are only an example, and the embodiments of the present application do not limit this.

[0085] In a possible implementation, the server can also extract and record the song preference information of the user account from the text features extracted by the feature extraction model from the historical conversation content of the user account. The song preference information can characterize the overall preference of the user account for songs. The song preference information can include at least one of song type, lyric description, song language, suitable scene of the song, singer name, song name, album name, and the movie / TV drama it belongs to, such as preference for Cantonese, sad songs, anime songs, etc. Correspondingly, the processing in step 304 above can be as follows:

[0086] Obtain the song preference information corresponding to the user account, and input the context information, extension information, and song preference information into the feature extraction model to obtain text features.

[0087] Similarly, the feature extraction model here can also be an LLM. Correspondingly, the operations performed by the LLM on the context information and song extension information in the above embodiments can also be performed on the song preference information.

[0088] See Figure 6 , which shows a schematic diagram of text feature extraction. Among them, RAG obtains extension information through a search engine, and the extension information, the song preference information of the user account, and the context information are input into the LLM together, and the LLM outputs text features.

[0089] Step 305: Determine the songs to be recommended based on the target text features, and display the relevant information of the songs to be recommended to the user account.

[0090] Among them, the relevant information of the song can include the audio of the song, the cover of the song, the lyrics of the song, etc.

[0091] In practice, the song library can store the audio, cover, lyrics, song attributes, etc. of the song. Among them, the song attributes can include the singer, the language of the song, the release year, the emotional type (such as sad songs, lively), the suitable scene, etc.

[0092] After obtaining the text features, the text features can be matched with the song attributes and lyrics of the songs in the song library to determine the M songs with the highest matching degree as the songs to be recommended, and obtain the relevant information of the songs to be recommended. The value of M can be configured according to actual needs. For example, M is set to 5.

[0093] The following exemplarily illustrates several matching relationships:

[0094] When performing the matching, the song name in the text features can be matched with the song names in the song library, the lyrics in the text features can be matched with the lyrics in the song library, the user emotion in the text features can be matched with the emotional type of the songs in the song library, the singer in the text features can be matched with the singers of the songs in the song library, and the suitable scene of the song in the text features can be matched with the suitable scene of the songs in the song library.

[0095] The above matching relationships are only one example, and matching can also be performed according to the music style, album, language, tempo, etc., which will not be elaborated here one by one.

[0096] The following exemplarily illustrates the calculation method of the matching degree:

[0097] For any song in the song library, calculate the number of matching items between the song and the text features, and calculate the ratio of the number of matching items to the number of features in the text features as the matching degree corresponding to the song.

[0098] For example, the text features include four features, namely: song title "Sunshine", lyrics "Hold this beam of light tightly", singer "XX", and style "popular". The song title of a song in the music library is "Sunshine", the lyrics include "Hold this beam of light tightly", the singer is "YY", and the style is "popular", then the song title, lyrics, and style are matching items, the number of matching items is 3, and the corresponding matching degree of the song is 3 / 4=0.75.

[0099] After obtaining the relevant information of the song, the relevant information is input into the NLG (Natural Language Generation) model to generate a reply to the conversation content, and the reply content is sent to the terminal, which is displayed in the dialog box. Figure 7 As shown, the reply content is "Here are some recommended songs you might like" and a list of recommended songs. The recommended song list can display information such as song title, singer, album, etc. Figure 7 The recommended song list only shows the song titles of the recommended songs.

[0100] All the above optional technical solutions can be arbitrarily combined to form optional embodiments of the present disclosure, and will not be described in detail here.

[0101] In the technical solution provided by the embodiment of the present application, the user inputs the conversation content in the dialog box. When the conversation content triggers the song recommendation behavior, the search engine is used to expand the conversation content, and the expanded information and the conversation content input by the user are used as the input of the feature extraction model. The feature extraction model extracts text features. The text features not only contain the features of the conversation content, but also contain the features of the extended information. Under the premise of combining the needs of the user, the extended recommendation can be made as much as possible. Then, according to the text features, the relevant information of the matching songs to be recommended is obtained in the music library, and based on the relevant information, the reply content to the conversation content input by the user is generated and displayed in the dialog box. It can be seen that in this solution, songs are recommended to the user in combination with the conversation content input by the user in real time, so that the recommended songs meet the current needs of the user.

[0102] Based on the same technical concept, the embodiment of the present application also provides a device for generating drum audio, which can be a computing device, such as Figure 8 As shown, the device includes an acquisition module 710, a search module 720 and a reply module 730, wherein:

[0103] The acquisition module 710 is used to acquire the conversation content currently input by the user account, wherein the currently input conversation content includes basic information of the song;

[0104] A search module 720, which is used to, when it is determined that the currently input conversation content belongs to the song recommendation intent, search for song extension information that matches the basic song information included in the currently input conversation content through a search engine;

[0105] A reply module 730, which is used to use the currently input conversation content and at least one historical conversation content associated with the currently input conversation content as context information; input the context information and the song extension information into a pre-trained feature extraction model to obtain target text features; the target text features are used to represent user emotions and / or song information; determine the songs to be recommended based on the target text features, and display relevant information about the songs to be recommended to the user account.

[0106] In a possible implementation, the acquisition module 710 is used to:

[0107] Input the conversation content currently input by the user account into an intent recognition model to obtain intent indication information output by the intent recognition model; where the intent indication information is used to indicate whether the currently input conversation content belongs to the song recommendation intent.

[0108] In a possible implementation, the search module 720 is used to:

[0109] Extract the basic song information in the currently input conversation content and input it into a Retrieval-Augmented Generation (RAG) model, so that the RAG model calls a search engine to query song search results that match the basic song information;

[0110] Select at least one song search result that meets the preset conditions from the song search results, and use the summary information of the at least one song search result as the song extension information.

[0111] In a possible implementation, the acquisition module 710 is further used to: acquire the song preference information corresponding to the user account; the song preference information is text features extracted by the feature extraction model according to the historical conversation content of the user account;

[0112] The step of inputting the context information and the song extension information into a pre-trained feature extraction model to obtain target text features includes: inputting the context information, the song extension information, and the song preference information into a pre-trained feature extraction model to obtain target text features.

[0113] In a possible implementation, the feature extraction model is a large language model; the reply module 730 is used to:

[0114] Generate a specific type of guiding statement, which is used to indicate that the input of the large language model is the context information and the song extension information, and to indicate that the output of the large language model is the target text feature of the specific type and the output format of the target text feature;

[0115] Input the guiding statement of the specific type into the pre-trained large language model to obtain the target text feature of the specific type and conforming to the output format.

[0116] In a possible implementation, the output format of the target text feature is the JSON format. The key of the target text feature in the JSON format is the identifier of the specific type, and the key of the target text feature in the JSON format is the text feature corresponding to the specific type.

[0117] In a possible implementation, the search module 720 is used for:

[0118] Input the currently input conversation content into the intent recognition model, and the intent recognition model outputs the target content type;

[0119] When the target content type is used to indicate that the currently input conversation content includes knowledge-based information, search for song extension information that matches the basic song information included in the currently input conversation content through a search engine.

[0120] In a possible implementation, the reply module 730 is used for:

[0121] Input the context information and the song extension information into the pre-trained feature extraction model;

[0122] Use the pre-trained feature extraction model to determine whether the context information contains information for characterizing the user's emotion. If it contains information for characterizing the user's emotion, extract the information for characterizing the user's emotion from the context information;

[0123] Use the pre-trained feature extraction model to determine whether the context information contains information for characterizing song information. If it contains information for characterizing song information, extract the information for characterizing song information from the context information;

[0124] Use the pre-trained feature extraction model to take the information for characterizing the user's emotion, the information for characterizing song information, and the target song information included in the song extension information as the target text feature.

[0125] In a possible implementation, the song information includes at least one of song type, lyric description, song language, suitable scene of the song, singer name, song name, and album name.

[0126] In the technical solution provided by the embodiments of the present application, when a user inputs conversation content in a dialog box, in the case where the conversation content triggers a song recommendation behavior, a search engine is used to expand the conversation content, and the expanded information and the conversation content input by the user are jointly used as the input of a feature extraction model. The feature extraction model extracts text features. These text features not only include the features of the conversation content but also the features of the expanded information, and can perform extended recommendations as much as possible on the premise of combining the user's needs. Furthermore, relevant information of the to-be-recommended songs that match the text features is obtained from the song library, and based on the relevant information, a reply content to the conversation content input by the user is generated and displayed in the dialog box. It can be seen that in this solution, songs are recommended in combination with the conversation content input by the user in real time, and the recommended songs meet the user's current needs.

[0127] It should be noted that when the song recommendation device provided in the above embodiments recommends songs, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be assigned to different functional modules according to needs, that is, the internal structure of the computing device is divided into different functional modules to complete all or part of the functions described above. In addition, the song recommendation device provided in the above embodiments and the method embodiments of song recommendation belong to the same concept, and the specific implementation process is detailed in the method embodiments and will not be elaborated here.

[0128] Figure 9 The block diagram of the structure of a terminal 600 provided by an exemplary embodiment of the present application is shown. The terminal 600 may be a portable mobile terminal, such as: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a notebook computer, or a desktop computer. The terminal 600 may also be referred to by other names such as user equipment, portable terminal, laptop terminal, desktop terminal, etc.

[0129] Generally, the terminal 600 includes: a processor 601 and a memory 602.

[0130] The processor 601 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 601 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 601 may also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 601 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 601 may further include an AI (Artificial Intelligence) processor, and this AI processor is used to process computational operations related to machine learning.

[0131] The memory 602 may include one or more computer-readable storage media, and this computer-readable storage media may be non-transitory. The memory 602 may also include high-speed random access memory, as well as non-volatile memory, such as one or more disk storage devices and flash storage devices.

[0132] In some embodiments, the terminal 600 may also optionally include: a peripheral device interface 603 and at least one peripheral device. The processor 601, the memory 602, and the peripheral device interface 603 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 603 through a bus, signal lines, or a circuit board. Specifically, the peripheral devices include at least one of a radio frequency circuit 604, a display screen 605, a camera assembly 606, an audio circuit 607, a positioning component 608, and a power supply 609.

[0133] The peripheral device interface 603 may be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 601 and the memory 602. In some embodiments, the processor 601, the memory 602, and the peripheral device interface 603 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 601, the memory 602, and the peripheral device interface 603 may be implemented on a separate chip or circuit board, and this embodiment does not limit this.

[0134] The radio frequency circuit 604 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 604 communicates with the communication network and other communication devices through electromagnetic signals. The radio frequency circuit 604 converts electrical signals into electromagnetic signals for transmission, or converts the received electromagnetic signals into electrical signals. Optionally, the radio frequency circuit 604 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and so on. The radio frequency circuit 604 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, various generations of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 604 may further include a circuit related to NFC (Near Field Communication), which is not limited in this application.

[0135] The display screen 605 is used to display a UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 605 is a touch display screen, the display screen 605 also has the ability to collect touch signals on or above the surface of the display screen 605. The touch signals can be input to the processor 601 as control signals for processing. At this time, the display screen 605 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, the display screen 605 can be one, disposed on the front panel of the terminal 600; in other embodiments, the display screen 605 can be at least two, respectively disposed on different surfaces of the terminal 600 or in a foldable design; in other embodiments, the display screen 605 can be a flexible display screen, disposed on the curved surface or the folding surface of the terminal 600. Even, the display screen 605 can be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 605 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0136] The camera component 606 is used to collect images or videos. Optionally, the camera component 606 includes a front camera and a rear camera. Generally, the front camera is disposed on the front panel of the terminal, and the rear camera is disposed on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth camera, a wide-angle camera, and a telephoto camera, so as to implement the function of background blurring by fusing the main camera and the depth camera, panoramic shooting by fusing the main camera and the wide-angle camera, and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera component 606 may further include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. The dual-color temperature flash refers to the combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.

[0137] The audio circuit 607 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals and input them to the processor 601 for processing, or input them to the radio frequency circuit 604 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively disposed at different parts of the terminal 600. The microphone can also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 601 or the radio frequency circuit 604 into sound waves. The speaker can be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 607 may further include a headphone jack.

[0138] The positioning component 608 is used to locate the current geographical location of the terminal 600 to implement navigation or LBS (Location Based Service). The positioning component 608 can be a positioning component based on GPS (Global Positioning System), Beidou system or Galileo system.

[0139] The power supply 609 is used to supply power to each component in the terminal 600. The power supply 609 can be alternating current, direct current, a disposable battery or a rechargeable battery. When the power supply 609 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. The wired rechargeable battery is a battery charged through a wired line, and the wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0140] In some embodiments, the terminal 600 further includes one or more sensors 610. The one or more sensors 610 include but are not limited to: an acceleration sensor 611, a gyroscope sensor 612, a pressure sensor 613, a fingerprint sensor 614, an optical sensor 615, and a proximity sensor 616.

[0141] The acceleration sensor 611 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established by the terminal 600. For example, the acceleration sensor 611 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 601 can control the display screen 605 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 611. The acceleration sensor 611 can also be used for collecting game or user movement data.

[0142] The gyroscope sensor 612 can detect the body direction and rotation angle of the terminal 600. The gyroscope sensor 612 can cooperate with the acceleration sensor 611 to collect the 3D actions of the user on the terminal 600. According to the data collected by the gyroscope sensor 612, the processor 601 can implement the following functions: motion sensing (such as changing the UI according to the user's tilting operation), image stabilization during shooting, game control, and inertial navigation.

[0143] The pressure sensor 613 can be disposed on the side frame of the terminal 600 and / or the lower layer of the display screen 605. When the pressure sensor 613 is disposed on the side frame of the terminal 600, it can detect the holding signal of the user on the terminal 600, and the processor 601 can perform left / right hand recognition or shortcut operations according to the holding signal collected by the pressure sensor 613. When the pressure sensor 613 is disposed on the lower layer of the display screen 605, the processor 601 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 605. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.

[0144] The fingerprint sensor 614 is used to collect the fingerprint of the user. The processor 601 can identify the user's identity according to the fingerprint collected by the fingerprint sensor 614, or the fingerprint sensor 614 can identify the user's identity according to the collected fingerprint. When the identity of the user is identified as a trusted identity, the processor 601 authorizes the user to perform relevant sensitive operations, and the sensitive operations include unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings, etc. The fingerprint sensor 614 can be disposed on the front, back, or side of the terminal 600. When there are physical buttons or manufacturer logos on the terminal 600, the fingerprint sensor 614 can be integrated with the physical buttons or manufacturer logos.

[0145] The optical sensor 615 is used to collect the ambient light intensity. In one embodiment, the processor 601 can control the display brightness of the display screen 605 according to the ambient light intensity collected by the optical sensor 615. Specifically, when the ambient light intensity is high, the display brightness of the display screen 605 is increased; when the ambient light intensity is low, the display brightness of the display screen 605 is decreased. In another embodiment, the processor 601 can also dynamically adjust the shooting parameters of the camera module 606 according to the ambient light intensity collected by the optical sensor 615.

[0146] The proximity sensor 616, also known as the distance sensor, is usually disposed on the front panel of the terminal 600. The proximity sensor 616 is used to collect the distance between the user and the front of the terminal 600. In one embodiment, when the proximity sensor 616 detects that the distance between the user and the front of the terminal 600 is gradually decreasing, the processor 601 controls the display screen 605 to switch from the lit state to the off state; when the proximity sensor 616 detects that the distance between the user and the front of the terminal 600 is gradually increasing, the processor 601 controls the display screen 605 to switch from the off state to the lit state.

[0147] Those skilled in the art can understand that Figure 9 the structure shown in

[0148] Figure 10 is a schematic structural diagram of a server provided by an embodiment of the present application. The server 1000 may vary greatly due to different configurations or performances, and may include one or more processors (central processing units, CPUs) 1001 and one or more memories 1002. Among them, at least one instruction is stored in the memory 1002, and the at least one instruction is loaded and executed by the processor 1001 to implement the methods provided by the above-mentioned various method embodiments. Of course, the computing device may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input / output. The computing device may also include other components for implementing the functions of the device, which will not be elaborated here.

[0149] In an exemplary embodiment, a computer-readable storage medium is further provided, such as a memory including instructions that can be executed by a processor in a terminal to complete the method in the above embodiment. The computer-readable storage medium may be non-transitory. For example, the computer-readable storage medium may be a ROM (Read-Only Memory), a RAM (Random Access Memory), a CD-ROM (Compact Disc Read-Only Memory), a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0150] In an exemplary embodiment, a computer program product is further provided. At least one instruction is stored in the computer program product, and the instruction is loaded and executed by a processor to implement the operations performed by the method for song recommendation as described above.

[0151] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals (including but not limited to signals transmitted between a user terminal and other devices, etc.) involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions. For example, the conversation content input by the user, the user's song listening records, favorite songs, and song comments involved in this application are all obtained under full authorization.

[0152] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiment can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk, or an optical disc, etc.

[0153] The above are only optional embodiments of this application, and are not intended to limit this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this application shall be included in the protection scope of this application.

Claims

1. A song recommendation method, characterized in that: The method comprises: Acquire the conversation content currently input by the user account, wherein the currently input conversation content includes basic information of the song; When it is determined that the currently input conversation content belongs to the song recommendation intention, searching for song extension information matching the song basic information included in the currently input conversation content through a search engine; using the currently input conversation content and at least one historical conversation content associated with the currently input conversation content as context information; Inputting the context information and the song extension information into a pre-trained feature extraction model to obtain target text features; the target text features are used to represent user emotions and / or song information; The songs to be recommended are determined based on the target text features, and relevant information of the songs to be recommended is displayed to the user account.

2. The method according to claim 1, characterized in that The obtaining of the conversation content currently input by the user account includes: The conversation content currently input by the user account is input into the intention recognition model to obtain the intention indication information output by the intention recognition model; wherein the intention indication information is used to indicate whether the conversation content currently input belongs to the song recommendation intention.

3. The method according to claim 1, characterized in that The searching for song extension information matching the basic song information included in the currently input conversation content through a search engine includes: Extracting basic song information from the currently input conversation content and inputting it into a search enhancement to generate a RAG model, so that the RAG model calls a search engine to query song search results that match the basic song information; At least one song search result that meets a preset condition is selected from the song search results, and summary information of the at least one song search result is used as the song extension information.

4. The method according to claim 1, characterized in that: The method further includes: obtaining song preference information corresponding to the user account; the song preference information is text features extracted by the feature extraction model based on historical conversation content of the user account; The step of inputting the context information and the song extension information into a pre-trained feature extraction model to obtain target text features includes: inputting the context information, the song extension information and the song preference information into a pre-trained feature extraction model to obtain target text features.

5. The method according to claim 1, characterized in that The feature extraction model is a large language model; the context information and the song extension information are input into a pre-trained feature extraction model to obtain target text features, including: Generate a specific type of guide sentence, the guide sentence is used to indicate that the input of the large language model is the context information and the song extension information, and is used to indicate that the output of the large language model is the specific type of target text feature and the output format of the target text feature; The specific type of guiding sentence is input into the pre-trained large language model to obtain target text features of the specific type that conform to the output format.

6. The method according to claim 5, characterized in that The output format of the target text feature is JSON format, the key of the target text feature in JSON format is the identifier of the specific type, and the key of the target text feature in JSON format is the text feature corresponding to the specific type.

7. The method according to claim 1, characterized in that The searching for song extension information matching the basic song information included in the currently input conversation content through a search engine includes: Inputting the currently input conversation content into an intention recognition model, and the intention recognition model outputs a target content type; In the case where the target content type is used to indicate that the currently input conversation content includes knowledge-related information, song extension information matching the song basic information included in the currently input conversation content is searched for through a search engine.

8. The method according to any one of claims 1 to 6, characterized in that The step of inputting the context information and the song extension information into a pre-trained feature extraction model to obtain target text features includes: Inputting the context information and the song extension information into a pre-trained feature extraction model; Determining whether the context information contains information for representing the user's emotions by using the pre-trained feature extraction model, and if the context information contains information for representing the user's emotions, extracting the information for representing the user's emotions from the context information; Determining whether the context information contains information for representing song information by using the pre-trained feature extraction model, and if the context information contains information for representing song information, extracting the information for representing song information from the context information; The information used to characterize the user's emotions, the information used to characterize the song information, and the target song information included in the song extension information are used as target text features through the pre-trained feature extraction model.

9. The method according to any one of claims 1 to 6, characterized in that: The song information includes at least one of the song type, lyrics description, song language, suitable scene of the song, singer name, song title, and album name.

10. A computing device, characterized in that: The computing device includes a processor and a memory, wherein the memory stores at least one instruction, and the instruction is loaded and executed by the processor to implement the operations performed by the song recommendation method as described in any one of claims 1 to 9.

11. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, which is loaded and executed by the processor to implement the operation performed by the song recommendation method as described in any one of claims 1 to 9.

12. A computer program product, characterized in that The computer program product stores at least one instruction, which is loaded and executed by a processor to implement the operations performed by the song recommendation method as described in any one of claims 1 to 9.