Method for adding songs in song list, computer equipment and storage medium

By calculating the audio and classification tag features of songs and using artificial intelligence to recommend candidate playlists, the inefficiency of users manually selecting and adding songs from multiple playlists is solved, achieving more efficient song management.

CN120631233APending Publication Date: 2025-09-12TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510716602.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

The more playlists users create in music apps, the less efficient it becomes to manually select and add songs, leading to decreased playlist management efficiency.

Method used

By calculating the audio features and classification label features of the target song, and using artificial intelligence technology to analyze the match between the song and the user's playlist, candidate playlists are automatically recommended to improve adding efficiency.

Benefits of technology

This reduces the number of steps users have to manually filter through a large number of playlists and improves the efficiency of adding songs to the target playlist.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120631233A_ABST
    Figure CN120631233A_ABST
Patent Text Reader

Abstract

The embodiment of the invention belongs to the technical field of artificial intelligence, and particularly relates to a method for adding songs in a song list, computer equipment and a storage medium. In the method, candidate song lists are selected from object song lists of a target object logging in a music application program based on the similarity between target music features of a target song and reference music features of a reference song, and the target music features are obtained by fusing audio features and classification tag features of the target song. And displaying the candidate song menus in the music application program, and when a selection confirmation request for a target song menu in the candidate song menus is received, adding the target song to the target song menu. In this way, the candidate song lists matched with the target song are displayed, it can be avoided that a user searches for the song lists needing to be added with the target song in all the created song lists, and then the efficiency of adding the songs into the song lists can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method for adding songs to a playlist, a computer device, and a storage medium. Background Art

[0002] When using music apps, users often create playlists based on their preferences and needs, and then add music to them based on their listening needs. For example, a user could add music they often listen to while exercising to one playlist and music they often listen to while driving to another. This way, users can choose the appropriate playlist for their listening needs.

[0003] As users spend more time using music apps, they create more and more playlists. When users need to add music to a playlist, they need to manually filter out the corresponding playlist from the large number of playlists they have created. This results in a low efficiency in adding music to playlists. Summary of the Invention

[0004] The embodiments of the present application provide a method, computer device, and storage medium for adding songs to a playlist, which can improve the efficiency of adding songs to a playlist in a music application. The corresponding technical solutions are as follows:

[0005] In a first aspect, a method for adding a song to a playlist is provided, the method comprising:

[0006] When a playlist addition request for a target song is received, which is triggered in a music application, a candidate playlist is selected from each target playlist of a target object logged into the music application based on a similarity between a target music feature of the target song and a reference music feature of a reference song, wherein the target music feature is obtained by fusing an audio feature and a classification label feature of the target song, the reference music feature is obtained by fusing an audio feature and a classification label feature of a reference song, and the reference song is a song in the target playlist;

[0007] Pushing the candidate playlist to the target object to display the candidate playlist in the music application;

[0008] When a selection confirmation request for a target song list in the candidate song list is received, the target song is added to the target song list.

[0009] Optionally, selecting a candidate playlist from each target playlist of a target object logging into the music application based on the similarity between the target music feature of the target song and the reference music feature of the reference song includes:

[0010] Determining the similarity between the target music feature and the reference music feature corresponding to each reference song in each object playlist;

[0011] Determining a first number of similar songs in each reference song in descending order of corresponding similarities;

[0012] According to the number of similar songs included in each target song list, a candidate song list is selected from each target song list.

[0013] Optionally, selecting a candidate playlist from each target playlist according to the number of similar songs included in each target playlist includes:

[0014] For each target song list, determining an average similarity between similar songs included in the target song list and the target song;

[0015] According to the average similarity corresponding to each object playlist, a candidate playlist is selected from each object playlist.

[0016] Optionally, before selecting a candidate playlist from each target playlist of a target object logged into the music application based on the similarity between the target music feature of the target song and the reference music feature of the reference song, the method further includes:

[0017] According to the target song identifier corresponding to the target song and the reference song identifier of the reference song, the similarity between the target song and the reference song is searched in the corresponding relationship, wherein the corresponding relationship stores the similarity between each song in the song library of the music application and the corresponding song identifier.

[0018] Optionally, the songs included in the corresponding relationship are songs in a song library of the music application whose corresponding popularity values ​​exceed a popularity threshold, and the method further includes:

[0019] If the similarity between the target song and the reference song is not found in the corresponding relationship, the target music features of the target song and the reference music features of the reference song are obtained, and the similarity between the target song and the reference song is calculated based on the target music features and the reference song features.

[0020] Optionally, selecting a candidate playlist from each target playlist of a target object logging into the music application based on the similarity between the target music feature of the target song and the reference music feature of the reference song includes:

[0021] Determining an average music feature of each reference song included in each target playlist based on the reference music features of each reference song included in each target playlist;

[0022] According to the similarity between the target music feature and the average music feature corresponding to each object playlist, a candidate playlist is selected from each object playlist.

[0023] Optionally, before selecting a candidate playlist from each target playlist of a target object logged into the music application based on the similarity between the target music feature of the target song and the reference music feature of the reference song, the method further includes:

[0024] Obtaining a classification label of the target song and extracting classification label features of the classification label;

[0025] Extracting target music features of the target song based on the audio data of the target song;

[0026] The classification label features and target music features corresponding to the target song are fused to obtain the target music features of the target song, and the fusion process includes weighted summation and feature splicing.

[0027] In a second aspect, a device for adding songs to a playlist is provided, the device comprising:

[0028] a selection module for, upon receiving a playlist addition request for a target song triggered in a music application, selecting a candidate playlist from each target playlist of a target object logged into the music application based on a similarity between a target music feature of the target song and a reference music feature of a reference song, wherein the target music feature is obtained by fusing an audio feature and a classification label feature of the target song, the reference music feature is obtained by fusing an audio feature and a classification label feature of a reference song, and the reference song is a song in the target playlist;

[0029] A push module, configured to push the candidate playlist to the target object so as to display the candidate playlist in the music application;

[0030] An adding module is used to add the target song to the target song list when a selection confirmation request for the target song list in the candidate song list is received.

[0031] Optionally, the selection module is used to determine the similarity between the target music feature and the reference music feature corresponding to each reference song in each object playlist; determine a first number of similar songs in each reference song in order of corresponding similarities from high to low; and select a candidate playlist in each object playlist based on the number of similar songs included in each object playlist.

[0032] Optionally, the selection module is used to determine, for each object playlist, an average similarity between similar songs included in the object playlist and the target song; and select a candidate playlist from each object playlist based on the average similarity corresponding to each object playlist.

[0033] Optionally, the above-mentioned device also includes a search module for searching the similarity between the target song and the reference song in a corresponding relationship based on the target song identifier corresponding to the target song and the reference song identifier of the reference song, wherein the corresponding relationship stores the similarity between each song in the song library of the music application and the corresponding song identifier.

[0034] Optionally, the songs included in the corresponding relationship are songs in the song library of the music application whose corresponding popularity values ​​exceed a popularity value threshold. The above-mentioned device also includes a calculation module for obtaining the target music features of the target song and the reference music features of the reference song when the similarity between the target song and the reference song is not found in the corresponding relationship, and calculating the similarity between the target song and the reference song based on the target music features and the reference song features.

[0035] Optionally, the selection module is used to determine the average music features of each reference song included in each object playlist based on the reference music features of each reference song included in each object playlist; and select candidate playlists from each object playlist based on the similarity between the target music features and the average music features corresponding to each object playlist.

[0036] Optionally, the above-mentioned device also includes a feature extraction module, which is used to obtain the classification label of the target song and extract the classification label features of the classification label; extract the target music features of the target song based on the audio data of the target song; and fuse the classification label features and target music features corresponding to the target song to obtain the target music features of the target song, and the fusion processing includes weighted summation and feature splicing.

[0037] In a third aspect, a computer device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction, which is loaded and executed by the processor to implement the operations performed by the method for adding songs to a playlist as described in the first aspect above.

[0038] In a fourth aspect, a computer-readable storage medium is provided, which stores at least one instruction, which is loaded and executed by a processor to implement the operations performed by the method for adding songs to a playlist as described in the first aspect above.

[0039] In a fifth aspect, a computer program product is provided, which stores at least one instruction, which is loaded and executed by a processor to implement the operations performed by the method for adding songs to a playlist as described in the first aspect above.

[0040] The beneficial effects of the technical solution provided by this application include at least:

[0041] In the technical solution provided by this application, when a user needs to add a target song to at least one of the last playlists created, the similarity between the target song and the songs included in the target playlist created by the user can be determined. Then, based on the similarity, a candidate playlist matching the target song is determined and displayed in the music application. In this way, the user only needs to select the target playlist to which the target song is to be added from the candidate playlists matching the target song. It can be seen that in the technical solution provided by this application, the user does not need to search through a large number of playlists to find the playlist to which the target song is to be added, which can improve the efficiency of adding songs to the playlist. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0043] Figure 1 This is a schematic diagram of an exemplary method of adding songs to a playlist provided in an embodiment of the present application;

[0044] Figure 2 This is a flow chart of a method for adding songs to a playlist provided in an embodiment of the present application;

[0045] Figure 3 This is a flow chart of a targeted processing of classification labels and audio signals provided by an embodiment of the present application;

[0046] Figure 4 This is a schematic diagram of a playlist display provided by an embodiment of the present application;

[0047] Figure 5 This is a schematic diagram of a playlist display provided by an embodiment of the present application;

[0048] Figure 6 is a schematic diagram of training a first encoder, a second encoder, and a multimodal feature fusion module provided in an embodiment of the present application;

[0049] Figure 7 This is a schematic diagram of the structure of a device for adding songs to a playlist provided in an embodiment of the present application;

[0050] Figure 8 A schematic structural diagram of a computer device provided by an exemplary embodiment of the present application is shown;

[0051] Figure 9 This is a structural diagram of a server provided in an embodiment of the present application. DETAILED DESCRIPTION

[0052] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0053] Song classification tags: These are keywords or phrases used to describe and categorize musical works. These tags can encompass a wide range of information, including but not limited to musical style, emotion, theme, instrumentation, rhythmic characteristics, and regional and cultural background, such as rock or folk. These tags can be manually compiled or generated using music information retrieval tools.

[0054] Embedding: is a technology that converts high-dimensional data (such as text, images or audio) into low-dimensional numerical vectors. By mapping complex feature information into a vector space, embedding can capture the semantic or feature similarity of the data, allowing machines to process, compare and search information more efficiently. For example, for the audio of a song, embedding can convert the complex information of the music (such as an audio signal) into a more concise, highly abstract numerical vector representation. This vector representation can capture and express the key features of music audio, such as melody, rhythm, harmony, etc. For the classification labels of songs, embedding has a larger space and can represent richer and more detailed music features. Among them, the distance between embeddings can reflect the degree of similarity between objects in the representation of each dimension.

[0055] With the development of Internet technology, users run music applications on their terminals, and playing songs through music applications has become one of the most common ways to play songs. Various music applications provide users with the function of creating playlists for their convenience. In one example, users can create different playlists for their own listening scenarios, such as creating a "fitness playlist", "work and study playlist", "driving playlist", etc. Users can selectively add songs to the created playlists, such as adding songs that they often listen to when exercising to the "fitness playlist". In this way, when exercising, users can choose to play songs in the "fitness playlist" in the music application, thereby realizing personalized music playback.

[0056] Figure 1This is a schematic diagram of an exemplary method of adding songs to a playlist provided in an embodiment of the present application. Figure 1 As shown, when the user is playing a song in the music application, an option to add a playlist may be displayed in the song playback interface. The user can click on this option to trigger the currently playing song to be added to the playlist. After the user clicks on this option, the music application can display multiple playlists created by the user. The display order of multiple playlists can be random, or they can be displayed in chronological order according to the time of creation. When there are too many playlists created by the user to be displayed in full, the user can use sliding and other methods to make the music application further display the playlists that are not currently displayed. The user can select the playlist to which the song needs to be added from the displayed playlists. For example, the user can trigger the currently playing song to be added to the corresponding playlist by clicking on the playlist.

[0057] As users spend more time using music apps, they create more and more playlists, often dozens or even hundreds of them. Consequently, when users need to add songs to a playlist, they must manually select the desired playlist from among these dozens or even hundreds of playlists. This results in a low efficiency in adding songs to playlists.

[0058] The present application provides a method for adding songs to a playlist. The method utilizes artificial intelligence technology to analyze the matching degree between the songs to be added and the playlists created by the user, intelligently analyzes the playlists that the user needs to add songs to, and prioritizes the display of the corresponding playlists. This avoids the user having to manually select the playlists that need to be added from a large number of playlists, thereby improving the efficiency of the user adding songs to the playlist.

[0059] Figure 2 This is a flow chart of a method for adding songs to a playlist provided by an embodiment of the present application. The method for adding songs to a playlist can be executed by a terminal, or by a server, or by a terminal and a server in cooperation. The terminal refers to a device that runs a music application, such as a mobile phone, tablet computer, personal computer (PC), and smart wearable device, and the server can refer to the background server of the music application, which can be a physical machine or a virtual instance implemented by virtual machine technology. See Figure 2 , the method comprising:

[0060] Step 201: When a playlist addition request for a target song is received, triggered in a music application, a candidate playlist is selected from each target playlist of a target object logged into the music application based on the similarity between the target music feature of the target song and the reference music feature of a reference song. The target music feature is obtained by fusing the audio feature and classification label feature of the target song, the reference music feature is obtained by fusing the audio feature and classification label feature of the reference song, and the reference song is a song in the target playlist.

[0061] The target object is the user's account logged into the music application, and the target playlist is the playlist created by the user after logging into the music application. The similarity between the target music features and the reference music features of the reference songs in each target playlist can be used to calculate the matching degree between the target song and each target playlist. Candidate playlists can refer to the corresponding playlists in the target playlists with a high matching degree.

[0062] When step 201 is executed by the terminal, the playlist addition request for the target song triggered by the user in the music application may refer to the playlist addition request for the target song triggered by the user after clicking the Add to Playlist option corresponding to the target song. When step 201 is executed by the server, the playlist addition request for the target song triggered by the user in the music application may refer to the playlist addition request for the target song sent by the terminal to the server after the user clicks the Add to Playlist option corresponding to the target song.

[0063] The following first introduces the generation of target music features of the target song.

[0064] The classification label features of the target song can be obtained by extracting the features of the classification label of the target song. In one example, the classification label of the target song can be obtained, and the text embedding corresponding to the classification label can be generated by a pre-trained first encoder to obtain the classification label features of the target song. The first encoder can be implemented using a word vector model, such as Word2Vec, GloVe, etc. The training of the first encoder can be seen below. Figure 6 The relevant content will not be introduced here.

[0065] The audio features of the target song may include the frequency domain features of the frequency domain signal corresponding to the audio of the target song and / or the time domain features of the corresponding time domain signal. For the frequency domain features, the audio signal of the target song may be first converted into a frequency domain signal. For example, a short-time Fourier transform is performed on the audio signal to obtain a frequency domain signal, and then the frequency domain features of the frequency domain signal are extracted, wherein the frequency domain features may include fine-grained features such as fundamental frequency, harmonics, spectrum envelope, and spectrum attenuation, which respectively reflect some characteristics of the music content, such as melody pitch, instrument characteristics, singing skills, emotional expression, etc. For the time domain features, the time-frequency domain features such as Mel-Frequency Cepstral Coefficients (MFCC) and Constant Q Transform of the audio signal of the target song may be extracted to obtain the dynamic relationship between music time and frequency.

[0066] In some embodiments, considering that the playing time of a song is long and the data volume of the audio signal is large, when calculating the audio features, you can choose to segment the audio signal of the target song, for example, only retain the audio signal of a chorus part. In addition, considering that the data volume of the frequency domain features and / or time domain features is still large, after obtaining the frequency domain features and / or time domain features, the frequency domain features and / or time domain features can be encoded by a second encoder to generate corresponding audio embeddings. Among them, the second encoder can be implemented by an audio encoder (Music-EnhancedRepresentation Transformer, MERT), a contrastive language-audio pretraining model (ContrastiveLanguage-Audio Pretraining, CLAP) or a jukebox model. The training of the second encoder can be seen below. Figure 6 The relevant content will not be introduced here.

[0067] In one example, after obtaining the classification label features and audio features corresponding to the target song, such as the text embedding of the classification label and the audio embedding of the audio signal, the classification label features and audio features corresponding to the target song can be fused to obtain the target music features of the target song. The fusion process includes weighted summation, feature concatenation, or feature fusion.

[0068] In implementation, the size of the text embedding and the size of the audio embedding can be the same, and the technician can set the first weight coefficient corresponding to the text embedding and the second weight coefficient corresponding to the audio embedding. In this way, the text embedding and the audio embedding can be weighted and summed by the first weight coefficient and the second weight coefficient to obtain the fused music embedding, that is, to obtain the target music features of the target song. Among them, the first weight coefficient and the second weight coefficient that the technician can set can be empirical values, or obtained through training, and their specific values ​​are not limited in the embodiments of this application.

[0069] Alternatively, in practice, the text embedding and the audio embedding may be concatenated to obtain the music embedding. The order in which the text embedding and the audio embedding are concatenated is not limited in this application.

[0070] Alternatively, in practice, the text embedding and audio embedding can be input into a pre-trained multimodal feature fusion module, and the text embedding and audio embedding can be fused to output the music embedding. The multimodal feature fusion module can be an artificial intelligence (AI) model for processing multimodal features. The training of the multimodal feature fusion module can be seen below. Figure 6 The relevant content will not be introduced here.

[0071] Figure 3 This is a flow chart of a target processing of classification labels and audio signals provided by an embodiment of the present application. Figure 3 As shown in the figure, after obtaining the target song's classification label, the classification label can be encoded using the first encoder to obtain a text embedding. After obtaining the target song's audio signal, the audio signal can be converted into a frequency domain signal and feature extracted to obtain the corresponding music features. This signal is then encoded using the second encoder to obtain an audio embedding. After obtaining the text and audio embeddings, they can be processed in a multimodal encoding fusion module to obtain a music embedding.

[0072] In an embodiment of the present application, when calculating the degree of matching between a song and an object playlist, the song's classification label and the song's audio signal can be combined to calculate the song's musical features. In this way, the musical features can reflect the style, emotion, theme, instrument, rhythm characteristics, etc. that can be reflected by the classification label, and can also reflect the melody pitch, instrument characteristics, singing skills, emotional expression, etc. of the audio signal. It can be seen that the audio features obtained in this way can reflect the characteristics of the song in multiple dimensions, can more accurately represent the song, and thus can improve the accuracy of subsequent calculations of the degree of matching with the object playlist.

[0073] In one achievable manner, the target song feature of the target song can be calculated by the terminal or server after receiving a request to add the target song to a playlist. Alternatively, the server can pre-store a correspondence between the song features and song identifiers (IDs) of each song in the song library of the music application. The server can search for the target song feature in this correspondence based on the target song ID, thereby improving the efficiency of obtaining the target song feature.

[0074] The following introduces the process of determining the candidate playlists for each object playlist based on the target music features and the reference music features of the reference songs in each object playlist.

[0075] When step 201 is executed by the terminal, the degree of matching between the target song and each object playlist created by the user can be calculated by the terminal, or the terminal can send a matching degree acquisition request to the server, and the server can calculate the degree of matching. When step 201 is executed by the server, the degree of matching between the target song and each object playlist created by the user can be calculated by the server. The following uses the example of the server calculating the degree of matching between the target song and each object playlist created by the user to describe the process of calculating the degree of matching between the target song and each object playlist in detail.

[0076] In one example, after obtaining the target music features of the target song, the server can determine the similarity between the target music features and the reference music features of the reference songs included in each object playlist, and then determine the matching degree of the target song with each object playlist based on the similarity. Similar to the target music features, the reference music features of each reference song can be calculated from the classification label features and audio features of the reference song. For the reference music features and the corresponding calculation process, please refer to the introduction to the target music features and the corresponding calculation process in step 201 above, which will not be repeated here. Similarly, the reference music features of the reference songs can also be found in the correspondence between the song features and the song ID in the above embodiment.

[0077] In the embodiment of the present application, at least two methods of determining candidate playlists may be provided as follows:

[0078] Method 1: The server may determine the matching degree between the target song and the target playlist based on the number of reference songs in the target playlist that have a high degree of similarity to the target song. The corresponding processing includes:

[0079] Step 2011: Determine the similarity between the target music feature and the reference music feature corresponding to each reference song in each object playlist.

[0080] During implementation, the server can calculate the similarity between the target song and each reference song based on the target music feature and the reference music feature corresponding to each reference song. In one example, the target music feature and the reference music feature are music embeddings, and the similarity between the target song and the reference song is obtained by calculating the distance between the music embeddings. In other words, the distance between the embeddings can indicate the similarity between the target song and the reference song. For example, the distance between the embeddings can be calculated using conventional algorithms such as Euclidean distance, cosine similarity, and Manhattan distance. The smaller the distance, the higher the similarity.

[0081] Since the audio of the song will not be updated, the similarity between songs will not change. Therefore, in order to avoid repeated calculation of the similarity between songs, the background server of the music application can pre-calculate the similarity between each song included in the song library, and store the song ID corresponding to the similarity of each pair of songs in the corresponding relationship. In this way, when it is necessary to calculate the similarity between the target song and the reference song, the similarity can be found in the corresponding relationship based on the song ID of the target song and the reference song. This can improve the efficiency of obtaining similarity, and thus improve the efficiency of calculating the matching degree between the song and the target playlist.

[0082] Optionally, in order to reduce the storage resources occupied by the corresponding relationship, the corresponding relationship can only store the similarity between songs whose popularity values ​​reach the popularity threshold. The popularity value can be obtained from the recent number of times the song has been played. For example, the similarity between songs with popularity values ​​in the top 1000 can be stored in the corresponding relationship. In this way, the storage resource occupation of the corresponding relationship can be reduced while meeting the needs of most users for song similarity calculation, and the efficiency of obtaining song similarity is further improved, thereby improving the efficiency of calculating the matching degree between songs and object playlists.

[0083] If the similarity between the target song and the reference song is not found in the corresponding relationship, the similarity between the target song and the reference song can be calculated based on the target music features of the target song and the reference music features of the reference song. The music features of each song can also be pre-stored in the server. In this way, even for a few songs for which similarity cannot be found in the comparison relationship, the music features can be quickly obtained to calculate the similarity.

[0084] Step 2012: Determine a first number of similar songs in each reference song in descending order of corresponding similarities.

[0085] After calculating the similarity between the target song and each reference song in each target playlist created by the user, the reference songs can be sorted in descending order of similarity, and the songs with the first number of similarities ranked first are selected as the similar songs. The specific value of the first number can be pre-set by a technician. Alternatively, the first number can be determined by the number of reference songs included in each target playlist created by the user. For example, the first number can be determined as one-third of the number of reference songs.

[0086] Step 2013: Select a candidate playlist from each target playlist based on the number of similar songs included in each target playlist.

[0087] Optionally, after determining the similar songs, the number of similar songs included in each target playlist can be determined. The greater the number of similar songs included in a target playlist, the more songs in the target playlist that are highly similar to the target song, and thus the higher the degree of match between the target song and the target playlist. Therefore, the number of similar songs included in each target playlist can be determined as the degree of match between the target song and the target playlist.

[0088] Considering that different similar songs have different similarities with the target song, this may also affect the matching degree between the target song and the target song. Therefore, the similarity between each similar song and the target song can be added to the calculation of the matching degree between the target song and the target song. In one example, for each target playlist, based on the number of similar songs included and the similarity between each similar song and the target song, the average similarity between the included similar songs and the target song is determined. The average similarity is used to indicate the matching degree between the target song and the target playlist.

[0089] In one example, if there are two or more object playlists that include the same number of similar songs, the average similarity between the similar songs in each object playlist and the target songs can be further calculated based on the similarity between the similar songs in each object playlist and the target songs, as well as the number of similar songs included in each object playlist, and then the object playlist with the highest average similarity can be determined as the object playlist with a higher degree of match.

[0090] In another example, the degree of match between the target song and each object playlist created by the user can be obtained by calculating the average similarity between similar songs in each object playlist and the target song. That is, the average similarity between similar songs in each object playlist and the target song is determined as the degree of match between each object playlist and the target song. In this case, if there are two or more object playlists with the same average similarity between similar songs and the target song, the object playlist containing more similar songs can be determined as the object playlist with a higher degree of match.

[0091] After determining the degree of match between the target song and the target playlist, the server or terminal can sort the target playlists created by the user in descending order of the degree of match, and determine the second number of target playlists with the highest order as candidate playlists. The value of the second number can be pre-set by a technician, such as 3, 5, etc., and the embodiment of the present application does not limit its specific value.

[0092] Method 2: The target song can be matched with the target playlist based on the playlist features of the target playlist and the target music features, wherein the playlist features of the target playlist can be the average music features corresponding to the reference music features of the reference songs included in the target playlist. The corresponding processing includes:

[0093] Step 201a: Determine the average music features of the reference songs included in each object playlist based on the reference music features of the reference songs included in each object playlist.

[0094] During implementation, the server may calculate or obtain the reference music features of each reference song in each target playlist. For example, the reference music features may be music embeddings. For the reference music features obtained for the reference songs, the multiple reference music features may be summed, and then the average music features of the multiple reference music features may be calculated based on the number of reference songs. For example, the embeddings corresponding to the multiple reference songs may be added together and then divided by the number of reference songs to obtain the average embedding. The average music features of the reference songs are the playlist features of the target playlist, that is, the average embedding is the playlist embedding.

[0095] Step 201b: Select a candidate playlist from each object playlist based on the similarity between the target music feature and the average music feature corresponding to each object playlist.

[0096] After obtaining the playlist features of each object playlist (i.e., the average music features of the reference songs), the server can calculate the similarity between the target music features of the target song and the corresponding average music features of each object playlist, and then determine the similarity as the matching degree between the target song and each object playlist.

[0097] In one example, the server can pre-calculate the playlist features corresponding to each target playlist and store them in association with the user ID and playlist ID. This way, each time the match between the target song and each target playlist needs to be calculated, the playlist features of the target playlist can be directly obtained based on the user ID and the target playlist ID. This eliminates the need to calculate the playlist features, thereby improving the efficiency of calculating the match between the target song and each target playlist. Furthermore, when new songs are added to each target playlist, the server can update the playlist features of the target playlist, thereby ensuring the accuracy of the playlist features.

[0098] After determining the matching degree between the target song and the object playlist, the server or terminal can sort the object playlists created by the user in descending order of matching degree, and determine the second-number object playlists with the highest ranking as candidate playlists.

[0099] Step 202: Push the candidate playlist to the target object so that the candidate playlist is displayed in the music application.

[0100] When step 202 is executed by the server, after determining the candidate playlist, the server may send the playlist to the terminal, and the terminal may display the received candidate playlist in the music application. When step 202 is executed by the terminal, after determining the candidate playlist, the terminal may display the received candidate playlist in the music application. The candidate playlist sent by the server may be the ID of the candidate playlist, and the candidate playlist displayed by the terminal is the playlist name of the candidate playlist.

[0101] Figure 4 This is a schematic diagram of a playlist display provided by an embodiment of the present application. For example, a user creates 9 playlists including playlist 1 to playlist 9. For example, after sorting the 10 playlists from high to low according to the matching degree, the corresponding order may be playlist 3, playlist 5, playlist 8, playlist 1, playlist 2, playlist 4, playlist 9, playlist 7, playlist 6. Assuming that the second number is set to 3, then in Figure 4In the example, playlist 3, playlist 5, and playlist 8 may be displayed. In other words, the user is most likely to add the target song to playlist 3, followed by playlist 5 and playlist 8. This way, the user does not need to search through multiple playlists to find the song to add, which can improve the efficiency of adding songs to playlists.

[0102] In one example, if the target song has been added to the target playlist created by the user, the playlist with the target song added can be displayed in front of other playlists and can be displayed in a different way from other playlists, and the user cannot select the playlist, and the user can be prompted that the target song has been added to the playlist. Figure 4 For example, Figure 5 This is a schematic diagram of another playlist display provided by an embodiment of the present application. Assuming that the target song has been added to playlist 1, playlist 1 can be displayed above playlists 3, 5, and 8, and displayed with dotted lines. This prevents users from adding songs to playlists repeatedly and allows the same song to be flexibly added to different playlists.

[0103] Step 203: When a selection confirmation request for a target song in the candidate song list is received, the target song is added to the target song list.

[0104] After the candidate playlists are displayed in the music application, the user can select the target playlist to which the target song needs to be added from the candidate playlists. Among them, when the terminal executes step 203, the selection confirmation request can be a request for the user to select the target playlist in the music application. In response to the request for the target playlist, the music application can add the target song to the target playlist and send a selection confirmation request to the server to instruct the server to synchronously add the target song to the target playlist. When the server executes step 203, the selection confirmation request can be a playlist confirmation request sent by the terminal to the server after the user selects the target playlist in the music application. In response to the selection confirmation request, the server can add the target song to the target playlist.

[0105] In an embodiment of the present application, when a user needs to add a target playlist to at least one of the last playlists created, the degree of match between the target audio features of the target song and each of the created playlists can be determined, and then candidate playlists can be displayed in the music application in descending order of match. In this way, the playlist with the higher the match, the more likely it is that the target song needs to be added to the playlist. In this way, the candidate playlist that best matches the target song is displayed first, so the user does not need to search through a large number of playlists to find the playlist to which the target song needs to be added, thereby improving the efficiency of adding songs to the playlist.

[0106] Figure 6Schematic diagram of training the first encoder, the second encoder and the multimodal feature fusion module provided in an embodiment of the present application. Figure 6 As shown in Figure 1, the sample classification label can be input into the first encoder, and the sample audio data after frequency domain conversion can be input into the second encoder. The first encoder can output the corresponding text embedding, and the second encoder can output the corresponding audio embedding. The text and audio embeddings can serve as input to the multimodal feature fusion module, which outputs the music embedding.

[0107] In one example, a technician can set a baseline music embedding corresponding to the sample classification label and sample audio data, then determine the loss between the music embedding output by the multimodal feature fusion module and the baseline music embedding. Based on this loss, the weight coefficients included in the first encoder, second encoder, and multimodal feature fusion module are then adjusted inversely. In another example, a technician can set a music embedding for a second sample song that has a high similarity to the first sample song corresponding to the sample classification label and sample audio data. The similarity between the music embedding of the first sample song and the music embedding of the second sample song can then be calculated. Based on the calculated similarity and the loss value for the baseline similarity between the first and second sample songs, the weight coefficients included in the first encoder, second encoder, and multimodal feature fusion module are then adjusted inversely. After training on a large number of sample songs in this manner, the outputs of the first encoder, second encoder, and multimodal feature fusion module gradually converge, thereby completing the training of the first encoder, second encoder, and multimodal feature fusion module.

[0108] In one feasible manner, the sample songs for training the first encoder, the second encoder and the multimodal feature fusion module can be selected from the music library of the music application. In this way, for each user using the music application, the trained first encoder, the second encoder and the multimodal feature fusion module can be used to calculate the music features corresponding to the song. In another feasible manner, the songs included in the playlist created by each user can be used as sample songs for training the first encoder, the second encoder and the multimodal feature fusion module. In this way, each user can be trained with the user's own first encoder, the second encoder and the multimodal feature fusion module, thereby improving the accuracy of calculating the matching degree between the song and the playlist for the user. In addition, in order to reduce cost, the exclusive first encoder, the second encoder and the multimodal feature fusion module can be trained only for users who create a number of playlists greater than the corresponding number threshold (such as users who create more than 30 playlists).

[0109] All the above optional technical solutions can be arbitrarily combined to form optional embodiments of the present disclosure, and will not be described in detail here.

[0110] Figure 7 This is a schematic diagram of the structure of a device for adding songs to a playlist provided in an embodiment of the present application. The device can be configured in the above-mentioned terminal and / or server, such as Figure 7 As shown, the device includes:

[0111] A selection module 710 is configured to, upon receiving a playlist addition request for a target song triggered in a music application, select a candidate playlist from each target playlist of a target object logged into the music application based on a similarity between a target music feature of the target song and a reference music feature of a reference song, wherein the target music feature is obtained by fusing an audio feature and a classification label feature of the target song, the reference music feature is obtained by fusing an audio feature and a classification label feature of a reference song, and the reference song is a song in the target playlist;

[0112] A push module 720 is configured to push the candidate playlist to the target object so as to display the candidate playlist in the music application;

[0113] The adding module 730 is used to add the target song to the target song list when receiving a selection confirmation request for the target song list in the candidate song list.

[0114] Optionally, the selection module 710 is used to determine the similarity between the target music feature and the reference music feature corresponding to each reference song in each object playlist; determine a first number of similar songs in each reference song in order of corresponding similarities from high to low; and select a candidate playlist in each object playlist based on the number of similar songs included in each object playlist.

[0115] Optionally, the selection module 710 is used to determine, for each object playlist, an average similarity between similar songs included in the object playlist and the target song; and select a candidate playlist from each object playlist based on the average similarity corresponding to each object playlist.

[0116] Optionally, the above-mentioned device also includes a search module for searching the similarity between the target song and the reference song in a corresponding relationship based on the target song identifier corresponding to the target song and the reference song identifier of the reference song, wherein the corresponding relationship stores the similarity between each song in the song library of the music application and the corresponding song identifier.

[0117] Optionally, the songs included in the corresponding relationship are songs in the song library of the music application whose corresponding popularity values ​​exceed a popularity value threshold. The above-mentioned device also includes a calculation module for obtaining the target music features of the target song and the reference music features of the reference song when the similarity between the target song and the reference song is not found in the corresponding relationship, and calculating the similarity between the target song and the reference song based on the target music features and the reference song features.

[0118] Optionally, the selection module 710 is used to determine the average music features of each reference song included in each object playlist based on the reference music features of each reference song included in each object playlist; and select a candidate playlist from each object playlist based on the similarity between the target music features and the average music features corresponding to each object playlist.

[0119] Optionally, the above-mentioned device also includes a feature extraction module, which is used to obtain the classification label of the target song and extract the classification label features of the classification label; extract the target music features of the target song based on the audio data of the target song; and fuse the classification label features and target music features corresponding to the target song to obtain the target music features of the target song, and the fusion processing includes weighted summation and feature splicing.

[0120] It should be noted that: the device for adding songs to a playlist provided in the above embodiment only uses the division of the above functional modules as an example to illustrate when adding songs to a playlist. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the device for adding songs to a playlist provided in the above embodiment and the method embodiment for adding songs to a playlist belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0121] Figure 8 FIG1 shows a schematic diagram of the structure of a computer device provided by an exemplary embodiment of the present application. The computer device may be an electronic device, that is, a terminal in the above embodiment, such as Figure 8As shown, electronic device 800 may be a portable mobile terminal, such as a smartphone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, or a desktop computer. Electronic device 800 may also be referred to as user equipment, a portable terminal, a laptop terminal, a desktop terminal, or other similar names.

[0122] Typically, the electronic device 800 includes a processor 801 and a memory 802 .

[0123] The processor 801 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 801 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 801 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 801 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 801 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.

[0124] The memory 802 may include one or more computer-readable storage media, which may be non-transitory. The memory 802 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 802 is used to store at least one instruction, which is used to be executed by the processor 801 to implement the method for adding songs to a playlist provided in the method embodiment of the present application.

[0125] In some embodiments, electronic device 800 may optionally include a peripheral device interface 803 and at least one peripheral device. The processor 801, memory 802, and peripheral device interface 803 may be connected via a bus or signal lines. Each peripheral device may be connected to peripheral device interface 803 via a bus, signal lines, or circuit boards. Specifically, the peripheral device may include at least one of a radio frequency circuit 804, a display screen 805, a camera assembly 806, an audio circuit 807, a positioning assembly 808, and a power supply 809.

[0126] The peripheral device interface 803 can be used to connect at least one I / O (Input / Output)-related peripheral device to the processor 801 and the memory 802. In some embodiments, the processor 801, the memory 802, and the peripheral device interface 803 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 801, the memory 802, and the peripheral device interface 803 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0127] The radio frequency circuit 804 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 804 communicates with communication networks and other communication devices via electromagnetic signals. The radio frequency circuit 804 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. Optionally, the radio frequency circuit 804 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The radio frequency circuit 804 can communicate with other terminals via at least one wireless communication protocol. Such wireless communication protocols include, but are not limited to, the World Wide Web, a metropolitan area network, an intranet, various generations of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 804 may also include circuits related to NFC (Near Field Communication), which is not limited in this application.

[0128] The display screen 805 is used to display a user interface (UI). This UI may include graphics, text, icons, videos, or any combination thereof. When the display screen 805 is a touch screen, it is also capable of collecting touch signals on or above the surface of the display screen 805. These touch signals can be input as control signals to the processor 801 for processing. In this case, the display screen 805 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there can be one display screen 805, located on the front panel of the electronic device 800. In other embodiments, there can be at least two display screens 805, located on different surfaces of the electronic device 800 or in a foldable design. In other embodiments, the display screen 805 can be a flexible display, located on a curved or foldable surface of the electronic device 800. Furthermore, the display screen 805 can be configured as a non-rectangular irregular shape, i.e., a special-shaped screen. The display screen 805 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0129] The camera assembly 806 is used to capture images or videos. Optionally, the camera assembly 806 includes a front camera and a rear camera. Typically, the front camera is arranged on the front panel of the terminal, and the rear camera is arranged on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize panoramic shooting and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera assembly 806 may also include a flash. The flash can be a monochrome temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.

[0130] The audio circuit 807 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals to be input into the processor 801 for processing, or input into the radio frequency circuit 804 to achieve voice communication. For the purpose of stereo sound collection or noise reduction, there can be multiple microphones, which are respectively arranged in different parts of the electronic device 800. The microphone can also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signals from the processor 801 or the radio frequency circuit 804 into sound waves. The speaker can be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signals into sound waves audible to humans, but also convert the electrical signals into sound waves inaudible to humans for purposes such as ranging. In some embodiments, the audio circuit 807 may also include a headphone jack.

[0131] Positioning component 808 is used to locate the current geographic location of electronic device 800 to implement navigation or LBS (Location Based Service). Positioning component 808 can be a positioning component based on the US GPS (Global Positioning System), China's Beidou system, or Russia's Galileo system.

[0132] Power supply 809 is used to power the various components of electronic device 800. Power supply 809 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 809 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, while a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0133] In some embodiments, the electronic device 800 further includes one or more sensors 810 , including but not limited to: an acceleration sensor 811 , a gyroscope sensor 812 , a pressure sensor 813 , a fingerprint sensor 814 , an optical sensor 815 , and a proximity sensor 816 .

[0134] The accelerometer 811 can detect the magnitude of acceleration along the three coordinate axes of the coordinate system established by the electronic device 800. For example, the accelerometer 811 can be used to detect the components of gravity acceleration along the three coordinate axes. The processor 801 can control the display screen 805 to display the user interface in a landscape or portrait view based on the gravity acceleration signal collected by the accelerometer 811. The accelerometer 811 can also be used to collect game or user motion data.

[0135] The gyroscope sensor 812 can detect the orientation and rotation angle of the electronic device 800. It can work in conjunction with the accelerometer 811 to capture the user's 3D movements of the electronic device 800. Based on the data collected by the gyroscope sensor 812, the processor 801 can implement the following functions: motion sensing (for example, changing the UI based on the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.

[0136] The pressure sensor 813 can be set on the side frame of the electronic device 800 and / or the lower layer of the display screen 805. When the pressure sensor 813 is set on the side frame of the electronic device 800, it can detect the user's grip signal of the electronic device 800, and the processor 801 performs left and right hand recognition or shortcut operations based on the grip signal collected by the pressure sensor 813. When the pressure sensor 813 is set on the lower layer of the display screen 805, the processor 801 controls the operable controls on the UI interface based on the user's pressure operation on the display screen 805. The operable controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.

[0137] The fingerprint sensor 814 is used to collect the user's fingerprint. The processor 801 identifies the user's identity based on the fingerprint collected by the fingerprint sensor 814, or the fingerprint sensor 814 identifies the user's identity based on the collected fingerprint. When the user's identity is identified as a trusted identity, the processor 801 authorizes the user to perform relevant sensitive operations, such as unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings. The fingerprint sensor 814 can be set on the front, back, or side of the electronic device 800. When a physical button or manufacturer logo is set on the electronic device 800, the fingerprint sensor 814 can be integrated with the physical button or manufacturer logo.

[0138] The optical sensor 815 is used to detect ambient light intensity. In one embodiment, the processor 801 can control the display brightness of the display screen 805 based on the ambient light intensity detected by the optical sensor 815. Specifically, when the ambient light intensity is high, the display brightness of the display screen 805 is increased; when the ambient light intensity is low, the display brightness of the display screen 805 is decreased. In another embodiment, the processor 801 can also dynamically adjust the shooting parameters of the camera assembly 806 based on the ambient light intensity detected by the optical sensor 815.

[0139] Proximity sensor 816, also known as a distance sensor, is typically located on the front panel of electronic device 800. Proximity sensor 816 is used to detect the distance between the user and the front of electronic device 800. In one embodiment, when proximity sensor 816 detects that the distance between the user and the front of electronic device 800 is gradually decreasing, processor 801 controls display screen 805 to switch from the screen-on state to the screen-off state. When proximity sensor 816 detects that the distance between the user and the front of electronic device 800 is gradually increasing, processor 801 controls display screen 805 to switch from the screen-off state to the screen-on state.

[0140] Those skilled in the art will understand that Figure 8 The structure shown in the figure does not constitute a limitation on the electronic device 800, and the electronic device 800 may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.

[0141] Figure 9 Schematic diagram of the structure of a computer device provided in an embodiment of the present application, which may be the server in the above embodiment. Figure 9 As shown, the server 900 may vary significantly due to different configurations or performances, and may include one or more processors (central processing units, CPUs) 901 and one or more memories 902, wherein the memories 902 store at least one instruction, which is loaded and executed by the processor 901 to implement the methods provided in the above-mentioned various method embodiments. Of course, the server may also have components such as a wired or wireless network interface, a keyboard, and input / output interfaces for input and output. The server may also include other components for implementing device functions, which will not be described in detail here.

[0142] The present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored on any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the method for adding songs to a playlist described in the above embodiment.

[0143] In an exemplary embodiment, a computer-readable storage medium is further provided, such as a memory including instructions, which can be executed by a processor in a terminal to complete the method of adding songs to a playlist in the above embodiment. The computer-readable storage medium can be non-transitory. For example, the computer-readable storage medium can be a ROM (Read-Only Memory), a RAM (Random Access Memory), a CD-ROM (Compact Disc Read-Only Memory), a magnetic tape, a floppy disk, an optical data storage device, etc.

[0144] It should be noted that the information involved in this application (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals (including but not limited to signals transmitted between user terminals and other devices, etc.) are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions. For example, the reference songs included in the playlist involved in this application and the calculation of the similarity between the playlist and the songs are all obtained with full authorization.

[0145] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.

[0146] It should be understood that the "multiple" mentioned in this article refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent three situations: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship. In addition, the step numbers described in this article only illustrate a possible execution sequence between the steps. In some other embodiments, the above steps may not be executed in the order of the numbers, such as two steps with different numbers are executed at the same time, or two steps with different numbers are executed in the opposite order to the diagram. The embodiments of the present application do not limit this.

[0147] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A method for adding songs to a playlist, characterized in that: The method comprises: When a playlist addition request for a target song is received, which is triggered in a music application, a candidate playlist is selected from each target playlist of a target object logged into the music application based on a similarity between a target music feature of the target song and a reference music feature of a reference song, wherein the target music feature is obtained by fusing an audio feature and a classification label feature of the target song, the reference music feature is obtained by fusing an audio feature and a classification label feature of a reference song, and the reference song is a song in the target playlist; Pushing the candidate playlist to the target object to display the candidate playlist in the music application; When a selection confirmation request for a target song list in the candidate song list is received, the target song is added to the target song list.

2. The method according to claim 1, characterized in that The selecting, based on the similarity between the target music feature of the target song and the reference music feature of the reference song, a candidate playlist from each target playlist of the target object logging into the music application program includes: Determining the similarity between the target music feature and the reference music feature corresponding to each reference song in each object playlist; Determining a first number of similar songs in each reference song in descending order of corresponding similarities; According to the number of similar songs included in each target song list, a candidate song list is selected from each target song list.

3. The method according to claim 2, characterized in that The selecting a candidate playlist from each target playlist according to the number of similar songs in each target playlist includes: For each target song list, determining an average similarity between similar songs included in the target song list and the target song; According to the average similarity corresponding to each object playlist, a candidate playlist is selected from each object playlist.

4. The method according to claim 2, characterized in that Before selecting a candidate playlist from each target playlist of a target object logged into the music application based on the similarity between the target music feature of the target song and the reference music feature of the reference song, the method further includes: According to the target song identifier corresponding to the target song and the reference song identifier of the reference song, the similarity between the target song and the reference song is searched in the corresponding relationship, wherein the corresponding relationship stores the similarity between each song in the song library of the music application and the corresponding song identifier.

5. The method according to claim 4, characterized in that The songs included in the corresponding relationship are songs in the song library of the music application whose corresponding popularity values ​​exceed the popularity value threshold, and the method further includes: If the similarity between the target song and the reference song is not found in the corresponding relationship, the target music features of the target song and the reference music features of the reference song are obtained, and the similarity between the target song and the reference song is calculated based on the target music features and the reference song features.

6. The method according to claim 1, characterized in that The selecting, based on the similarity between the target music feature of the target song and the reference music feature of the reference song, a candidate playlist from each target playlist of the target object logging into the music application program includes: Determining an average music feature of each reference song included in each target playlist based on the reference music features of each reference song included in each target playlist; According to the similarity between the target music feature and the average music feature corresponding to each object playlist, a candidate playlist is selected from each object playlist.

7. The method according to any one of claims 1 to 6, characterized in that Before selecting a candidate playlist from each target playlist of a target object logged into the music application based on the similarity between the target music feature of the target song and the reference music feature of the reference song, the method further includes: Obtaining a classification label of the target song and extracting classification label features of the classification label; Extracting target music features of the target song based on the audio data of the target song; The classification label features and target music features corresponding to the target song are fused to obtain the target music features of the target song, and the fusion process includes weighted summation and feature splicing.

8. A computer device, characterized in that: The computer device includes a processor and a memory, wherein the memory stores at least one instruction, and the instruction is loaded and executed by the processor to implement the operations performed by the method for adding songs to a playlist as described in any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that The storage medium stores at least one instruction, which is loaded and executed by the processor to implement the operations performed by the method for adding songs to a playlist as described in any one of claims 1 to 7.

10. A computer program product, characterized in that The computer program product stores at least one instruction, which is loaded and executed by a processor to implement the operations performed by the method for adding songs to a playlist as described in any one of claims 1 to 7.