Song list generation method and device and storage medium
By analyzing the song sentimentality and preference characteristics associated with user operation behavior, selecting songs with similar characteristics in the song library to generate recommended song lists, solving the problem of inefficient playlist generation in the existing technology and achieving efficient personalized playlist recommendations.
Patent Information
- Application Number
- CN202510002899.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, the generation of playlists requires manual combination of operators, which is inefficient and difficult to meet the needs of personalized recommendations.
By determining the songs associated with the operation behavior of the user account and their emotionality, select songs that meet the target type of emotions, extract their preference characteristics, and select songs with similar characteristics in the song library to generate a recommended song list.
It realizes that the target type emotional playlists can be automatically generated without the participation of operators, and improves the efficiency of playlist generation.
Smart Images

Figure CN119988669A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of audio technology, and in particular to a method, device and storage medium for generating a playlist. Background Art
[0002] A playlist is a collection of songs with common characteristics. Recommending a playlist to users is an important way to recommend songs. At present, playlists are generally generated by operators who combine songs according to the attributes of the songs and the attributes of the users. For example, for users born in the 1990s, songs that were released after 2000 and were popular at the time are obtained to form a playlist of "Exclusive Memories of the Post-90s". The playlist generated by the above solution requires manual work by operators, which is very inefficient. Summary of the invention
[0003] The embodiment of the present application provides a method, device and storage medium for generating drum audio, which can automatically generate a playlist recommended to different users, with higher efficiency. The technical solution is as follows:
[0004] In a first aspect, a method for generating a playlist is provided, the method comprising:
[0005] Determine at least one song associated with the operation behavior of the user account, and determine the emotion degree of each song in the at least one song in the target emotion type;
[0006] In a case where it is determined according to the emotion level of the at least one song that the user account meets the song list recommendation condition corresponding to the target type emotion, a target song that meets the screening condition of the target type emotion is selected from the at least one song according to the emotion level of the at least one song;
[0007] Determine the preference characteristics of the target song under the target type emotion;
[0008] Select songs in the music library whose emotional preference characteristics and the preference characteristics of the target song under the target type emotion meet the preset similarity conditions to generate a recommended song list corresponding to the user account.
[0009] In a second aspect, a model training method is provided, the method comprising:
[0010] Obtain sample songs and the emotion labels corresponding to each sample song in the music library;
[0011] Extracting audio features of each of the sample songs;
[0012] Inputting the audio features of the sample song into the emotion classification model to be trained to obtain a predicted emotion degree, wherein the predicted emotion degree is used to indicate the degree of a specific target type of emotion presented by the sample song predicted by the emotion classification model to be trained;
[0013] The emotion classification model to be trained is trained according to the predicted emotion degree and the emotion label corresponding to the sample song to obtain a trained emotion classification model.
[0014] In a third aspect, a model training method is provided, the method comprising:
[0015] Obtaining a favorite playlist of at least one user account in the music library;
[0016] For each of the collection song lists, determine the target emotion level corresponding to each song in the collection song list, and select sample songs whose corresponding target emotion levels meet the emotion condition of the second target type from the collection song list to form a sample song list corresponding to the collection song list;
[0017] For each of the sample playlists, every two sample songs in the sample playlist form a positive sample pair, and a sample song in the positive sample pair and another song randomly obtained from the music library form a negative sample pair corresponding to the positive sample pair, wherein the other song is a song outside the sample playlist;
[0018] Based on each of the positive sample pairs and the negative sample pairs corresponding to the positive sample pairs, the emotion preference model to be trained is trained to obtain a trained emotion preference model.
[0019] In a fourth aspect, a device for generating a playlist is provided, the device comprising:
[0020] A determination module is used to determine at least one song associated with the operation behavior of the user account and determine the emotional degree of each song in the at least one song in the target type emotion; when it is determined that the user account meets the song list recommendation condition corresponding to the target type emotion according to the emotional degree of the at least one song, select a target song that meets the screening condition of the target type emotion from the at least one song according to the emotional degree of the at least one song; and determine the preference characteristics of the target song under the target type emotion;
[0021] A generation module is used to select songs in the music library whose preference characteristics meet preset similarity conditions with the preference characteristics of the target song under the target type emotion, so as to generate a recommended song list corresponding to the user account.
[0022] In a fifth aspect, a model training device is provided, the device comprising:
[0023] An acquisition module is used to acquire sample songs and the emotion label corresponding to each sample song in the music library;
[0024] Extracting audio features of each of the sample songs;
[0025] Inputting the audio features of the sample song into the emotion classification model to be trained to obtain a predicted emotion degree, wherein the predicted emotion degree is used to indicate the degree of a specific target type of emotion presented by the sample song predicted by the emotion classification model to be trained;
[0026] The training module is used to train the emotion classification model to be trained according to the predicted emotion degree and the emotion label corresponding to the sample song to obtain a trained emotion classification model.
[0027] In a sixth aspect, a device for model training is provided, the device comprising:
[0028] An acquisition module is used to acquire the favorite song list of at least one user account in the music library; for each of the favorite song lists, determine the target emotion level corresponding to each song in the favorite song list, and screen out sample songs whose corresponding target emotion levels meet the emotion condition of the second target type in the favorite song list to form a sample song list corresponding to the favorite song list; for each of the sample song lists, form a positive sample pair with every two sample songs in the sample song list, and form a negative sample pair corresponding to the positive sample pair with a sample song in the positive sample pair and another song randomly acquired from the music library, wherein the other song is a song outside the sample song list;
[0029] The training module is used to train the emotion preference model to be trained based on each positive sample pair and the negative sample pair corresponding to the positive sample pair to obtain a trained emotion preference model.
[0030] In the seventh aspect, a computing device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction, and the instruction is loaded and executed by the processor to implement the operations performed by the playlist generation method described in the first aspect or the operations performed by the model training method described in the second aspect or the operations performed by the model training method described in the third aspect.
[0031] In the eighth aspect, a computer-readable storage medium is provided, wherein the storage medium stores at least one instruction, which is loaded and executed by a processor to implement the operations performed by the playlist generation method described in the first aspect or the operations performed by the model training method described in the second aspect or the operations performed by the model training method described in the third aspect.
[0032] In the ninth aspect, a computer program product is provided, wherein at least one instruction is stored in the computer program product, and the instruction is loaded and executed by a processor to implement the operations performed by the playlist generation method described in the first aspect or the operations performed by the model training method described in the second aspect or the operations performed by the model training method described in the third aspect.
[0033] The beneficial effects of the technical solution provided by this application are:
[0034] In the method for generating a playlist provided in an embodiment of the present application, for a target user, the degree of the target type emotion presented by the songs that the user has historically operated is first determined. Then, based on the degree of the target type emotion presented by the songs that the user has operated, it is determined whether a playlist related to the target type emotion needs to be recommended to the user. When it is determined that a playlist related to the target type emotion needs to be recommended to the user, songs that match the target type emotion presented by the user's historical operations are selected from the music library to generate a recommended playlist. It can be seen that the entire playlist generation process does not require the participation of operators, and a playlist of the target type emotion that meets the user's preferences is automatically generated, and the playlist generation efficiency is higher. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0036] Figure 1 This is a flow chart of a method for generating a song list provided in an embodiment of the present application;
[0037] Figure 2 This is a flow chart of a method for training a sentiment classification model provided in an embodiment of the present application;
[0038] Figure 3 This is a flow chart of a method for training an emotion preference model provided in an embodiment of the present application;
[0039] Figure 4 This is a flowchart of a method for generating a song list provided in an embodiment of the present application;
[0040] Figure 5 This is a schematic diagram of the structure of a device for generating a song list provided in an embodiment of the present application;
[0041] Figure 6 is a schematic diagram of the structure of a terminal provided in an embodiment of the present application;
[0042] Figure 7 It is a structural diagram of a server provided in an embodiment of the present application. DETAILED DESCRIPTION
[0043] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0044] An embodiment of the present application provides a method for generating a playlist, which can be implemented by a computing device. The computing device can be a terminal or a server. The terminal can be a mobile phone, a tablet computer, a desktop computer, a laptop computer, etc., and the server can be a single server, a server cluster, etc.
[0045] The implementation scenario of the embodiment of the present application is described below. The terminal may be installed with a music application, and the user may log in to the target user account in the music application. The user may search for the target type emotional song list in the music application, for example, searching for a sad song list. Furthermore, the terminal sends a request to the server to obtain the target type emotional song list, and the target user account may be carried in the request. After receiving the request, the server obtains the target type emotional song list corresponding to the target user account and returns it to the terminal. Alternatively, there is a song list recommendation interface in the music application, and at least one type of emotional option is displayed in the song list recommendation interface, such as sadness, happiness, excitement, etc. The user can select any type of emotional option. After detecting the user's selection operation on the target type emotional option, the terminal sends a request to the server to obtain the target type emotional song list, and the target user account may be carried in the request. After receiving the request, the server obtains the target type emotional song list corresponding to the target user account and returns it to the terminal.
[0046] In the above example, the server obtains the target type emotional song list corresponding to the target user account. It can be to obtain the target type emotional song list corresponding to the target user account that has been generated, or it can be to generate the target type emotional song list corresponding to the target user account in real time. In either case, the song list generation method provided in the embodiment of the present application can be used to achieve it.
[0047] In the method for generating a song list provided in an embodiment of the present application, for a target user, the degree of the target type emotion presented by the songs that the user has historically operated is first determined, and the target type emotion can be sad, happy, passionate, etc. Then, based on the degree of the target type emotion presented by the songs that the user has operated, it is determined whether it is necessary to recommend a song list related to the target type emotion to the target user. When it is determined that a song list related to the target type emotion needs to be recommended to the target user, the target type emotion preference features of the target songs that the user has operated are first extracted, and then the songs to be recommended that match the target type emotion preference features are searched in the music library. Finally, a song list consisting of the target type emotion songs to be recommended is generated. It can be seen that the entire song list generation process does not require the participation of operators, and a song list of target type emotions that meets the user's preferences is automatically generated, and the song list generation efficiency is higher.
[0048] The following is an explanation of the drum audio generation method provided in the embodiment of the present application. The method can be implemented by a computing device. Figure 1 , the processing flow of the method may include the following steps:
[0049] Step 101: Determine the emotion level of at least one song associated with the operation behavior of the user account.
[0050] The user account may be an account of any user registered in the music application, and the emotion level is used to indicate the degree of the target type emotion presented by the corresponding song, and the target type emotion may be sadness, happiness, excitement, etc.
[0051] In implementation, the computing device determines at least one song associated with the operation behavior of the user account, wherein the operation behavior may be one or more of the operations of finishing, collecting, sharing, commenting, clicking, etc. Then, for the determined songs, the emotion level corresponding to each song is determined respectively.
[0052] Exemplarily, the computing device may query a database for all songs that have been played by the target user account, and then, for each song that has been played by the target user account, determine the emotion level corresponding to the song.
[0053] The following is an explanation of the method for determining the emotion level of a song:
[0054] The computing device first extracts the audio features of the song, and then inputs the extracted audio features into the pre-trained emotion classification model, which outputs the emotion level corresponding to the song. The emotion level corresponding to the song is the confidence level of the song presenting the target type of emotion, which can be used to indicate the degree to which the song presents the target type of emotion.
[0055] There are many methods for extracting audio features of songs, such as MFCC (Mel-Frequency Cepstral Coefficients) algorithm, neural network model, FT (Fourier Transform), STFT (Short-Time Fourier Transform), FFT (Fast Fourier Transform), etc.
[0056] Combine the following Figure 2 Explain the training of the sentiment classification model:
[0057] First, sample songs and the emotion tags corresponding to each sample song are obtained in the music library. Specifically, each song in the music library has an emotion tag corresponding to it. The emotion tag can be annotated by relevant personnel to indicate the emotion presented by the song. For each emotion tag, multiple songs corresponding to the emotion tag can be obtained in the music library as sample songs.
[0058] Then, for each sample song, the audio features of the sample song are extracted. Specifically, there can be multiple extraction methods, such as MFCC algorithm, neural network model, FT, STFT, FFT, etc. In addition, the method for extracting the audio features of the sample song here needs to be the same as the method for extracting the audio features of the song when determining the target emotion of the song in the above step 101.
[0059] Then, the audio features of the sample song are input into the emotion classification model to be trained to obtain a predicted emotion degree, which is used to indicate the degree of a target type of emotion presented by the sample song predicted by the emotion classification model to be trained.
[0060] Next, the loss value is calculated based on the predicted sentiment and the sentiment label corresponding to the sample song, and then the initial sentiment classification model is trained based on the loss value. After the training stop condition is reached, the training is stopped to obtain the sentiment classification model.
[0061] Step 102: When it is determined that the user account meets the playlist recommendation condition corresponding to the target type emotion based on the emotion level of at least one song, a target song that meets the screening condition of the target type emotion is selected from at least one song based on the emotion level of at least one song.
[0062] In implementation, the computing device determines whether the user account meets the song list recommendation conditions corresponding to the target type emotion based on the emotion level of each song in at least one song obtained in step 101. Specifically, the computing device can calculate the target emotion level of the user account based on the emotion level of each song in at least one song obtained in step 101. If the target emotion level of the target user account reaches the first threshold, it is determined that the user account meets the song list recommendation conditions corresponding to the target type emotion. If the target emotion level of the user account does not reach the first threshold, it is determined that the user account does not meet the song list recommendation conditions corresponding to the target type emotion, wherein the first threshold can be configured by relevant personnel according to actual needs. For example, for the following calculation method one, the first threshold can be 0.9, and for the following calculation method two, the first threshold can be 12. There are multiple methods for calculating the target emotion score corresponding to a user account, and two of them are listed below for illustration:
[0063] Calculation method 1: Calculate the average value of the sentiment level of each song in the at least one song mentioned above as the target sentiment level of the user account. The calculation process can be shown in the following formula (1):
[0064]
[0065] Among them, score is the target sentiment of the user account, J is the number of songs with at least one song mentioned above, and emo_score j is the emotion degree corresponding to the jth song.
[0066] Calculation method 2: For each of the at least one song mentioned above, multiply the emotion level of the song by the playing time of the song to obtain the weighted emotion level corresponding to the song, where the playing time of the song refers to the total time the user account plays the song, and the unit can be minutes. Then calculate the average value of each weighted emotion level as the target emotion level of the user account. The calculation process can be shown in the following formula (2):
[0067]
[0068] Where score is the target sentiment score of the user account, J is the number of songs with at least one of the above songs, and dur j is the playing time of the jth song, emo_score j is the emotion level of the jth song.
[0069] When it is determined that the user account meets the song list recommendation condition of the target type emotion, a target song that meets the screening condition of the target type emotion is selected from the songs associated with the operation behavior of the user account. There may be multiple screening conditions, and three of them are listed below as examples:
[0070] Condition 1: Select the song with the highest emotion from the songs associated with the operation behavior of the user account as the target song that meets the screening condition of the target type emotion.
[0071] Condition 2: Select the song with the highest weighted sentiment from the songs associated with the operation behaviors of the user account as the target song that meets the screening condition of the target type sentiment.
[0072] Condition 3: Select N songs with the highest weighted sentiment (or sentiment) from the songs associated with the operation behavior of the user account, and randomly select a song from these N songs as the target song that meets the screening condition of the target type sentiment. Wherein, N is a positive integer, and the value can be configured by relevant personnel according to needs, for example, N = 5.
[0073] Step 103: Determine the preference characteristics of the target song under the target type emotion.
[0074] In implementation, after determining the target song, the audio features of the target song are first extracted, and then the audio features of the target song are input into a pre-trained emotion preference model, and the emotion preference model outputs the preference features of the target song under the target type emotion.
[0075] There are many methods for extracting the audio features of the target song, such as MFCC algorithm, neural network model, FT, STFT, FFT, etc.
[0076] The sentiment preference model may be a resnet (Residual Network) model, a DenseNet (Densely Connected Convolutional Networks), or the like.
[0077] Combine the following Figure 3 Explain the training of the sentiment preference model:
[0078] S301. Obtain a favorite song list of at least one user account in a music library.
[0079] In implementation, the computing device may select all or part of the user accounts from the registered accounts, and for each user account, obtain at least one favorite playlist of the user account.
[0080] S302. For each collection song list, determine the emotional level corresponding to each song in the collection song list, filter out sample songs whose corresponding emotional levels meet the target type emotional conditions, and form the sample songs filtered out from the collection song list into a sample song list corresponding to the collection song list.
[0081] In implementation, for each favorite song list determined in step S301, the audio features of each song in the favorite song list can be extracted first, and the audio features of each song can be input into the above-mentioned trained emotion classification model to obtain the emotion level corresponding to each song in the favorite song list. If the emotion level corresponding to the song is greater than the second threshold, it is determined that the song meets the target type emotion condition, and the song is added as a sample song to the sample song list corresponding to the favorite song list. The second threshold can be configured by relevant personnel according to actual needs. For example, the second threshold can be 0.8.
[0082] In a possible implementation, after obtaining the sample playlist, the sample playlist can be screened according to the number of sample songs included in the sample playlist. For example, if the number of sample songs included in the sample playlist is less than the third threshold, the sample playlist is removed. The third threshold can be configured by relevant personnel according to actual needs. For example, the second threshold can be 5.
[0083] S303. For each sample playlist, every two sample songs in the sample playlist form a positive sample pair, and a sample song in the positive sample pair and another song randomly obtained from the music library form a negative sample pair corresponding to the positive sample pair, wherein the other songs are songs outside the sample playlist.
[0084] In implementation, for each sample playlist, every two sample songs in the sample playlist are taken as a positive sample pair. For example, if the sample playlist includes song 1, song 2, and song 3, song 1 and song 2 are taken as a positive sample pair, song 1 and song 3 are taken as a positive sample pair, and song 2 and song 3 are taken as a positive sample pair.
[0085] For each positive sample pair, a song in the positive sample pair and another song randomly obtained from the music library form a negative sample pair corresponding to the positive sample pair.
[0086] S304: Based on each positive sample pair and the negative sample pair corresponding to the positive sample pair, the emotion preference model to be trained is trained to obtain a trained emotion preference model.
[0087] In the implementation, each positive sample pair and the negative sample pair corresponding to the positive sample pair are taken as a group of samples. For each group of samples, the audio features of each song in the group of samples are extracted respectively, and the audio features of each song in the group of samples are input into the emotion preference model to be trained respectively. The emotion preference model to be trained outputs the estimated emotion preference features of each song. Then, for each group of samples, the estimated emotion preference features of each song in the group of samples are input into the loss function to obtain the loss value. Then, according to the loss value, the emotion preference model to be trained is trained. When the training stop condition is met, the training is stopped to obtain the emotion preference model.
[0088] The loss function can be expressed as follows:
[0089]
[0090] Among them, Loss is the loss value, N is the number of sample groups used in one round of training, max() is to take the maximum value, dis() is to calculate the distance between the two, and a i and p i are the audio features corresponding to the two songs in the positive sample pair in the i-th group of samples used in this round of training, n i are the audio features corresponding to the two songs in the negative sample pair in the i-th group of samples used in this round of training. margin is a parameter used to control the size of Loss. The value of margin can be configured by relevant personnel according to actual needs.
[0091] Step 104: Select songs in the music library whose preference characteristics and the preference characteristics of the target song under the target type emotion meet the preset similarity conditions to generate a recommended song list corresponding to the user account.
[0092] In implementation, for songs in the music library, the above-mentioned emotional preference model can be used to obtain the preference features of the songs under the target type of emotion in advance and store them. After obtaining the preference features of the target song under the target type of emotion in the above-mentioned step 103, the similarity between the preference features of each song in the music library under the target type of emotion and the preference features of the target song under the target type of emotion can be calculated respectively, and the songs whose calculated similarity meets the preset similarity condition are used as songs to be recommended. The similarity condition can be that the similarity is greater than a fourth threshold, wherein the fourth threshold can be configured by relevant personnel according to actual needs, for example, the fourth threshold value is 0.9.
[0093] All or part of the songs to be recommended determined will form a playlist corresponding to the user account.
[0094] For example, the number of songs in the playlist can be preconfigured, and here, the number of songs in the playlist is recorded as N. If the number of songs to be recommended is greater than N, then N songs to be recommended are selected from them to form a recommended playlist corresponding to the user account. When selecting, N songs to be recommended can be randomly selected, or N songs to be recommended whose preference characteristics under the target type emotion are most similar to the preference characteristics of the target song under the target type emotion can be selected. If the number of songs to be recommended is less than or equal to N, all the songs to be recommended that have been determined to be recommended will form a recommended playlist corresponding to the user account.
[0095] All the above optional technical solutions can be arbitrarily combined to form optional embodiments of the present disclosure, and will not be described in detail here.
[0096] In the method for generating a playlist provided in an embodiment of the present application, for a target user, the degree of the target type emotion presented by the songs that the user has historically operated is first determined. Then, based on the degree of the target type emotion presented by the songs that the user has operated, it is determined whether a playlist related to the target type emotion needs to be recommended to the user. When it is determined that a playlist related to the target type emotion needs to be recommended to the user, songs that match the target type emotion presented by the user's historical operations are selected from the music library to generate a recommended playlist. It can be seen that the entire playlist generation process does not require the participation of operators, and a playlist of the target type emotion that meets the user's preferences is automatically generated, and the playlist generation efficiency is higher.
[0097] See also Figure 4 In the method for generating a playlist provided in the embodiment of the present application, the audio features of the songs associated with the operation behavior of the user account can be used to obtain the emotional level of each song through the emotional classification model, and then, according to the emotional level of each song, the target song that meets the screening conditions is selected. Then, the audio features of the target song are input into the emotional preference model to obtain the preference features of the target song under the target type emotion. Finally, the target song is used as a seed, and more songs to be recommended are matched in the music library according to the preference features of the target song.
[0098] Based on the same technical concept, the embodiment of the present application also provides a device for generating a playlist, which can be a computing device, such as Figure 5 As shown, the device includes a determination module 510 and a generation module 520, wherein:
[0099] The determination module 510 is used to determine the emotional degree of each of at least one song associated with the operation behavior of the user account, wherein the emotional degree is used to indicate the degree of a specific target type of emotion presented by the corresponding song; when it is determined that the user account meets the song list recommendation condition corresponding to the target type of emotion according to the emotional degree of the at least one song, select a target song that meets the screening condition of the target type of emotion from the at least one song according to the emotional degree of the at least one song; and determine the preference characteristics of the target song under the target type of emotion;
[0100] The generation module 520 is used to select songs in the music library whose emotional preference characteristics and the preference characteristics of the target song under the target type emotion meet the preset similarity conditions, so as to generate a recommended song list corresponding to the user account.
[0101] In one possible implementation, the determination module 510 is used to: for each song of at least one song associated with the operation behavior of the user account, extract the audio features of the song, and input the audio features into a pre-trained sentiment classification model to obtain the sentiment degree corresponding to the song.
[0102] In a possible implementation, the determination module 510 is also used to: determine the target emotion level of the user account based on the emotion level of each song in at least one song associated with the operation behavior of the user account; if the target emotion level of the user account reaches the target emotion level threshold, determine that the target user account meets the playlist recommendation conditions corresponding to the target type emotion.
[0103] In one possible implementation, the determination module 510 is used to: multiply the emotional level of each song in at least one song associated with the operation behavior of the user account by the playback time of each song to obtain the weighted emotional level corresponding to each song; calculate the average value of the weighted emotional levels of each song in the at least one song to obtain the target emotional level of the target account.
[0104] In a possible implementation, the determination module 510 is used to select a song with the largest weighted emotion among the at least one song as a target song that meets the screening condition of the target type emotion.
[0105] In a possible implementation, the determination module 510 is used to extract the audio features of the target song, and input the audio features of the target song into a pre-trained emotion preference model to obtain the preference features of the target song under the target type emotion.
[0106] In the method for generating a playlist provided in an embodiment of the present application, for a target user, the degree of the target type emotion presented by the songs that the user has historically operated is first determined. Then, based on the degree of the target type emotion presented by the songs that the user has operated, it is determined whether a playlist related to the target type emotion needs to be recommended to the user. When it is determined that a playlist related to the target type emotion needs to be recommended to the user, songs that match the target type emotion presented by the user's historical operations are selected from the music library to generate a recommended playlist. It can be seen that the entire playlist generation process does not require the participation of operators, and a playlist of the target type emotion that meets the user's preferences is automatically generated, and the playlist generation efficiency is higher.
[0107] It should be noted that: the device for generating a playlist provided in the above embodiment only uses the division of the above functional modules as an example when generating a playlist. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computing device is divided into different functional modules to complete all or part of the functions described above. In addition, the device for generating a playlist provided in the above embodiment belongs to the same concept as the method embodiment for generating a playlist. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0108] Figure 6 The structure block diagram of a terminal 600 provided by an exemplary embodiment of the present application is shown. The terminal 600 may be a portable mobile terminal, such as a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer or a desktop computer. The terminal 600 may also be called a user device, a portable terminal, a laptop terminal, a desktop terminal or other names.
[0109] Typically, the terminal 600 includes a processor 601 and a memory 602 .
[0110] The processor 601 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 601 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 601 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 601 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 601 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.
[0111] The memory 602 may include one or more computer-readable storage media, which may be non-transitory. The memory 602 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices.
[0112] In some embodiments, the terminal 600 may further optionally include: a peripheral device interface 603 and at least one peripheral device. The processor 601, the memory 602 and the peripheral device interface 603 may be connected via a bus or a signal line. Each peripheral device may be connected to the peripheral device interface 603 via a bus, a signal line or a circuit board. Specifically, the peripheral device includes: at least one of a radio frequency circuit 604, a display screen 605, a camera assembly 606, an audio circuit 607, a positioning assembly 608 and a power supply 609.
[0113] In some embodiments, the terminal 600 further includes one or more sensors 610 , including but not limited to: an acceleration sensor 611 , a gyroscope sensor 612 , a pressure sensor 613 , a fingerprint sensor 614 , an optical sensor 615 , and a proximity sensor 616 .
[0114] Those skilled in the art will understand that Figure 6The structure shown in the figure does not constitute a limitation on the terminal 600, and the terminal 600 may include more or less components than those shown in the figure, or combine some components, or adopt a different component arrangement.
[0115] Figure 7 1 is a schematic diagram of the structure of a server provided in an embodiment of the present application. The server 1000 may have relatively large differences due to different configurations or performances, and may include one or more processors (central processing units, CPU) 1001 and one or more memories 1002, wherein the memory 1002 stores at least one instruction, and the at least one instruction is loaded and executed by the processor 1001 to implement the methods provided in the above-mentioned various method embodiments. Of course, the computing device may also have components such as a wired or wireless network interface, a keyboard, and an input and output interface for input and output, and the computing device may also include other components for implementing device functions, which will not be described in detail here.
[0116] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions, and the instructions can be executed by a processor in a terminal to complete the method for generating a playlist in the above embodiment. The computer-readable storage medium can be non-transitory. For example, the computer-readable storage medium can be a ROM (Read-Only Memory), a RAM (Random Access Memory), a CD-ROM (Compact Disc Read-Only Memory), a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0117] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals (including but not limited to signals transmitted between user terminals and other devices, etc.) involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions. For example, the relevant information of songs operated by users and the relevant information of songs in the music library involved in this application are all obtained with full authorization.
[0118] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.
[0119] The above description is only an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A method for generating a song list, characterized in that: The method comprises: Determine at least one song associated with the operation behavior of the user account, and determine the emotion degree of each song in the at least one song under the target emotion type; In a case where it is determined according to the emotion level of the at least one song that the user account meets the song list recommendation condition corresponding to the target type emotion, a target song that meets the screening condition of the target type emotion is selected from the at least one song according to the emotion level of the at least one song; Determine the preference characteristics of the target song under the target type emotion; Select songs in the music library whose preference characteristics and the preference characteristics of the target song under the target type emotion meet the preset similarity conditions to generate a recommended song list corresponding to the user account.
2. The method according to claim 1, characterized in that Determining the emotion level of each song in the at least one song under the target emotion type includes: For each song of at least one song associated with the operation behavior of the user account, the audio features of the song are extracted, and the audio features are input into a pre-trained emotion classification model to obtain the emotion degree corresponding to the song.
3. The method according to claim 1, characterized in that Before selecting a target song that satisfies the screening condition of the target type emotion from the at least one song according to the emotion level of the at least one song, the method further includes: Determining a target emotion level of the user account according to the emotion level of each song in at least one song associated with the operation behavior of the user account; If the target emotion level of the user account reaches the target emotion level threshold, it is determined that the target user account meets the song list recommendation condition corresponding to the target type emotion.
4. The method according to claim 3, characterized in that The step of determining the target emotion level of the user account according to the emotion level of each song in at least one song associated with the operation behavior of the user account includes: Multiplying the emotion level of each song in at least one song associated with the operation behavior of the user account by the playing time of each song to obtain a weighted emotion level corresponding to each song; An average value of the weighted sentiment levels of each song in the at least one song is calculated to obtain a target sentiment level of the target account.
5. The method according to claim 4, characterized in that The step of selecting a target song that meets the target type emotion screening condition from the at least one song according to the emotion level of the at least one song comprises: A song with the largest weighted emotion among the at least one song is selected as a target song that meets the screening condition of the target type emotion.
6. The method according to claim 1, characterized in that Determining the preference characteristics of the target song under the target type emotion includes: The audio features of the target song are extracted, and the audio features of the target song are input into a pre-trained emotion preference model to obtain the preference features of the target song under the target type emotion.
7. A method for model training, characterized in that: The method comprises: Obtain sample songs and the emotion labels corresponding to each sample song in the music library; Extracting audio features of each of the sample songs; Inputting the audio features of the sample song into the emotion classification model to be trained to obtain a predicted emotion degree, wherein the predicted emotion degree is used to indicate the degree of the target type emotion presented by the sample song predicted by the emotion classification model to be trained; The emotion classification model to be trained is trained according to the predicted emotion degree and the emotion label corresponding to the sample song to obtain a trained emotion classification model.
8. A method for model training, characterized in that: The method comprises: Obtaining a favorite playlist of at least one user account in the music library; For each of the collection song lists, determine the emotional degree corresponding to each song in the collection song list, and select sample songs whose corresponding emotional degrees meet the target type emotional conditions in the collection song list to form a sample song list corresponding to the collection song list; For each of the sample playlists, every two sample songs in the sample playlist form a positive sample pair, and a sample song in the positive sample pair and another song randomly obtained from the music library form a negative sample pair corresponding to the positive sample pair, wherein the other song is a song outside the sample playlist; Based on each of the positive sample pairs and the negative sample pairs corresponding to the positive sample pairs, the emotion preference model to be trained is trained to obtain a trained emotion preference model.
9. The method according to claim 8, characterized in that The number of songs included in each of the sample playlists is greater than or equal to a preset threshold.
10. A computing device, characterized in that The computing device includes a processor and a memory, wherein the memory stores at least one instruction, and the instruction is loaded and executed by the processor to implement the operation performed by the method according to any one of claims 1 to 9.
11. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, which is loaded and executed by the processor to implement the operation performed by the method according to any one of claims 1 to 9.
12. A computer program product, characterized in that The computer program product stores at least one instruction, which is loaded and executed by a processor to implement the operations performed by the method according to any one of claims 1 to 9.