Song recommendation model training method, song recommendation method, device and medium

By splitting audio clips from song audio training sets and calculating contrast loss, the audio features are independently learned, and the problem of low accuracy of the song recommendation model in the prior art is solved, achieving more efficient song recommendation.

CN115062225BActive Publication Date: 2025-07-18TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210741571.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-28
Publication Date
2025-07-18
Estimated Expiration
2042-06-28

AI Technical Summary

Technical Problem

During the training process of the existing song recommendation model, the audio characteristics and user behavior characteristics are mixed, resulting in low recommendation accuracy, especially when new songs are promoted, it cannot be accurately recommended to the appropriate user group.

Method used

By selecting song audio from the song audio training set, dividing it into multiple audio clips, using the initial model to extract embedded features, and computing the comparison loss of the same song and different song audio clips, performing parameter adjustments, forming a song recommendation model, and learning audio features independently.

Benefits of technology

It improves the performance of the song recommendation model, improves the accuracy of song recommendation, and can better recommend new songs based on the characteristics of the song itself, solving the "cold start" problem.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115062225B_ABST
    Figure CN115062225B_ABST
Patent Text Reader

Abstract

The present application discloses a method for training a song recommendation model, a song recommendation method, a device, and a medium, including: selecting different song audios from a song audio training set; splitting the song audios into multiple audio segments; inputting the audio segments into an initial model so that the initial model extracts embedding features of the audio segments; calculating a contrastive loss based on the embedding features of different audio segments of the same song audio and the embedding features of audio segments of different song audios; adjusting the parameters of the initial model based on the contrastive loss; when the training completion condition is satisfied, determining the initial model with adjusted parameters as a song recommendation model. It can improve the performance of the song recommendation model and thus improve the accuracy of song recommendation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of song recommendation, and particularly to a method for training a song recommendation model, a song recommendation method, a device, and a medium. Background Art

[0002] Song recommendation is one of the important channels for users to listen to songs and also an important way to promote new songs. The goal of the recommendation algorithm is to analyze user behavior and the correlation between songs, speculate on users' music preferences, and push appropriate songs to appropriate people. As the main body of the recommended content, song audio contains rich information and the correlation between songs, which is an important research object of the song recommendation system.

[0003] Currently, existing song audio recommendation models are usually trained based on user behavior. For example, a binary classification model of whether a user likes or dislikes a song. During the training process, the extracted features include not only audio information but also user behavior information, which will mix the learning of audio features and user behavior, affecting the performance of the entire recommendation model. In summary, during the implementation of the present invention, the inventors have at least found the problem of low accuracy in song recommendation in the prior art. Summary of the Invention

[0004] In view of this, the purpose of this application is to provide a method for training a song recommendation model, a song recommendation method, a device, and a medium, which can improve the performance of the song recommendation model and thus improve the accuracy of song recommendation. The specific solutions are as follows:

[0005] In a first aspect, this application discloses a method for training a song recommendation model, including:

[0006] Select different song audios from a song audio training set;

[0007] Cut the song audio into multiple audio segments;

[0008] Input the audio segments into an initial model so that the initial model extracts the embedding features of the audio segments;

[0009] Calculate a contrastive loss based on the embedding features of different audio segments of the same song audio and the embedding features of audio segments of different song audios; wherein, during the training process, the contrastive loss decreases as the spatial distance between the embedding features of different audio segments of the same song audio decreases and the spatial distance between the embedding features of audio segments of different song audios increases;

[0010] Adjust the parameters of the initial model based on the contrastive loss;

[0011] When the training completion condition is satisfied, the initial model with adjusted parameters is determined as the song recommendation model.

[0012] Optionally, calculating the contrast loss based on the embedding features of different audio segments of the same song audio and the embedding features of audio segments of different song audios includes:

[0013] Determining target embedding feature pairs based on the embedding features of different audio segments of the same song audio;

[0014] Calculating the spatial distance of the target embedding feature pairs to obtain a first spatial distance;

[0015] Calculating the spatial distance between the embedding features of audio segments of different song audios to obtain a second spatial distance;

[0016] Calculating the contrast loss based on the first spatial distance and the second spatial distance.

[0017] Optionally, both the first spatial distance and the second spatial distance are cosine distances or both are Euclidean distances.

[0018] Optionally, inputting the audio segment into the initial model includes:

[0019] Directly inputting the audio segment into the initial model;

[0020] Or, extracting the spectral features of the audio segment and inputting the spectral features into the initial model.

[0021] Optionally, before inputting the audio segment into the initial model, it further includes:

[0022] For each audio segment of the song audio, determining a target data augmentation processing method from multiple preset data augmentation processing methods;

[0023] Performing data augmentation processing on the corresponding audio segment based on the target data augmentation processing method.

[0024] Optionally, the preset data augmentation processing methods include noise addition processing, random cropping processing, speed change processing, pitch change processing, and time-domain flipping processing.

[0025] In a second aspect, the present application discloses a song recommendation method, including:

[0026] Obtaining the song audio to be recommended and the target song audio; wherein, the target song audio is the song audio liked by the user;

[0027] Segmenting both the song audio to be recommended and the target song audio into multiple target audio segments;

[0028] Input the target audio clip into the song recommendation model so that the song recommendation model extracts the embedding features of the target audio clip; wherein, the song recommendation model is obtained based on the aforementioned song recommendation model training method;

[0029] Calculate the similarity between the song audio to be recommended and the target song audio based on the embedding features;

[0030] Recommend the song audio to be recommended to the corresponding user terminal based on the similarity.

[0031] Optionally, the calculating the similarity between the song audio to be recommended and the target song audio based on the embedding features includes:

[0032] Calculate the spatial distance between the embedding features of the target audio clip pairs to obtain a third spatial distance; wherein, the target audio clip pairs are feature pairs composed of the target audio clip of the song audio to be recommended and the target audio clip of the target song audio;

[0033] Calculate the similarity between the song audio to be recommended and the target song audio based on all the third spatial distances corresponding to the target audio clip pairs.

[0034] In a third aspect, the present application discloses an electronic device, including a memory and a processor, wherein:

[0035] The memory is used to store a computer program;

[0036] The processor is used to execute the computer program to implement the aforementioned song recommendation model training method, and / or, the aforementioned song recommendation method.

[0037] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the aforementioned song recommendation model training method, and / or, the aforementioned song recommendation method.

[0038] It can be seen that in this application, different song audios are selected from the song audio training set, and then the song audios are segmented into multiple audio segments. Then, the audio segments are input into the initial model so that the initial model can extract the embedding features of the audio segments, and calculate the contrast loss based on the embedding features of different audio segments of the same song audio and the embedding features of audio segments of different song audios. Among them, during the training process, the contrast loss decreases as the spatial distance between the embedding features of different audio segments of the same song audio decreases, and as the spatial distance between the embedding features of audio segments of different song audios increases. Then, the parameters of the initial model are adjusted based on the contrast loss. When the training completion condition is met, the initial model with adjusted parameters is determined as the song recommendation model. In this way, during the training process of the song recommendation model, the similarity between different audio segments of the same song audio and the difference between audio segments of different song audios can be learned, so that the trained model can extract better embedding features representing song audios, improve the performance of the song recommendation model, and further improve the accuracy of song recommendation. Description of the Drawings

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.

[0040] Figure 1 Schematic diagram of the hardware composition framework applicable to a song recommendation model training method and / or a song recommendation method provided by an embodiment of the present application;

[0041] Figure 2 Schematic diagram of the hardware composition framework applicable to another song recommendation model training method and / or a song recommendation method provided by an embodiment of the present application;

[0042] Figure 3 Schematic flowchart of a song recommendation model training method provided by an embodiment of the present application;

[0043] Figure 4 Schematic diagram of the layer structure of a specific SampleCNN module provided by an embodiment of the present application;

[0044] Figure 5 Schematic diagram of a specific song recommendation model training provided by an embodiment of the present application;

[0045] Figure 6 Schematic flowchart of a song recommendation method provided by an embodiment of the present application;

[0046] Figure 7 This is a flowchart for calculating the similarity between a specific song audio to be recommended and a target song audio provided by an embodiment of the present application. Detailed implementation manners

[0047] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0048] Song recommendation is one of the important channels for users to listen to songs, and it is also an important way for promoting and publicizing new songs. For some new songs, before they are promoted, it is impossible to infer which user groups they are suitable for based on historical user behavior. Therefore, it is only possible to analyze based on the audio information of the songs themselves and push the new songs to suitable user groups to solve the "cold start" problem of new songs. However, the current song audio recommendation systems are all trained based on user behavior, such as a binary classification model of whether a user likes or dislikes. The audio embedding features extracted in this way contain not only audio information but also user behavior information. When solving the "cold start" problem of songs, some invalid user information will be introduced and it is impossible to well represent the similarities and differences between audios.

[0049] A song audio recommendation scheme process in the prior art mainly includes: 1. First, obtain two parts of data sets: a batch of song data sets containing user's song listening behaviors, including songs collected by the user, recently fully played songs, favorite songs, collected songs, etc. A batch of song audios that the user likes / dislikes. 2. Use an audio-independent, user behavior-based song recommendation system to extract the embedding representing the user portrait from the first data set, and this embedding represents the user's preferences. 3. For the second batch, the data set of user's liked / disliked songs containing audios, extract the embedding from the list of songs the user likes and calculate the mean value to represent the audio of the user's favorite songs, and extract the embedding from the list of songs the user dislikes and calculate the mean value to represent the audio of the user's disliked songs. 4. The loss function calculates the cosine distance between the embedding features of the user's liked / disliked song audios and the user portrait embedding. The training objective of the network is to make the feature of the user's favorite song closer to the user portrait embedding and make the feature of the user's disliked song farther from the user portrait embedding. 5. In the prediction stage, input the song audio to be recommended into the network to extract the embedding feature, compare the distance between the embedding feature and the user portrait embedding, and determine whether the current song is liked by the user through a set threshold. This scheme combines audio analysis and user portrait analysis, and integrates the recommendation based on user behavior and the recommendation based on audio. Although this approach can achieve certain results, there are the following defects: First, when directly integrating user behavior analysis and audio analysis, the user behavior characteristics will affect the extraction of audio features, and the differences in user portraits will mislead the model to learn the differences in audio features; Second, when the user's historical assets are insufficient, a complete user portrait cannot be obtained, and the similarity calculated between the audio embedding and the user portrait embedding is unreliable. To solve the above problems, this case does not introduce user portrait information and only learns the correlation between audios to provide a more pure audio feature for the recommendation system. In summary, in the process of implementing the present invention, the inventor has at least found that there is a problem of low accuracy in song recommendation in the prior art. For this reason, this application provides a song recommendation model training scheme, which can improve the performance of the song recommendation model and thus improve the accuracy of song recommendation.

[0050] For ease of understanding, first introduce the hardware composition framework used in the song recommendation model training method and / or the corresponding scheme of the song recommendation method provided by the embodiments of this application. Please refer to Figure 1 , Figure 1Schematic diagram of the hardware composition framework applicable to a song recommendation model training method and / or a song recommendation method provided by an embodiment of the present application. The electronic device 100 may include a processor 101 and a memory 102, and may further include one or more of a multimedia component 103, an information input / output (I / O) interface 104, and a communication component 105.

[0051] Among them, the processor 101 is used to control the overall operation of the electronic device 100 to complete all or part of the steps in the song recommendation model training method and / or the song recommendation method; the memory 102 is used to store various types of data to support the operation of the electronic device 100. These data may include, for example, instructions for any application or method operating on the electronic device 100, and application-related data. The memory 102 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc. In this embodiment, the memory 102 stores at least programs and / or data for implementing the following functions:

[0052] Select different song audios from the song audio training set;

[0053] Segment the song audio into multiple audio segments;

[0054] Input the audio segments into the initial model so that the initial model extracts the embedding features of the audio segments;

[0055] Calculate a contrastive loss based on the embedding features of different audio segments of the same song audio and the embedding features of audio segments of different song audios; wherein, during the training process, the contrastive loss decreases as the spatial distance between the embedding features of different audio segments of the same song audio decreases and the spatial distance between the embedding features of audio segments of different song audios increases;

[0056] Adjust the parameters of the initial model based on the contrastive loss;

[0057] When the training completion condition is met, the initial model with adjusted parameters is determined as the song recommendation model.

[0058] The multimedia component 103 may include a screen and an audio component. The screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in the memory 102 or sent through the communication component 105. The audio component further includes at least one speaker for outputting audio signals. The I / O interface 104 provides an interface between the processor 101 and other interface modules, and the other interface modules may be a keyboard, a mouse, buttons, etc. These buttons may be virtual buttons or physical buttons. The communication component 105 is used for wired or wireless communication between the electronic device 100 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, or 4G, or a combination of one or more of them. Accordingly, the communication component 105 may include: a Wi-Fi component, a Bluetooth component, and an NFC component.

[0059] The electronic device 100 may be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components, and is used to execute the song recommendation model training method and / or the song recommendation method.

[0060] Of course, Figure 1 The structure of the illustrated electronic device 100 does not limit the electronic device in the embodiments of the present application. In practical applications, the electronic device 100 may include more or fewer components than Figure 1 those illustrated, or combine certain components.

[0061] It can be understood that the number of electronic devices is not limited in the embodiments of the present application, and multiple electronic devices may cooperate to complete the song recommendation model training method and / or the song recommendation method. In a possible implementation manner, please refer toFigure 2 , Figure 2 This is a schematic diagram of the hardware composition framework applicable to another song recommendation model training method and / or song recommendation method provided by the embodiments of the present application. As can be seen from Figure 2 , this hardware composition framework may include: a first electronic device 11 and a second electronic device 12, which are connected through a network 13.

[0062] In the embodiments of the present application, the hardware structures of the first electronic device 11 and the second electronic device 12 may refer to Figure 1 the electronic device 100 in. That is, it can be understood that there are two electronic devices 100 in this embodiment, and the two perform data interaction. Further, the form of the network 13 is not limited in the embodiments of the present application, that is, the network 13 may be a wireless network (such as WIFI, Bluetooth, etc.) or a wired network.

[0063] Among them, the first electronic device 11 and the second electronic device 12 may be the same type of electronic device. For example, both the first electronic device 11 and the second electronic device 12 are servers; they may also be different types of electronic devices. For example, the first electronic device 11 may be a smart phone or other intelligent terminal, and the second electronic device 12 may be a server. In a possible implementation manner, a server with strong computing power may be used as the first electronic device 11 and the second electronic device 12 to improve the model training efficiency. In a possible implementation manner, a server with strong computing power may be used as the second electronic device 12 to improve the data processing efficiency and reliability, and thus improve the processing efficiency of song recommendation. At the same time, a smart phone with low cost and wide application range is used as the first electronic device 11 to realize the interaction between the second electronic device 12 and the user. It can be understood that the interaction process may be: the smart phone collects user behavior data and transmits it to the server, the server obtains the song audio liked by the user based on the user behavior data, and obtains the song audio to be recommended, and segments both the song audio to be recommended and the target song audio into multiple target audio segments; inputs the target audio segments into the song recommendation model so that the song recommendation model extracts the embedding features of the target audio segments; calculates the similarity between the song audio to be recommended and the target song audio based on the embedding features; and recommends the song audio to be recommended to the corresponding smart phone based on the similarity.

[0064] Referring to Figure 3 shown, the embodiments of the present application disclose a song recommendation model training method, including:

[0065] Step S11: Select different song audios from the song audio training set.

[0066] In a specific implementation manner, two different song audios may be selected from the song audio training set.

[0067] Step S12: Segment the song audio into multiple audio segments.

[0068] In a specific implementation, the selected song audio can be evenly segmented into multiple audio segments according to a preset segmentation length.

[0069] Moreover, for each audio segment of the song audio, determine a target data augmentation processing method from multiple preset data augmentation processing methods; perform data augmentation processing on the corresponding audio segment based on the target data augmentation processing method. Among them, the preset data augmentation processing methods include but are not limited to noise addition processing, random cropping processing, speed change processing, pitch change processing, and time-domain flipping processing.

[0070] That is to say, for all audio segments of each song audio, data augmentation processing can be performed according to the determined target data augmentation processing method, for example, speed change processing is performed uniformly. Among them, the target data augmentation processing method can be randomly determined. It can be understood that through data augmentation processing, the robustness of the model can be increased.

[0071] Step S13: Input the audio segment into the initial model so that the initial model extracts the embedding features of the audio segment.

[0072] In one implementation, the audio segment can be directly input into the initial model.

[0073] In another implementation, the spectral features of the audio segment can be extracted and the spectral features can be input into the initial model.

[0074] Moreover, in one implementation, the initial model can include a SampleCNN (i.e., Sample Convolution Neural Network, sample convolutional neural network) module and a fully connected module. The enhanced audio segment is first input into the SampleCNN module and then passes through the fully connected module to extract the embedding features. For example, see Figure 4 as shown Figure 4 which is a schematic diagram of a specific layer structure of the SampleCNN module provided by this application. Taking the first row as an example, explain Figure 4The structure in it, conv3-128 is a 3-layer convolutional network with 128 neurons in each layer, a stride of 3. The specific data dimension after convolution is 19,683×128, and the number of parameters in the current network layer is 512. The difference between SampleCNN and the ordinary convolutional network is that SampleCNN uses strided convolution to gradually compress audio features without extracting spectral features, ensuring the integrity of audio information. Of course, in some other embodiments, other network structures can also be used to replace SampleCNN.

[0075] Step S14: Calculate the contrastive loss based on the embedding features of different audio segments of the same song audio and the embedding features of audio segments of different song audios. Among them, during the training process, the contrastive loss decreases as the spatial distance between the embedding features of different audio segments of the same song audio decreases and the spatial distance between the embedding features of audio segments of different song audios increases;

[0076] In a specific implementation, the target embedding feature pair can be determined based on the embedding features of different audio segments of the same song audio; calculate the spatial distance of the target embedding feature pair to obtain the first spatial distance; calculate the spatial distance between the embedding features of audio segments of different song audios to obtain the second spatial distance; calculate the contrastive loss based on the first spatial distance and the second spatial distance. And, the distance sum can be calculated based on all the second spatial distances, and the comparison loss can be calculated based on the first spatial distance and the second spatial distance.

[0077] In one implementation, both the first spatial distance and the second spatial distance are cosine distances. In another implementation, both the first spatial distance and the second spatial distance are Euclidean distances.

[0078] Furthermore, the calculation formula of the comparison loss can be as follows:

[0079]

[0080] where, l i,j represents the comparison loss calculated based on two song audios, the base of the logarithm is 2, z i and z j represent the embedding features of different audio segments of the same song audio, that is, zi and zj form the target embedding feature pair, sim(z i , z j ) represents the cosine distance between z i and z j , τ is a constant that controls the change range of the loss function, z i and z k represent the embedding features of audio segments of different song audios. It is 1 when k≠i, N represents the number of segments of a song audio, exp represents taking the power of e, and it is ensured that each term is greater than 1. That is, the denominator is the power of e for each second spatial distance, obtaining the corresponding power-of-e result, and then summing up the power-of-e results of all second spatial distances.

[0081] Moreover, during the training process, a batch (i.e., batch processing) can be two different song audios, then the comparison loss for one round of iteration is l i,j , or it can be N groups of data, each group of data including two different song audios, then the comparison loss for one round of iteration is N*l i,j . N is a positive integer greater than 1.

[0082] It can be understood that the learning objective of the network is to make different segments of the same song audio after data augmentation closer in the embedding space, and the segments of different song audios farther apart in the embedding space.

[0083] For example, as shown in Figure 5 , the embodiments of the present application disclose a specific schematic diagram for training a song recommendation model. Two different song audios are selected from the training song audios and respectively cut into n sub-segments, and then data augmentation processing is performed on these sub-segments. The data augmentation methods include adding noise, random cropping, speed change, pitch change, and time-domain flipping. The augmented segments are input into SampleCNN, and then embedding features are extracted through a fully connected module, and then the comparison loss is calculated based on the embedding features. In this way, using a self-supervised learning method, the deep features of the audio are extracted through the SampleCNN network, making different segments of the same song closer in the embedding feature space, and different songs farther apart in the embedding features. Without any prior information related to songs and users, the learned features can better represent the characteristics of the songs themselves. When this characteristic is applied to the recommendation of new songs, it can better describe the similarity relationship between new songs and other songs, providing better audio representation for the song recommendation system.

[0084] Step S15: Adjust the parameters of the initial model based on the comparison loss.

[0085] Step S16: When the training completion condition is met, determine the initial model with adjusted parameters as the song recommendation model.

[0086] In one implementation, when the comparison loss is less than or equal to a preset loss threshold, it is determined that the training completion condition is met.

[0087] It can be seen that in the embodiment of the present application, different song audios are selected from the song audio training set, and then the song audios are segmented into multiple audio segments, and then the audio segments are input into the initial model, so that the initial model extracts the embedding features of the audio segments, and calculates the contrastive loss based on the embedding features of different audio segments of the same song audio and the embedding features of audio segments of different song audios. Among them, during the training process, the contrastive loss decreases as the spatial distance between the embedding features of different audio segments of the same song audio decreases, and as the spatial distance between the embedding features of audio segments of different song audios increases. Then, the parameters of the initial model are adjusted based on the contrastive loss. When the training completion condition is met, the initial model with adjusted parameters is determined as the song recommendation model. In this way, during the training process of the song recommendation model, the similarity between different audio segments of the same song audio and the difference between audio segments of different song audios can be learned, so that the trained model can extract embedding features that better represent song audios, improve the performance of the song recommendation model, and further improve the accuracy of song recommendations.

[0088] See Figure 6 As shown, the embodiment of the present application discloses a song recommendation method, including:

[0089] Step S21: Obtain the song audio to be recommended and the target song audio; wherein, the target song audio is the song audio liked by the user.

[0090] In a specific implementation manner, the song audio liked by the user can be determined based on the user's historical behavior data. For example, the song audio with historical play, collection, sharing and other behaviors can be determined as the song audio liked by the user.

[0091] Step S22: Segment both the song audio to be recommended and the target song audio into multiple target audio segments.

[0092] Step S23: Input the target audio segments into the song recommendation model, so that the song recommendation model extracts the embedding features of the target audio segments.

[0093] Among them, the song recommendation model is obtained based on the song recommendation model training method disclosed in the foregoing embodiment, and will not be elaborated here.

[0094] Step S24: Calculate the similarity between the song audio to be recommended and the target song audio based on the embedding features.

[0095] In a specific embodiment, the spatial distance between the embedding features of the target audio segment pair can be calculated to obtain a third spatial distance; wherein, the target audio segment pair is a feature pair composed of the target audio segment of the song audio to be recommended and the target audio segment of the target song audio; based on the third spatial distances corresponding to all the target audio segment pairs, the similarity between the song audio to be recommended and the target song audio is calculated.

[0096] In one embodiment, the mean value of the third spatial distances corresponding to all the target audio segment pairs can be calculated to obtain the similarity between the song audio to be recommended and the target song audio.

[0097] In another embodiment, the maximum value among the third spatial distances corresponding to all the target audio segment pairs can be calculated to obtain the similarity between the song audio to be recommended and the target song audio.

[0098] Step S25: Recommend the song audio to be recommended to the corresponding user terminal based on the similarity.

[0099] In one embodiment, if the similarity is higher than a preset similarity threshold, the song audio to be recommended can be recommended to the corresponding user terminal.

[0100] For example, referring to Figure 7 As shown, an embodiment of the present application discloses a specific flowchart for calculating the similarity between the song audio to be recommended and the target song audio. The song audio to be recommended and the target song audio are respectively segmented into n segments, and then input into the song recommendation model to obtain the embedding features of each segment. The spatial distance between the embedding features of the segment pairs is calculated, and then the distance mean value is calculated to obtain the similarity between the song audio to be recommended and the target song audio. That is, in the inference stage, by traversing all the segments between the two song audios, the segments are input into the network to extract embedding features, the cosine distances of all the segment pairs are calculated and the mean value is taken, and the similarity between the two songs can be obtained. In this way, during the model training process, without any user information, the model is trained in a self-supervised learning manner, and the obtained model can provide purer and richer audio information, separating the user behavior analysis and the audio feature learning, and better playing the respective roles of the two modules, so that the recommendation accuracy can be greatly improved.

[0101] It can be seen that in the embodiment of the present application, the song audio to be recommended and the target song audio are obtained. The target song audio is the song audio liked by the user. Both the song audio to be recommended and the target song audio are segmented into multiple target audio segments, and then the target audio segments are input into the song recommendation model trained by the method disclosed in the foregoing embodiment, so that the song recommendation model extracts the embedding features of the target audio segments, and then calculates the similarity between the song audio to be recommended and the target song audio based on the embedding features. Finally, the song audio to be recommended is recommended to the corresponding user terminal based on the similarity. In this way, the similarity between the song audio to be recommended and the song audio liked by the user can be calculated, and songs can be accurately recommended to the user.

[0102] Next, taking a certain music APP as an example, the technical solution of the present application will be described.

[0103] The background server of this music APP determines a song audio training set based on the music library of this music APP for training the song recommendation model. During the training process, different song audios are selected from the song audio training set; the song audios are segmented into multiple audio segments; the audio segments are input into the initial model so that the initial model extracts the embedding features of the audio segments; the contrast loss is calculated based on the embedding features of different audio segments of the same song audio and the embedding features of the audio segments of different song audios; the parameters of the initial model are adjusted based on the contrast loss; when the training completion condition is met, the initial model with adjusted parameters is determined as the song recommendation model. The new song "Singing in Youth" by Mao Buyi is determined as the song audio to be recommended. Based on the collected user behavior data of this music APP, the song audio liked by the user is determined. For example, if the user has shared the song "Blue and White Porcelain" by Zhou Chuanxiong, then first both "Singing in Youth" and "Blue and White Porcelain" are segmented into multiple audio segments, and then input into the song recommendation model. The song recommendation model extracts the embedding features of each audio segment, and then calculates the similarity between "Singing in Youth" and "Blue and White Porcelain". The similarity between "Singing in Youth" and multiple song audios liked by the user can be calculated. Whether to recommend "Singing in Youth" to this user is determined based on the calculated multiple similarities. If it is determined to recommend "Singing in Youth" to this user, "Singing in Youth" is pushed to this music APP installed in the user's client for display.

[0104] Next, the computer-readable storage medium provided by the embodiment of the present application will be introduced. The computer-readable storage medium described below can be mutually referred to the song recommendation model training method and / or the song recommendation method described above.

[0105] The present application also provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the above-described song recommendation model training method and / or song recommendation method are implemented.

[0106] The computer-readable storage medium may include: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.

[0107] The various embodiments in this specification are described in a progressive manner. The key point of each embodiment is the difference from other embodiments. For the same or similar parts between the embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method part.

[0108] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in the form of hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0109] The steps of the methods or algorithms described in combination with the embodiments disclosed in this article can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0110] Finally, it should also be noted that in this article, relationships such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "including", "comprising" or any other variant are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.

[0111] The above has introduced in detail the method for training a song recommendation model, the song recommendation method, the device and the medium provided by this application. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation on this application.

Claims

1. A method for training a song recommendation model, characterized in that Including: Selecting different song audios from a song audio training set; Segmenting the song audios into multiple audio segments; Inputting the audio segments into an initial model so that the initial model extracts embedding features of the audio segments; Calculating a contrastive loss based on the embedding features of different audio segments of the same song audio and the embedding features of audio segments of different song audios; wherein, during the training process, the contrastive loss decreases as the spatial distance between the embedding features of different audio segments of the same song audio decreases and as the spatial distance between the embedding features of audio segments of different song audios increases; Adjusting the parameters of the initial model based on the contrastive loss; When the training completion condition is satisfied, determining the initial model with adjusted parameters as a song recommendation model; Wherein, calculating the contrastive loss based on the embedding features of different audio segments of the same song audio and the embedding features of audio segments of different song audios includes: determining a target embedding feature pair based on the embedding features of different audio segments of the same song audio; calculating the spatial distance of the target embedding feature pair to obtain a first spatial distance; calculating the spatial distance between the embedding features of audio segments of different song audios to obtain a second spatial distance; calculating the contrastive loss based on the first spatial distance and the second spatial distance.

2. The method for training a song recommendation model according to claim 1, wherein Both the first spatial distance and the second spatial distance are cosine distances or both are Euclidean distances.

3. The method for training a song recommendation model according to claim 1, wherein The inputting the audio segments into the initial model includes: Directly inputting the audio segments into the initial model; Or, extracting spectral features of the audio segments and inputting the spectral features into the initial model.

4. The method for training a song recommendation model according to any one of claims 1 to 3, characterized in that, Before inputting the audio segments into the initial model, it further includes: For each audio segment of each song audio, determining a target data augmentation processing method from multiple preset data augmentation processing methods; Performing data augmentation processing on the corresponding audio segment based on the target data augmentation processing method.

5. The method for training a song recommendation model according to claim 4, wherein The preset data augmentation processing methods include noise addition processing, random cropping processing, speed change processing, pitch change processing, and time domain flipping processing.

6. A song recommendation method, characterized in that, Including: Obtaining a song audio to be recommended and a target song audio; wherein, the target song audio is a song audio liked by the user; Segmenting both the song audio to be recommended and the target song audio into multiple target audio segments; Inputting the target audio segments into the song recommendation model so that the song recommendation model extracts embedding features of the target audio segments; wherein, the song recommendation model is obtained based on the song recommendation model training method according to any one of claims 1 to 5; Calculating the similarity between the song audio to be recommended and the target song audio based on the embedding features; Recommending the song audio to be recommended to the corresponding user terminal based on the similarity.

7. The song recommendation method according to claim 6, wherein The calculating the similarity between the song audio to be recommended and the target song audio based on the embedding features includes: Calculate the spatial distance between the embedding features of the target audio segment pair to obtain a third spatial distance; wherein, the target audio segment pair is a feature pair composed of the target audio segment of the to-be-recommended song audio and the target audio segment of the target song audio; Calculate the similarity between the to-be-recommended song audio and the target song audio based on the third spatial distance corresponding to all the target audio segment pairs.

8. An electronic device, characterized in that, It includes a memory and a processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program to implement the song recommendation model training method according to any one of claims 1 to 5, and / or the song recommendation method according to any one of claims 6 to 7.

9. A computer-readable storage medium, characterized in that, For storing computer programs, wherein the computer program, when executed by the processor, implements the song recommendation model training method according to any one of claims 1 to 5, and / or the song recommendation method according to any one of claims 6 to 7.

Citation Information

Patent Citations

  • Method and system for recommending songs

    CN102654859A

  • Song recommendation method and apparatus

    CN107885745A