Audio recommendation model training method, device, medium and equipment
By training the audio recommendation model, using the user's actual and simulated playback probability to optimize the model parameters, the problem of low music exposure is solved, and a more comprehensive music recommendation and user preferences are achieved.
Patent Information
- Application Number
- CN202110003789.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-04
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-01-04
AI Technical Summary
The prior art has the problem of low music exposure when recommending music, and it is impossible to effectively recommend music that lacks user behavior.
By obtaining the actual probability and simulation probability of the user playing candidate music after playing the target music, determining the loss function and optimizing the model parameters, the audio recommendation model is trained to improve the exposure of music.
It improves the exposure of music and can more fully meet users' listening needs. The recommended music takes into account the characteristics of the music itself and reflects the user's true preferences.
Smart Images

Figure CN114722236B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to a method and device for training an audio recommendation model, as well as a computer-readable storage medium and an electronic device for implementing the method. Background Art
[0002] In the music playing scenario, music applications generally provide users with playlists (free listening, music streaming, playlist recommendations, etc.) to facilitate users to listen. With the development of artificial intelligence technology, user preferences are increasingly taken into consideration in the process of determining recommended playlists.
[0003] In the related art, the music sequence of the user is obtained, and the music representation is indirectly obtained by using ideas such as Word2vec, and then the music recommendation model is trained to recommend music to the user. However, although the songs recommended by the music recommendation scheme can characterize the user's listening behavior, since the user behavior is generally reflected in the top popular music, the songs recommended by the scheme generally cannot include music without user behavior.
[0004] It can be seen that the audio recommendation solution provided by the related technology has the problem of low music exposure rate.
[0005] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention
[0006] The purpose of the present disclosure is to provide a training method and device for an audio recommendation model, an electronic device, and a computer-readable storage medium, thereby improving the exposure rate of music to a certain extent.
[0007] According to a first aspect of the present disclosure, a training method for an audio recommendation model is provided, the method comprising: obtaining M groups of sample data, each group of sample data comprising target music and N candidate music after a user plays the target music, where M and N are positive integers; obtaining the actual probability of a user playing the i-th candidate music in the j-th group after playing the target music in the j-th group, and obtaining an actual probability distribution of the music in the j-th group, where j is a positive integer not greater than M; determining, based on the model parameters to be optimized of the audio recommendation model, a simulated probability of a user playing the i-th candidate music in the j-th group after playing the target music in the j-th group, and obtaining a simulated probability distribution of the music in the j-th group; and determining a loss function according to the actual probability distribution and the simulated probability distribution, optimizing the model parameters based on the loss function, and obtaining a trained audio recommendation model. .
[0008] According to a second aspect of the present disclosure, a training device for an audio recommendation model is provided, comprising: a sample acquisition module, an actual probability determination module, a simulated probability determination module, and a model parameter optimization module.
[0009] Among them, the above-mentioned sample acquisition module is configured to: obtain M groups of sample data, each group of sample data includes target music and N candidate music after the user plays the above-mentioned target music, M and N are positive integers; the above-mentioned simulation probability determination module is configured to: obtain the actual probability of the user playing the i-th candidate music in the above-mentioned j-th group after playing the target music in the j-th group, and obtain the actual probability distribution of the music in the above-mentioned j-th group, j is a positive integer not greater than M; the above-mentioned actual probability determination module is configured to: determine the simulated probability of the user playing the i-th candidate music in the above-mentioned j-th group after playing the target music in the above-mentioned j-th group based on the model parameters to be optimized of the above-mentioned audio recommendation model, and obtain the simulated probability distribution of the music in the above-mentioned j-th group; the above-mentioned model parameter optimization module is configured to: determine the loss function according to the above-mentioned actual probability distribution and the above-mentioned simulated probability distribution, optimize the above-mentioned model parameters based on the above-mentioned loss function, and obtain the trained audio recommendation model.
[0010] In an exemplary embodiment of the present disclosure, based on the aforementioned embodiment, the above-mentioned actual probability determination module is specifically configured to: collect a large number of users' listening behavior sequences regarding the target music in the above-mentioned j-th group; in the above-mentioned listening behavior sequences, count the N matching music that are in the same calculation window as the target music in the above-mentioned j-th group; and obtain the number of occurrences of the i-th candidate music, and normalize the number of occurrences of the i-th candidate music to obtain the actual probability of the above-mentioned user playing the i-th candidate music after playing the target music in the above-mentioned j-th group.
[0011] In an exemplary embodiment of the present disclosure, based on the above-mentioned embodiment, the above-mentioned simulation probability determination module includes: a feature extraction submodule and a normalization submodule.
[0012] Among them, the above-mentioned feature extraction submodule is configured to: perform feature extraction on the target music in the above-mentioned j-th group based on the first parameter of the above-mentioned audio recommendation model, and obtain the target feature vector corresponding to the above-mentioned target music; the above-mentioned normalization unit is configured to: determine the simulated probability of the user playing the i-th candidate music in the above-mentioned j-th group after playing the target music in the above-mentioned j-th group based on the second parameter of the above-mentioned audio recommendation model and the above-mentioned target feature vector.
[0013] In an exemplary embodiment of the present disclosure, based on the aforementioned embodiment, the above-mentioned normalization submodule is specifically configured as follows: according to the second parameter of the above-mentioned audio recommendation model and the above-mentioned target feature vector, the simulation probabilities corresponding to the above-mentioned N candidate music are exponentially normalized to obtain the simulation probability that the user plays the i-th candidate music in the above-mentioned j-th group after playing the target music in the above-mentioned j-th group.
[0014] In an exemplary embodiment of the present disclosure, based on the above-mentioned embodiment, the feature extraction submodule includes: a feature sequence acquisition unit and a feature extraction unit.
[0015] Among them, the above-mentioned feature sequence acquisition unit is configured to: obtain P frequency domain feature sequences corresponding to the above-mentioned target music, P is an integer greater than 1; the above-mentioned feature extraction unit: based on the first parameter of the above-mentioned audio recommendation model, performs feature extraction on the above-mentioned P frequency domain feature sequences to obtain the target feature vector corresponding to the above-mentioned target music.
[0016] In an exemplary embodiment of the present disclosure, based on the aforementioned embodiment, the feature sequence acquisition unit is specifically configured to: divide the target music into at least P audio segments according to a time domain window; perform time-frequency conversion on multiple sampling points belonging to the same audio segment to obtain frequency domain sequences corresponding to the P audio segments; and sample the P frequency domain sequences respectively to obtain P frequency domain feature sequences corresponding to the P audio segments respectively.
[0017] In an exemplary embodiment of the present disclosure, based on the aforementioned embodiment, the feature extraction unit is specifically configured to: call the cascaded S-layer convolutional neural network layer in the audio recommendation model to perform feature extraction on the x-th frequency domain feature sequence based on the first parameter of the S-layer convolutional neural network layer, and obtain S convolution feature vectors corresponding to the x-th frequency domain feature sequence, where x is an integer not greater than P; and, concatenate the S convolution feature vectors corresponding to the x-th frequency domain feature sequence to obtain a fragment feature vector corresponding to the x-th frequency domain feature sequence, and concatenate the fragment feature vectors corresponding to the P frequency domain feature sequences to obtain a target feature vector corresponding to the target music.
[0018] In an exemplary embodiment of the present disclosure, based on the aforementioned embodiment, the feature extraction unit is further specifically configured as follows: after splicing the segment feature vectors corresponding to the P frequency domain feature sequences, pooling is performed on the spliced segment feature vectors, and the vector after pooling is used as the target feature vector corresponding to the target music.
[0019] In an exemplary embodiment of the present disclosure, based on the aforementioned embodiment, the training device of the audio recommendation model further includes: a sub-model pre-training module.
[0020] The sub-model pre-training module is configured as follows: based on the first parameter of the S-layer convolutional neural network layer, feature extraction is performed on the frequency domain feature sequence of the pre-training sample audio to obtain a sample feature vector and a positive sample feature vector that has a time domain before-after relationship with the sample feature vector; a negative sample feature vector that has no time domain before-after relationship with the sample feature vector is obtained; and a triplet loss function is determined based on the sample feature vector, the positive sample feature vector and the negative sample feature vector, and the first parameter is optimized based on the triplet loss function.
[0021] In an exemplary embodiment of the present disclosure, based on the aforementioned embodiment, the above-mentioned model parameter optimization module is specifically configured to: determine the relative entropy function according to the above-mentioned actual probability distribution and the above-mentioned simulated probability distribution; and optimize the second parameter of the above-mentioned audio recommendation model by calculating the minimum value of the above-mentioned relative entropy function.
[0022] In an exemplary embodiment of the present disclosure, based on the aforementioned embodiment, the feature extraction unit is further specifically configured to: input the yth audio vector of the frequency domain feature sequence into the yth node of the gated recurrent unit, where y is a positive integer not greater than P; based on the first parameter, determine the yth hidden state according to the yth audio vector and the y-1th hidden state, where the y-1th hidden state is output by the y-1th node; and use the P hidden states output by the P nodes in sequence as a target feature vector corresponding to the target music.
[0023] In an exemplary embodiment of the present disclosure, based on the aforementioned embodiment, the above-mentioned model parameter optimization module is also specifically configured to: determine the relative entropy function according to the above-mentioned actual probability distribution and the above-mentioned simulated probability distribution; and optimize the first parameter and the second parameter of the above-mentioned audio recommendation model by calculating the minimum value of the above-mentioned relative entropy function.
[0024] According to a third aspect of the present disclosure, an audio recommendation method based on artificial intelligence is provided, comprising: obtaining music currently played by a target user, and inputting the music into a trained audio recommendation model; performing feature extraction on the music according to a first parameter of the audio recommendation model to obtain a feature vector corresponding to the music; predicting the probability that the next music played by the target user after playing the current music is the i-th candidate music according to a second parameter of the audio recommendation model and the feature vector, and obtaining a predicted probability distribution about N candidate music; and determining a recommended music list according to the predicted probability distribution about the N candidate music.
[0025] According to a fourth aspect of the present disclosure, an audio recommendation device based on artificial intelligence is provided, comprising: an acquisition module, a first processing module, a second processing module and a recommendation module.
[0026] Among them, the above-mentioned acquisition module is configured to: acquire the music currently played by the target user, and input the above-mentioned music into the trained audio recommendation model; the above-mentioned first processing module is configured to: perform feature extraction on the above-mentioned music according to the first parameter of the above-mentioned audio recommendation model, and obtain the feature vector corresponding to the above-mentioned music; the above-mentioned second processing module is configured to: predict the probability that the next music played by the above-mentioned target user after playing the current music is the i-th candidate music according to the second parameter of the above-mentioned audio recommendation model and the above-mentioned feature vector, and obtain the predicted probability distribution of the N candidate music; and the above-mentioned recommendation module is configured to: determine the recommended music list according to the above-mentioned predicted probability distribution of the N candidate music.
[0027] According to a fifth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the training method of the audio recommendation model described in any embodiment of the first aspect above is implemented, and the training method of the audio recommendation model described in any embodiment of the second aspect above is implemented.
[0028] According to a sixth aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the training method of the audio recommendation model described in any embodiment of the first aspect above, and execute the training method of the audio recommendation model described in any embodiment of the second aspect above, by executing the executable instructions.
[0029] According to a seventh aspect of the present disclosure, a computer program product or a computer program is provided, the computer program product or the computer program including computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the training method of the audio recommendation model provided in each of the above embodiments.
[0030] The exemplary embodiments of the present disclosure may have some or all of the following beneficial effects:
[0031] In an AI-based audio recommendation solution provided in an example embodiment of the present disclosure, the following two aspects are taken into consideration during the training of the audio recommendation model, namely, the actual probability of the user playing the i-th candidate music after playing the target music and the simulated probability of the model output during the training process. The loss function of the model is further determined based on the two probability distributions. This solution trains the model based on the distribution of user listening to songs, so that the music recommended by the trained audio recommendation model takes into account both the characteristics of the music itself and the characterization of the user's preferences, which can reflect the user's true preference for music and is conducive to improving the accuracy of audio recommendations.
[0032] Among them, the above-mentioned target music and candidate music can be any music, that is, not limited to popular music or long-tail music. In other words, the music recommended to the user can not only be popular music but also long-tail music. In this way, more comprehensive music can be recommended to users, which is easier to meet the user's listening needs and is conducive to improving the exposure rate of music. Specifically, on the one hand, the music recommended according to the target music currently played by the user may include long-tail music, so that the types of music recommended to the user are more comprehensive. On the other hand, when the user is currently playing long-tail music (that is, the above-mentioned target music is long-tail music), the above-mentioned audio recommendation model can also effectively predict the recommended music for the user. It can be seen that this solution can meet the audio recommendation needs of niche listening users who like long-tail music, that is, for users with different listening preferences, their favorite music can be recommended.
[0033] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification are used to explain the principles of the present disclosure. Obviously, the accompanying drawings described below are only some embodiments of the present disclosure, and for ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without creative work.
[0035] Figure 1 A system architecture diagram schematically shows an exemplary application environment of an audio recommendation model training and device to which the embodiments of the present disclosure can be applied.
[0036] Figure 2 The flowchart of the audio recommendation method based on artificial intelligence according to an embodiment of the present disclosure is schematically shown.
[0037] Figure 3 The figure schematically shows an audio recommendation scenario diagram based on artificial intelligence according to an embodiment of the present disclosure.
[0038] Figure 4 The flowchart of the method for training an audio recommendation model according to an embodiment of the present disclosure is schematically shown.
[0039] Figure 5 The following schematically shows a flow chart of a method for training an audio recommendation model according to another embodiment of the present disclosure.
[0040] Figure 6 A schematic flow chart of a method for acquiring actual probability distribution according to an embodiment of the present disclosure is shown.
[0041] Figure 7 A flowchart of a method for determining simulation probabilities of candidate music according to an embodiment of the present disclosure is shown.
[0042] Figure 8 A schematic flow chart of a method for extracting a target feature vector according to an embodiment of the present disclosure is shown.
[0043] Fig. 9 A flowchart of a method for training a feature extraction sub-model according to an embodiment of the present disclosure is shown.
[0044] Fig.10 A schematic diagram of a method for extracting a target feature vector according to another embodiment of the present disclosure is shown.
[0045] Fig.11 A schematic flow chart of a method for extracting a target feature vector according to another embodiment of the present disclosure is shown.
[0046] Fig.12 The structure of a training device for an audio recommendation model according to an embodiment of the present disclosure is schematically shown.
[0047] Fig.13 The structure diagram of an audio recommendation device based on artificial intelligence according to an embodiment of the present disclosure is schematically shown.
[0048] Fig.14 A schematic diagram of the structure of a computer system suitable for implementing an electronic device of an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0049] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as being limited to the examples set forth herein; on the contrary, these embodiments are provided so that the present disclosure will be more comprehensive and complete, and the concepts of the example embodiments are fully conveyed to those skilled in the art. The described features, structures, or characteristics may be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure may be practiced while omitting one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, known technical solutions are not shown or described in detail to avoid obscuring various aspects of the present disclosure.
[0050] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and their repeated description will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0051] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.
[0052] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0053] The key technologies of speech technology include automatic speech recognition technology (ASR), text-to-speech technology (TTS) and voiceprint recognition technology. Enabling computers to listen, see, speak and feel is the future development direction of human-computer interaction, among which speech has become one of the most promising human-computer interaction methods in the future.
[0054] Machine Learning (ML) is a multi-disciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.
[0055] The solution provided by the embodiments of the present disclosure involves technologies such as machine learning and voice computing of artificial intelligence, which are specifically described by the following embodiments:
[0056] The AI-based audio recommendation scenario provided by this technical solution can be: in the mature stage of audio applications, a full range of music that enriches user behavior can be obtained, so that user preferences can be portrayed in more dimensions, making it easier to recommend music that users like.
[0057] In the process of audio recommendation, audio representation is crucial. The so-called audio representation is to use a vector to represent a piece of music. This vector can be used to calculate the similarity of songs. For example, in the "similar songs" scenario, the vector can be used as a feature to serve the upstream sorting model to improve the accuracy of the model. Audio representation can also serve user portraits to more finely characterize user preferences.
[0058] In the related technologies of audio representation, there is a method of processing and extracting the audio sequence of the music itself to make the extracted audio representation. This audio representation does not need to consider user behavior and is more suitable for the cold start of the music app (i.e., when there is no user data). However, this music representation only reflects the characteristics of the music, but not the characteristics of user behavior, and cannot characterize user preferences.
[0059] However, another related technology uses the music sequence listened to by the user and the idea of Word2vec to indirectly obtain the representation of the music, and then recommends music to the user. As mentioned above, the songs recommended by this solution generally cannot include long-tail music that lacks user behavior. That is, there is a problem of low exposure rate of long-tail music. At the same time, users cannot obtain long-tail music from the recommended music provided by the related technology, resulting in failure to fully meet user needs.
[0060] Among them, the above-mentioned long-tail music refers to music with a long-tail effect (English name: Long Tail Effect). In the long-tail effect, "head" and "tail" are two statistical terms. The protruding part in the middle of the normal curve is called the "head", and the relatively flat parts on both sides are called the "tail". From the perspective of user demand, most of the demands will be concentrated in the head, which we can call popular, and the demands distributed in the tail are personalized, scattered and small-scale demands. This part of differentiated and small-scale demands will form a long "tail" on the demand curve, and the so-called long-tail effect lies in its quantity. Adding up all non-popular markets will form a market larger than the popular market.
[0061] For text and video, we can appropriately abandon the long-tail part and conduct modeling to achieve content recommendation. This is because long-tail text and long-tail videos are rarely read and browsed by users, and the quality is not necessarily high, or there is serious homogeneity (such as plagiarism, etc.). But for music, there is no particularly fixed standard to measure the "quality" of music. Because music is relatively abstract, and the difference between popular music and unpopular music of the same style is not that big, or it is difficult to describe the specific difference (unlike videos, the difference between "good-looking" and "ugly" videos of the same style is very large, and the same is true for articles. The writing style of good articles and poor-quality articles is far apart).
[0062] It can be seen that long-tail music has corresponding value, so its exposure should be increased, while also meeting the listening needs of different users.
[0063] In view of the problem that the long-tail music has low exposure rate and cannot fully meet the user's listening needs in the related art, this technical solution determines the audio representation based on the user's listening distribution. The specific implementation plan will be described in the following examples.
[0064] Figure 1 A system architecture diagram schematically shows an exemplary application environment in which a speech recognition method and device according to an embodiment of the present disclosure can be applied.
[0065] like Figure 1 As shown, the system architecture 100 may include one or more of terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc. The terminal devices 101, 102, 103 may be various electronic devices with display screens, including but not limited to desktop computers, portable computers, smart phones, tablet computers, etc. It should be understood that Figure 1 The number of terminal devices, networks and servers in the example is only illustrative. Any number of terminal devices, networks and servers may be provided as required. For example, the server 105 may be a server cluster composed of multiple servers.
[0066] The speech recognition method provided in the embodiment of the present disclosure is generally executed by the server 105, and accordingly, the speech recognition device is generally set in the server 105. However, it is easy for those skilled in the art to understand that the speech recognition method provided in the embodiment of the present disclosure can also be executed by the terminal devices 101, 102, and 103, and accordingly, the speech recognition device can also be set in the terminal devices 101, 102, and 103, which is not particularly limited in this exemplary embodiment.
[0067] For example, in an exemplary embodiment, the terminal devices 101, 102, and 103 may send M groups of sample data, each group of sample data including the target music and N candidate music after the user plays the target music, to the server 105. Moreover, the terminal devices 101, 102, and 103 may also obtain the actual probability of the user playing the i-th candidate music in the group after playing the target music in the j-th group and send it to the server 105. The server 105 determines the simulated probability of the user playing the i-th candidate music in the group after playing the target music in the j-th group based on the model parameters to be optimized of the audio recommendation model, and obtains the simulated probability distribution of the music in the group. Then, the server 105 determines the loss function based on the above actual probability distribution and the above simulated probability distribution, optimizes the model parameters based on the loss function, and obtains the trained audio recommendation model.
[0068] Exemplarily, the server 105 can also send the audio recommendation model to the terminal devices 101 , 102 , and 103 , so that the user can directly obtain the recommended playlist through the terminal devices 101 , 102 , and 103 .
[0069] The following is a detailed description of an embodiment of a method for training an audio recommendation model provided by the present technical solution and an embodiment of a use scenario of the trained audio recommendation model. First, an embodiment of a use scenario of the trained audio recommendation model is introduced. Figure 2 The flowchart of the audio recommendation method based on artificial intelligence according to an embodiment of the present disclosure is schematically shown. Figure 2 , the embodiment shown in the figure includes:
[0070] Step S210, obtaining the music currently played by the target user, and inputting the music into the trained audio recommendation model; Step S220, performing feature extraction on the music according to the first parameter of the audio recommendation model, and obtaining a feature vector corresponding to the music; Step S230, predicting the probability that the next music played by the target user after playing the current music is the i-th candidate music according to the second parameter of the audio recommendation model and the feature vector, and obtaining a predicted probability distribution about the N candidate music; and, Step S240, determining a recommended music list according to the predicted probability distribution about the N candidate music.
[0071] In an exemplary embodiment, the target user is any user corresponding to the music currently being played. The music can be the music currently being played in any music application, or the music being played on a web page, and can be popular music or long-tail music. Figure 3 , "Song A" is currently playing in the terminal's application or web page.
[0072] Exemplarily, first obtain the "song A" currently being played by the target user and input it into the trained audio recommendation model to process "song A" based on the model parameters of the model, wherein the specific implementation of the relevant processing will be specifically described in the following embodiments. And based on the output of the model, the pre-stored probability of the target user playing the next piece of music is obtained to obtain the predicted probability distribution of multiple pieces of music. Further, the above-mentioned recommended music list is determined according to the size of the probability value in the above-mentioned predicted probability distribution.
[0073] Among them, the probability value is the probability predicted by the above audio recommendation model that the user will play other songs after the current song A. Therefore, the position of the song in the recommended music list can be determined in descending order of probability values. Figure 3 For example, the probability values corresponding to song A1, song A2, song A3... song An are from large to small, so the "recommended song list" can be determined, and the recommended song list includes: song A1, song A2, song A3... song An in sequence.
[0074] Exemplary, reference Figure 3 After the user finishes listening to the currently playing song A or before the current song A is finished playing, the user can play the song recommended by the system. For example, the user touches "Song A2" in the "Recommended Playlist" to switch to the system recommended song A2, so as to listen to the music that suits his or her preferences conveniently and quickly.
[0075] The following embodiment introduces the training process of the above audio recommendation model:
[0076] For example, Figure 4 A flowchart of a method for training an audio recommendation model according to an embodiment of the present disclosure is schematically shown. Figure 4 The training process of the audio recommendation model is generally introduced. For each set of sample data: on the one hand, the feature extraction layer 410 of the audio recommendation model is used to extract features of the target music in the set. Specifically, multiple audio features corresponding to the target music in the set are obtained to obtain an audio feature sequence, such as a frequency domain feature sequence G'1 to G' P ; Encode the audio feature sequence to obtain the feature vector corresponding to the target music (referred to as "target feature vector H"). Furthermore, through the normalization layer 420 of the audio recommendation model and the above-mentioned target feature vector H, predict the probability of the user playing the next music (the candidate music in the group), and obtain the simulated probability distribution Q' about each candidate music in the same group. On the other hand, obtain the actual probability of the user playing each candidate music in the group after playing the target music in the group, and obtain the actual probability distribution Q about each candidate music in the group. And determine the loss function 430 through the above-mentioned simulated probability distribution Q' and the actual probability distribution Q, and train the audio recommendation model according to the loss function.
[0077] For example, Figure 5 A flowchart of a method for training an audio recommendation model according to another embodiment of the present disclosure is schematically shown. Figure 5 For a detailed introduction to this solution, refer to Figure 5 , the embodiment shown in the figure includes steps S510 to S540.
[0078] In step S510, M groups of sample data are obtained, each group of sample data includes target music and N candidate music after the user plays the target music, and M and N are positive integers.
[0079] Among them, the above-mentioned target music is any song, and the candidate music in the same group of samples is the music that the user may listen to after listening to the target music in the group. Through the audio recommendation model, the probability of the user playing the next music (the candidate music in the same group) can be predicted based on the music currently played by the user (the above-mentioned target music), and further, the simulated probability distribution of each candidate music in the same group is obtained. As mentioned above, the above-mentioned target music and candidate music can both include popular music and long-tail music, and the style type of music is not limited. For example, the above-mentioned target music / candidate music can be folk music, rap, Chinese style, electronic music, rock, etc.
[0080] In step S520, the actual probability of the user playing the i-th candidate music in the j-th group after playing the target music in the j-th group is obtained, and the actual probability distribution of the music in the j-th group is obtained, where j is a positive integer not greater than M.
[0081] In an exemplary embodiment, as a specific implementation of step S520, Figure 6 A schematic flow chart of a method for obtaining an actual probability distribution according to an embodiment of the present disclosure is shown, comprising:
[0082] Step S521, collecting the listening behavior sequences of a large number of users regarding the target music in the j-th group; step S522, in the listening behavior sequence, counting the N matching music in the same calculation window as the target music in the j-th group; and step S523, obtaining the number of occurrences of the i-th candidate music, and normalizing the number of occurrences of the i-th candidate music, to obtain the actual probability of the user playing the i-th candidate music after playing the target music in the j-th group.
[0083] Exemplarily, a large number of listening behavior sequences of users are collected, and N matching music in the listening behavior sequences that are in the same calculation window C as the target music X are counted; the number of occurrences of the i-th matching music is obtained, and the number of occurrences of the i-th matching music is normalized to obtain the actual probability qi of the user playing the i-th matching music after playing the target music, and the actual probability distribution Q about the N matching music is obtained = [q1, q2, q3, ...q N ].
[0084] For example, the listening sequences of three users are obtained, and their listening behavior sequences are {S5, S7, S1, S8, S3, S2}, {S5, S2, S1, S9, S3, S2}, {S1, S4, S6}, where S1, S2, ... represent different songs, and the calculation window size is set to C = 2. For music S2 (target music), the target window related to the target music in the listening behavior sequence of the first user is: (S3, S2), the target window related to the target music in the listening behavior sequence of the second user is: [(S5, S2), (S3, S2)], and the listening behavior sequence of the third user is not related to the target music. It can be seen that the music in the same calculation window as the target music S2 is {S3, S5, S1}, and the number of occurrences of {S3, S5, S1} is {2, 1, 1} respectively. After normalization, the probability distribution of {S3, S5, S1} is {1 / 2, 1 / 4, 1 / 4}. That is, after the user plays the target music S2: the probability of the user playing S3 is 1 / 2, the probability of playing S5 is 1 / 4, the probability of playing S1 is 1 / 4, and the probability of playing S1, S4, S6, S7, S8, S9 is 0, thus obtaining a complete actual probability distribution Q = [1 / 4, 1 / 2, 0, 1 / 4, 0, 0, 0, 0] for the candidate music [S1, S3, S4, S5, S6, S7, S8, S9].
[0085] It should be noted that the more user listening behavior sequences there are, the higher the prediction accuracy of the trained audio recommendation model. In other words, this technical solution is suitable for scenarios where more user behavior data can be obtained, and the model is trained based on massive user behavior data to achieve the technical effect that the model output can depict the user's true listening preferences.
[0086] Continue to refer Figure 5 In step S530, based on the model parameters to be optimized of the audio recommendation model, the simulated probability that the user plays the i-th candidate music in the j-th group after playing the target music in the j-th group is determined, and the simulated probability distribution of the music in the j-th group is obtained.
[0087] In an exemplary embodiment, Figure 7FIG. 5 is a flowchart showing a method for determining the simulation probability of candidate music according to an embodiment of the present disclosure, which can be used as a specific implementation of step S530. Figure 7 The embodiment shown in the figure includes steps S531 and S532.
[0088] In step S531, feature extraction is performed on the target music in the j-th group based on the first parameter of the audio recommendation model to obtain a target feature vector corresponding to the target music.
[0089] In an exemplary embodiment, Figure 8 FIG. 4 is a flow chart showing a method for extracting a target feature vector according to an embodiment of the present disclosure. Figure 8 , including steps S810 to S850. Steps S810 to S830 are used to obtain multiple frequency domain feature sequences corresponding to the target music, and steps S840 and S850 are used to extract features from the multiple frequency domain feature sequences of the target music based on the first parameter of the audio recommendation model. Specifically:
[0090] In step S810, the target music is divided into at least P audio segments according to a time domain window, where P is an integer greater than 1; in step S820, time-frequency conversion is performed on multiple sampling points belonging to the same audio segment to obtain frequency domain sequences corresponding to the P audio segments; and, in step S830, the P frequency domain sequences are sampled respectively to obtain P frequency domain feature sequences corresponding to the P audio segments respectively.
[0091] It should be noted that audio signals are expressed in two dimensions, time domain and frequency domain, and the feature sequence corresponding to the target music can be either a time domain feature sequence of the target music or a frequency domain feature sequence of the target music. In this embodiment, the frequency domain feature sequence is used as an example for description.
[0092] For example, the target music is sampled in the time dimension to obtain a discrete time sequence containing multiple sampled signals. Then, the discrete time sequence is grouped to obtain multiple audio clips. For example, the target music is sampled in the time dimension, for example, an audio signal T is sampled every 0.1s. k (k is a positive integer less than or equal to n), and the discrete time series T1~T n , where each T k The value represents the size of the audio at that sampling point. Then, the above discrete time series are combined according to a fixed time period (for example, 3s). As mentioned above, the time period length is 3s and the sampling interval is 0.1s, then each audio segment contains 3s / 0.1s=30 values. For example, the discrete time series T1~T 30As a group, and denoted as G1, the discrete time series T 31 ~T 60 As a group, and recorded as G2, and so on to get multiple audio clips: G1 ~ G P .
[0093] Furthermore, performing time-frequency conversion on the above audio clips will obtain the frequency domain feature sequence of the target music. For example, performing time-frequency conversion on each audio clip, exemplarily, frequency domain conversion is achieved through FFT (fast Fourier transform), MFCC (Mel Frequency Cepstrum Coefficient) or DFT (Discrete Fourier Transform), to obtain the frequency signal corresponding to each group of audio clips, representing the different frequency distributions contained in each audio clip. Furthermore, the frequency signal corresponding to each group of audio clips is sampled, for example, once every 10hz, to obtain a discrete frequency sequence F1~F n Assuming that the upper and lower limits of the frequency are 0 to f, the number of each frequency sequence is f / 10, and each audio segment G x (x is a positive integer not greater than P) can be expressed as a frequency sequence of f / 10. For music, some parts of the music have a heavy bass, so the corresponding time series G x The mid- and low-frequency values are very large, and some parts have very high treble, so the corresponding time series G x The mid-high frequency values are very large. Assume there are P G x , then we get a Pxn matrix, which can be used as the P frequency domain feature sequences corresponding to the above target music.
[0094] As an exemplary embodiment of determining the target feature vector according to the target music frequency domain feature sequence:
[0095] The P frequency domain feature sequences are extracted based on the pre-trained feature extraction sub-model to obtain the target feature vector corresponding to the target music. The pre-trained feature extraction sub-model includes a cascade of S convolutional neural network layers. Fig. 9 FIG. 1 is a flow chart of a method for training a feature extraction sub-model according to an embodiment of the present disclosure. Fig. 9 , the method comprising:
[0096] Step S910, based on the first parameter of the S-layer convolutional neural network layer, feature extraction is performed on the frequency domain feature sequence of the pre-trained sample audio to obtain a sample feature vector and a positive sample feature vector that has a time domain before and after relationship with the sample feature vector; Step S920, negative sample feature vectors that have no time domain before and after relationship with the sample feature vector are obtained; and, Step S930, a triplet loss function is determined based on the sample feature vector, the positive sample feature vector and the negative sample feature vector, and the first parameter is optimized based on the triplet loss function.
[0097] The purpose of the triplet loss is to make the first distance between the feature expression of the sample feature vector Anchor and the positive sample feature vector Positive belonging to the audio positive sample pair as small as possible through iterative optimization, and the second distance between the feature expression of the sample feature vector Anchor and the negative sample feature vector Negative belonging to the audio negative sample pair as large as possible. When the first distance and the second distance meet the preset requirements respectively, the model parameters of the feature extraction sub-model obtain the optimal value.
[0098] Specifically, the temporal causal relationship between the sample feature vector Anchor and the positive sample feature vector Positive means that the end of the sample feature vector Anchor is connected to the beginning of the positive sample feature vector Positive in the time domain, or the beginning of the sample feature vector Anchor is connected to the end of the positive sample feature vector Positive in the time domain. For example, the sample feature vector Anchor and the positive sample feature vector Positive correspond to the first and second bars of the same song respectively.
[0099] Correspondingly, there is no temporal relationship between the sample feature vector Anchor and the negative sample feature vector Negative in the same audio. Specifically: the end of the sample feature vector Anchor is not connected to the beginning of the negative sample feature vector Negative in the time domain, or the beginning of the sample feature vector Anchor is not connected to the end of the negative sample feature vector Negative in the time domain. For example: the sample feature vector Anchor and the negative sample feature vector Negative come from different songs.
[0100] Furthermore, based on the above pre-trained feature extraction sub-model, feature extraction is performed on the frequency domain feature sequence of the target music. Figure 8In step S840, the cascaded S-layer convolutional neural network layer in the audio recommendation model is called to perform feature extraction on the x-th frequency domain feature sequence based on the first parameter of the S-layer convolutional neural network layer to obtain S convolution feature vectors corresponding to the x-th frequency domain feature sequence, where x is an integer not greater than P.
[0101] In this embodiment, through multiple cascaded convolutional neural network layers, information of different granularities of the target music can be extracted to obtain better audio representation effects. Among them, the cascaded S-layer convolutional neural network layer includes: the first convolutional neural network layer to the S-th convolutional neural network layer, and the output of the previous neural network layer is used as the input of the next neural network. In this embodiment, the higher the convolutional neural network layer, the larger the convolution kernel size and the longer the step size, so that the dimension of the output convolution feature vector is smaller and the granularity of the audio representation is coarser.
[0102] Exemplarily, the convolution kernel size of the previous convolutional neural network layer is smaller than the convolution kernel size of the next convolutional neural network layer. Exemplarily, the step size of the previous convolutional neural network layer is smaller than the step size of the next convolutional neural network layer.
[0103] It should be noted that the convolution kernel size and step size are parameters of the convolution neural network layer, which can be used to control the size of the feature vector of the segment sample output by the convolution neural network layer. The convolution neural network layer may include several convolution kernels. When the convolution kernel is working, it will regularly scan the feature sequence corresponding to the audio segment sample, perform matrix element multiplication and summation in the receptive field, and superimpose the deviation. Specifically, the size of the convolution kernel determines the size of the receptive field. The step size defines the distance between the positions of the convolution kernel when it scans the feature sequence twice adjacently. For example, when the step size is 1, the convolution kernel will scan the elements of the feature sequence one by one.
[0104] In step S850, the S convolution feature vectors corresponding to the x-th frequency domain feature sequence are concatenated to obtain the segment feature vector corresponding to the x-th frequency domain feature sequence, and the segment feature vectors corresponding to the P frequency domain feature sequences are concatenated to obtain the target feature vector corresponding to the target music.
[0105] In an exemplary embodiment, the feature extraction submodel includes 4 cascaded convolutional neural network layers. For the xth frequency domain feature sequence, four convolutional feature vectors t1, t2, t3 and t4 are obtained respectively through the above convolutional neural network layers. Splicing is performed in the order of the convolutional neural network layers to obtain the segment feature vector {t1, t2, t3, t4} corresponding to the xth frequency domain feature sequence. Similarly, the segment feature vectors corresponding to P frequency domain feature sequences are obtained and spliced to obtain the target feature vector corresponding to the target music.
[0106] In an exemplary embodiment, the spliced segment feature vectors may be pooled, and the pooled vectors are used as the target feature vectors corresponding to the target music. Pooling the spliced segment feature vectors may reduce the dimension of the vectors to compress the amount of data and parameters, thereby reducing overfitting.
[0107] As another exemplary embodiment of determining the target feature vector according to the target music frequency domain feature sequence:
[0108] refer to Fig.10 A schematic diagram of a method for extracting a target feature vector according to another embodiment of the present disclosure is shown.
[0109] The frequency domain feature sequence G'1~G' corresponding to the above target music P As the input of the encoding layer Encoder 1010. Further, in Encoder 1010, the frequency domain feature sequence is encoded based on the encoding equation containing the parameters to be optimized, so as to process the frequency domain feature sequence containing multiple variables into a vector [h1,h2,h3,...,h P ], and the above target feature vector H is obtained.
[0110] Specifically, for the above frequency domain feature sequences G'1~G' P The Encoder 1010 that performs encoding processing can be a Recurrent Neural Network (RNN), a Long Short-Term Memory (LSTM) neural network, or a variant of LSTM, a Gated Recurrent Unit (GRU) neural network, or a bidirectional Long Short-Term Memory (Bi-directional LSTM, Bi LSTM) network.
[0111] refer to Fig.10 , using the GRU network as the encoding layer 1010, the frequency domain feature sequence G'1~G' P The embodiment of encoding processing is described below. That is, the GRU network is used to identify the speech domain feature sequence G'1~G' P Encode to obtain the corresponding encoded hidden state sequence (i.e., the target feature vector corresponding to the target music). Fig.10 , the frequency domain feature sequence [G'1, G'2, G'3, ..., G' P ] is input into the coding layer 1010, so as to encode the frequency domain feature sequence based on the coding equation containing the parameters to be optimized. A hidden state sequence can be obtained through the coding function of the coding layer 1010: H = [h1, h2, h3, ..., hP ].
[0112] Specifically, Fig.11 FIG. 2 is a flow chart of a method for extracting a target feature vector according to another embodiment of the present disclosure. Fig.11 ,include:
[0113] Step S1110, input the yth audio vector of the frequency domain feature sequence into the yth node of the gated recurrent unit, where y is a positive integer not greater than P; Step S1120, based on the first parameter, determine the yth hidden state according to the yth audio vector and the y-1th hidden state, where the y-1th hidden state is output by the y-1th node; and, Step S1130, use the P hidden states output by P nodes in sequence as a target feature vector corresponding to the target music.
[0114] It should be noted that the method of determining the target feature vector corresponding to the target music is not limited to the above two methods, and other acquisition methods may also be used, which are not limited here.
[0115] Continue to refer Figure 7 In step S532, based on the second parameter of the audio recommendation model and the target feature vector, the simulated probability of the user playing the i-th candidate music in the j-th group after playing the target music in the j-th group is determined.
[0116] In an exemplary embodiment, continue to refer to Figure 4 , based on the normalization parameter to be optimized (i.e., the second parameter) of the normalization layer 420 in the model and the target feature vector H, the simulated probability corresponding to each candidate music in the group is subjected to exponential normalization processing, and the simulated probability of the user playing the i-th candidate music in the group after playing the target music in the j-th group is obtained. Specifically, a specific implementation of the normalization processing is as follows:
[0117]
[0118] Among them, p i is the simulated probability that the user plays the i-th candidate music in the group after playing the target music in the i-th group, w i is the normalized parameter corresponding to the i-th candidate music, and H is the target feature vector corresponding to the target music.
[0119] Thus, the normalization layer outputs the simulated probability of each candidate music, and obtains the simulated probability distribution Q'=[p1,p2,p3,...p N ].
[0120] Continue to refer Figure 5In step S540, a loss function is determined according to the actual probability distribution and the simulated probability distribution, and the model parameters are optimized based on the loss function to obtain a trained audio recommendation model.
[0121] In an exemplary embodiment, continue to refer to Figure 4 , according to the actual probability distribution Q = [q1, q2, q3, ... q N ] and simulated probability distribution Q'=[p1,p2,p3,……p N ] Determine the relative entropy function as the loss function 430. The specific formula is as follows; further, by calculating the minimum value of the relative entropy function, the first parameter and the second parameter of the above-mentioned audio recommendation model are optimized.
[0122]
[0123] Exemplarily, the cross entropy loss of the actual probability distribution and the simulated probability distribution is calculated, and the above first parameter and the second parameter are optimized at any time, so as to complete an iterative process. Through multiple rounds of iterative optimization, the model parameters are optimized to meet the preset model evaluation indicators.
[0124] In an exemplary embodiment, the iteratively optimized audio recommendation model is evaluated by one or more of the following model evaluation indicators: accuracy, recall, and receiver operating characteristic curve (ROC) area under the curve (AUC) (a model evaluation indicator specifically used to evaluate the predictive value of the model; Area Under Curve). Specifically:
[0125] Exemplarily, after iterative optimization through training samples, the audio recommendation model after iterative optimization (referred to as "the model to be tested") is tested through test samples, and the test results of the model to be tested are verified using at least one test indicator, and the audio sharp reduction model that meets the test indicator is tested. Figure 2 The illustrated embodiment predicts a recommended music list based on the current music being played by the user.
[0126] In an exemplary embodiment, the specific way of testing the model to be tested may be:
[0127] First, the descriptive features of the test samples are input into the model to be tested, and the output data of the model are as follows: true positive TP, true negative TN, false negative FN and false positive FP. Among them, TP is the number of samples that are still in the positive class after being judged as the positive class in the test sample set by the model to be tested, TN is the number of samples that are still in the negative class after being judged as the negative class in the test sample set by the model to be tested, FN is the number of samples that are still in the negative class after being judged as the positive class in the test sample set by the model to be tested, and FP is the number of samples that are still in the positive class after being judged as the negative class in the test sample set by the model to be tested. Positive and negative classes refer to the two categories manually annotated for the samples in the first part, that is, if a sample is manually annotated to belong to a specific class, then the sample belongs to the positive class, and samples that do not belong to the specific class belong to the negative class.
[0128] Secondly, the test results of the model to be tested are calculated according to the true positive TP, true negative TN, false negative FN and false positive FP.
[0129] In the exemplary embodiment, the test indicators are introduced by taking accuracy and recall as examples. Specifically:
[0130] The precision p and recall r are calculated according to the following two formulas respectively;
[0131] p=TP / (TP+FP)
[0132] r=TP / (TP+FN)
[0133] If the setting conditions corresponding to the test indicators are: if the accuracy test result is greater than p' (preset value), the accuracy setting conditions are met, otherwise the accuracy setting conditions are not met, and if the recall test result is greater than r' (preset value), the recall setting conditions are met, otherwise the recall setting conditions are not met.
[0134] In an exemplary embodiment, when the test results meet the set conditions corresponding to the test indicators, the model to be tested can be used as a prediction model for predicting a list of music to be recommended based on the music currently played by the user; when the test results do not meet the set conditions, the above-mentioned model to be tested continues to iteratively optimize until the test results of the model to be tested meet the set conditions.
[0135] In an exemplary embodiment, when judging whether the test results meet the set conditions corresponding to the test indicators, only the accuracy or the recall rate can be used as the test indicators, that is, the accuracy / recall rate only needs to meet the set conditions; and both the accuracy and the recall rate can be used as test indicators at the same time, that is, the accuracy and the recall rate only need to meet the set conditions.
[0136] It should be noted that the specific test method is formulated according to actual needs and is not limited to the above accuracy and / or recall rate as test indicators.
[0137] In an exemplary embodiment, the test indicator may also be AUC. Specifically:
[0138] In an exemplary embodiment, the false positive rate FPR and the true positive rate TPR are determined using the following two formulas:
[0139] FPR=FP / (FP+TN)
[0140] TPR=TP / (TP+FN)
[0141] Furthermore, the receiver operating characteristic curve (ROC curve for short) is plotted with FPR as the horizontal coordinate and TPR as the vertical coordinate. Among them, the ROC curve is the characteristic curve of each indicator obtained, which is used to show the relationship between each indicator, and further calculates the area under the ROC curve AUC. The ROC curve is the characteristic curve of each indicator obtained, which is used to show the relationship between each indicator. AUC is the area under the ROC curve. The larger the AUC, the higher the predictive value of the model, and the AUC can be used to test the model to be tested. And when the evaluation result is that the AUC value meets the preset threshold, the obtained model can be used to predict the recommended music list based on the music currently played by the user. That is, given a target music, the model can predict the probability distribution of other music that the user will listen to next after listening to the target music.
[0142] In an exemplary embodiment, after obtaining a music recommendation model that meets the prediction model evaluation index, Figure 7 right Figure 2 The following is explained in the embodiment shown:
[0143] In step S220, the audio feature sequence (e.g., frequency domain feature sequence) of the music currently played by the target user is obtained based on the first parameter of the trained audio recommendation model, and the feature vector is encoded (refer to the embodiment corresponding to step S531). Furthermore, in step S230, the probability that the next music played by the user is the i-th candidate music is predicted based on the second parameter of the trained audio recommendation model and the above-mentioned feature vector (refer to the embodiment corresponding to step S532), and the predicted probability distribution of the N candidate music is obtained. Finally, the following is obtained: Figure 3 The "Recommended Playlist" shown.
[0144] Among them, the music in the "recommended playlist" integrates the user's listening preferences determined based on big data, and each candidate music can be sorted from high to low according to the degree of preference (determined by probability).
[0145] Those skilled in the art will appreciate that all or part of the steps to implement the above embodiments are implemented as a computer program executed by a processor (including a CPU and a GPU). When the computer program is executed by the processor, the above functions defined by the above method provided by the present disclosure are performed. The program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk or an optical disk, etc.
[0146] In addition, it should be noted that the above figures are only schematic illustrations of the processes included in the method according to the exemplary embodiment of the present disclosure, and are not intended to be limiting. It is easy to understand that the processes shown in the above figures do not indicate or limit the time sequence of these processes. In addition, it is also easy to understand that these processes can be performed synchronously or asynchronously, for example, in multiple modules.
[0147] The following is an introduction to the training device of the audio recommendation model provided by this technical solution:
[0148] This example embodiment provides a training device for an audio recommendation model. Fig.12 As shown, the training device 1200 of the audio recommendation model includes: a sample acquisition module 1201 , an actual probability determination module 1202 , a simulated probability determination module 1203 and a model parameter optimization module 1204 .
[0149] Among them, the above-mentioned sample acquisition module 1201 is configured to: obtain M groups of sample data, each group of sample data includes target music and N candidate music after the user plays the target music, M and N are positive integers; the above-mentioned simulation probability determination module 1202 is configured to: obtain the actual probability of the user playing the i-th candidate music in the j-th group after playing the target music in the j-th group, and obtain the actual probability distribution of the music in the j-th group, j is a positive integer not greater than M; the above-mentioned actual probability determination module 1203 is configured to: determine the simulated probability of the user playing the i-th candidate music in the j-th group after playing the target music in the j-th group based on the model parameters to be optimized of the audio recommendation model, and obtain the simulated probability distribution of the music in the j-th group; the above-mentioned model parameter optimization module 1204 is configured to: determine the loss function according to the actual probability distribution and the simulated probability distribution, optimize the model parameters based on the loss function, and obtain the trained audio recommendation model.
[0150] In an exemplary embodiment of the present disclosure, based on the aforementioned embodiment, the above-mentioned actual probability determination module 1202 is specifically configured to: collect a large number of users' listening behavior sequences regarding the target music in the j-th group; in the listening behavior sequence, count the N matching music in the same calculation window as the target music in the j-th group; and obtain the number of occurrences of the i-th candidate music, and normalize the number of occurrences of the i-th candidate music to obtain the actual probability of the user playing the i-th candidate music after playing the target music in the j-th group.
[0151] In an exemplary embodiment of the present disclosure, based on the aforementioned embodiment, the simulation probability determination module 1203 includes: a feature extraction submodule 12031 and a normalization submodule 12032 .
[0152] Among them, the above-mentioned feature extraction submodule 12031 is configured to: perform feature extraction on the target music in the j-th group based on the first parameter of the audio recommendation model, and obtain the target feature vector corresponding to the target music; the above-mentioned normalization unit 12032 is configured to: determine the simulated probability of the user playing the i-th candidate music in the j-th group after playing the target music in the j-th group based on the second parameter of the audio recommendation model and the target feature vector.
[0153] In an exemplary embodiment of the present disclosure, based on the aforementioned embodiment, the above-mentioned normalization submodule 12032 is specifically configured as: according to the second parameter of the audio recommendation model and the target feature vector, the simulation probabilities corresponding to the N candidate music are exponentially normalized to obtain the simulation probability that the user plays the i-th candidate music in the j-th group after playing the target music in the j-th group.
[0154] In an exemplary embodiment of the present disclosure, based on the above-mentioned embodiment, the feature extraction submodule 12031 includes: a feature sequence acquisition unit 311 and a feature extraction unit 312 .
[0155] Among them, the above-mentioned feature sequence acquisition unit 311 is configured to: obtain P frequency domain feature sequences corresponding to the target music, P is an integer greater than 1; the above-mentioned feature extraction unit 312: based on the first parameter of the audio recommendation model, perform feature extraction on the P frequency domain feature sequences to obtain the target feature vector corresponding to the target music.
[0156] In an exemplary embodiment of the present disclosure, based on the aforementioned embodiment, the feature sequence acquisition unit 311 is specifically configured to: divide the target music into at least P audio segments according to a time domain window; perform time-frequency conversion on multiple sampling points belonging to the same audio segment to obtain frequency domain sequences corresponding to the P audio segments; and sample the P frequency domain sequences respectively to obtain P frequency domain feature sequences corresponding to the P audio segments respectively.
[0157] In an exemplary embodiment of the present disclosure, based on the aforementioned embodiment, the feature extraction unit 312 is specifically configured to: call the cascaded S-layer convolutional neural network layer in the audio recommendation model to perform feature extraction on the x-th frequency domain feature sequence based on the first parameter of the S-layer convolutional neural network layer, and obtain S convolutional feature vectors corresponding to the x-th frequency domain feature sequence, where x is an integer not greater than P; and, concatenate the S convolutional feature vectors corresponding to the x-th frequency domain feature sequence to obtain a fragment feature vector corresponding to the x-th frequency domain feature sequence, and concatenate the fragment feature vectors corresponding to the P frequency domain feature sequences to obtain a target feature vector corresponding to the target music.
[0158] In an exemplary embodiment of the present disclosure, based on the aforementioned embodiment, the feature extraction unit 312 is further specifically configured as: after splicing the segment feature vectors corresponding to the P frequency domain feature sequences, performing pooling processing on the spliced segment feature vectors, and using the vector after pooling processing as the target feature vector corresponding to the target music.
[0159] In an exemplary embodiment of the present disclosure, based on the aforementioned embodiment, the training device 1200 of the audio recommendation model further includes: a sub-model pre-training module 1205.
[0160] The sub-model pre-training module 1205 is configured to: extract features from the frequency domain feature sequence of the pre-training sample audio based on the first parameter of the S-layer convolutional neural network layer to obtain a sample feature vector and a positive sample feature vector that has a time domain before-after relationship with the sample feature vector; obtain a negative sample feature vector that has no time domain before-after relationship with the sample feature vector; and determine a triplet loss function based on the sample feature vector, the positive sample feature vector and the negative sample feature vector, and optimize the first parameter based on the triplet loss function.
[0161] In an exemplary embodiment of the present disclosure, based on the aforementioned embodiment, the above-mentioned model parameter optimization module 1204 is specifically configured to: determine a relative entropy function according to the actual probability distribution and the simulated probability distribution; and optimize the second parameter of the audio recommendation model by calculating the minimum value of the relative entropy function.
[0162] In an exemplary embodiment of the present disclosure, based on the aforementioned embodiment, the feature extraction unit 312 is further specifically configured to: input the yth audio vector of the frequency domain feature sequence into the yth node of the gated recurrent unit, where y is a positive integer not greater than P; based on the first parameter, determine the yth hidden state according to the yth audio vector and the y-1th hidden state, where the y-1th hidden state is output by the y-1th node; and use the P hidden states output by the P nodes in sequence as a target feature vector corresponding to the target music.
[0163] In an exemplary embodiment of the present disclosure, based on the aforementioned embodiment, the above-mentioned model parameter optimization module 1204 is also specifically configured to: determine a relative entropy function according to the actual probability distribution and the simulated probability distribution; and optimize the first parameter and the second parameter of the audio recommendation model by calculating the minimum value of the relative entropy function.
[0164] The specific details of each module or unit in the above-mentioned audio recommendation model training device have been described in detail in the corresponding audio recommendation model training method, so they will not be repeated here.
[0165] The following is an introduction to the audio recommendation device based on artificial intelligence provided by this technical solution:
[0166] This example embodiment provides an audio recommendation device based on artificial intelligence. Fig.13 As shown, the audio recommendation device 1300 based on artificial intelligence includes: an acquisition module 1301, a first processing module 1302, a second processing module 1303 and a recommendation module 1304.
[0167] Among them, the above-mentioned acquisition module 1301 is configured to: acquire the music currently played by the target user, and input the music into the trained audio recommendation model; the above-mentioned first processing module 1302 is configured to: extract features of the music according to the first parameter of the audio recommendation model, and obtain the feature vector corresponding to the music; the above-mentioned second processing module 1303 is configured to: predict the probability that the next music played by the target user after playing the current music is the i-th candidate music according to the second parameter of the audio recommendation model and the feature vector, and obtain the predicted probability distribution of the N candidate music; and the above-mentioned recommendation module 1304 is configured to: determine the recommended music list according to the predicted probability distribution of the N candidate music.
[0168] The specific details of each module or unit in the above-mentioned artificial intelligence-based audio recommendation device have been described in detail in the corresponding artificial intelligence-based audio recommendation method, so they will not be repeated here.
[0169] Fig.14 A schematic diagram of the structure of a computer system suitable for implementing an electronic device of an embodiment of the present invention is shown.
[0170] It should be noted that Fig.14 The computer system 1400 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.
[0171] like Fig.14 As shown, the computer system 1400 includes a processor 1401, wherein the processor 1401 may include: a graphics processing unit (GPU), a central processing unit (CPU), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1402 or a program loaded from a storage part 1408 to a random access memory (RAM) 1403. In RAM 1403, various programs and data required for system operation are also stored. The processor (GPU / CPU) 1401, ROM 1402, and RAM 1403 are connected to each other via a bus 1404. An input / output (I / O) interface 1405 is also connected to the bus 1404.
[0172] The following components are connected to the I / O interface 1405: an input section 1406 including a keyboard, a mouse, etc.; an output section 1407 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1408 including a hard disk, etc.; and a communication section 1409 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1409 performs communication processing via a network such as the Internet. A drive 1410 is also connected to the I / O interface 1405 as needed. A removable medium 1411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1410 as needed so that a computer program read therefrom is installed into the storage section 1408 as needed.
[0173] In particular, according to an embodiment of the present disclosure, the process described below with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 1409, and / or installed from the removable medium 1411. When the computer program is executed by the processor (GPU / CPU) 1401, various functions defined in the system of the present application are executed. In some embodiments, the computer system 1400 may also include an AI (Artificial Intelligence) processor for processing computing operations related to machine learning.
[0174] It should be noted that the computer-readable storage medium shown in the embodiment of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by an instruction execution system, device or device or used in combination with it. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, wherein a computer-readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable storage medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0175] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0176] The units involved in the embodiments described in the present disclosure may be implemented by software or hardware, and the units described may also be arranged in a processor. The names of these units do not, in some cases, constitute limitations on the units themselves.
[0177] As another aspect, the present application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiment; or may exist independently without being assembled into the electronic device. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed by an electronic device, the electronic device implements the method described in the above embodiment.
[0178] For example, the electronic device can implement Figure 5 As shown in: Step S510, obtaining M groups of sample data, each group of sample data includes target music and N candidate music after the user plays the target music, M and N are positive integers; Step S520, obtaining the actual probability of the user playing the i-th candidate music in the j-th group after playing the target music in the j-th group, and obtaining the actual probability distribution of the music in the j-th group, j is a positive integer not greater than M; Step S530, based on the model parameters to be optimized of the audio recommendation model, determining the simulated probability of the user playing the i-th candidate music in the j-th group after playing the target music in the j-th group, and obtaining the simulated probability distribution of the music in the j-th group; and Step S540, determining the loss function according to the actual probability distribution and the simulated probability distribution, optimizing the model parameters based on the loss function, and obtaining the trained audio recommendation model.
[0179] For another example, the electronic device can implement the various steps shown in other figures.
[0180] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into multiple modules or units to be embodied.
[0181] Through the description of the above implementation, it is easy for those skilled in the art to understand that the example implementation described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the implementation of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the implementation of the present disclosure.
[0182] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.
[0183] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A training method for an audio recommendation model, characterized in that: The method comprises: Obtain M groups of sample data, each group of sample data includes a target music and N candidate music after the user plays the target music, where M and N are positive integers; Obtaining the actual probability of a user playing the i-th candidate music in the j-th group after playing the target music in the j-th group, and obtaining the actual probability distribution of the music in the j-th group, where j is a positive integer not greater than M; the actual probability of playing the i-th candidate music in the j-th group is obtained by normalizing the number of times the i-th candidate music appears in multiple target windows, where the target window is a window related to the target music in the j-th group determined from a large number of users' listening behavior sequences with respect to the target music in the j-th group; Based on the model parameters to be optimized of the audio recommendation model, determining the simulated probability that the user plays the i-th candidate music in the j-th group after playing the target music in the j-th group, and obtaining a simulated probability distribution of the music in the j-th group; A loss function is determined according to the actual probability distribution and the simulated probability distribution, and the model parameters are optimized based on the loss function to obtain a trained audio recommendation model.
2. The method according to claim 1, characterized in that Obtaining the actual probability of the user playing the i-th candidate music in the j-th group after playing the target music in the j-th group includes: Collecting a large number of users' listening behavior sequences for the target music in the jth group; In the song listening behavior sequence, statistics are collected on N matching music pieces that are in the same calculation window as the target music piece in the jth group; The number of times the i-th candidate music appears is obtained, and the number of times the i-th candidate music appears is normalized to obtain the actual probability that the user plays the i-th candidate music after playing the target music in the j-th group.
3. The method according to claim 1, characterized in that Determining the simulated probability that the user plays the i-th candidate music in the j-th group after playing the target music in the j-th group based on the model parameters to be optimized of the audio recommendation model includes: Extracting features of the target music in the j-th group based on the first parameter of the audio recommendation model to obtain a target feature vector corresponding to the target music; Based on the second parameter of the audio recommendation model and the target feature vector, a simulated probability of the user playing the i-th candidate music in the j-th group after playing the target music in the j-th group is determined.
4. The method according to claim 3, characterized in that Determining a simulated probability that the user plays the i-th candidate music in the j-th group after playing the target music in the j-th group based on the second parameter of the audio recommendation model and the target feature vector includes: According to the second parameter of the audio recommendation model and the target feature vector, the simulation probabilities corresponding to the N candidate music are exponentially normalized to obtain the simulation probability that the user plays the i-th candidate music in the j-th group after playing the target music in the j-th group.
5. The method according to claim 3, characterized in that: Extracting features of the target music in the j-th group based on the first parameter of the audio recommendation model includes: Obtain P frequency domain feature sequences corresponding to the target music, where P is an integer greater than 1; Based on the first parameter of the audio recommendation model, feature extraction is performed on the P frequency domain feature sequences to obtain a target feature vector corresponding to the target music.
6. The method according to claim 5, characterized in that Obtaining P frequency domain feature sequences corresponding to the target music, including: Dividing the target music into at least P audio segments according to the time domain window; Performing time-frequency conversion on multiple sampling points belonging to the same audio clip to obtain frequency domain sequences corresponding to P audio clips; The frequency domain sequences corresponding to the P audio clips are sampled respectively, and P frequency domain feature sequences corresponding to the P audio clips are obtained.
7. The method according to claim 5, characterized in that Based on the first parameter of the audio recommendation model, feature extraction is performed on the P frequency domain feature sequences, including: Calling the cascaded S-layer convolutional neural network layer in the audio recommendation model to perform feature extraction on the x-th frequency domain feature sequence based on the first parameter of the S-layer convolutional neural network layer, and obtaining S convolution feature vectors corresponding to the x-th frequency domain feature sequence, where x is an integer not greater than P; The S convolution feature vectors corresponding to the x-th frequency domain feature sequence are concatenated to obtain the segment feature vector corresponding to the x-th frequency domain feature sequence, and the segment feature vectors corresponding to the P frequency domain feature sequences are concatenated to obtain the target feature vector corresponding to the target music.
8. The method according to claim 7, characterized in that After splicing the segment feature vectors corresponding to the P frequency domain feature sequences, the method further includes: The concatenated segment feature vectors are pooled, and the vectors after the pooling process are used as the target feature vectors corresponding to the target music.
9. The method according to claim 7 or 8, characterized in that: The method further comprises: Extracting features from the frequency domain feature sequence of the pre-trained sample audio based on the first parameter of the S-layer convolutional neural network layer to obtain a sample feature vector and a positive sample feature vector that has a time domain before and after relationship with the sample feature vector; Obtaining a negative sample feature vector that has no temporal context relationship with the sample feature vector; A triplet loss function is determined based on the sample feature vector, the positive sample feature vector, and the negative sample feature vector, and the first parameter is optimized based on the triplet loss function.
10. The method according to claim 9, characterized in that Determining a loss function according to the actual probability distribution and the simulated probability distribution, and optimizing the model parameters based on the loss function, comprising: Determining a relative entropy function based on the actual probability distribution and the simulated probability distribution; The second parameter of the audio recommendation model is optimized by calculating the minimum value of the relative entropy function.
11. The method according to claim 5, characterized in that Based on the first parameter of the audio recommendation model, feature extraction is performed on the P frequency domain feature sequences, including: Inputting the yth audio vector of the frequency domain feature sequence into the yth node of the gated cycle unit, where y is a positive integer not greater than P; Based on the first parameter, determine the yth hidden state according to the yth audio vector and the y-1th hidden state, where the y-1th hidden state is output by the y-1th node; The P hidden states outputted by the P nodes in sequence are used as a target feature vector corresponding to the target music.
12. The method according to claim 11, characterized in that Determining a loss function according to the actual probability distribution and the simulated probability distribution, and optimizing the model parameters based on the loss function, comprising: Determining a relative entropy function based on the actual probability distribution and the simulated probability distribution; By calculating the minimum value of the relative entropy function, the first parameter and the second parameter of the audio recommendation model are optimized.
13. A training device for an audio recommendation model, characterized in that: The device comprises: The sample acquisition module is configured to: acquire M groups of sample data, each group of sample data includes a target music and N candidate music after the user plays the target music, where M and N are positive integers; The actual probability determination module is configured to: obtain the actual probability of a user playing the i-th candidate music in the j-th group after playing the target music in the j-th group, and obtain the actual probability distribution of the music in the j-th group, where j is a positive integer not greater than M; the actual probability of playing the i-th candidate music in the j-th group is obtained by normalizing the number of times the i-th candidate music appears in multiple target windows, and the target window is a window related to the target music in the j-th group determined from the listening behavior sequence of a large number of users with respect to the target music in the j-th group; A simulation probability determination module is configured to: determine the simulation probability that the user plays the i-th candidate music in the j-th group after playing the target music in the j-th group based on the model parameters to be optimized of the audio recommendation model, and obtain a simulation probability distribution of the music in the j-th group; The parameter optimization module is configured to: determine a loss function according to the actual probability distribution and the simulated probability distribution, optimize the model parameters based on the loss function, and obtain a trained audio recommendation model.
14. A computer-readable storage medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method according to any one of claims 1 to 12.
15. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; The processor is configured to perform the method of any one of claims 1 to 12 by executing the executable instructions.
16. A computer program product, characterized in that The invention comprises a computer program carried on a computer-readable storage medium, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 12 is implemented.
Citation Information
Patent Citations
User portrait data processing method and device and storage medium
CN110162698A