A neural network model training method, a song mining method and device

By using a pre-trained neural network model, based on feature extraction and similarity calculation of target songs and similar historical songs, the interactive information of target songs is predicted, and the problem of poor song mining performance in the prior art is solved, and the high-accurate song mining results are achieved.

CN112231511BActive Publication Date: 2025-06-13TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011124015.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-20
Publication Date
2025-06-13
Estimated Expiration
2040-10-20

AI Technical Summary

Technical Problem

The prior art has poor performance in song mining and is difficult to accurately meet user needs.

Method used

By obtaining the target song and historical songs with similar conditions, using a pre-trained neural network model, the interactive information of the target song is predicted based on the feature extraction and similarity calculation of the historical song and the target song, thereby determining the song mining results.

Benefits of technology

It realizes that without manually defining song value indicators, improves the accuracy of song mining, makes the mining results consistent with the actual needs of users, and improves the song mining performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112231511B_ABST
    Figure CN112231511B_ABST
Patent Text Reader

Abstract

The present application discloses a song mining method, device, equipment and computer-readable storage medium, which includes: obtaining a target song to be mined; obtaining historical songs that meet the similarity conditions with the target song; determining first target input information based on the historical songs and second target input information based on the target song; transmitting the first target input information and the second target input information to a pre-trained neural network model, and obtaining target interaction information of the target song output by the pre-trained neural network model; and determining a mining result of the target song based on the target interaction information. Since the target interaction information is the information of the interaction between the user and the target song predicted by the pre-trained neural network model, the present application realizes predicting the target interaction information according to the target song and the historical songs similar to the target song, and the mining result is consistent with the actual needs of the user for the song, and the song mining performance is good.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of information processing technology, and more specifically, to a method for training a neural network model, a method and device for song mining. Background Art

[0002] Currently, with the development of communication technology and the popularity of music, more and more songs are produced by singers, and more and more songs can be browsed by users. However, users have limited energy and it is difficult to find the songs that meet their own needs among numerous songs. Therefore, it is necessary to mine songs to obtain songs that meet user needs. For example, song value indicators can be established and, with the help of deep learning methods, songs that meet user needs can be mined based on the song value indicators. However, the above solutions require manual definition of song value indicators, and the finally mined songs may not be the songs that meet user needs, resulting in poor song mining performance.

[0003] In summary, how to improve the performance of song mining is an urgent problem to be solved by those skilled in the art in the current field. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a method for training a neural network model, a method and device for song mining, a device, a device and a computer-readable storage medium, which can effectively improve the performance of song mining. The specific solutions are as follows:

[0005] In a first aspect, the present application discloses a song mining method, including:

[0006] Obtain a target song to be mined;

[0007] Obtain historical songs that meet the similarity condition with the target song;

[0008] Determine first target input information based on the historical songs, and determine second target input information based on the target song;

[0009] Transmit the first target input information and the second target input information to a pre-trained neural network model, and obtain the target interaction information of the target song output by the pre-trained neural network model;

[0010] Determine the mining result of the target song based on the target interaction information;

[0011] Wherein, the target interaction information is used to represent the result of the interaction between the user and the target song.

[0012] Optionally, the pre-trained neural network model outputs the target interaction information based on the first target input information and the second target input information, including:

[0013] Extract features from the first target input information based on a learnable CNN network structure to obtain first target CNN features;

[0014] Extract features from the second target input information based on the learnable CNN network structure to obtain second target CNN features;

[0015] Calculate the similarity between the first target CNN features and the second target CNN features, and determine the target interaction information based on the similarity.

[0016] Optionally, the calculating the similarity between the first target CNN features and the second target CNN features, and determining the target interaction information based on the similarity includes:

[0017] Calculate the dot product value of the first target CNN features and the second target CNN features;

[0018] Concatenate the dot product value and the first target CNN features into a long feature;

[0019] Based on the long feature, calculate the similarity between the target song and each song in the historical songs through a fully connected layer;

[0020] Use the similarity as the feature weight value of the corresponding song in the historical songs, and perform weighted summation on the first target CNN features based on the feature weight value to obtain the comprehensive feature of the historical songs;

[0021] Concatenate the comprehensive feature and the second target CNN features together to obtain a concatenated feature;

[0022] Classify the concatenated feature through a fully connected layer to obtain the target interaction information of the target song.

[0023] In a second aspect, the present application discloses a song mining device, including:

[0024] A target song acquisition module, configured to acquire a target song to be mined;

[0025] A historical song acquisition module, configured to acquire historical songs that meet the similarity condition with the target song;

[0026] A target input information acquisition module, configured to determine first target input information based on the historical songs, and determine second target input information based on the target song;

[0027] A target interaction information acquisition module, configured to transmit the first target input information and the second target input information to a pre-trained neural network model, and acquire the target interaction information of the target song output by the pre-trained neural network model;

[0028] A mining result determination module, configured to determine a mining result of the target song based on the target interaction information;

[0029] Wherein, the target interaction information is used to characterize the result of the interaction between the user and the target song.

[0030] In a third aspect, the present application discloses a method for training a neural network model, including:

[0031] Obtaining sample songs with known interaction information;

[0032] Dividing a training set from the sample songs;

[0033] In the training set, selecting a training song set and a first song set that satisfies a similarity condition with the training song set;

[0034] Determining first training input information based on the first song set, and determining second training input information based on the training song set;

[0035] Taking the first training input information and the second training input information as inputs of an initial neural network model, training the initial neural network model, and inputting the known interaction information of the training song set and the predicted interaction information of the training song set output by the initial neural network model into a preset loss function to obtain a loss value;

[0036] Judging whether the loss value converges. If the loss value does not converge, adjusting the network parameters of the initial neural network model according to the loss value, and returning to execute the step of selecting a training song set and a first song set that satisfies a similarity condition with the training song set in the training set; if the loss value converges, completing the training of the initial neural network model to obtain a pre-trained neural network model for song mining based on the pre-trained neural network model.

[0037] Optionally, after obtaining the pre-trained neural network model, it further includes:

[0038] Dividing a validation set from the sample songs;

[0039] Performing performance evaluation on the pre-trained neural network model based on the validation set to obtain a performance evaluation result;

[0040] Judging whether the performance evaluation result meets a preset requirement;

[0041] If the performance evaluation result meets the preset requirement, allowing the application of the pre-trained neural network model;

[0042] If the performance evaluation result does not meet the preset requirements, change the training strategy and continue to train the pre-trained neural network model.

[0043] Optionally, the performance evaluation of the pre-trained neural network model based on the validation set to obtain a performance evaluation result includes:

[0044] In the validation set, select a validation song set and a second song set that satisfies the similarity condition with the validation song set;

[0045] Determine first validation input information based on the second song set, and determine second validation input information based on the validation song set;

[0046] Use the first validation input information and the second validation input information as inputs to the pre-trained neural network model, and obtain predicted interaction information output by the pre-trained neural network model;

[0047] Determine the performance evaluation result based on the predicted interaction information and the known interaction information of the validation song set.

[0048] Optionally, the determination of the first training input information based on the first song set and the second training input information based on the training set includes:

[0049] Determine the target Mel spectrogram feature of the first song set, and determine the target Mel spectrogram feature of the first song set as the first training input information;

[0050] Determine the target Mel spectrogram feature of the training set, and determine the target Mel spectrogram feature of the training set as the second training input information.

[0051] Optionally, the determination process of the target Mel spectrogram feature of the song includes:

[0052] Perform short-time Fourier transform on the audio of the song to obtain a short-time Fourier transform result;

[0053] Perform Mel spectrogram coefficient conversion on the short-time Fourier transform result to obtain an initial Mel spectrogram feature;

[0054] Truncate the initial Mel spectrogram feature to obtain the target Mel spectrogram feature of the song.

[0055] In a fourth aspect, the present application provides a training device for a neural network model, including:

[0056] A sample song acquisition module, configured to acquire sample songs with known interaction information;

[0057] A training set division module, configured to divide a training set from the sample songs;

[0058] A first song set selection module, configured to select a training song set and a first song set that meets a similarity condition with the training song set from the training set;

[0059] A training input information determination module, configured to determine first training input information based on the first song set and determine second training input information based on the training song set;

[0060] A loss value acquisition module, configured to use the first training input information and the second training input information as inputs to an initial neural network model, train the initial neural network model, and input the known interaction information of the training song set and the predicted interaction information of the training song set output by the initial neural network model into a preset loss function to obtain a loss value;

[0061] An adjustment module, configured to determine whether the loss value converges. If the loss value does not converge, adjust the network parameters of the initial neural network model according to the loss value, and prompt the first song set selection module to execute the step of selecting a training song set and a first song set that meets a similarity condition with the training song set from the training set; if the loss value converges, complete the training of the initial neural network model to obtain a pre-trained neural network model for song mining based on the pre-trained neural network model.

[0062] In a fifth aspect, the present application discloses an electronic device, including:

[0063] A memory, configured to store a computer program;

[0064] A processor, configured to execute the computer program to implement the training method of the neural network model or the song mining method as described in any one of the above.

[0065] In a sixth aspect, the present application discloses a computer-readable storage medium, configured to store a computer program, and when the computer program is executed by a processor, implement the training method of the neural network model or the song mining method as described in any one of the above.

[0066] A song mining method provided by the present application first obtains a target song to be mined; obtains historical songs that meet the similarity conditions with the target song; then determines first target input information based on the historical songs and second target input information based on the target song; and transmits the first target input information and the second target input information to a pre-trained neural network model to obtain target interaction information of the target song output by the pre-trained neural network model; finally, determines the mining result of the target song based on the target interaction information. Since the target interaction information is the information of the interaction between the user and the target song predicted by the pre-trained neural network model, the present application realizes predicting the target interaction information according to the target song and the historical songs similar to the target song. In this process, there is no need to manually define song value indicators, and since the mining result is adapted to the target interaction information, and the target interaction information reflects the user's demand for the target song, the mining result is consistent with the actual demand of the user for the song, and the song mining performance is good. A song mining device, an electronic device, and a computer-readable storage medium provided by the present application also solve the corresponding technical problems. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0068] Figure 1 It is a schematic diagram of the system framework applicable to the song mining solution provided by the present application;

[0069] Figure 2 It is a flowchart of a song mining method provided by an embodiment of the present application;

[0070] Figure 3 It is a training flowchart of the neural network model in the present application;

[0071] Figure 4 It is another training flowchart of the neural network model in the present application;

[0072] Figure 5 It is a determination flowchart of the target Mel spectrogram feature in the present application;

[0073] Figure 6 It is a flowchart of the neural network model determining the target interaction information in the present application;

[0074] Figure 7 It is a deployment schematic diagram of the neural network model;

[0075] Figure 8Structural schematic diagram of a song mining device provided by this application;

[0076] Figure 9 Structural schematic diagram of a training device for a neural network model provided by this application;

[0077] Figure 10 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment. Detailed implementation manners

[0078] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0079] Currently, with the development of communication technology and the popularity of music, more and more songs are produced by singers, and more and more songs can be browsed by users. However, users have limited energy and it is difficult to find the songs that meet their own needs among numerous songs. Therefore, it is necessary to mine songs to obtain songs that meet user needs. For example, song value indicators can be established, and with the help of deep learning methods, songs that users like can be mined based on the song value indicators. However, the above solutions require manual definition of song value indicators, and the manually defined song value indicators are prone to deviate from the true needs of users for songs. For example, if the song value indicators are defined manually according to the quality of songs, although good or bad songs can be mined, the finally mined songs may not be the songs that users need, and the mining accuracy is poor. To overcome the above technical problems, this application provides a song mining solution, which can improve the accuracy of the song mining method.

[0080] In the song mining solution of this application, the specific system framework can be referred to Figure 1 As shown, it may specifically include: a background server 01 and a number of user terminals 02 that establish a communication connection with the background server 01.

[0081] In this application, the background server 01 is used to execute the steps of the song mining method, including obtaining the target song to be mined; obtaining the historical songs that meet the similarity conditions with the target song; determining the first target input information based on the historical songs and the second target input information based on the target song; transmitting the first target input information and the second target input information to a pre-trained neural network model, and obtaining the target interaction information of the target song output by the pre-trained neural network model; determining the mining result of the target song based on the target interaction information; where the target interaction information is used to represent the result of the interaction between the user and the target song.

[0082] Furthermore, the background server 01 may also be provided with a target song database, a historical song database, a target interaction information database, and a mining result database. Among them, the target song database is used to store the target songs to be mined, the historical song database is used to store the songs with known interaction information, the target interaction information database is used to store the interaction information output by the pre-trained neural network model, and the mining result database is used to store the finally obtained mining results. It can be understood that after the target songs are mined by the song mining solution of this application, the target songs can be transferred from the target song database to the historical song database, so that the target songs can be used as historical songs to mine new target songs later.

[0083] Of course, this application can also set the above-mentioned target song database and other databases in a third-party business server, and the above-mentioned business server can specifically collect the data uploaded by the business side. In this way, when the background server 01 needs to use the corresponding data, it can obtain the corresponding data by sending a corresponding data call request to the above-mentioned business server, such as obtaining historical songs by sending a historical song call request to the above-mentioned business server.

[0084] In this application, the background server 01 can respond to the song mining requests of one or more user terminals 02. It can be understood that the song mining requests initiated by different user terminals 02 in this application can be on-demand requests for the same song or on-demand requests for different songs.

[0085] Figure 2 It is a flowchart of a song mining method provided by an embodiment of this application. See Figure 2 As shown, the song mining method includes:

[0086] Step S11: Obtain the target song to be mined.

[0087] In this embodiment, the target song refers to a song with unknown interaction information, which may be a newly released song by a singer, or a song that has been released by the singer but has not been listened to by users or has been listened to by few users, etc.

[0088] It can be understood that since the interaction information of the target song is unknown, it is impossible to determine whether the target song meets the user's needs. If the target song is not mined, on the one hand, it is easy to miss the songs that meet the user's needs, and on the other hand, it is easy to reduce the value of the songs. Therefore, in this application, it is necessary to obtain the target song to be mined and mine the target song.

[0089] It should be noted that the interaction information of the song is used to characterize the result of the user's interaction with the song. For example, the interaction information can be the popularity information of the song, the moment when the user plays the song, etc., and its type can be determined according to actual needs. In addition, the audio data volume of the song can be reduced to ensure the operation efficiency of the song mining method, that is, the format of the target song and other songs can be MP3 (Moving Picture Experts Group Audio Layer III).

[0090] Step S12: Obtain historical songs that meet the similarity conditions with the target song.

[0091] In this embodiment, after obtaining the target song to be mined, instead of directly or through deep learning methods to determine the song value features of the target song, it is necessary to first obtain historical songs that meet the similarity conditions with the target song. Since the similarity conditions are used to judge whether the historical song is similar to the target song, and the historical song is a song with known interaction information, the historical songs obtained in this application are those that are similar to the target song and have known interaction information. In this way, the target interaction information of the target song can be predicted with the help of the historical songs similar to the target song.

[0092] It should be noted that the similarity conditions for judging whether a historical song is similar to the target song can be determined according to actual needs. For example, the similarity conditions can be one or more of language similarity conditions, genre similarity conditions, singer similarity conditions, era similarity conditions, etc.

[0093] Step S13: Determine the first target input information based on the historical songs and determine the second target input information based on the target song.

[0094] Step S14: Transmit the first target input information and the second target input information to a pre-trained neural network model, and obtain the target interaction information of the target song output by the pre-trained neural network model.

[0095] In this embodiment, since there is no artificially defined song value index, there is no need to extract specific value index information from the song. Then, this application can use a neural network model to predict the target interaction information of the target song. For example, after obtaining historical songs that meet the similarity conditions with the target song, the first target input information is determined based on the historical songs, the second target input information is determined based on the target song, the first target input information and the second target input information are transmitted to a pre-trained neural network model, and the target interaction information of the target song output by the pre-trained neural network model is obtained. The first target input information and the second target input information are the input information when the neural network model predicts the interaction information of the target song. The first target input information is also the song input information corresponding to the historical song. The first target input information can be directly the historical song itself, or the corresponding song features obtained after feature extraction of the historical song, etc.; the second target input information is also the song input information corresponding to the target song. The second target input information can be directly the target song itself, or the corresponding song features obtained after feature extraction of the target song, etc.; this application does not make specific limitations here.

[0096] It should be noted that since the interaction information of the historical song is known, if this application determines the first target input information based on the historical song and transmits the first target input information to the pre-trained neural network model, the interaction information of the historical song can be used to constrain the target interaction information of the target song, which can ensure the authenticity of the target interaction information.

[0097] Step S15: Determine the mining result of the target song based on the target interaction information; where the target interaction information includes the result of the interaction between the user and the target song predicted by the pre-trained neural network model.

[0098] In this embodiment, after obtaining the target interaction information by using the pre-trained neural network model, only the predicted interaction information between the user and the target song is obtained. At this time, the mining result of the target song cannot be determined. Therefore, after obtaining the target interaction information, it is also necessary to determine the mining result of the target song based on the target interaction information. Specifically, the target interaction information can be directly used as the mining result, or the target interaction information can be extended to obtain the mining result. For example, the corresponding level of the song is determined according to the target interaction information, and this level is used as the mining result, etc. In this process, according to the specific value of the interaction information output by the pre-trained neural network model and the level definition of the song, the corresponding relationship between the target interaction information and the corresponding level of the song can be determined, and then the corresponding level corresponding to the target interaction information can be determined by using this corresponding relationship.

[0099] It should be noted that the type of the mining result in this application can be determined according to the scenario applied in this application and the type of interaction information. For example, if the song mining method provided in this application is applied to the scenario of popular song mining, the type of interaction information can be limited to song popularity information, and the type of song popularity information can be the song completion rate, the song play volume, the song click volume, the song repost volume, etc. In this process, after obtaining the target song popularity information, the target song popularity information can be directly used as the mining result, or the popular level of the target song can be judged according to the target song popularity information, and this popular level can be used as the mining result of the target song, etc. In addition, in this scenario, the promotion method of the target song or the corresponding signing information of the target song can be further set according to the mining result. For example, if the mining result indicates that the popular level of the target song is relatively high, the target song can be promoted in the way of a headline song or in the name of a high-quality song, and the promotion intensity of the target song can be increased. The target song can also be signed, the author of the target song can be signed, and the author of the target song can be invited to join the music platform, etc.

[0100] If the song mining method provided in this application is applied to the scenario of recommending songs to be played at different time periods for users, the type of interaction information can be limited to the moment when the user plays the song. In this process, the moment when the user plays the target song obtained can be directly determined as the mining result, or the moment when the user plays the target song can be compared with the time period when the user usually listens to songs to determine the mining result. For example, if the predicted moment when the user plays the target song is 5 pm, and the user usually listens to songs at 6 pm, the information indicating that the playing moment of the target song is 6 pm can be determined as the mining result. In this scenario, if the user listens to songs at 6 pm later, the user will hear the target song, enriching the user's song listening experience.

[0101] In addition, it should be further noted that the users served by the song mining method provided in this application can be individual users or group users, etc., and this application does not make specific limitations here.

[0102] A song mining method provided by this application first obtains a target song to be mined; obtains historical songs that meet similar conditions with the target song; then determines first target input information based on the historical songs and second target input information based on the target song; and transmits the first target input information and the second target input information to a pre-trained neural network model to obtain target interaction information of the target song output by the pre-trained neural network model; finally, determines a mining result of the target song based on the target interaction information. Since the target interaction information is the information of the interaction between the user and the target song predicted by the pre-trained neural network model, this application realizes predicting the target interaction information according to the target song and the historical songs similar to the target song. In this process, there is no need to manually define song value indicators, and since the mining result is adapted to the target interaction information, and the target interaction information reflects the user's demand for the target song, the mining result is consistent with the actual demand of the user for the song, and the song mining performance is good.

[0103] Figure 3 It is a training flow chart of the neural network model in this application.

[0104] In a song mining method provided by an embodiment of this application, before transmitting the first target input information and the second target input information to a pre-trained neural network model, the neural network model can also be trained, and the training process can include:

[0105] Step S201: Obtain sample songs with known interaction information.

[0106] In this embodiment, during the training of the neural network model, sample songs with known interaction information can be obtained first, so as to subsequently train the neural network model with the sample songs.

[0107] It can be understood that in a specific application scenario, during the process of obtaining sample songs with known interaction information, the interaction records between users and songs can be collected first. For example, the corresponding song interaction records can be collected through a music software, and then the sample songs and the interaction information of the sample songs can be determined according to the interaction records between users and songs. During this process, the songs played by users in the music software can be used as sample songs, and the interaction records between users and sample songs can be converted into corresponding interaction information. For example, if the interaction record between a user and a sample song is that the user has played the sample song three times, then the interaction information of this sample song can be that the number of plays is three. Another example is that if the interaction record between a user and a sample song is that the user starts playing the sample song at 8:00 in the morning, then the interaction information of this sample song can be that the playing time is 8:00 in the morning, etc. It should be noted that the information of the sample songs can be determined according to actual needs. For example, the sample songs can include song IDs, song information, singer information, etc. The type of interaction information can also be determined according to actual needs. For example, the interaction information can be song popularity information, the time when the user plays the song, etc.

[0108] Step S202: Divide a training set from the sample songs.

[0109] In this embodiment, after obtaining the sample songs with known interaction information, a training set can be divided from the sample songs. Since the training set is obtained by dividing the sample songs, the interaction information of the training set is also known. In this way, the initial neural network model can be trained using the training set subsequently.

[0110] Step S203: Select a training song set and a first song set that meets the similarity condition with the training song set from the training set.

[0111] In this embodiment, because in the process of the neural network model outputting the target interaction information, the participation of historical songs similar to the target song is required, and the interaction information of the historical songs is needed to constrain the interaction information of the target song. Therefore, during the training process of the neural network model, a song set similar to the training song set needs to be determined. That is, after dividing a training set from the sample songs, a training song set and a first song set that meets the similarity condition with the training song set can be selected from the training set. Since the first song set is also obtained by dividing the training set, the interaction information of the first song set is also known. In this way, the first song set can be used to cooperate with the training song set to train the neural network model subsequently.

[0112] It should be noted that the quantities of the sample songs, the training song set, and the first song set can be determined according to actual needs, and the present application does not make specific limitations here. For example, 80% of the sample songs can be used as the training set; then in the training set, 256 songs can be selected as the training song set each time, and 5000 songs outside the training song set can be selected as the first song set each time, etc.

[0113] Step S204: Determine first training input information based on the first song set, and determine second training input information based on the training song set.

[0114] Step S205: Use the first training input information and the second training input information as the input of the initial neural network model, train the initial neural network model, and input the known interaction information of the training song set and the predicted interaction information of the training song set output by the initial neural network model into a preset loss function to obtain a loss value.

[0115] In this embodiment, after determining the training song set and the first song set, the first training input information can be determined based on the first song set, the second training input information can be determined based on the training song set, and the first training input information and the second training input information are used as the input of the initial neural network model to train the initial neural network model. Then, the known interaction information of the training song set and the predicted interaction information of the training song set output by the initial neural network model are input into a preset loss function to obtain a loss value, so as to adjust the network parameters of the initial neural network model according to the loss value subsequently.

[0116] It should be noted that the first training input information and the second training input information are the input information for training the neural network model; the first training input information is also the song input information corresponding to the first song set. The first training input information can be directly the first song set itself, or the corresponding song features obtained after feature extraction of the first song set, etc.; the second training input information is also the song input information corresponding to the training song set. The second training input information can be directly the training song set itself, or the corresponding song features obtained after feature extraction of the training song set, etc. This application does not make specific limitations here. In addition, the type of the loss function can be determined according to actual needs. For example, the loss function can be an L2 loss function, a PairweiseLoss loss function, a cross-entropy loss function, etc.

[0117] Step S206: Determine whether the loss value converges. If the loss value does not converge, execute Step S207; if the loss value converges, execute Step S208.

[0118] Step S207: Adjust the network parameters of the initial neural network model according to the loss value, and return to Step S203.

[0119] Step S208: Complete the training of the initial neural network model to obtain a pre-trained neural network model.

[0120] In this embodiment, after obtaining the loss value, it is necessary to determine whether the training of the initial neural network model is completed by judging whether the loss value converges. If the loss value does not converge, it indicates that the initial neural network model does not meet the training standard. At this time, it is necessary to adjust the network parameters of the initial neural network model according to the loss value and return to execute step S203 to start a new round of training process for the initial neural network model. It should be noted that the training song set selected when executing step S203 again can be different from the previously selected training song set to achieve a new training effect. If the loss value converges, it indicates that the initial neural network model has met the training standard, and at this time, a pre-trained neural network model can be obtained.

[0121] It should be noted that the type of the neural network model applied in this application can be determined according to actual needs. For example, the neural network model can be a convolutional neural network, a recurrent neural network, a Transfomer, etc.

[0122] It can be seen that in this embodiment, the initial neural network model is trained through the interaction information, the training song set, and the first song set similar to the training song set. Thus, a pre-trained neural network model can be quickly obtained, and there is no need to manually define the song value index during the training process, avoiding the situation where the mining result does not meet the user's needs due to manual definition.

[0123] Figure 4 This is another training flowchart of the neural network model in this application.

[0124] In a song mining method provided by an embodiment of this application, before transmitting the first target input information and the second target input information to the pre-trained neural network model, the neural network model can also be trained, and the training process can include:

[0125] Step S201: Obtain sample songs with known interaction information.

[0126] Step S202: Divide a training set from the sample songs.

[0127] Step S203: In the training set, select a training song set and a first song set that meet the similarity condition with the training song set.

[0128] Step S204: Determine the first training input information based on the first song set, and determine the second training input information based on the training song set.

[0129] Step S205: Use the first training input information and the second training input information as the inputs of the initial neural network model, train the initial neural network model, and input the known interaction information of the training song set and the predicted interaction information of the training song set output by the initial neural network model into a preset loss function to obtain a loss value.

[0130] Step S206: Determine whether the loss value converges. If the loss value does not converge, execute Step S207; if the loss value converges, execute Step S208.

[0131] Step S207: Adjust the network parameters of the initial neural network model according to the loss value, and return to Step S203.

[0132] Step S208: Complete the training of the initial neural network model to obtain a pre-trained neural network model, and execute Step S209.

[0133] Step S209: Divide a validation set from the sample songs, and use the validation set to evaluate the performance of the pre-trained neural network model. If the performance evaluation result meets the preset requirements, allow the application of the pre-trained neural network model. If the performance evaluation result does not meet the preset requirements, change the training strategy and continue to train the neural network model to obtain a pre-trained neural network model that allows application.

[0134] In this embodiment, during the training of the neural network model, after training the initial neural network model with the training set, the first song set, and the interaction information of the training song set, it may directly obtain a neural network model that meets the requirements, or it may not obtain a neural network model that meets the requirements. At this time, it is necessary to verify and adjust the trained neural network model obtained by training with the training set, the first song set, and the interaction information of the training song set to obtain a pre-trained neural network model that meets the requirements. During this process, for the convenience of operation, a validation set can be directly divided from the sample songs. It should be noted that the validation set and the training set should be different song sets, that is, there are no overlapping songs in the validation set and the training set.

[0135] In practical applications, during the process of evaluating the performance of the pre-trained neural network model with the validation set, a validation song set and a second song set that meet the similarity conditions with the validation song set can be selected from the validation set first; determine the first validation input information based on the second song set, and determine the second validation input information based on the validation song set; use the first validation input information and the second validation input information as the inputs of the trained neural network model, and obtain the predicted interaction information output by the trained neural network model; determine the performance evaluation result based on the predicted interaction information and the known interaction information of the validation song set.

[0136] It should be noted that the quantity of the validation set, the selected validation song set in the validation set, and the quantity of songs in the second song set can be determined according to actual needs. For example, 20% of the remaining sample songs can be used as the validation set; in the validation set, 256 songs are selected as the validation song set each time, and 5000 songs are selected as the second song set, etc. In addition, the first validation input information and the second validation input information are the input information for validating the neural network model; the first validation input information is also the song input information corresponding to the second song set. The first validation input information can directly be the second song set itself, or the corresponding song features obtained after feature extraction of the second song set, etc.; the second validation input information is also the song input information corresponding to the validation song set. The second validation input information can directly be the validation song set itself, or the corresponding song features obtained after feature extraction of the validation song set, etc.; the present application does not make specific limitations here. Furthermore, if the trained neural network model meets the requirements, the gap between the predicted interaction information and the known interaction information of the validation song set will not be too large. Therefore, the performance of the trained neural network model can be evaluated based on the predicted interaction information and the known interaction information of the validation song set. At this time, the preset requirement can be that the gap between the predicted interaction information and the known interaction information of the validation song set is within a preset range, etc.

[0137] It can be seen that in this embodiment, only by dividing the training set and the validation set from the sample songs can the training and adjustment of the neural network model be completed with the help of the training set and the validation set, quickly obtaining a pre-trained neural network model that meets the requirements, accelerating the training efficiency of the neural network model, and further accelerating the song mining efficiency.

[0138] In a song mining method provided by an embodiment of the present application, although the corresponding song can be directly used as the input information of the neural network model, due to the large amount of data in the song itself, it will increase the operation burden of the neural network model. To avoid this situation, the song information of the corresponding song can be extracted, and the extracted song information can be used as the input of the neural network model. The extracted song information can be the corresponding Mel Spectrogram (Mel) feature of the song. That is, the process of determining the first training input information based on the first song set and the second training input information based on the training song set can be specifically to determine the target Mel Spectrogram feature of the first song set and determine the target Mel Spectrogram feature of the first song set as the first training input information; determine the target Mel Spectrogram feature of the training song set and determine the target Mel Spectrogram feature of the training song set as the second training input information. Correspondingly, subsequently, it is necessary to determine the target Mel Spectrogram feature of the second song set as the first verification input information, determine the target Mel Spectrogram feature of the verification song set as the second verification input information, determine the target Mel Spectrogram feature of the historical song as the first target input information, and determine the target Mel Spectrogram feature of the target song as the second target input information.

[0139] Figure 5 It is a flowchart for determining the target Mel Spectrogram feature in the present application.

[0140] In the song mining method provided by an embodiment of the present application, the process of determining the target Mel Spectrogram feature of the song may include the following steps:

[0141] Step S31: Perform a short-time Fourier transform on the audio of the song to obtain a short-time Fourier transform result.

[0142] In this embodiment, in the process of determining the target Mel Spectrogram feature of the song, a short-time Fourier transform can be first performed on the audio of the song to obtain a short-time Fourier transform result. The short-time Fourier transform process can be determined according to actual needs. For example, the time window parameter W of the short-time Fourier transform 1 can be 1024, and the size can be R T*F etc., where R represents a real number, T represents time, and F represents the frequency domain; and the style of the short-time Fourier transform result can be determined according to actual needs.

[0143] Step S32: Perform a Mel Spectrogram coefficient conversion on the short-time Fourier transform result to obtain an initial Mel Spectrogram feature.

[0144] In this embodiment, if the short-time Fourier transform result is directly used as the input of the neural network model, the input information of the neural network model will still be relatively large. To further reduce the data volume of the input information of the neural network model, after obtaining the short-time Fourier transform result, the short-time Fourier transform result can be further subjected to Mel-frequency cepstral coefficient conversion to obtain the initial Mel-frequency cepstral features, and the structure of the initial Mel-frequency cepstral features can be R T*F , where R represents the real number, T represents time, and F represents the frequency domain.

[0145] Step S33: Truncate the initial Mel-frequency cepstral features to obtain the target Mel-frequency cepstral features of the song.

[0146] In this embodiment, if the initial Mel-frequency cepstral features are directly used as the input information of the neural network model, since the initial Mel-frequency cepstral features may carry invalid information and the data volume of the initial Mel-frequency cepstral features may be relatively large, in order to further reduce the data volume of the input information of the neural network model, the initial Mel-frequency cepstral features can be truncated to obtain the target Mel-frequency cepstral features of the song, and the target Mel-frequency cepstral features are used as the input information of the neural network model.

[0147] It should be noted that truncating the initial Mel-frequency cepstral features means truncating the initial Mel-frequency cepstral features according to the time information. For example, if the time length of the initial Mel-frequency cepstral features is 3 minutes and the time length of the required target Mel-frequency cepstral features is 2 minutes, then the initial Mel-frequency cepstral features can be truncated at the 2-minute mark to obtain the target Mel-frequency cepstral features. At this time, the target Mel-frequency cepstral features are R 5167*F .

[0148] It can be seen that in this embodiment, by determining the target Mel-frequency cepstral features of the song and using the target Mel-frequency cepstral features of the song as the input information of the neural network model, the data volume of the input information of the neural network model can be reduced, the running efficiency of the neural network model can be accelerated, and thus the efficiency of song mining can be improved.

[0149] In the song mining method provided by the embodiment of the present application, the process of the pre-trained neural network model outputting the target interaction information based on the first target input information and the second target input information can be specifically as follows: Feature extraction is performed on the first target input information based on a learnable CNN (Convolutional Neural Networks) network structure to obtain the first target CNN features; Feature extraction is performed on the second target input information based on a learnable CNN network structure to obtain the second target CNN features; Calculate the similarity between the first target CNN features and the second target CNN features, and determine the target interaction information based on the similarity.

[0150] That is, in this embodiment, the neural network model can extract the CNN features of historical songs and the CNN features of target songs, and determine the target interaction information according to the similarity between the CNN features of the target song and the CNN features of the historical songs, further enhancing the constraint of the historical songs on the target interaction information. Since the interaction information of the historical songs is known and is real interaction information, the target interaction information can be made more real.

[0151] It should be noted that since the learnable CNN network structure does not require manual definition of similar criteria and degree calculation methods, the CNN features output by the learnable CNN network structure are all automatically learned by the network structure. Moreover, high-quality songs have similarities, and low-quality songs also have commonalities. Therefore, in this embodiment, the learnable CNN network structure can output CNN features representing the commonalities of songs without manual definition, facilitating the subsequent neural network model to predict interaction information based on the CNN features. That is, the CNN features output by the learnable CNN network structure carry features for predicting interaction information, and the songs can be converted into a data type that is easy for the neural network model to process with the help of the CNN features, facilitating the subsequent neural network model to predict interaction information of songs based on the CNN features. In addition, the first target CNN feature is the CNN feature of the historical song, the second target CNN feature is the CNN feature of the target song, and the learnable CNN network structure can be a part of the neural network model.

[0152] Figure 6 It is a flowchart for the neural network model in this application to determine the target interaction information.

[0153] In the song mining method provided by the embodiment of this application, the process of the neural network model determining the target interaction information may include the following steps:

[0154] Step S41: Extract features from the first target input information based on the learnable CNN network structure to obtain the first target CNN feature.

[0155] Step S42: Extract features from the second target input information based on the learnable CNN network structure to obtain the second target CNN feature.

[0156] In this embodiment, the neural network model can first extract features from the first target input information and the second target input information based on the learnable CNN network structure to obtain the corresponding CNN features, so as to subsequently determine the target interaction information according to the CNN features of the target song and the CNN features of the historical song.

[0157] Step S43: Calculate the dot product value of the first target CNN feature and the second target CNN feature.

[0158] Step S44: Concatenate the dot product value and the first target CNN feature into a long feature.

[0159] Step S45: Based on the long feature, calculate the similarity between the target song and each song in the historical songs through a fully connected layer.

[0160] In this embodiment, in order to quickly calculate the similarity value, the neural network model can calculate the dot product value between the first target CNN feature and the second target CNN feature, concatenate the dot product value and the first target CNN feature into a long feature, and then based on the long feature, calculate the similarity between the target song and each song in the historical songs through a fully connected layer. Correspondingly, a connection layer and a fully connected layer need to be built in the neural network model. Of course, the structure of the neural network model can also be enriched according to actual needs. For example, the neural network model can also include attention, softmax, etc.

[0161] It should be noted that the calculation method of the similarity can be determined according to actual needs. For example, the similarity can be determined through cosine distance, L1 distance, L2 distance, etc.

[0162] Step S46: Use the similarity as the feature weight value of the corresponding song in the historical songs, and based on the feature weight value, perform weighted summation on the first target CNN feature to obtain the comprehensive feature of the historical songs.

[0163] Step S47: Concatenate the comprehensive feature and the second target CNN feature together to obtain a concatenated feature.

[0164] Step S48: Classify the concatenated feature through a fully connected layer to obtain the target interaction information of the target song.

[0165] In this embodiment, in the process of the neural network model determining the target interaction information based on the similarity, the similarity can be used as the feature weight value of the corresponding song in the historical songs, and based on the feature weight value, perform weighted summation on the first target CNN feature to obtain the comprehensive feature of the historical songs. Concatenate the comprehensive feature and the second target CNN feature together to obtain a concatenated feature, and classify the concatenated feature through a fully connected layer to obtain the target interaction information of the target song.

[0166] It should be noted that in a specific application scenario, since the historical songs are existing songs, the CNN features of the historical songs can be pre-saved, and when needed, the CNN features of the historical songs can be directly input into the neural network model, so that the neural network model can obtain the CNN features of the historical songs without any operation, which improves the processing efficiency of the neural network model. Please refer to Figure 7 , Figure 7 for a deployment schematic diagram of the neural network model. In Figure 7Among them, Seed features represent the CNN features of historical songs, audio backbone represents the network for extracting the CNN features of target songs, Seed Feature Fusion represents the network for calculating similarity and outputting target interaction information, Concat represents the connection layer, FC represents the fully connected layer; Pooling represents pooling; Out Product represents the dot product value; similarity represents similarity; MelSpectrum represents the extraction of target Mel spectrum features; ConvBlock represents the convolutional block; Average Pooling represents average pooling.

[0167] It can be seen that in this embodiment, the neural network model can determine the similarity between the target song and the historical song through the CNN features, and can increase the coupling degree between the similarity and the target interaction information through weighted summation, concatenation, and classification, improving the accuracy of the neural network model to determine the target interaction information according to the similarity.

[0168] For the sake of easy understanding, the song mining method provided by this application will now be described in combination with the popular song mining scenario. Assuming that the interaction information relied on during the song mining process is the song completion rate, the song mining method provided by this application may include the following steps:

[0169] Obtain the specific information of each song being played from the original massive user play records. The specific information includes song ID, user ID, play duration, song duration, etc.;

[0170] Take all the songs in the user play records as sample songs, download the sample songs in the MP3 format, and calculate the completion rate of the sample songs; and the completion rate is the ratio of the number of complete plays of the song to the number of effective plays. The number of complete plays can be the number of plays with a play percentage greater than or equal to 90%, and the number of effective plays can be the number of plays with a play percentage greater than or equal to 30%, etc.;

[0171] Randomly divide the sample songs into a training set and a validation set in a ratio of 8:2;

[0172] In the training set, select a training song set and a first song set that meet the similarity conditions with the training song set. For example, take 256 songs in the training set as the training song set, and take 5000 songs in the training set other than the training song set as the first song set;

[0173] Determine the target Mel spectrum features of the first song set as the first training input information, and determine the target Mel spectrum features of the training song set as the second training input information. The acquisition process of the target Mel spectrum features refers to the above embodiment. At this time, the size of the second training input information is R M*1*5167*F, the value of M is 5000, and the size of the first training input information is R B*1*5167*F , the value of B is 256;

[0174] Take the first training input information and the second training input information as the input of the initial neural network model, train the initial neural network model, and input the known completion rate of the training song set and the predicted completion rate of the training song set output by the initial neural network model into a preset loss function to obtain a loss value;

[0175] Determine whether the loss value converges. If the loss value does not converge, adjust the network parameters of the initial neural network model according to the loss value, and return to execute the step of selecting a training song set and a first song set that meets the similarity condition with the training song set in the training set; if the loss value converges, complete the training of the initial neural network model to obtain a pre-trained neural network model, and execute subsequent steps;

[0176] Evaluate the performance of the pre-trained neural network model using the validation set. If the performance evaluation result meets the preset requirements, allow the application of the pre-trained neural network model. If the performance evaluation result does not meet the preset requirements, change the training strategy and continue to train the neural network model to obtain a pre-trained neural network model that allows application

[0177] After obtaining the pre-trained neural network model, determine the completion rate threshold TH for distinguishing songs of different qualities according to the predicted completion rate of the validation set and the known completion rate of the validation set by the pre-trained neural network model;

[0178] According to Figure 7 Deploy the pre-trained neural network model as shown;

[0179] Obtain the target song to be mined;

[0180] Obtain historical songs that meet the similarity condition with the target song;

[0181] Determine the CNN features of the historical song as the first target input information, and determine the target song as the second target input information;

[0182] Transmit the first target input information and the second target input information to the pre-trained neural network model, and obtain the target completion rate of the target song output by the pre-trained neural network model;

[0183] Based on the target completion rate and the completion rate threshold, determine the mining result after mining the quality of the target song;

[0184] Determine the promotion method and corresponding signing information of the target song according to the mining result.

[0185] SeeFigure 8 As shown in the figure, an embodiment of the present application also correspondingly discloses a song mining device, which is applied to a background server and includes:

[0186] A target song acquisition module 11, configured to acquire a target song to be mined;

[0187] A historical song acquisition module 12, configured to acquire historical songs that meet similar conditions with the target song;

[0188] A target input information acquisition module 13, configured to determine first target input information based on the historical songs and determine second target input information based on the target song;

[0189] A target interaction information acquisition module 14, configured to transmit the first target input information and the second target input information to a pre-trained neural network model, and acquire target interaction information of the target song output by the pre-trained neural network model;

[0190] A mining result determination module 15, configured to determine a mining result of the target song based on the target interaction information;

[0191] Wherein, the target interaction information is used to represent the result of the interaction between the user and the target song.

[0192] It can be seen that in the present application, first, the target song to be mined is acquired; historical songs that meet similar conditions with the target song are acquired; then, the first target input information is determined based on the historical songs, and the second target input information is determined based on the target song; and the first target input information and the second target input information are transmitted to a pre-trained neural network model to acquire the target interaction information of the target song output by the pre-trained neural network model; finally, the mining result of the target song is determined based on the target interaction information. Since the target interaction information is the information of the interaction between the user and the target song predicted by the pre-trained neural network model, the present application realizes predicting the target interaction information according to the target song and the historical songs similar to the target song. In this process, there is no need to manually define song value indicators, and since the mining result is adapted to the target interaction information, and the target interaction information reflects the user's demand for the target song, the mining result is consistent with the user's actual demand for the song, and the song mining performance is good.

[0193] A CNN feature determination module, configured to perform feature extraction on the first target input information based on a learnable CNN network structure to obtain a first target CNN feature; perform feature extraction on the second target input information based on a learnable CNN network structure to obtain a second target CNN feature;

[0194] A target interaction information determination module, configured to calculate the similarity between the first target CNN feature and the second target CNN feature, and determine the target interaction information based on the similarity.

[0195] In some specific embodiments, the target interaction information determination module may specifically be configured to: calculate the dot product value of the first target CNN feature and the second target CNN feature; concatenate the dot product value and the first target CNN feature into a long feature; based on the long feature, calculate the similarity between the target song and each song in the historical songs through a fully connected layer; use the similarity as the feature weight value of the corresponding song in the historical songs, and perform weighted summation on the first target CNN feature based on the feature weight value to obtain the comprehensive feature of the historical songs; concatenate the comprehensive feature and the second target CNN feature together to obtain a concatenated feature; classify the concatenated feature through a fully connected layer to obtain the target interaction information of the target song.

[0196] See Figure 9 As shown, an embodiment of the present application also correspondingly discloses a training device for a neural network model, which is applied to a background server and includes:

[0197] A sample song acquisition module 111, configured to acquire sample songs with known interaction information;

[0198] A training set division module 112, configured to divide a training set from the sample songs;

[0199] A first song set selection module 113, configured to select a training song set and a first song set that satisfies a similarity condition with the training song set from the training set;

[0200] A training input information determination module 114, configured to determine first training input information based on the first song set and second training input information based on the training song set;

[0201] A loss value acquisition module 115, configured to use the first training input information and the second training input information as inputs to an initial neural network model, train the initial neural network model, and input the known interaction information of the training song set and the predicted interaction information of the training song set output by the initial neural network model into a preset loss function to obtain a loss value;

[0202] An adjustment module 116, configured to determine whether the loss value converges. If the loss value does not converge, adjust the network parameters of the initial neural network model according to the loss value, and prompt the first song set selection module to execute the step of selecting a training song set and a first song set that satisfies a similarity condition with the training song set from the training set; if the loss value converges, complete the training of the initial neural network model to obtain a pre-trained neural network model for song mining based on the pre-trained neural network model.

[0203] In some specific embodiments, the training device of the neural network model may further include:

[0204] A validation set partitioning module, configured to partition a validation set from the sample songs after the adjustment module obtains a pre-trained neural network model;

[0205] A performance evaluation module, configured to perform a performance evaluation on the pre-trained neural network model based on the validation set to obtain a performance evaluation result;

[0206] A judgment module, configured to judge whether the performance evaluation result meets a preset requirement; if the performance evaluation result meets the preset requirement, the pre-trained neural network model is allowed to be applied; if the performance evaluation result does not meet the preset requirement, the training strategy is changed and the pre-trained neural network model is continued to be trained.

[0207] In some specific embodiments, the performance evaluation module may include:

[0208] A song set selection unit, configured to select a validation song set and a second song set that meets the similarity condition with the validation song set from the validation set;

[0209] A validation input information determination unit, configured to determine first validation input information based on the second song set and determine second validation input information based on the validation song set;

[0210] A predicted interaction information acquisition unit, configured to use the first validation input information and the second validation input information as inputs of the pre-trained neural network model to obtain predicted interaction information output by the pre-trained neural network model;

[0211] A performance evaluation result determination unit, configured to determine the performance evaluation result based on the predicted interaction information and the known interaction information of the validation song set.

[0212] In some specific embodiments, the training input information determination module may include:

[0213] A first training input information determination unit, configured to determine the target Mel spectrogram features of the first song set and determine the target Mel spectrogram features of the first song set as the first training input information;

[0214] A second training input information determination unit, configured to determine the target Mel spectrogram features of the training set and determine the target Mel spectrogram features of the training set as the second training input information.

[0215] In some specific embodiments, the training device of the neural network model may include a target Mel spectrogram determination module, which is configured to perform short-time Fourier transform on the audio of the song to obtain a short-time Fourier transform result; perform Mel spectrogram coefficient conversion on the short-time Fourier transform result to obtain an initial Mel spectrogram feature; and truncate the initial Mel spectrogram feature to obtain the target Mel spectrogram feature of the song.

[0216] Furthermore, an embodiment of the present application also provides an electronic device. Figure 10 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment, and the content in the figure should not be considered as any limitation to the scope of use of the present application.

[0217] Figure 10 It is a schematic structural diagram of an electronic device 20 provided by an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the song mining method or the neural network model training method disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be a server.

[0218] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed on it here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitation is made here.

[0219] In addition, as a carrier for resource storage, the memory 22 may be a read-only memory, a random access memory, a magnetic disk, or an optical disc, etc., and the resources stored thereon may include an operating system 221, a computer program 222, and video data 223, etc., and the storage method may be short-term storage or permanent storage.

[0220] Among them, the operating system 221 is used to manage and control each hardware device on the electronic device 20 and the computer program 222, so as to implement the operation and processing of the massive video data 223 in the memory 22 by the processor 21. It can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the song mining method or neural network model training method executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 may further include a computer program that can be used to complete other specific tasks. The data 223 may include various video data collected by the electronic device 20.

[0221] Further, an embodiment of the present application also discloses a storage medium in which a computer program is stored. When the computer program is loaded and executed by a processor, the steps of the song mining method or neural network model training method disclosed in any of the foregoing embodiments are implemented.

[0222] The computer-readable storage medium involved in the present application includes random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium well-known in the technical field.

[0223] For the description of the relevant parts in the song mining device, electronic device and computer-readable storage medium provided in the embodiments of the present application, please refer to the corresponding detailed description in the song mining method provided in the embodiments of the present application, which will not be elaborated here. In addition, the parts of the above technical solutions provided in the embodiments of the present application that are consistent with the corresponding technical solutions in the prior art in terms of implementation principles are not described in detail to avoid excessive elaboration.

[0224] It should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including the element.

[0225] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A song mining method, characterized in that, it includes: Obtain a target song to be mined, where the target song is a song with unknown interaction information; Obtain historical songs that meet the similarity conditions with the target song, and the historical songs are songs with known interaction information; Determine the first target input information based on the historical songs, and determine the second target input information based on the target song; Transmit the first target input information and the second target input information to a pre-trained neural network model, and obtain the target interaction information of the target song output by the pre-trained neural network model, and the interaction information of the historical songs is used to constrain the target interaction information of the target song; Determine the mining result of the target song based on the target interaction information; wherein, the target interaction information is used to characterize the result of the user's interaction with the target song; wherein, the pre-trained neural network model outputs the target interaction information based on the first target input information and the second target input information, including: Extract features from the first target input information based on a learnable CNN network structure to obtain first target CNN features; Extract features from the second target input information based on the learnable CNN network structure to obtain second target CNN features; Calculate the similarity between the first target CNN feature and the second target CNN feature, and determine the target interaction information based on the similarity.

2. The method according to claim 1, characterized in that, The calculating the similarity between the first target CNN feature and the second target CNN feature, and determining the target interaction information based on the similarity includes: Calculate the dot product value of the first target CNN feature and the second target CNN feature; Concatenate the dot product value and the first target CNN feature into a long feature; Based on the long feature, calculate the similarity between the target song and each song in the historical songs through a fully connected layer; Use the similarity as the feature weight value of the corresponding song in the historical songs, and perform weighted summation on the first target CNN feature based on the feature weight value to obtain the comprehensive feature of the historical songs; Concatenate the comprehensive feature and the second target CNN feature together to obtain a concatenated feature; Classify the concatenated feature through a fully connected layer to obtain the target interaction information of the target song.

3. A song mining device, characterized in that, it includes: A target song acquisition module, configured to acquire a target song to be mined, where the target song is a song with unknown interaction information; A historical song acquisition module, configured to acquire historical songs that meet the similarity conditions with the target song, and the historical songs are songs with known interaction information; A target input information acquisition module, configured to determine the first target input information based on the historical songs, and determine the second target input information based on the target song; A target interaction information acquisition module, configured to transmit the first target input information and the second target input information to a pre-trained neural network model, and acquire target interaction information of the target song output by the pre-trained neural network model, and the interaction information of the historical song is used to constrain the target interaction information of the target song; A mining result determination module, configured to determine a mining result of the target song based on the target interaction information; wherein, the target interaction information is used to characterize the result of the user's interaction with the target song; wherein, the pre-trained neural network model includes: A CNN feature determination module, configured to perform feature extraction on the first target input information based on a learnable CNN network structure to obtain a first target CNN feature; perform feature extraction on the second target input information based on the learnable CNN network structure to obtain a second target CNN feature; A target interaction information determination module, configured to calculate a similarity between the first target CNN feature and the second target CNN feature, and determine the target interaction information based on the similarity.

4. A training method for a neural network model, characterized in that, comprising: Obtain sample songs with known interaction information; Divide a training set from the sample songs; In the training set, select a training song set and a first song set that satisfies a similarity condition with the training song set; Determine first training input information based on the first song set, and determine second training input information based on the training song set; Use the first training input information and the second training input information as inputs of an initial neural network model, train the initial neural network model, and input the known interaction information of the training song set and the predicted interaction information of the training song set output by the initial neural network model into a preset loss function to obtain a loss value; Determine whether the loss value converges. If the loss value does not converge, adjust the network parameters of the initial neural network model according to the loss value, and return to execute the step of selecting a training song set and a first song set that satisfies a similarity condition with the training song set in the training set; if the loss value converges, complete the training of the initial neural network model to obtain a pre-trained neural network model; wherein, the pre-trained neural network model outputs target interaction information based on first target input information and second target input information, including: Perform feature extraction on the first target input information based on a learnable CNN network structure to obtain a first target CNN feature; the first target input information is determined based on a historical song; Perform feature extraction on the second target input information based on the learnable CNN network structure to obtain a second target CNN feature; the second target input information is determined based on a target song; Calculate a similarity between the first target CNN feature and the second target CNN feature, and determine the target interaction information based on the similarity.

5. The method according to claim 4, characterized in that, After obtaining the pre-trained neural network model, the following steps are further included: Dividing a validation set from the sample songs; Evaluating the performance of the pre-trained neural network model based on the validation set to obtain a performance evaluation result; Judging whether the performance evaluation result meets a preset requirement; If the performance evaluation result meets the preset requirement, allowing the application of the pre-trained neural network model; If the performance evaluation result does not meet the preset requirement, changing the training strategy and continuing to train the pre-trained neural network model.

6. The method according to claim 5, wherein, The step of evaluating the performance of the pre-trained neural network model based on the validation set to obtain a performance evaluation result includes: In the validation set, selecting a validation song set and a second song set that satisfies the similarity condition with the validation song set; Determining first validation input information based on the second song set and second validation input information based on the validation song set; Using the first validation input information and the second validation input information as the input of the pre-trained neural network model, and obtaining predicted interaction information output by the pre-trained neural network model; Determining the performance evaluation result based on the predicted interaction information and the known interaction information of the validation song set.

7. The method according to claim 4, wherein, The step of determining first training input information based on the first song set and second training input information based on the training set includes: Determining the target Mel spectrogram feature of the first song set and using the target Mel spectrogram feature of the first song set as the first training input information; Determining the target Mel spectrogram feature of the training set and using the target Mel spectrogram feature of the training set as the second training input information.

8. The method according to claim 7, wherein, The process of determining the target Mel spectrogram feature of the song includes: Performing short-time Fourier transform on the audio of the song to obtain a short-time Fourier transform result; Performing Mel spectrogram coefficient conversion on the short-time Fourier transform result to obtain an initial Mel spectrogram feature; Truncating the initial Mel spectrogram feature to obtain the target Mel spectrogram feature of the song.

9. A training device for a neural network model, wherein, It includes: A sample song acquisition module for acquiring sample songs with known interaction information; A training set division module for dividing a training set from the sample songs; A first song set selection module for selecting a training song set and a first song set that satisfies the similarity condition with the training song set in the training set; A training input information determination module for determining first training input information based on the first song set and second training input information based on the training song set; A loss value acquisition module, configured to use the first training input information and the second training input information as inputs to an initial neural network model, train the initial neural network model, and input the known interaction information of the training song set and the predicted interaction information of the training song set output by the initial neural network model into a preset loss function to obtain a loss value; An adjustment module, configured to determine whether the loss value converges. If the loss value does not converge, adjust the network parameters of the initial neural network model according to the loss value, and prompt the first song set selection module to execute the step of selecting a training song set and a first song set that satisfies a similarity condition with the training song set in the training set; if the loss value converges, complete the training of the initial neural network model to obtain a pre-trained neural network model for song mining based on the pre-trained neural network model; Wherein, the pre-trained neural network model outputs target interaction information based on first target input information and second target input information, including: Performing feature extraction on the first target input information based on a learnable CNN network structure to obtain first target CNN features; the first target input information is determined based on historical songs; Performing feature extraction on the second target input information based on the learnable CNN network structure to obtain second target CNN features; the second target input information is determined based on target songs; Calculating the similarity between the first target CNN features and the second target CNN features, and determining the target interaction information based on the similarity; the target interaction information is used to characterize the result of the user's interaction with the target song.

Citation Information

Patent Citations

  • Convolution neural network-based music recommending system and method

    CN108595550A

  • Song list recommendation method and device and storage medium

    CN108984731A