Song recognition method, computer device and storage medium

The method uses a trained audio separation model to enhance song recognition by separating vocals and instrumental tracks, addressing the challenge of identifying new recordings and cover songs, and improving recognition accuracy.

CN116189707BActive Publication Date: 2025-07-15TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310120143.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-02
Publication Date
2025-07-15
Estimated Expiration
2043-02-02

AI Technical Summary

Technical Problem

It is difficult to accurately identify new covers or new original music works in the prior art, and the traditional song fingerprint matching method has poor effect when facing cover adapted songs.

Method used

The trained audio separation model separates the person's voice and accompaniment audio of the song to be identified, and combines the basic frequency sequence and audio feature similarity to filter out songs that meet the preset similarity conditions from the song library, and uses an identification method that combines the basic frequency similarity and feature similarity.

Benefits of technology

It improves the accuracy of song recognition, can effectively deal with noise interference, accurately recognize vocals and accompaniment, adapt to covers and accompaniment adaptations, and improves the recognition effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116189707B_ABST
    Figure CN116189707B_ABST
Patent Text Reader

Abstract

The present application relates to a song recognition method, a computer device, and a storage medium. The method includes: inputting a song to be recognized into a trained audio separation model to obtain a human voice audio and an accompaniment audio corresponding to the song to be recognized; the trained audio separation model is trained by using human voice samples, accompaniment samples, and mixed music samples; the mixed music samples contain human voice samples and accompaniment samples; obtaining a fundamental frequency sequence of the human voice audio, and respectively determining the fundamental frequency similarity between the fundamental frequency sequence and the fundamental frequency sequences of each song in the song library; obtaining the audio features of the accompaniment audio, and respectively determining the feature similarity between the audio features and the audio features of each song in the song library; screening out songs that meet the preset similarity condition from the song library according to the fundamental frequency similarity and the feature similarity, and using them as the song recognition result corresponding to the song to be recognized. Using this method can improve the song recognition effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of audio technology, and in particular, to a song recognition method, a computer device, a storage medium, and a computer program product. Background Art

[0002] Song recognition by listening is an important technology in the music field. However, many song recognition technologies are relatively accurate in recognizing the original works, even with a certain amount of background noise. It is difficult to accurately recognize newly covered or newly original music works.

[0003] In traditional technologies, the correspondence of songs is often recognized through song fingerprints. However, the fingerprints of new music works are not stored in the database in time, and there are often certain differences between the covered and adapted songs and the original songs. As a result, the recognition effect of songs by using the traditional method in song fingerprint matching is poor. Summary of the Invention

[0004] Based on this, it is necessary to provide a song recognition method, a computer device, a computer-readable storage medium, and a computer program product that can improve the recognition effect of songs for the above technical problems.

[0005] In a first aspect, the present application provides a song recognition method. The method includes:

[0006] Input the song to be recognized into the trained audio separation model to obtain the corresponding human voice audio and accompaniment audio of the song to be recognized; the trained audio separation model is trained by using human voice samples, accompaniment samples, and mixed music samples; the mixed music samples contain the human voice samples and the accompaniment samples;

[0007] Obtain the fundamental frequency sequence of the human voice audio, and respectively determine the fundamental frequency similarity between the fundamental frequency sequence and the fundamental frequency sequences of each song in the song library;

[0008] Obtain the audio features of the accompaniment audio, and respectively determine the feature similarity between the audio features and the audio features of each song in the song library;

[0009] According to the fundamental frequency similarity and the feature similarity, screen out the songs that meet the preset similarity conditions from the song library as the song recognition result corresponding to the song to be recognized.

[0010] In one of the embodiments, the trained audio separation model is trained in the following manner:

[0011] Perform Fourier transforms on the vocal sample, the accompaniment sample, and the mixed music sample respectively to obtain the vocal sample audio of the vocal sample, the accompaniment sample audio of the accompaniment sample, and the mixed music sample audio of the mixed music sample;

[0012] Input the mixed music sample audio into the vocal separation model and the accompaniment separation model in the audio separation model to be trained respectively, and obtain the separated vocal audio and the separated accompaniment audio corresponding to the mixed music sample audio;

[0013] According to the distance between the separated vocal audio and the vocal sample audio, and the distance between the separated accompaniment audio and the accompaniment sample audio, obtain the audio separation loss function of the audio separation model to be trained;

[0014] According to the audio separation loss function, perform iterative training on the audio separation model to be trained to obtain the trained audio separation model.

[0015] In one embodiment, obtaining the audio features of the accompaniment audio and respectively determining the feature similarity between the audio features and the audio features of each song in the song library includes:

[0016] Input the accompaniment audio into the trained audio feature extraction model to obtain the audio features of the accompaniment audio; the trained audio feature extraction model is trained by the original song, the positive cover song, and the negative cover song; the positive cover song is a cover song of the original song; the negative cover song is a cover song of a song other than the original song;

[0017] Respectively perform similarity conversion on the distances between the audio features of the accompaniment audio and the audio features of each song in the song library to obtain the feature similarity between the audio features of the accompaniment audio and the audio features of each song in the song library.

[0018] In one embodiment, the trained audio feature extraction model is trained in the following manner:

[0019] Input the original song, the positive cover song, and the negative cover song into the audio feature extraction model to be trained to obtain the audio features of the original song, the audio features of the positive cover song, and the audio features of the negative cover song;

[0020] According to the distance between the audio features of the original song and the audio features of the positive cover song, and the distance between the audio features of the original song and the audio features of the negative cover song, obtain the feature extraction loss function of the audio feature extraction model to be trained;

[0021] Iteratively train the audio feature extraction model to be trained according to the feature extraction loss function to obtain the trained audio feature extraction model.

[0022] In one embodiment, obtaining the fundamental frequency sequence of the human voice audio includes:

[0023] Determine the pitch period of the human voice audio;

[0024] Perform autocorrelation processing on the human voice audio according to the pitch period to obtain the fundamental frequency sequence of the human voice audio.

[0025] In one embodiment, determining the fundamental frequency similarity between the fundamental frequency sequence and the fundamental frequency sequences of each song in the song library respectively includes:

[0026] Perform dynamic time warping processing on the fundamental frequency sequence of the human voice audio and the fundamental frequency sequences of each song in the song library to obtain the sequence distance between the fundamental frequency sequence of the human voice audio and the fundamental frequency sequences of each song in the song library;

[0027] Perform similarity conversion on each of the sequence distances to obtain the fundamental frequency similarity.

[0028] In one embodiment, according to the fundamental frequency similarity and the feature similarity, screening out the songs that meet the preset similarity condition from the song library as the song recognition result corresponding to the song to be recognized includes:

[0029] Obtain the first importance parameter corresponding to the fundamental frequency similarity and the second importance parameter corresponding to the feature similarity;

[0030] Perform fusion processing on each of the fundamental frequency similarities and the first importance parameter, and each of the feature similarities and the second importance parameter to obtain the target similarity between the song to be recognized and each song in the song library;

[0031] According to the target similarity of each song in the song library, screen out the song with the highest target similarity from the song library as the song recognition result corresponding to the song to be recognized.

[0032] In a second aspect, the present application also provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0033] Input the song to be recognized into the trained audio separation model to obtain the corresponding vocal audio and accompaniment audio of the song to be recognized; the trained audio separation model is trained by vocal samples, accompaniment samples and mixed music samples; the mixed music samples contain the vocal samples and the accompaniment samples;

[0034] Obtain the fundamental frequency sequence of the vocal audio, and respectively determine the fundamental frequency similarity between the fundamental frequency sequence and the fundamental frequency sequences of each song in the song library;

[0035] Obtain the audio features of the accompaniment audio, and respectively determine the feature similarity between the audio features and the audio features of each song in the song library;

[0036] According to the fundamental frequency similarity and the feature similarity, screen out the songs that meet the preset similarity conditions from the song library as the song recognition result corresponding to the song to be recognized.

[0037] In a third aspect, the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0038] Input the song to be recognized into the trained audio separation model to obtain the corresponding vocal audio and accompaniment audio of the song to be recognized; the trained audio separation model is trained by vocal samples, accompaniment samples and mixed music samples; the mixed music samples contain the vocal samples and the accompaniment samples;

[0039] Obtain the fundamental frequency sequence of the vocal audio, and respectively determine the fundamental frequency similarity between the fundamental frequency sequence and the fundamental frequency sequences of each song in the song library;

[0040] Obtain the audio features of the accompaniment audio, and respectively determine the feature similarity between the audio features and the audio features of each song in the song library;

[0041] According to the fundamental frequency similarity and the feature similarity, screen out the songs that meet the preset similarity conditions from the song library as the song recognition result corresponding to the song to be recognized.

[0042] In a fourth aspect, the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0043] Input the song to be recognized into the trained audio separation model to obtain the corresponding vocal audio and accompaniment audio of the song to be recognized; the trained audio separation model is trained with vocal samples, accompaniment samples, and mixed music samples; the mixed music samples contain the vocal samples and the accompaniment samples;

[0044] Obtain the fundamental frequency sequence of the vocal audio, and respectively determine the fundamental frequency similarity between the fundamental frequency sequence and the fundamental frequency sequences of each song in the song library;

[0045] Obtain the audio features of the accompaniment audio, and respectively determine the feature similarity between the audio features and the audio features of each song in the song library;

[0046] According to the fundamental frequency similarity and the feature similarity, screen out the songs that meet the preset similarity conditions from the song library as the song recognition result corresponding to the song to be recognized.

[0047] The above song recognition method, computer device, storage medium, and computer program product input the song to be recognized into the trained audio separation model to obtain the corresponding vocal audio and accompaniment audio of the song to be recognized, realizing the separation of the human voice and accompaniment in the song to be recognized. At the same time, the noise other than the human voice and accompaniment in the song to be recognized is also separated, avoiding the interference of noise on the recognition process, thereby improving the recognition accuracy of the song. Among them, the trained audio separation model is trained with vocal samples, accompaniment samples, and mixed music samples; the mixed music samples contain vocal samples and accompaniment samples; furthermore, obtain the fundamental frequency sequence of the vocal audio, and respectively determine the fundamental frequency similarity between the fundamental frequency sequence and the fundamental frequency sequences of each song in the song library; obtain the audio features of the accompaniment audio, and respectively determine the feature similarity between the audio features and the audio features of each song in the song library, so that both the fundamental frequency similarity of the vocal audio and the feature similarity of the accompaniment audio are effectively processed; according to the fundamental frequency similarity and the feature similarity, screen out the songs that meet the preset similarity conditions from the song library as the song recognition result corresponding to the song to be recognized, enabling the song recognition result to combine the factors of fundamental frequency similarity and feature similarity for song recognition, and being able to avoid the influence of the differences in a single aspect of the song on the recognition effect, thereby further improving the recognition effect of the song. Description of the Drawings

[0048] Figure 1 It is a schematic flowchart of the song recognition method in an embodiment;

[0049] Figure 2 It is a schematic flowchart of the training method of the trained audio separation model in an embodiment;

[0050] Figure 3Schematic diagram of the principle of the training method of the trained audio separation model in an embodiment;

[0051] Figure 4 Schematic diagram of the principle of the training process of the audio feature extraction model in an embodiment;

[0052] Figure 5 Schematic diagram of the principle of iteratively training the audio feature extraction model to be trained according to the feature extraction loss function in an embodiment;

[0053] Figure 6 Schematic diagram of the principle of performing dynamic time warping processing on the fundamental frequency sequence of the human voice audio and the fundamental frequency sequences of each song in the song library in an embodiment;

[0054] Figure 7 Schematic diagram of the flow of the song recognition method in another embodiment;

[0055] Figure 8 Schematic diagram of the flow of the song recognition method in yet another embodiment;

[0056] Figure 9 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0057] In order to make the objectives, technical solutions, and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0058] In one embodiment, as Figure 1 shown, a song recognition method is provided. In this embodiment, it is exemplified that the method is applied to a terminal. It can be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. Among them, the terminal can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. In this embodiment, the method includes the following steps:

[0059] Step S101, input the song to be recognized into the trained audio separation model to obtain the human voice audio and the accompaniment audio corresponding to the song to be recognized; the trained audio separation model is trained by using human voice samples, accompaniment samples, and mixed music samples; the mixed music samples contain human voice samples and accompaniment samples.

[0060] Among them, the song to be recognized refers to the song for which the corresponding original singer's song needs to be recognized; the song to be recognized includes the vocals of the singer and the accompaniment of the song. The vocal audio refers to the audio data of the vocals separated from the song. The accompaniment audio refers to the audio data of the accompaniment separated from the song. The audio data can be a spectrogram in the audio field.

[0061] Among them, the audio separation model refers to a model used to separate the vocal audio and the accompaniment audio of a song. The audio separation model can be a model composed of a vocal separation model and an accompaniment separation model, and of course, the audio separation model can also be changed accordingly according to the change of the song recognition requirements. Specifically, the terminal obtains the song to be recognized that needs to be recognized; then performs a Fourier transform on the song to be recognized to obtain the audio to be recognized corresponding to the song to be recognized; and then inputs the audio to be recognized into the vocal separation model and the accompaniment separation model in the trained audio separation model respectively to obtain the vocal audio corresponding to the audio to be recognized output by the vocal separation model, and the accompaniment audio corresponding to the audio to be recognized output by the accompaniment separation model. It can be understood that in the case where the terminal directly obtains the audio to be recognized that needs to be recognized, there is no need to perform a Fourier transform on the audio to be recognized, and the audio to be recognized can be directly input into the vocal separation model and the accompaniment separation model in the trained audio separation model.

[0062] Furthermore, the terminal can also perform an inverse Fourier transform on the vocal audio and the accompaniment audio respectively, and then obtain the vocals corresponding to the vocal audio (i.e., the vocals of the singer of the song to be recognized), and the accompaniment corresponding to the accompaniment audio (i.e., the accompaniment of the song to be recognized). That is to say, the terminal can first convert the song to be recognized into the audio to be recognized in audio format, and then input the audio to be recognized into the trained audio separation model to separate the vocal audio and the accompaniment audio through the trained audio separation model, and then restore the vocal audio and the accompaniment audio to the vocals and accompaniment in song format through the inverse transform, thereby realizing the accurate separation of the vocals and the accompaniment from the song to be recognized.

[0063] Step S102, obtain the fundamental frequency sequence of the vocal audio, and respectively determine the fundamental frequency similarity between the fundamental frequency sequence and the fundamental frequency sequences of each song in the song library.

[0064] Among them, the fundamental frequency refers to the lowest frequency in the vocal audio, and the fundamental frequency is used to determine the pitch of the sound. The fundamental frequency sequence refers to the data composed of multiple fundamental frequencies; for example, first extract a fundamental frequency from each segment (or each frame) of the vocal audio, and then combine the fundamental frequencies of all segments (or all frames) of the vocal audio to obtain the fundamental frequency sequence.

[0065] Specifically, the terminal can obtain the fundamental frequency sequences of each song in the song library by performing fundamental frequency extraction processing on the vocal data of each song in the song library respectively. Then the terminal obtains the fundamental frequency sequences corresponding to the vocal audio of each song in the song library. Among them, the vocal audio of the songs in the song library can be the same as the vocal samples used to train the audio separation model, or can be richer than the vocal samples. The terminal performs fundamental frequency extraction processing on the vocal audio obtained in step S101 above to obtain the fundamental frequency sequence of the vocal audio. Then, the fundamental frequency sequence of the vocal audio is respectively subjected to sequence similarity processing with the fundamental frequency sequences of each song in the song library to obtain the fundamental frequency similarity between the fundamental frequency sequence and the fundamental frequency sequences of each song in the song library.

[0066] Step S103: Obtain the audio features of the accompaniment audio, and respectively determine the feature similarity between the audio features and the audio features of each song in the song library.

[0067] Specifically, the terminal can obtain the audio features of each song in the song library by performing feature extraction processing on the accompaniment data of each song in the song library respectively. Then the terminal obtains the audio features corresponding to the accompaniment data of each song in the song library. It can be understood that the accompaniment data in the song library can be the same as the accompaniment samples used to train the audio separation model, or can be richer than the accompaniment samples. The terminal performs feature extraction processing on the accompaniment audio obtained in step S101 above to obtain the audio features of the accompaniment audio. Then, the audio features of the accompaniment audio are respectively subjected to feature similarity processing with the audio features of each song in the song library to obtain the feature similarity between the audio features and the audio features of each song in the song library.

[0068] It should be noted that the above step S102 and the above step S103 can be executed in parallel, or step S102 can be executed first and then step S103. Of course, step S103 can also be executed first and then step S102. Here, the execution order of the above step S102 and the above step S103 is not limited.

[0069] Step S104: According to the fundamental frequency similarity and the feature similarity, screen out the songs in the song library that meet the preset similarity conditions as the song recognition result corresponding to the song to be recognized.

[0070] Among them, the preset similarity condition refers to the determination condition set for the fundamental frequency similarity and the feature similarity between the songs in the song library and the song to be recognized. The song recognition result refers to the original song corresponding to the song to be recognized obtained by recognition.

[0071] Specifically, after the terminal obtains the fundamental frequency similarity and feature similarity between the song to be recognized and each song in the song library, it can obtain a preset similarity condition, and then determine the target song in the song library that meets the preset similarity condition as the song recognition result corresponding to the song to be recognized.

[0072] In practical applications, the preset similarity condition can be set to the highest similarity of any one of the fundamental frequency similarity or the feature similarity unilaterally. The preset similarity condition can also be set to the highest target similarity obtained by fusing the fundamental frequency similarity or the feature similarity. The preset similarity condition can be set according to the actual situation.

[0073] In the above song recognition method, by inputting the song to be recognized into the trained audio separation model, the human voice audio and the accompaniment audio corresponding to the song to be recognized are obtained, realizing the separation of the human voice and the accompaniment in the song to be recognized. At the same time, the noise other than the human voice and the accompaniment in the song to be recognized is also separated, avoiding the interference of the noise on the recognition process, thereby improving the recognition accuracy of the song. Among them, the trained audio separation model is trained by using human voice samples, accompaniment samples and mixed music samples. The mixed music sample contains human voice samples and accompaniment samples. Furthermore, the fundamental frequency sequence of the human voice audio is obtained, and the fundamental frequency similarity between the fundamental frequency sequence and the fundamental frequency sequences of each song in the song library is determined respectively. The audio features of the accompaniment audio are obtained, and the feature similarity between the audio features and the audio features of each song in the song library is determined respectively, so that the fundamental frequency similarity of the human voice audio and the feature similarity of the accompaniment audio are both effectively processed. According to the fundamental frequency similarity and the feature similarity, the song that meets the preset similarity condition is selected from the song library as the song recognition result corresponding to the song to be recognized, so that the song recognition result can combine the two factors of the fundamental frequency similarity and the feature similarity for song recognition, and can avoid the influence of the differences in a single aspect of the song on the recognition effect, thereby further improving the recognition effect of the song.

[0074] In one embodiment, as Figure 2 shown, the trained audio separation model is trained in the following manner:

[0075] Step S201, perform Fourier transform on the human voice sample, the accompaniment sample and the mixed music sample respectively to obtain the human voice sample audio of the human voice sample, the accompaniment sample audio of the accompaniment sample, and the mixed music sample audio of the mixed music sample.

[0076] Among them, the human voice sample refers to the human voice audio data set used to train the audio separation model. The accompaniment sample refers to the accompaniment audio data set used to train the audio separation model. The mixed music sample refers to the song data set obtained by fusing the human voice audio and the accompaniment audio.

[0077] Specifically, the terminal obtains a mixed music sample, as well as the corresponding vocal sample and accompaniment sample of the mixed music sample; then, Fourier transforms are respectively performed on the mixed music sample, the corresponding vocal sample and accompaniment sample of the mixed music sample to obtain the spectrogram of the vocal sample, the spectrogram of the accompaniment sample, and the spectrogram of the mixed music sample. The terminal can use the spectrogram of the vocal sample as the vocal sample audio of the vocal sample, the spectrogram of the accompaniment sample as the accompaniment sample audio of the accompaniment sample, and the spectrogram of the mixed music sample as the mixed music sample audio of the mixed music sample.

[0078] Step S202: Input the mixed music sample audio into the vocal separation model and the accompaniment separation model in the audio separation model to be trained respectively, to obtain the separated vocal audio and the separated accompaniment audio corresponding to the mixed music sample audio.

[0079] Among them, the separated vocal audio refers to the audio data of the vocals separated from the mixed music sample audio, which is the same as the vocal audio obtained in the above step S101. The separated accompaniment audio refers to the audio data of the background accompaniment separated from the mixed music sample audio, which is the same as the accompaniment audio obtained in the above step S101.

[0080] Step S203: According to the distance between the separated vocal audio and the vocal sample audio, and the distance between the separated accompaniment audio and the accompaniment sample audio, obtain the audio separation loss function of the audio separation model to be trained.

[0081] Step S204: According to the audio separation loss function, perform iterative training on the audio separation model to be trained to obtain the trained audio separation model.

[0082] In practical applications, the audio separation model can be implemented through the Spleeter network, and the vocal separation model and the accompaniment separation model can be respectively implemented through networks with the unet structure.

[0083] Specifically, the terminal inputs the mixed music sample audio into the vocal separation model and the accompaniment separation model in the audio separation model to be trained respectively. The vocal separation model outputs the separated vocal audio corresponding to the mixed music sample audio, and the accompaniment separation model outputs the separated accompaniment audio corresponding to the mixed music sample audio. The terminal obtains the distance between the separated vocal audio and the vocal sample audio. It can perform Manhattan distance processing on the separated vocal audio and the vocal sample audio, so that the terminal obtains the Manhattan distance between the separated vocal audio and the vocal sample audio. The terminal obtains the distance between the separated accompaniment audio and the accompaniment sample audio. It can perform Manhattan distance processing on the separated accompaniment audio and the accompaniment sample audio to obtain the Manhattan distance between the separated accompaniment audio and the accompaniment sample audio. Furthermore, the terminal performs mean processing on the distance between the separated vocal audio and the vocal sample audio, and the distance between the separated accompaniment audio and the accompaniment sample audio to obtain the distance mean. According to the distance mean, an audio separation loss function of the audio separation model to be trained is constructed. Based on the audio separation loss function, the audio separation model to be trained is iteratively updated until the training end condition is reached. The audio separation model that reaches the training end condition is used as the trained audio separation model.

[0084] Figure 3 It is a schematic diagram of the principle of the training method of the trained audio separation model, as Figure 3 shown. In practical applications, the terminal can obtain the mixed music with accompaniment and vocals, as well as the corresponding pure accompaniment music and pure vocal music of the mixed music. Then the terminal performs Fourier transform on the mixed music, the corresponding pure accompaniment music and pure vocal music of the mixed music respectively to obtain the spectrogram corresponding to the mixed music, the spectrogram corresponding to the pure accompaniment music, and the spectrogram corresponding to the pure vocal music. The terminal inputs the spectrogram corresponding to the mixed music into the vocal separation model and the accompaniment separation model in the audio separation model to be trained respectively, and obtains the spectrogram of the vocals separated from the spectrogram of the mixed music and the spectrogram of the accompaniment separated from the spectrogram of the mixed music. The terminal determines the first Manhattan distance between the separated vocal spectrogram and the spectrogram corresponding to the pure vocal music, and determines the second Manhattan distance between the separated accompaniment spectrogram and the spectrogram corresponding to the pure accompaniment music. Mean processing is performed on the first Manhattan distance and the second Manhattan distance to obtain the distance mean. An audio separation loss function is constructed according to the distance mean. Based on the audio separation loss function, the audio separation model to be trained is iteratively updated to obtain the trained audio separation model.

[0085] In this embodiment, the mixed music sample audio of the mixed music sample is respectively input into the vocal separation model and the accompaniment separation model in the audio separation model to be trained, and the separated vocal audio and the separated accompaniment audio corresponding to the mixed music sample audio are obtained; furthermore, according to the distance between the separated vocal audio and the vocal sample audio, and the distance between the separated accompaniment audio and the accompaniment sample audio, an audio separation loss function is constructed to iteratively train the audio separation model to be trained, so that the separated vocal audio output by the audio separation model can continuously approach the vocal sample audio, and at the same time the separated accompaniment audio can continuously approach the accompaniment sample audio. Finally, a trained audio separation model is obtained. Since the trained audio separation model only needs to separate the vocal audio and the accompaniment audio from the input song to be recognized or the mixed music sample audio, the noise in the song to be recognized or the mixed music sample audio is also removed during the separation process, thus ensuring the reliability of the subsequent song recognition process and improving the recognition accuracy of the song.

[0086] In one embodiment, in the above step S103, the audio features of the accompaniment audio are obtained, and the feature similarity between the audio features and the audio features of each song in the song library is determined respectively, which specifically includes the following contents: The accompaniment audio is input into the trained audio feature extraction model to obtain the audio features of the accompaniment audio; The trained audio feature extraction model is trained by the original song, the positive cover song and the negative cover song; The positive cover song is a cover song of the original song; The negative cover song is a cover song of a song other than the original song; The distance between the audio features of the accompaniment audio and the audio features of each song is respectively subjected to similarity conversion to obtain the feature similarity between the audio features of the accompaniment audio and the audio features of each song.

[0087] Among them, the audio feature extraction model refers to a model used to extract audio features from audio data.

[0088] Specifically, after the terminal obtains the trained audio feature extraction model, it can also obtain the audio features of each song in the song library. Then, the accompaniment audio of each song in the song library is input into the trained audio feature extraction model to obtain the audio features corresponding to the accompaniment audio of each song in the song library (for easy distinction from the audio features of the song to be recognized, it can be called the song library audio features). Furthermore, the terminal obtains the accompaniment audio processed in the above step S101 and the trained audio feature extraction model; the terminal inputs the accompaniment audio into the trained audio feature extraction model to obtain the audio features of the accompaniment audio. The terminal performs feature similarity processing on the audio features of the accompaniment audio and each song library audio feature. It can be to perform cosine distance processing on the audio features of the accompaniment audio and each song library audio feature pairwise to obtain the cosine distance between the audio features of the accompaniment audio and each song library audio feature; perform similarity conversion on the cosine distance to obtain the cosine similarity between the audio features of the accompaniment audio and each song library audio feature, then the terminal obtains the feature similarity between the audio features of the accompaniment audio and each song library audio feature.

[0089] In this embodiment, by inputting the accompaniment audio into the trained audio feature extraction model, the audio features of the accompaniment audio are obtained; furthermore, similarity conversion is performed on the distances between the audio features of the accompaniment audio and the audio features of each song in the song library respectively, to obtain the feature similarity between the audio features of the accompaniment audio and the audio features of each song in the song library. By reasonably determining the feature similarity between the accompaniment audio and the accompaniment audio of each song in the song library through audio features, song recognition of the song to be recognized is also realized from the perspective of accompaniment through the feature similarity, improving the recognition accuracy of the song.

[0090] In one embodiment, the trained audio feature extraction model is trained in the following manner: The original songs, positive cover songs, and negative cover songs are input into the audio feature extraction model to be trained to obtain the audio features of the original songs, the audio features of the positive cover songs, and the audio features of the negative cover songs; according to the distance between the audio features of the original songs and the audio features of the positive cover songs, and the distance between the audio features of the original songs and the audio features of the negative cover songs, the feature extraction loss function of the audio feature extraction model to be trained is obtained; according to the feature extraction loss function, the audio feature extraction model to be trained is iteratively trained to obtain the trained audio feature extraction model.

[0091] Among them, a positive cover song refers to a song that is covered or covered after adaptation based on the accompaniment music, singing style, and rhythm of the original song. A negative cover song refers to a song that is covered or covered after adaptation based on a song other than the original song. For example, assume there are two original songs, A and B. Song a is covered based on A, and song b is covered based on B. Then the positive cover song of A is a, and the negative cover song of A is b; similarly, the positive cover song of B is b, and the negative cover song of B is a.

[0092] Among them, the feature extraction loss function refers to the loss function used to train the audio feature extraction model.

[0093] Specifically, the terminal obtains the original song, as well as the positive cover song of the original song and the negative cover song of the original song; takes the original song, the positive cover song, and the negative cover song as a group and inputs them into the audio feature extraction model to be trained, obtaining the audio features of the original song (for easy distinction from other audio features, it can be called the original audio features), the audio features of the positive cover song (for easy distinction from other audio features, it can be called the positive audio features), and the audio features of the negative cover song (for easy distinction from other audio features, it can be called the negative audio features).

[0094] Figure 4 It is a schematic diagram of the principle of the training process of the audio feature extraction model, as Figure 4 shown. The terminal can take the original song, the positive cover song, and the negative cover song as a group and input them into the audio feature extraction model to be trained. Among them, the audio feature extraction model can be implemented by a convolutional network based on the residual structure (Residual Network, ResNet), and the audio feature extraction model can also be implemented by a convolutional network based on the Incerion structure. Of course, the audio feature extraction model can also be implemented by a convolutional network based on the Visual Geometry Group (VGG) structure. Here, the network structure of the audio feature extraction model is not limited. The feature similarity refers to an index used to measure the similarity between the audio features of the accompaniment audio and the audio features of each song. The pooling layer in the audio feature extraction model to be trained outputs the original audio features, the positive audio features, and the negative audio features. Then, the terminal obtains the feature extraction loss function of the audio feature extraction model to be trained according to the distance between the audio features of the original song and the audio features of the positive cover song, and the distance between the audio features of the original song and the audio features of the negative cover song.

[0095] The terminal performs cosine distance processing on the original audio features and the positive audio features to obtain the first cosine distance between the original audio features and the positive audio features; performs cosine distance processing on the original audio features and the negative audio features to obtain the second cosine distance between the original audio features and the negative audio features; and then constructs a feature extraction loss function for the audio feature extraction model to be trained based on the first cosine distance and the second cosine distance. Among them, the feature extraction loss function L can be represented by the following formula:

[0096] L = max(d(a, p) - d(a, n) + margin, 0)

[0097] Among them, a represents the audio features of the original song; p represents the audio features of the positive cover song; n represents the audio features of the negative cover song; d() represents the cosine distance function, then d(a, p) represents the first cosine distance, and d(a, n) represents the second cosine distance; margin represents the function interval, which is used to control that the inter-class distance of the audio features of these three types of songs, namely the original song, the positive cover song, and the negative cover song, is greater than the threshold value.

[0098] Figure 5 FIG. is a schematic diagram of the principle of iteratively training the audio feature extraction model to be trained according to the feature extraction loss function. As Figure 5 shown, through the feature extraction loss function, the audio feature extraction model to be trained is iteratively trained, so that the distance between the original audio features output by the audio feature extraction model and the positive audio features becomes closer, and at the same time, the distance between the original audio features output by the audio feature extraction model and the negative audio features becomes more distant; when the training end condition is met, the terminal obtains the trained audio feature extraction model.

[0099] In this embodiment, various forms of positive cover songs of the original song and negative cover songs of the original song are used as the training data set of the audio feature extraction model to be trained, and the distance between the original audio features and the positive audio features and the distance between the original audio features and the negative audio features are used to train the audio feature extraction model to be trained, so that during the iterative training process of the audio feature extraction model to be trained, the original audio features output by it are closer to the positive audio features, and the original audio features output by it are also more distant from the negative audio features. Therefore, the trained audio feature extraction model can accurately identify the original song corresponding to the song. Even for the cover song adapted from the original song, the trained audio feature extraction model also has a high recognition accuracy.

[0100] In one embodiment, step S102 of obtaining the fundamental frequency sequence of the human voice audio specifically includes the following: determining the pitch period of the human voice audio; performing autocorrelation processing on the human voice audio according to the pitch period to obtain the fundamental frequency sequence of the human voice audio.

[0101] Among them, the pitch period refers to the smallest positive period of the signal waveform of the audio.

[0102] Specifically, the terminal divides the human voice audio to obtain multiple audio segments of the human voice audio; the terminal determines the fundamental frequency of each audio segment according to the pitch period of each audio segment. It can estimate the pitch period of each audio segment through the YIN algorithm to obtain the fundamental frequency of each audio segment; it can also estimate the pitch period of each audio segment through the autocorrelation function to obtain the fundamental frequency of each audio segment; and then combine the fundamental frequencies of all audio segments to obtain the fundamental frequency sequence of the human voice audio. In practical applications, the terminal can translate the original signal waveform of the audio segment by τ periods and determine the degree of coincidence between the translated signal waveform and the original signal waveform; according to the degree of coincidence between each translated signal waveform and the original signal waveform, select the target signal waveform with the highest degree of coincidence with the original waveform signal from all the translated signal waveforms, and use the translation period τ corresponding to the target signal waveform as the pitch period of this audio segment; according to the pitch period, determine the fundamental frequency of this audio segment; combine the fundamental frequencies of all audio segments into the fundamental frequency sequence of the human voice audio. Among them, the degree of coincidence between the translated signal waveform and the original signal waveform can be measured by the following formula:

[0103]

[0104] Among them, x i represents the value of the signal waveform x of the audio segment at time point i; τ represents the translation period, which is used to represent the period of the signal waveform x at time point t. When the coincidence degree r t (τ) takes the maximum value, the fundamental frequency at time point t is obtained as 1 / τ.

[0105] Furthermore, the terminal can also obtain the fundamental frequency sequences of each song in the song library, and respectively perform fundamental frequency extraction processing on the human voice audio of each song in the song library in the same way as obtaining the fundamental frequency sequence of the human voice audio in this embodiment, then the terminal obtains the fundamental frequency sequence corresponding to the human voice audio of each song in the song library.

[0106] In this embodiment, the fundamental period of the human voice audio is determined; then, based on the fundamental period, autocorrelation processing is performed on the human voice audio to obtain the fundamental frequency sequence of the human voice audio, realizing the reasonable acquisition of the fundamental frequency sequence corresponding to the human voice audio of the song to be recognized and the fundamental frequency sequence corresponding to the human voice audio of each song in the song library. Thus, the fundamental frequency sequence of each song in the song library can be used as a processing basis to perform the subsequent fundamental frequency similarity determination step.

[0107] In one embodiment, in step S102 above, the fundamental frequency similarities between the fundamental frequency sequence and the fundamental frequency sequences of each song in the song library are determined respectively, specifically including the following: performing dynamic time warping processing on the fundamental frequency sequence of the human voice audio and the fundamental frequency sequences of each song in the song library to obtain the sequence distance between the fundamental frequency sequence of the human voice audio and the fundamental frequency sequences of each song in the song library; performing similarity conversion on each sequence distance to obtain the fundamental frequency similarity.

[0108] Among them, the sequence distance refers to the shortest distance between the fundamental frequency sequence of the human voice audio and the fundamental frequency sequences of each song in the song library. The fundamental frequency similarity is an index used to measure the similarity between the fundamental frequency sequence of the human voice audio and the fundamental frequency sequences of each song in the song library.

[0109] Specifically, after the terminal obtains the fundamental frequency sequence of the human voice audio, it can obtain the fundamental frequency sequences of each song in the song library (for the convenience of distinguishing from the fundamental frequency sequence of the human voice audio, it can be called the song library fundamental frequency sequence); then the terminal performs pairwise dynamic time warping (DTW) processing on the fundamental frequency sequence and each song library fundamental frequency sequence to determine all the similarity points between the fundamental frequency sequence of the human voice audio and each song library fundamental frequency sequence, and obtain the sum of the distances between all the similarity points, which is used as the sequence distance between the fundamental frequency sequence of the human voice audio and each song library fundamental frequency sequence; among them, the distance between similarity points can be the Euclidean distance. The terminal inputs each sequence distance into the similarity conversion model in turn to obtain the fundamental frequency similarity corresponding to each sequence distance. Among them, the similarity conversion model can be represented by the following formula:

[0110]

[0111] Among them, S k represents the fundamental frequency similarity between the fundamental frequency sequence of the human voice audio and the fundamental frequency sequence of the kth song in the song library; d represents the sequence distance between the fundamental frequency sequence of the human voice audio and each song library fundamental frequency sequence.

[0112] Figure 6 is a schematic diagram of the principle of performing dynamic time warping processing on the fundamental frequency sequence of the human voice audio and the fundamental frequency sequences of each song in the song library. Figure 6 The solid line located above in represents the fundamental frequency sequence of the human voice audio. Figure 5The solid line located below in the figure represents the fundamental frequency sequence of a certain song in the music library. The similarity points between the fundamental frequency sequence and the fundamental frequency sequence of the music library are connected by a dotted line. The sum of the Euclidean distances between the similarity points of the fundamental frequency sequence and the similarity points of the fundamental frequency sequence of the music library is obtained through dynamic time warping, and the sequence distance between the fundamental frequency sequence and the fundamental frequency sequence of the music library is obtained. Then, the sequence distance is converted into a fundamental frequency similarity using a similarity conversion model, and the terminal obtains the fundamental frequency similarity between the fundamental frequency sequence and the fundamental frequency sequences of each song in the music library.

[0113] In this embodiment, by performing dynamic time warping processing on the fundamental frequency sequence of the human voice audio and the fundamental frequency sequences of each song in the music library, the sequence distance between the fundamental frequency sequence of the human voice audio and the fundamental frequency sequences of each song in the music library is obtained. Then, each sequence distance is converted into a fundamental frequency similarity. Through the fundamental frequency similarity between the human voice audio and the fundamental frequency sequences of each song in the music library, song recognition of the song to be recognized is realized from the perspective of the human voice, and the recognition accuracy of the song is improved.

[0114] In one embodiment, in step S104 above, according to the fundamental frequency similarity and the feature similarity, songs that meet the preset similarity condition are screened out from the music library as the song recognition result corresponding to the song to be recognized, which specifically includes the following content: obtaining a first importance parameter corresponding to the fundamental frequency similarity and a second importance parameter corresponding to the feature similarity; respectively performing fusion processing on each fundamental frequency similarity and the first importance parameter, and each feature similarity and the second importance parameter to obtain the target similarity between the song to be recognized and each song in the music library; according to the target similarity of each song in the music library, the song with the highest target similarity is screened out from the music library as the song recognition result corresponding to the song to be recognized.

[0115] Among them, the importance parameter refers to a parameter used to describe the importance of the similarity.

[0116] The terminal can combine the fundamental frequency similarity and the feature similarity to determine the song recognition result corresponding to the song to be recognized. Specifically, the terminal respectively obtains a first importance parameter corresponding to the fundamental frequency similarity and a second importance parameter corresponding to the feature similarity. The terminal fuses according to the first importance parameter and each fundamental frequency similarity, and the second importance parameter and each feature similarity to obtain the target similarity between the song to be recognized and each song in the music library. It can be to perform weighted processing on each fundamental frequency similarity and the first importance parameter, and each feature similarity and the second importance parameter in turn, then the terminal obtains the target similarity between the song to be recognized and each song in the music library. Furthermore, the terminal screens out the song with the highest target similarity from the music library as the song recognition result corresponding to the song to be recognized.

[0117] Illustrating by way of example, the target similarity can be calculated by the following formula:

[0118] num k = μS k +(1 - μ)δ k

[0119] Among them, num k represents the target similarity between the song to be recognized and the k-th song in the song library; S k represents the fundamental frequency similarity between the fundamental frequency sequence of the human voice audio and the fundamental frequency sequence of the k-th song in the song library; δ k represents the feature similarity between the audio features of the accompaniment audio and the audio features of the k-th song in the song library; μ represents the first importance parameter, and 1 - μ represents the second importance parameter.

[0120] The terminal can also determine the song recognition result corresponding to the song to be recognized according to any one of the fundamental frequency similarity and the feature similarity. For example, the terminal can screen out the song with the highest fundamental frequency similarity from the song library according to the fundamental frequency similarity of each song in the song library as the song recognition result corresponding to the song to be recognized. For another example, the terminal can also screen out the song with the highest feature similarity from the song library according to the feature similarity of each song in the song library as the song recognition result corresponding to the song to be recognized.

[0121] In this embodiment, by screening out the song recognition result corresponding to the song to be recognized from the song library according to the fundamental frequency similarity and the feature similarity between the song to be recognized and each song in the song library, the song recognition result can combine both the fundamental frequency similarity and the feature similarity for song recognition, and also has a high recognition accuracy when facing works with human voice humming and works with adapted accompaniment, thus greatly improving the song recognition effect.

[0122] In one embodiment, as Figure 7 shown, another song recognition method is provided. Taking this method applied to the terminal as an example for illustration, it includes the following steps:

[0123] Step S701, input the song to be recognized into the trained audio separation model to obtain the human voice audio and the accompaniment audio corresponding to the song to be recognized.

[0124] Among them, the trained audio separation model is trained by using human voice samples, accompaniment samples, and mixed music samples; the mixed music samples contain human voice samples and accompaniment samples.

[0125] Step S702, determine the pitch period of the human voice audio; according to the pitch period, perform autocorrelation processing on the human voice audio to obtain the fundamental frequency sequence of the human voice audio.

[0126] Step S703: Perform dynamic time warping processing on the fundamental frequency sequence of the human voice audio and the fundamental frequency sequences of each song in the song library to obtain the sequence distances between the fundamental frequency sequence of the human voice audio and the fundamental frequency sequences of each song in the song library; perform similarity conversion on each sequence distance to obtain the fundamental frequency similarity.

[0127] Step S704: Input the accompaniment audio into the trained audio feature extraction model to obtain the audio features of the accompaniment audio.

[0128] Among them, the trained audio feature extraction model is trained by the original songs, positive cover songs, and negative cover songs; the positive cover songs are the cover songs of the original songs; the negative cover songs are the cover songs of the songs other than the original songs.

[0129] Step S705: Perform similarity conversion on the distances between the audio features of the accompaniment audio and the audio features of each song in the song library respectively to obtain the feature similarities between the audio features of the accompaniment audio and the audio features of each song in the song library.

[0130] Step S706: Obtain the first importance parameter corresponding to the fundamental frequency similarity and the second importance parameter corresponding to the feature similarity.

[0131] Step S707: Perform fusion processing on each fundamental frequency similarity and the first importance parameter, and each feature similarity and the second importance parameter respectively to obtain the target similarities between the song to be recognized and each song in the song library.

[0132] Step S708: According to the target similarities of each song in the song library, screen out the song with the highest target similarity from the song library as the song recognition result corresponding to the song to be recognized.

[0133] The above song recognition method can achieve the following beneficial effects: By inputting the song to be recognized into the trained audio separation model, the corresponding human voice audio and accompaniment audio of the song to be recognized are obtained, realizing the separation of the human voice and accompaniment in the song to be recognized. At the same time, it can separate the noise in the song to be recognized other than the human voice and accompaniment, avoiding the interference of noise on the recognition process, thereby improving the recognition accuracy of the song. The trained audio separation model is obtained by training with human voice samples, accompaniment samples, and mixed music samples. The mixed music samples contain human voice samples and accompaniment samples. Furthermore, the fundamental frequency sequence of the human voice audio is obtained, and the fundamental frequency similarity between the fundamental frequency sequence and the fundamental frequency sequences of each song in the song library is determined respectively. The audio features of the accompaniment audio are obtained, and the feature similarity between the audio features and the audio features of each song in the song library is determined respectively, so that the fundamental frequency similarity of the human voice audio and the feature similarity of the accompaniment audio are both effectively processed. According to the fundamental frequency similarity and the feature similarity, the songs that meet the preset similarity conditions are selected from the song library as the song recognition results corresponding to the song to be recognized, enabling the song recognition results to combine the two factors of fundamental frequency similarity and feature similarity for song recognition, and being able to avoid the influence of the differences in a single aspect of the song on the recognition effect, thereby further improving the recognition effect of the song.

[0134] To more clearly illustrate the song recognition method provided by the embodiments of the present disclosure, the above song recognition method will be specifically described below with a specific embodiment. As Figure 8 shown, another song recognition method is provided, which can be applied to a terminal and specifically includes the following contents:

[0135] The terminal obtains the song to be recognized that needs to be recognized for the song, inputs the song to be recognized into the trained audio separation model, and obtains the human voice audio and the accompaniment audio corresponding to the song to be recognized. The terminal performs fundamental frequency extraction processing on the human voice audio to obtain the fundamental frequency sequence of the human voice audio, and performs dynamic time warping processing pairwise on the fundamental frequency sequence of the human voice audio and each fundamental frequency sequence in the human voice fundamental frequency library to obtain the fundamental frequency similarity between the fundamental frequency sequence of the human voice audio and each fundamental frequency sequence in the human voice fundamental frequency library; wherein, the human voice fundamental frequency library is constructed by the fundamental frequency sequences of each song in the song library. At the same time, the terminal inputs the accompaniment audio into the trained audio feature extraction model to obtain the audio features of the accompaniment audio, and performs cosine processing pairwise on the audio features of the accompaniment audio and each audio feature in the accompaniment feature library to obtain the cosine similarity between the audio features of the accompaniment audio and each audio feature in the accompaniment feature library; wherein, the accompaniment feature library is constructed by the audio features of each song in the song library. Furthermore, the terminal obtains the first importance parameter corresponding to the fundamental frequency similarity and the second importance parameter corresponding to the feature similarity; performs weighted processing on each fundamental frequency similarity and the first importance parameter, and each feature similarity and the second importance parameter to obtain the target similarity between the song to be recognized and each song in the song library; takes the song with the highest target similarity in the song library as the song recognition result corresponding to the song to be recognized.

[0136] In this embodiment, the human voice audio and the accompaniment audio in the song to be recognized are separated by the trained audio separation model, so as to process the human voice audio and the accompaniment audio in a split manner. And during the separation process, by separating the noise other than the human voice and the accompaniment in the song to be recognized, the interference of the noise on the recognition of the song to be recognized is avoided; in addition, various rich audio features such as the genre and melody of the accompaniment can be parsed through the trained audio feature extraction model, the fundamental frequency sequence of the singing voice of the singer can be parsed through the fundamental frequency extraction processing on the human voice audio, and the song is recognized by combining these two aspects of the audio features and the fundamental frequency sequence, so that this method can not only accurately recognize the songs with human voice and accompaniment, but also accurately recognize the cover works of human voice humming and the cover works after the accompaniment is adapted, thereby greatly improving the application range of the above song recognition method.

[0137] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0138] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as Figure 9 shown. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a song recognition method. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device may be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0139] Those skilled in the art can understand that Figure 9 the structure shown in

[0140] is only a block diagram of a part of the structure related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different component layout.

[0141] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0142] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0143] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0144] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., and are not limited thereto.

[0145] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0146] The above-described embodiments merely represent several implementation manners of the present application, and the description thereof is relatively specific and detailed. However, it should not be construed as a limitation to the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A song recognition method, characterized in that, The method includes: Input the song to be recognized into the trained audio separation model to obtain the corresponding human voice audio and accompaniment audio of the song to be recognized; the trained audio separation model is trained by using human voice samples, accompaniment samples and mixed music samples; the mixed music samples contain the human voice samples and the accompaniment samples; Obtain the fundamental frequency sequence of the human voice audio, and respectively determine the fundamental frequency similarity between the fundamental frequency sequence and the fundamental frequency sequences of each song in the song library; Obtain the audio features of the accompaniment audio, and respectively determine the feature similarity between the audio features and the audio features of each song in the song library; Perform fusion processing on each of the fundamental frequency similarities and the first importance parameters corresponding to the fundamental frequency similarities, and each of the feature similarities and the second importance parameters corresponding to the feature similarities, to obtain the target similarity between the song to be recognized and each song in the song library; the first importance parameter and the second importance parameter are parameters used to characterize the importance of describing the similarity; the second importance parameter is obtained according to the difference between the preset value and the first importance parameter; According to the target similarity of each song in the song library, screen out the song with the highest target similarity from the song library as the song recognition result corresponding to the song to be recognized.

2. The method according to claim 1, characterized in that, The trained audio separation model is trained in the following manner: Perform Fourier transform on the human voice sample, the accompaniment sample and the mixed music sample respectively to obtain the human voice sample audio of the human voice sample, the accompaniment sample audio of the accompaniment sample, and the mixed music sample audio of the mixed music sample; Input the mixed music sample audio into the human voice separation model and the accompaniment separation model in the audio separation model to be trained respectively, to obtain the separated human voice audio and the separated accompaniment audio corresponding to the mixed music sample audio; According to the distance between the separated human voice audio and the human voice sample audio, and the distance between the separated accompaniment audio and the accompaniment sample audio, obtain the audio separation loss function of the audio separation model to be trained; According to the audio separation loss function, perform iterative training on the audio separation model to be trained to obtain the trained audio separation model.

3. The method according to claim 1, characterized in that, The obtaining the audio features of the accompaniment audio, and respectively determining the feature similarity between the audio features and the audio features of each song in the song library includes: Input the accompaniment audio into the trained audio feature extraction model to obtain the audio features of the accompaniment audio; the trained audio feature extraction model is trained by using original singer songs, positive cover songs and negative cover songs; the positive cover songs are cover songs of the original singer songs; the negative cover songs are cover songs of songs other than the original singer songs; Perform similarity conversion on the distance between the audio features of the accompaniment audio and the audio features of each song in the song library respectively, to obtain the feature similarity between the audio features of the accompaniment audio and the audio features of each song in the song library.

4. The method according to claim 3, characterized in that The trained audio feature extraction model is obtained through the following method: Input the original song, the positive cover song, and the negative cover song into the audio feature extraction model to be trained, and obtain the audio features of the original song, the audio features of the positive cover song, and the audio features of the negative cover song; According to the distance between the audio features of the original song and the audio features of the positive cover song, and the distance between the audio features of the original song and the audio features of the negative cover song, obtain the feature extraction loss function of the audio feature extraction model to be trained; According to the feature extraction loss function, perform iterative training on the audio feature extraction model to be trained to obtain the trained audio feature extraction model.

5. The method according to claim 1, wherein The obtaining of the fundamental frequency sequence of the human voice audio includes: Determine the pitch period of the human voice audio; According to the pitch period, perform autocorrelation processing on the human voice audio to obtain the fundamental frequency sequence of the human voice audio.

6. The method according to claim 1, characterized in that The respectively determining the fundamental frequency similarity between the fundamental frequency sequence and the fundamental frequency sequences of each song in the song library includes: Perform dynamic time warping processing on the fundamental frequency sequence of the human voice audio and the fundamental frequency sequences of each song in the song library to obtain the sequence distance between the fundamental frequency sequence of the human voice audio and the fundamental frequency sequences of each song in the song library; Perform similarity conversion on each of the sequence distances to obtain the fundamental frequency similarity.

7. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Song recognition method and device

    CN112270929A