Method, server, program product and storage medium for song recognition

By extracting accompaniment and separating human voices from the audio to be identified, and obtaining and fusing matching degrees with different weights, the problem of human voice noise affecting song recognition in noisy environments is solved, thus improving the accuracy and reliability of song recognition.

CN119360889BActive Publication Date: 2025-11-04TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411475925.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-22
Publication Date
2025-11-04
Estimated Expiration
2044-10-22

AI Technical Summary

Technical Problem

In noisy environments, the audio captured by the microphone contains human voice noise, which affects the accuracy of song recognition and leads to a decrease in the accuracy of song identification.

Method used

By extracting the accompaniment and separating the human voice from the audio to be identified, the first and second accompaniment audios are obtained respectively, their features are extracted, and the matching degree with different weights is fused to identify the song.

Benefits of technology

It effectively reduces the impact of human voice noise on song recognition, improves the accuracy and reliability of song recognition, and avoids the risk of failure when only acquiring accompaniment audio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119360889B_ABST
    Figure CN119360889B_ABST
Patent Text Reader

Abstract

The application provides a song recognition method, server, computer program product and storage medium, and relates to the technical field of audio. The method comprises the following steps: extracting accompaniment from to-be-recognized audio to obtain first accompaniment audio of the to-be-recognized audio; separating vocals and accompaniment of the to-be-recognized audio to obtain second accompaniment audio of the to-be-recognized audio; extracting features from the first accompaniment audio to obtain first accompaniment audio features; extracting features from the second accompaniment audio to obtain second accompaniment audio features; matching the first accompaniment audio features with audio features of songs in a song library to obtain first matching degrees of the songs; matching the second accompaniment audio features with the audio features of the songs to obtain second matching degrees of the songs; and combining the first matching degrees and the second matching degrees to identify whether the songs are matching songs corresponding to the to-be-recognized audio. The application can effectively reduce the influence of vocal noise in a noisy speech environment on song recognition.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of audio technology, in particular to a song recognition method, a server, a computer program product and a storage medium. BACKGROUND

[0002] The song recognition function is a function possessed by most music applications. When a user hears a favorite song, the user can recognize the relevant information of the song through the song recognition function of the music application. At present, the song recognition function is usually implemented by the following method: audio is collected through a microphone, audio features are extracted, the audio features of songs in a song library are matched, and the relevant information of the song with the highest matching degree is returned. However, in a noisy environment, the microphone will collect ambient sound when collecting the audio of the song, which will affect the accuracy of song recognition. SUMMARY

[0003] The embodiments of the present application provide a song recognition method, a server, a program product and a storage medium, which can effectively reduce the influence of human voice noise in a noisy speaking environment on the accuracy of song recognition. The technical solution is as follows:

[0004] In a first aspect, a song recognition method is provided, and the method comprises:

[0005] extracting accompaniment from the to-be-recognized audio to obtain first accompaniment audio of the to-be-recognized audio; and separating vocals and accompaniment from the to-be-recognized audio to obtain second accompaniment audio of the to-be-recognized audio;

[0006] extracting features from the first accompaniment audio to obtain first accompaniment audio features; and extracting features from the second accompaniment audio to obtain second accompaniment audio features;

[0007] matching the first accompaniment audio features with audio features of a song in a song library to obtain a first matching degree of the song; and matching the second accompaniment audio features with the audio features of the song to obtain a second matching degree of the song;

[0008] combining the first matching degree and the second matching degree of the song to identify whether the song is a matching song corresponding to the to-be-recognized audio.

[0009] In a possible implementation, the combining the first matching degree and the second matching degree of the song to identify whether the song is a matching song corresponding to the to-be-recognized audio comprises:

[0010] fusing the first matching degree and the second matching degree of the song to obtain a fused matching degree of the song;

[0011] The song with the highest fusion matching degree in the song library is taken as the matching song of the audio to be identified.

[0012] In a possible implementation, the fusing the first matching degree and the second matching degree of the song to obtain the fusion matching degree of the song comprises:

[0013] The first matching degree and the second matching degree are fused based on a first weight corresponding to the first matching degree and a second weight corresponding to the second matching degree to obtain the fusion matching degree of the song.

[0014] In a possible implementation, the identifying whether the song is taken as the matching song corresponding to the audio to be identified in combination with the first matching degree and the second matching degree of the song comprises:

[0015] The greater matching degree between the first matching degree of the song and the second matching degree of the song is taken as the target matching degree of the song.

[0016] The song with the highest target matching degree in the song library is taken as the matching song of the audio to be identified.

[0017] In a possible implementation, before the extracting the accompaniment of the audio to be identified to obtain the first accompaniment audio of the audio to be identified and separating the vocals and the accompaniment of the audio to be identified to obtain the second accompaniment audio of the audio to be identified, the method further comprises:

[0018] Extracting an audio feature of the audio to be identified, and matching the audio feature with an audio feature of a song in the song library.

[0019] In a case where no song satisfying a matching condition exists in the song library, the extracting the accompaniment of the audio to be identified to obtain the first accompaniment audio of the audio to be identified and separating the vocals and the accompaniment of the audio to be identified to obtain the second accompaniment audio of the audio to be identified are performed.

[0020] In a possible implementation, before the extracting the accompaniment of the audio to be identified to obtain the first accompaniment audio of the audio to be identified and separating the vocals and the accompaniment of the audio to be identified to obtain the second accompaniment audio of the audio to be identified, the method further comprises:

[0021] Performing vocals noise detection on the audio to be identified to obtain a detection result.

[0022] In a case where the detection result meets the noise condition, the steps of performing accompaniment extraction on the to-be-identified audio to obtain first accompaniment audio of the to-be-identified audio, and performing voice and accompaniment separation on the to-be-identified audio to obtain second accompaniment audio of the to-be-identified audio are performed.

[0023] In a possible implementation, the method further includes:

[0024] In a case where the detection result does not meet the noise condition, audio features of the to-be-identified audio are extracted, and the audio features are matched with audio features of songs in the song library.

[0025] The song in the song library that meets the matching condition is taken as a matching song corresponding to the to-be-identified audio.

[0026] In a second aspect, a device for song identification is provided, and the device includes:

[0027] The obtaining module is configured to perform accompaniment extraction on the to-be-identified audio to obtain first accompaniment audio of the to-be-identified audio, perform voice and accompaniment separation on the to-be-identified audio to obtain second accompaniment audio of the to-be-identified audio, perform feature extraction on the first accompaniment audio to obtain first accompaniment audio features, and perform feature extraction on the second accompaniment audio to obtain second accompaniment audio features.

[0028] The matching module is configured to match the first accompaniment audio features with audio features of songs in a song library to obtain first matching degrees of the songs, match the second accompaniment audio features with the audio features of the songs to obtain second matching degrees of the songs, and combine the first matching degrees and the second matching degrees of the songs to identify whether the songs are matching songs corresponding to the to-be-identified audio.

[0029] In a possible implementation, the matching module is configured to:

[0030] fuse the first matching degrees and the second matching degrees of the songs to obtain a fused matching degree of the songs.

[0031] The song in the song library that has the highest fused matching degree is taken as a matching song of the to-be-identified audio.

[0032] In a possible implementation, the matching module is configured to:

[0033] fuse the first matching degrees and the second matching degrees of the songs based on a first weight corresponding to the first matching degrees and a second weight corresponding to the second matching degrees to obtain a fused matching degree of the songs.

[0034] In a possible implementation, the matching module is configured to:

[0035] taking the greater one of the first matching degree of the song and the second matching degree of the song as a target matching degree of the song;

[0036] taking the song with the highest target matching degree in the song library as a matching song of the audio to be identified.

[0037] In a possible implementation, the matching module is further configured to:

[0038] extracting an audio feature of the audio to be identified, and matching the audio feature with an audio feature of a song in the song library;

[0039] in a case where no song satisfying a matching condition exists in the song library, performing the steps of extracting an accompaniment of the audio to be identified to obtain a first accompaniment audio of the audio to be identified, and separating a vocal and an accompaniment of the audio to be identified to obtain a second accompaniment audio of the audio to be identified.

[0040] In a possible implementation, the apparatus further includes a detection module configured to:

[0041] performing vocal noise detection on the audio to be identified to obtain a detection result;

[0042] The obtaining module is configured to:

[0043] in a case where the detection result satisfies a noise condition, performing the steps of extracting an accompaniment of the audio to be identified to obtain a first accompaniment audio of the audio to be identified, and separating a vocal and an accompaniment of the audio to be identified to obtain a second accompaniment audio of the audio to be identified.

[0044] In a possible implementation, the matching module is further configured to:

[0045] in a case where the detection result does not satisfy the noise condition, extracting an audio feature of the audio to be identified, and matching the audio feature with an audio feature of a song in the song library;

[0046] taking a song satisfying a matching condition in the song library as a matching song corresponding to the audio to be identified.

[0047] In a third aspect, a server is provided, the terminal includes a processor and a memory, the memory stores at least one instruction, the instruction is loaded and executed by the processor to implement the operations performed by the method for identifying a song according to the first aspect and any possible implementation of the first aspect.

[0048] In a fourth aspect, a computer-readable storage medium is provided, and the storage medium stores at least one instruction, which is loaded and executed by a processor to implement the operations performed by the song recognition method according to the first aspect and any possible implementation of the first aspect.

[0049] In a fifth aspect, a computer program product is provided, and the computer program product stores at least one instruction, which is loaded and executed by a processor to implement the operations performed by the song recognition method according to the first aspect and any possible implementation of the first aspect.

[0050] The technical scheme provided in the present application has the beneficial effects that: in the technical scheme provided in the present application, after obtaining the to-be-recognized audio, the accompaniment audio of the to-be-recognized audio is obtained, and then the accompaniment audio features and the audio features of the songs in the song library are matched. Since the accompaniment audio does not contain environmental sound such as human voice, even if the to-be-recognized audio contains human voice noise, the extracted accompaniment audio does not contain human voice noise, so that the influence of the collected human voice noise on song recognition can be effectively avoided. In addition, in the technical scheme provided in the present application, two different technical schemes of accompaniment voice separation and accompaniment direct extraction are used when the accompaniment audio is obtained, which can avoid the problem of song recognition failure caused by the failure of obtaining the accompaniment audio when a single scheme is used to obtain the accompaniment audio, and effectively improves the reliability of the scheme. BRIEF DESCRIPTION OF DRAWINGS

[0051] In order to more clearly illustrate the technical schemes in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0052] Figure 1 is a song recognition method flowchart provided by an embodiment of the present application;

[0053] Figure 2 is an interface schematic diagram of a music application program provided by an embodiment of the present application;

[0054] Figure 3 is an accompaniment acquisition method schematic diagram provided by an embodiment of the present application;

[0055] Figure 4 is a song recognition method flowchart provided by an embodiment of the present application;

[0056] Figure 5 is a song recognition method flowchart provided by an embodiment of the present application;

[0057] Figure 6 is a flowchart of a song recognition method provided by an embodiment of the present application;

[0058] Figure 7 is a structural schematic diagram of a song recognition device provided by an embodiment of the present application;

[0059] Figure 8 is a structural schematic diagram of a terminal provided by an embodiment of the present application;

[0060] Figure 9 is a structural schematic diagram of a server provided by an embodiment of the present application. DETAILED DESCRIPTION

[0061] To make the objectives, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the drawings.

[0062] The present application provides a song recognition method, which can be implemented by a server. The terminal can be a user-side device such as a mobile phone, a tablet computer, a notebook computer, a desktop computer, etc. The server can be a single server or a server cluster. The music application program or applet installed in the terminal can provide a song recognition function. The server can be a background server of the application program or applet.

[0063] In a possible implementation scenario of the present application, the user can start the song recognition function of the music application program installed in the terminal. Then, the terminal collects the audio in the external environment through the microphone and sends the collected audio directly or after processing to the server. Then, the server can recognize the audio by executing the song recognition method provided by the present application and return the relevant information of the target song to the terminal.

[0064] The song recognition method provided by the present application will be described below. Referring to Figure 1 , the method can include the following steps:

[0065] Step 201, obtaining audio to be recognized.

[0066] In implementation, when the user hears a song played by a device and wants to know the relevant information of the song but does not know the song title, the user can open the music application program or applet installed in the terminal and open the song recognition function. Referring to Figure 2 , a music application program interface is shown, in which a song recognition function option is displayed. The user can click the song recognition function option, the music application program jumps to the song recognition interface, and the terminal starts collecting the audio in the external environment through the microphone.

[0067] Then, the terminal collects external audio through a microphone as to-be-recognized audio. At this time, if someone is speaking in the environment, the to-be-recognized audio collected by the terminal will be mixed with human voice noise. After the to-be-recognized audio is collected, the terminal sends a song recognition request to the server, wherein the to-be-recognized audio collected is carried in the song recognition request.

[0068] Correspondingly, after receiving the song recognition request sent by the terminal, the server parses the song recognition request to obtain the to-be-recognized audio carried therein.

[0069] Step 202, obtaining accompaniment audio of the to-be-recognized audio.

[0070] In implementation, for the to-be-recognized audio, the server obtains the accompaniment audio of the to-be-recognized audio. The scheme for obtaining the accompaniment audio can be various, and some of them are exemplarily listed below for description:

[0071] Scheme one:

[0072] The to-be-recognized audio is subjected to accompaniment extraction to obtain the accompaniment audio of the to-be-recognized audio. Specifically, the to-be-recognized audio is input into a pre-trained accompaniment extraction model, and the accompaniment extraction model outputs the accompaniment audio of the to-be-recognized audio. The accompaniment extraction model is a neural network model.

[0073] Scheme two:

[0074] The to-be-recognized audio is subjected to human voice and accompaniment separation to obtain human voice audio and accompaniment audio of the to-be-recognized audio. Specifically, the to-be-recognized audio is input into a pre-trained voice-accompaniment separation model, and the voice-accompaniment separation model outputs the human voice audio and the accompaniment audio of the to-be-recognized audio. The voice-accompaniment separation model is a neural network model.

[0075] Scheme three:

[0076] The to-be-recognized audio is subjected to accompaniment extraction to obtain first accompaniment audio of the to-be-recognized audio, and the to-be-recognized audio is subjected to human voice and accompaniment separation to obtain second accompaniment audio of the to-be-recognized audio. Specifically, as shown in Figure 3 the to-be-recognized audio is input into pre-trained accompaniment extraction model and voice-accompaniment separation model respectively, the accompaniment extraction model inputs the first accompaniment audio of the to-be-recognized audio, and the voice-accompaniment separation model outputs the second accompaniment audio and human voice audio of the to-be-recognized audio.

[0077] In this scheme three, the above scheme one and scheme two are combined, and two sets of audio obtaining schemes are used at the same time. This design has good complementarity and also has a disaster recovery mechanism. Even if one of the models is out of time and fails, the accompaniment audio can also be obtained through the other model.

[0078] Step 203, performing feature extraction on the accompaniment audio to obtain accompaniment audio features.

[0079] In implementation, after obtaining the accompaniment audio of the audio to be identified, the server extracts features from the accompaniment audio to obtain accompaniment audio features.

[0080] The feature extraction method can include Chroma Features extraction, Landmark audio fingerprint extraction, etc. Correspondingly, the extracted accompaniment audio features are different corresponding to different feature extraction methods. For example, Chroma Features extraction is performed on the accompaniment audio to obtain Chroma Features of the accompaniment audio as accompaniment audio features. For another example, Landmark audio fingerprint extraction is performed on the accompaniment audio to obtain Landmark audio fingerprint of the accompaniment audio as accompaniment audio features. The present application embodiment does not limit the specific feature extraction method, and only Landmark audio fingerprint extraction is exemplarily described below.

[0081] A spectrogram of the accompaniment audio of the audio to be identified is generated, where the horizontal axis of the spectrogram represents time, the vertical axis represents frequency, and the vertical axis represents energy. In the spectrogram of the accompaniment audio, local peaks of time-frequency points are extracted, and the time-frequency points corresponding to the local peaks are recorded, for example, time-frequency point 1 (t1, f1), where t1 represents the time of time-frequency point 1 and f1 represents the frequency of time-frequency point 1. Then, for each local peak, the peaks within a certain range around the local peak are found, and the time-frequency points corresponding to the peaks within the certain range are recorded. Then, for each local peak and the time-frequency points corresponding to the peaks within a certain range around the local peak, information is simplified and recorded in the form of frequency and time difference, and the simplified information is taken as the features of the signals within a certain range around the local peak. For example, the local peak corresponds to time-frequency point 1 (t1, f1), and the peaks within a certain range correspond to time-frequency point 2 (t2, f2). The information is simplified and the simplified information is recorded as (f1, f2, t1-t2) as the features of the signals within a certain range around the local peak. For each feature of the signals within a certain range around the local peak, a hash value can be further calculated to obtain a fingerprint of the signals within a certain range around the local peak, which can also be referred to as an index. Finally, the fingerprint of the signals within a certain range around each local peak of the accompaniment audio is taken as the Landmark audio fingerprint of the accompaniment audio.

[0082] Step 204, matching the accompaniment audio features with the audio features of the songs in the song library, and taking the songs satisfying the matching condition as the matching songs corresponding to the audio to be identified.

[0083] In implementation, after obtaining the accompaniment audio features, matching the accompaniment audio features and the audio features of the songs in the song library can obtain the matching degrees of the to-be-identified audio and each song in the song library. Then, the songs reaching the matching degree threshold are determined, and the song with the highest matching degree is selected from the songs reaching the matching degree threshold, and the song is taken as the matching song corresponding to the to-be-identified audio. The matching degree threshold can be configured by the technical personnel according to actual needs.

[0084] In a possible implementation, if the first accompaniment audio is obtained by accompaniment extraction in step 202, and the second accompaniment audio is obtained by sound accompaniment separation, then the audio features of the first accompaniment audio and the audio features of the second accompaniment audio are calculated respectively in step 203.

[0085] Correspondingly, in step 204, the audio features of the first accompaniment audio and the audio features of each song in the song library can be matched to obtain the first matching degrees of the to-be-identified audio and each song in the song library. The audio features of the first accompaniment audio and the audio features of each song in the song library are matched to obtain the second matching degrees of the to-be-identified audio and each song in the song library. Then, the song meeting the matching condition is selected as the matching song corresponding to the to-be-identified audio. In this matching case, there can be multiple methods for selecting the song meeting the matching condition, and some of the methods are exemplarily described below.

[0086] Method one:

[0087] For each song in the song library, the first matching degree and the second matching degree of the song are fused to obtain a fused matching degree. Then, the song with the fused matching degree reaching the matching degree threshold is determined, and the song with the highest fused matching degree is selected from the songs with the fused matching degrees reaching the matching degree threshold, and the song is taken as the matching song corresponding to the to-be-identified audio.

[0088] The above method of fusing the first matching degree and the second matching degree is shown in the following formula (1):

[0089] M=α×m1+β×m2 (1)

[0090] Wherein, M is the fusion matching degree, m1 is the first matching degree, m2 is the second matching degree, a is the first weight, b is the second weight, a+b=1, and the specific values of a and b can be configured by the technician according to the actual demand. For example, in the case that the performance of the accompaniment extraction model is better than that of the vocal accompaniment separation model, the first weight can be configured to be higher than the second weight, such as the first weight being 0.6 and the second weight being 0.4. Similarly, in the case that the performance of the vocal accompaniment extraction model is better than that of the accompaniment extraction model, the second weight can be configured to be higher than the first weight, such as the first weight being 0.4 and the second weight being 0.6. In the case that the performance of the vocal accompaniment extraction model is equivalent to that of the accompaniment extraction model, the first weight and the second weight can be configured to be the same value, such as the first weight being 0.5 and the second weight being 0.5.

[0091] Method two:

[0092] For each song in the song library, compare the first matching degree and the second matching degree of the song, and take the larger matching degree as the target matching degree of the song. Then, determine the songs whose target matching degrees reach the matching degree threshold, and select the song with the highest target matching degree from the songs whose target matching degrees reach the matching degree threshold, and take the song with the highest target matching degree as the matching song corresponding to the to-be-identified audio.

[0093] Method three:

[0094] Determine the songs in the song library whose matching degrees reach the matching degree threshold, and select the song with the highest matching degree from the songs whose matching degrees reach the matching degree threshold, and take the song as the matching song corresponding to the to-be-identified audio.

[0095] The song with the highest matching degree is exemplarily described as follows:

[0096] For example, the first matching degree of the first song is 96%, and the second matching degree is 95%, the first matching degree of the second song is 94%, and the second matching degree is 97%, wherein the second matching degree of the second song is the highest among all matching degrees, and the second song is the song with the highest matching degree.

[0097] Method four:

[0098] Determine the songs in the song library whose first matching degrees reach the matching degree threshold, and select the song with the highest first matching degree from the songs whose first matching degrees reach the matching degree threshold, and take the song as the first matching song corresponding to the to-be-identified audio. And determine the songs in the song library whose second matching degrees reach the matching degree threshold, and select the song with the highest second matching degree from the songs whose second matching degrees reach the matching degree threshold, and take the song as the second matching song corresponding to the to-be-identified audio.

[0099] After the matching song corresponding to the to-be-identified audio is determined, relevant information of the matching song can be further acquired, and then the relevant information of the matching song is sent to the terminal. The relevant information can include a song name, a singer, lyrics, and the like. In the case of the fourth method, if the first matching song and the second matching song are different songs, the relevant information of the first matching song and the second matching song can be sent to the terminal, and the user can select from the first matching song and the second matching song.

[0100] In a possible implementation, referring to Figure 4 After the to-be-identified audio is acquired in step 201, the following steps can be performed first:

[0101] Step 301: Feature extraction is performed on the to-be-identified audio to obtain audio features of the to-be-identified audio.

[0102] In implementation, after the to-be-identified audio is acquired, feature extraction can be performed on the to-be-identified audio first to obtain audio features of the to-be-identified audio.

[0103] The method of feature extraction can include Chroma Features extraction and Landmark audio fingerprint extraction. Correspondingly, the audio features of the to-be-identified audio extracted according to different feature extraction methods are different. For example, Chroma Features extraction is performed on the to-be-identified audio, and Chroma Features of the to-be-identified audio can be obtained as the audio features of the to-be-identified audio. For another example, Landmark audio fingerprint extraction is performed on the to-be-identified audio, and a Landmark audio fingerprint of the to-be-identified audio can be obtained as the audio features of the to-be-identified audio. The specific feature extraction method is not limited in the embodiments of the present application.

[0104] Step 302: The audio features of the to-be-identified audio are matched with the audio features of the songs in the song library.

[0105] Step 303: It is determined whether there is a song in the song library that meets the matching condition.

[0106] In implementation, the processing of steps 302 and 303 is similar to the processing of step 204, which is not described herein again.

[0107] Step 304: In the case where there is no song in the song library that meets the matching condition, step 202 is continued.

[0108] Step 305: In the case where there is a song in the song library that meets the matching condition, the song that meets the matching condition is taken as the matching song of the to-be-identified audio.

[0109] In a possible implementation, referring to Figure 5After the audio to be recognized is obtained in step 201, the following steps can be performed first:

[0110] Step 401, human voice noise detection is performed on the audio to be recognized to obtain a detection result.

[0111] In implementation, after the audio to be recognized is obtained, the audio to be recognized can be input into a human voice noise detection model, and the human voice noise detection model outputs the detection result. The detection result can be a confidence degree of the audio to be recognized having human voice noise.

[0112] The human voice noise detection model used here can be a neural network model, which can be trained before use to learn the characteristics of human voice noise, so as to perform the task of human voice noise detection.

[0113] Step 402, it is determined whether the detection result meets a noise condition.

[0114] In implementation, if the confidence degree of the audio to be recognized having human voice noise reaches a confidence degree threshold, it is determined that the detection result meets the noise condition, and if the confidence degree of the audio to be recognized having human voice noise does not reach the confidence degree threshold, it is determined that the detection result does not meet the noise condition. The confidence degree reaching the confidence degree threshold can be that the confidence degree is greater than the confidence degree threshold, or that the confidence degree is greater than or equal to the confidence degree threshold. The confidence degree threshold can be configured by a technician according to actual needs.

[0115] Step 403, in the case where the detection result meets the noise condition, step 202 is performed.

[0116] Step 404, in the case where the detection result does not meet the noise condition, feature extraction is performed on the audio to be recognized to obtain audio features of the audio to be recognized.

[0117] Step 405, the audio features of the audio to be recognized and the audio features of the songs in the song library are matched.

[0118] Step 406, it is determined whether there is a song in the song library that meets a matching condition.

[0119] Step 407, in the case where there is no song in the song library that meets the matching condition, step 202 is continued to be performed.

[0120] Step 408, in the case where there is a song in the song library that meets the matching condition, the song that meets the matching condition is taken as a matching song of the audio to be recognized.

[0121] In implementation, the processing of steps 404-407 above is the same as or similar to the processing of steps 301-305 above, and will not be described here.

[0122] All the optional technical solutions described above can be combined to form optional embodiments of the present disclosure, which will not be described again here.

[0123] In the technical solution provided in the embodiments of the present application, after obtaining the to-be-identified audio, the accompaniment audio of the to-be-identified audio is extracted, and then the accompaniment audio features and the audio features of the songs in the song library are matched. Since the accompaniment audio does not contain human voice, even if the to-be-identified audio contains human voice noise, the extracted accompaniment audio does not contain human voice noise, so that the influence of the collected human voice noise on song recognition can be effectively avoided.

[0124] In the above-mentioned process of obtaining the accompaniment audio of the to-be-identified audio, if only one obtaining scheme is used, in the case that the obtaining scheme fails, the current song recognition will fail, and the stability is poor. Based on this, in the technical solution provided in the embodiments of the present application, multiple schemes for obtaining accompaniment audio can be combined to avoid this problem, thereby improving the stability and reliability of the song recognition system. The following will be combined Figure 6 Another method for song recognition provided in the embodiments of the present application will be described. Referring to Figure 6 The method can include the following processing steps:

[0125] Step 501, extracting accompaniment from the to-be-identified audio to obtain the first accompaniment audio of the to-be-identified audio, and separating the human voice and the accompaniment of the to-be-identified audio to obtain the second accompaniment audio of the to-be-identified audio. In the implementation, the to-be-identified audio is input into a pre-trained accompaniment extraction model, and the accompaniment extraction model outputs the accompaniment audio of the to-be-identified audio. The to-be-identified audio is input into a pre-trained voice-accompaniment separation model, and the voice-accompaniment separation model outputs the human voice audio and the accompaniment audio of the to-be-identified audio.

[0126] Step 502, extracting features from the first accompaniment audio to obtain the first accompaniment audio features, and extracting features from the second accompaniment audio to obtain the second accompaniment audio features. In the implementation, after obtaining the first accompaniment audio and the second accompaniment audio of the to-be-identified audio, the server extracts features from the first accompaniment audio to obtain the first accompaniment audio features, and extracts features from the second accompaniment audio to obtain the second accompaniment audio features.

[0127] The method of feature extraction can include Chroma Features extraction, Landmark audio fingerprint extraction, etc. Correspondingly, the extracted accompaniment audio features are different corresponding to different feature extraction methods. For example, Chroma Features extraction is performed on the accompaniment audio, and Chroma Features of the accompaniment audio can be obtained as accompaniment audio features. For another example, Landmark audio fingerprint extraction is performed on the accompaniment audio, and Landmark audio fingerprint of the accompaniment audio can be obtained as accompaniment audio features.

[0128] Step 503, match the first accompaniment audio features and the audio features of the songs in the song library to obtain a first matching degree of the songs, and match the second accompaniment audio features and the audio features of the songs to obtain a second matching degree of the songs. The matching degree can be a value of similarity, distance, etc. measuring the degree of similarity, for example, Euclidean distance, cosine similarity, Euclidean distance, etc. The present application is not limited to this.

[0129] Step 504, combine the first matching degree and the second matching degree of the songs to identify whether the songs are the matching songs corresponding to the to-be-identified audio. In implementation, the implementation of this step 504 can have multiple methods, some of which are exemplarily listed below for description:

[0130] Method one:

[0131] For each song in the song library, the first matching degree and the second matching degree of the song are fused to obtain a fused matching degree. Then, the song whose fused matching degree reaches a matching degree threshold is determined, and among the songs whose fused matching degrees reach the matching degree threshold, the song with the highest fused matching degree is selected as the matching song corresponding to the to-be-identified audio. The above method of fusing the first matching degree and the second matching degree is shown in the following formula (1):

[0132] M=α×m1+β×m2 (1)

[0133] Wherein, M is the fusion matching degree, m1 is the first matching degree, m2 is the second matching degree, a is the first weight, b is the second weight, a+b=1, and the specific values of a and b can be configured by the technician according to the actual demand. For example, in the case that the performance of the accompaniment extraction model is better than that of the vocal accompaniment separation model, the first weight can be configured to be higher than the second weight, such as the first weight being 0.6 and the second weight being 0.4. Similarly, in the case that the performance of the vocal accompaniment extraction model is better than that of the accompaniment extraction model, the second weight can be configured to be higher than the first weight, such as the first weight being 0.4 and the second weight being 0.6. In the case that the performance of the vocal accompaniment extraction model is equivalent to that of the accompaniment extraction model, the first weight and the second weight can be configured to be the same value, such as the first weight being 0.5 and the second weight being 0.5.

[0134] Method two:

[0135] For each song in the song library, compare the first matching degree and the second matching degree of the song, and take the larger matching degree as the target matching degree of the song. Then, determine the songs whose target matching degrees reach the matching degree threshold, and select the song with the highest target matching degree from the songs whose target matching degrees reach the matching degree threshold, and take the song with the highest target matching degree as the matching song corresponding to the to-be-identified audio.

[0136] Method three:

[0137] Determine the songs in the song library whose matching degrees in at least one of the first matching degree and the second matching degree reach the matching degree threshold, and select the song with the highest matching degree from the songs whose matching degrees reach the matching degree threshold, and take the song as the matching song corresponding to the to-be-identified audio. The selection of the song with the highest matching degree is exemplarily described as follows: for example, the first matching degree of a first song is 96%, the second matching degree is 95%, the first matching degree of a second song is 94%, and the second matching degree is 97%, wherein the second matching degree of the second song is the highest among all matching degrees, and the second song is the song with the highest matching degree.

[0138] Method four:

[0139] Determine the songs in the song library whose first matching degrees reach the matching degree threshold, and select the song with the highest matching degree from the songs whose first matching degrees reach the matching degree threshold, and take the song as the matching song corresponding to the to-be-identified audio. And determine the songs in the song library whose second matching degrees reach the matching degree threshold, and select the song with the highest matching degree from the songs whose second matching degrees reach the matching degree threshold, and take the song as the matching song corresponding to the to-be-identified audio.

[0140] In Figure 6The method shown employs two audio acquisition schemes simultaneously, which is highly complementary and also has a disaster recovery mechanism. Even if one model times out and fails, the accompaniment audio can still be obtained through the other model.

[0141] Based on the same technical concept, this application also provides a song recognition device, which can be applied to a server (see [link]). Figure 7 The device may include an acquisition module 510 and a matching module 520, wherein:

[0142] The acquisition module 510 is used to extract accompaniment from the audio to be identified to obtain a first accompaniment audio of the audio to be identified; to separate the human voice and accompaniment from the audio to be identified to obtain a second accompaniment audio of the audio to be identified; to extract features from the first accompaniment audio to obtain first accompaniment audio features; and to extract features from the second accompaniment audio to obtain second accompaniment audio features.

[0143] The matching module 520 is used to match the first accompaniment audio features with the audio features of songs in the music library to obtain a first matching degree of the song; match the second accompaniment audio features with the audio features of the song to obtain a second matching degree of the song; and combine the first matching degree and the second matching degree of the song to identify whether the song is a matching song corresponding to the audio to be identified.

[0144] In one possible implementation, the matching module 520 is used to:

[0145] The first matching degree and the second matching degree of the song are fused to obtain the fused matching degree of the song;

[0146] The song with the highest matching degree in the music library is selected as the matching song for the audio to be identified.

[0147] In one possible implementation, the matching module 520 is used to:

[0148] Based on the first weight corresponding to the first matching degree and the second weight corresponding to the second matching degree, the first matching degree and the second matching degree are fused to obtain the fused matching degree of the song.

[0149] In one possible implementation, the matching module 520 is used to:

[0150] The larger of the first matching degree and the second matching degree of the song is taken as the target matching degree of the song;

[0151] The song with the highest target matching degree in the music library is selected as the matching song for the audio to be identified.

[0152] In a possible implementation, the matching module 520 is further configured to:

[0153] extract an audio feature of the to-be-identified audio, and match the audio feature with an audio feature of a song in the song library;

[0154] In a case where there is no song in the song library that meets the matching condition, perform the steps of extracting accompaniment from the to-be-identified audio to obtain first accompaniment audio of the to-be-identified audio, and separating vocals from the to-be-identified audio to obtain second accompaniment audio of the to-be-identified audio.

[0155] In a possible implementation, the apparatus further includes a detection module configured to:

[0156] perform vocals noise detection on the to-be-identified audio to obtain a detection result;

[0157] The obtaining module 510 is configured to:

[0158] In a case where the detection result meets the noise condition, perform the steps of extracting accompaniment from the to-be-identified audio to obtain first accompaniment audio of the to-be-identified audio, and separating vocals from the to-be-identified audio to obtain second accompaniment audio of the to-be-identified audio.

[0159] In a possible implementation, the matching module 520 is further configured to:

[0160] In a case where the detection result does not meet the noise condition, extract an audio feature of the to-be-identified audio, and match the audio feature with an audio feature of a song in the song library;

[0161] take the song in the song library that meets the matching condition as a matching song corresponding to the to-be-identified audio.

[0162] In the technical scheme provided in the embodiments of the present application, after obtaining the to-be-identified audio, the accompaniment audio of the to-be-identified audio is extracted, and then the accompaniment audio feature is matched with the audio feature of a song in the song library. Because the accompaniment audio does not contain vocals, even if the to-be-identified audio contains vocals noise, the extracted accompaniment audio does not contain vocals noise, so that the influence of the collected vocals noise on song recognition can be effectively avoided.

[0163] It should be noted that the device for song recognition provided in the above embodiment is only used for example to divide the above functions, and in actual application, the above functions can be distributed to different function modules according to needs, that is, the internal structure of the server is divided into different function modules to complete all or part of the above described functions. In addition, the device for song recognition provided in the above embodiment and the method for song recognition provided in the above embodiment belong to the same concept, and the specific implementation process is described in the method embodiment, which will not be repeated here.

[0164] Figure 8 A structure block diagram of a terminal 600 provided by an example embodiment of the present application is shown. The terminal 600 can be a portable mobile terminal, such as a smart phone, a tablet computer, an MP3 (Moving Picture Experts Group Audio Layer III) player, an MP4 (Moving Picture Experts Group Audio Layer IV) player, a notebook computer, or a desktop computer. The terminal 600 can also be referred to as a user equipment, a portable terminal, a laptop terminal, a desktop terminal, or other names.

[0165] Generally, the terminal 600 includes a processor 601 and a memory 602.

[0166] The processor 601 can include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 601 can be implemented in at least one of a hardware form of a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), and a PLA (Programmable Logic Array). The processor 601 can also include a main processor and a coprocessor. The main processor is a processor for processing data in an awake state, also referred to as a CPU (Central Processing Unit). The coprocessor is a low-power processor for processing data in a standby state. In some embodiments, the processor 601 can be integrated with a GPU (Graphics Processing Unit) for rendering and drawing content to be displayed on a display screen. In some embodiments, the processor 601 can also include an AI (Artificial Intelligence) processor for processing machine learning-related computing operations.

[0167] The memory 602 can include one or more computer-readable storage media. The computer-readable storage media can be non-transitory. The memory 602 can also include high-speed random access memory and can include non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 602 stores at least one instruction for execution by the processor 601 to implement a method of generating a master audio according to an embodiment of the method provided in the present application.

[0168] In some embodiments, the terminal 600 can further optionally include a peripheral device interface 603 and at least one peripheral device. The processor 601, the memory 602, and the peripheral device interface 603 can be connected through a bus or a signal line. Each peripheral device can be connected to the peripheral device interface 603 through a bus, a signal line, or a circuit board. Specifically, the peripheral device includes at least one of a radio frequency circuit 604, a display screen 605, a camera component 606, an audio circuit 607, a positioning component 608, and a power supply 609.

[0169] The peripheral device interface 603 can be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 601 and the memory 602. In some embodiments, the processor 601, the memory 602, and the peripheral device interface 603 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 601, the memory 602, and the peripheral device interface 603 can be implemented on a separate chip or circuit board, and the present embodiment does not limit this.

[0170] The radio frequency circuit 604 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 604 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 604 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. Optionally, the radio frequency circuit 604 includes an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and the like. The radio frequency circuit 604 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to the World Wide Web, a metropolitan area network, an intranet, various generations of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 604 can also include NFC (Near Field Communication) related circuitry, which is not limited by the present application.

[0171] The display screen 605 is configured to display a UI (User Interface). The UI can include graphics, text, icons, video, and any combination thereof. When the display screen 605 is a touch display screen, the display screen 605 is further configured to capture touch signals on or above the surface of the display screen 605. The touch signals can be input to the processor 601 as control signals for processing. In this case, the display screen 605 can also be configured to provide virtual buttons and / or virtual keyboard, also known as soft buttons and / or soft keyboard. In some embodiments, the display screen 605 can be one, disposed on the front panel of the terminal 600; in other embodiments, the display screen 605 can be at least two, respectively disposed on different surfaces of the terminal 600 or in a folding design; in other embodiments, the display screen 605 can be a flexible display screen, disposed on a curved surface or a folding surface of the terminal 600. Even, the display screen 605 can also be disposed in an irregular shape, i.e., a special-shaped screen. The display screen 605 can be made of LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode), etc.

[0172] The camera assembly 606 is configured to capture images or videos. Optionally, the camera assembly 606 includes a front camera and a rear camera. Typically, the front camera is disposed on the front panel of the terminal, and the rear camera is disposed on the back of the terminal. In some embodiments, the rear camera is at least two, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, to realize the background blur function by fusing the main camera and the depth-of-field camera, the panoramic shooting and VR (Virtual Reality) shooting function by fusing the main camera and the wide-angle camera, or other fusion shooting functions. In some embodiments, the camera assembly 606 can further include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. The dual-color temperature flash refers to the combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.

[0173] The audio circuit 607 can include a microphone and a speaker. The microphone is used to collect sound waves of a user and an environment, and convert the sound waves into an electrical signal input to the processor 601 for processing, or input to the radio frequency circuit 604 to realize voice communication. For the purpose of stereo sound collection or noise reduction, the microphone can be multiple, and arranged at different parts of the terminal 600. The microphone can also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert an electrical signal from the processor 601 or the radio frequency circuit 604 into sound waves. The speaker can be a conventional diaphragm speaker, or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, not only can it convert an electrical signal into a sound wave audible to humans, but also can convert an electrical signal into an inaudible sound wave to humans for ranging purposes, etc. In some embodiments, the audio circuit 607 can also include a headphone jack.

[0174] The positioning component 608 is used to position the current geographic location of the terminal 600 to realize navigation or LBS (Location Based Service). The positioning component 608 can be a GPS (Global Positioning System) based positioning component, a Beidou system based positioning component, or a Galileo system based positioning component.

[0175] The power supply 609 is used to supply power to various components in the terminal 600. The power supply 609 can be an alternating current, a direct current, a disposable battery, or a rechargeable battery. When the power supply 609 includes a rechargeable battery, the rechargeable battery can be a wired charging battery or a wireless charging battery. The wired charging battery is a battery charged through a wired line, and the wireless charging battery is a battery charged through a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0176] In some embodiments, the terminal 600 further includes one or more sensors 610. The one or more sensors 610 include, but are not limited to, an acceleration sensor 611, a gyroscope sensor 612, a pressure sensor 613, a fingerprint sensor 614, an optical sensor 615, and a proximity sensor 616.

[0177] The acceleration sensor 611 can detect the acceleration magnitude in three coordinate axes of the coordinate system established by the terminal 600. For example, the acceleration sensor 611 can be used to detect the components of the gravitational acceleration in three coordinate axes. The processor 601 can control the display screen 605 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 611. The acceleration sensor 611 can also be used for game or user motion data collection.

[0178] The gyroscope sensor 612 can detect the orientation and rotation angle of the terminal 600. The gyroscope sensor 612, in conjunction with the accelerometer sensor 611, can collect 3D motion data from the user on the terminal 600. Based on the data collected by the gyroscope sensor 612, the processor 601 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.

[0179] The pressure sensor 613 can be disposed on the side bezel of the terminal 600 and / or on the lower layer of the display screen 605. When the pressure sensor 613 is disposed on the side bezel of the terminal 600, it can detect the user's grip signal on the terminal 600, and the processor 601 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 613. When the pressure sensor 613 is disposed on the lower layer of the display screen 605, the processor 601 can control the operable controls on the UI interface based on the user's pressure operation on the display screen 605. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.

[0180] The fingerprint sensor 614 is used to collect the user's fingerprint. The processor 601 identifies the user's identity based on the fingerprint collected by the fingerprint sensor 614, or the fingerprint sensor 614 identifies the user's identity based on the collected fingerprint. When the user's identity is identified as trusted, the processor 601 authorizes the user to perform relevant sensitive operations, including unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings. The fingerprint sensor 614 can be located on the front, back, or side of the terminal 600. When the terminal 600 has physical buttons or a manufacturer's logo, the fingerprint sensor 614 can be integrated with the physical buttons or manufacturer's logo.

[0181] An optical sensor 615 is used to collect ambient light intensity. In one embodiment, the processor 601 can control the display brightness of the display screen 605 based on the ambient light intensity collected by the optical sensor 615. Specifically, when the ambient light intensity is high, the display brightness of the display screen 605 is increased; when the ambient light intensity is low, the display brightness of the display screen 605 is decreased. In another embodiment, the processor 601 can also dynamically adjust the shooting parameters of the camera assembly 606 based on the ambient light intensity collected by the optical sensor 615.

[0182] The proximity sensor 616, also known as a distance sensor, is typically mounted on the front panel of the terminal 600. The proximity sensor 616 is used to detect the distance between the user and the front of the terminal 600. In one embodiment, when the proximity sensor 616 detects that the distance between the user and the front of the terminal 600 is gradually decreasing, the processor 601 controls the display screen 605 to switch from a screen-on state to a screen-off state; when the proximity sensor 616 detects that the distance between the user and the front of the terminal 600 is gradually increasing, the processor 601 controls the display screen 605 to switch from a screen-off state to a screen-on state.

[0183] Those skilled in the art will understand that Figure 8 The structure shown does not constitute a limitation on terminal 600, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0184] Figure 9 This is a schematic diagram of a computing device 1000 provided in an embodiment of this application. The computing device 1000 can vary significantly due to differences in configuration or performance. It may include one or more central processing units (CPUs) 1001 and one or more memories 1002. The memories 1002 store at least one instruction, which is loaded and executed by the processors 1001 to implement the methods provided in the various method embodiments described above. Of course, the computing device may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The computing device may also include other components for implementing device functions, which will not be elaborated upon here.

[0185] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions that can be executed by a processor in a terminal to complete the song recognition method described above. This computer-readable storage medium can be non-transitory. For example, the computer-readable storage medium can be ROM (Read-Only Memory), RAM (Random Access Memory), CD-ROM (Compact Disc Read-Only Memory), magnetic tape, floppy disk, and optical data storage devices, etc.

[0186] In an exemplary embodiment, a computer program product is also provided, which stores at least one instruction that is loaded and executed by a processor to perform the operations performed by the song recognition method described above.

[0187] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals (including but not limited to signals transmitted between the user terminal and other devices) involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the audio to be identified and the audio characteristics of songs in the music library involved in this application were obtained with full authorization.

[0188] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0189] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for song recognition, characterized in that, The method includes: The accompaniment is extracted from the audio to be identified to obtain the first accompaniment audio of the audio to be identified; the human voice and accompaniment are separated from the audio to be identified to obtain the second accompaniment audio of the audio to be identified. Feature extraction is performed on the first accompaniment audio to obtain the first accompaniment audio features; feature extraction is performed on the second accompaniment audio to obtain the second accompaniment audio features; The first accompaniment audio feature is matched with the audio features of songs in the music library to obtain the first matching degree of the song; the second accompaniment audio feature is matched with the audio features of the song to obtain the second matching degree of the song. By combining the first matching degree and the second matching degree of the song, it is determined whether the song is the matching song corresponding to the audio to be identified.

2. The method according to claim 1, characterized in that, The step of combining the first matching degree and the second matching degree of the song to identify whether the song is a matching song corresponding to the audio to be identified includes: The first matching degree and the second matching degree of the song are fused to obtain the fused matching degree of the song; The song with the highest matching degree in the music library is selected as the matching song for the audio to be identified.

3. The method according to claim 2, characterized in that, The process of fusing the first matching degree and the second matching degree of the song to obtain the fused matching degree of the song includes: Based on the first weight corresponding to the first matching degree and the second weight corresponding to the second matching degree, the first matching degree and the second matching degree are fused to obtain the fused matching degree of the song.

4. The method according to claim 1, characterized in that, The step of combining the first matching degree and the second matching degree of the song to identify whether the song is a matching song corresponding to the audio to be identified includes: The larger of the first matching degree and the second matching degree of the song is taken as the target matching degree of the song; The song with the highest target matching degree in the music library is selected as the matching song for the audio to be identified.

5. The method according to any one of claims 1-4, characterized in that, Before extracting the accompaniment from the audio to be identified to obtain a first accompaniment audio, and separating the vocals and accompaniment from the audio to be identified to obtain a second accompaniment audio, the method further includes: Extract the audio features of the audio to be identified, and match the audio features with the audio features of songs in the music library; If no song that meets the matching conditions exists in the music library, the steps of extracting the accompaniment from the audio to be identified to obtain the first accompaniment audio of the audio to be identified, and separating the vocals and accompaniment from the audio to be identified to obtain the second accompaniment audio of the audio to be identified are performed.

6. The method according to any one of claims 1-4, characterized in that, Before extracting the accompaniment from the audio to be identified to obtain a first accompaniment audio, and separating the vocals and accompaniment from the audio to be identified to obtain a second accompaniment audio, the method further includes: Human voice noise detection is performed on the audio to be identified, and the detection result is obtained; If the detection result meets the noise condition, the following steps are performed: extracting the accompaniment from the audio to be identified to obtain the first accompaniment audio of the audio to be identified; separating the human voice and accompaniment from the audio to be identified to obtain the second accompaniment audio of the audio to be identified.

7. The method according to claim 6, characterized in that, The method further includes: If the detection result does not meet the noise condition, the audio features of the audio to be identified are extracted, and the audio features are matched with the audio features of songs in the music library. Songs in the music library that meet the matching conditions are used as the matching songs corresponding to the audio to be identified.

8. A server, characterized in that, The server includes a processor and a memory, the memory storing at least one instruction that is loaded and executed by the processor to perform the operations of the song recognition method as described in any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction, which is loaded and executed by a processor to perform the operations of the song recognition method as described in any one of claims 1 to 7.

10. A computer program product, characterized in that, The computer program product stores at least one instruction, which is loaded and executed by a processor to perform the operations of the song recognition method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Language recognition method and device, model training method and device, and facility

    CN110853618A

  • Chorus voice separation method, computer equipment and storage medium

    CN116524949A