Audio recognition method, computer device and computer program product
The audio recognition method enhances song identification by comparing beat level differences between target and candidate songs, effectively filtering out irrelevant content and improving accuracy.
Patent Information
- Application Number
- CN202310244837.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-07
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-03-07
AI Technical Summary
When existing cover recognition technology has speech sounds or noises in the target audio, it is easy to misidentify them as song content, affecting the accuracy of the original song recall results.
By obtaining the number of beat points of the target audio and candidate song audio, calculating the level difference, and comparing it with the preset difference threshold, we determine whether the candidate song is the original song audio, and avoiding false filtering caused by direct detection of the number of beat points.
It improves the accuracy of the original song recall results, avoids erroneous recognition caused by noise or speaking, and improves the accuracy of the cover recognition process.
Smart Images

Figure CN116417012B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio technology, and in particular, to an audio recognition method, a computer device, and a computer program product. Background Art
[0002] With the development of computer technology, it has become increasingly widespread to use cover song recognition technology to search for cover songs or the original songs associated with original works.
[0003] In the related art, the cover song recognition technology can perform matching based on information such as the lyrics or pitch of the cover audio, and regard the song with similar lyrics content or pitch sequence as the original song of the cover song. However, when there is speech or other noise in the audio to be recognized, the speech or other noise will be wrongly regarded as the song content, and the song with similar lyrics content or pitch sequence will be regarded as the original song of this audio, affecting the accuracy of the recall result of the original song in the cover song recognition process. Summary of the Invention
[0004] Based on this, it is necessary to provide an audio recognition method, a computer device, and a computer program product that can improve the accuracy of the recall result of the original song for the above technical problems.
[0005] In a first aspect, this application provides an audio recognition method. The method includes:
[0006] Obtain candidate song audios obtained after performing song matching on the target audio;
[0007] Determine the number-of-beats level corresponding to the beats in the target audio, and determine the number-of-beats level corresponding to the beats in the candidate song audio;
[0008] Obtain the level difference between the number-of-beats level of the target audio and the number-of-beats level of the candidate song audio;
[0009] If the level difference is less than a preset difference threshold, determine the candidate song audio as the original song audio of the target audio;
[0010] If the level difference is greater than or equal to the preset difference threshold, determine that the recall of the original song audio of the target audio fails.
[0011] In one embodiment, the determining the number-of-beats level corresponding to the beats in the target audio, and determining the number-of-beats level corresponding to the beats in the candidate song audio includes:
[0012] Divide the target audio into multiple target audio segments with a preset duration, and obtain the number-of-beats level corresponding to the beats in each target audio segment; and,
[0013] Divide the candidate song audio into multiple candidate song audio segments of the preset duration, and obtain the beat quantity levels corresponding to the beats in each candidate song audio segment.
[0014] In one embodiment, the obtaining of the level difference between the beat quantity level of the target audio and the beat quantity level of the candidate song audio includes:
[0015] Determine the segment level difference between the beat quantity level of each target audio segment and the beat quantity level of the corresponding candidate song audio segment;
[0016] Determine the level difference between the beat quantity levels of the target audio and the candidate song audio according to the multiple segment level differences.
[0017] In one embodiment, the determining of the beat quantity level corresponding to the beats in the target audio and the determining of the beat quantity level corresponding to the beats in the candidate song audio include:
[0018] Input the audio features corresponding to the target audio into the trained beat information recognition model to obtain the beat quantity level of the target audio output by the beat information recognition model; and,
[0019] Input the audio features corresponding to the candidate song audio into the beat information recognition model to obtain the beat quantity level of the candidate song audio output by the beat information recognition model.
[0020] In one embodiment, the beat information recognition model is trained through the following steps:
[0021] Obtain multiple sample audios including speaker corpus audio and / or noise audio;
[0022] Perform supervised training on the beat information recognition model to be trained based on the multiple sample audios and the beat quantity level labels of each sample audio;
[0023] When the training end condition is satisfied, obtain the trained beat information recognition model.
[0024] In one embodiment, the obtaining of the candidate song audio obtained after song matching for the target audio includes:
[0025] Determine the lyrics of the target audio, and determine the lyric similarity between the lyrics of the target audio and the lyrics of each song audio in the audio library;
[0026] If there is a song audio with the largest lyric similarity that is greater than the first similarity threshold, then determine this song audio as the candidate song audio for the target audio.
[0027] In one embodiment, after determining the lyric similarities between the lyrics of the target audio and the lyrics of multiple song audios in the audio library, it further includes:
[0028] If the lyric similarities between the lyrics of the target audio and the lyrics of each song audio in the audio library are all less than the first similarity threshold, then determine the melody similarity between the melody of the target audio and the melody of each song audio in the audio library;
[0029] If there is a song audio with the largest melody similarity that is greater than the second similarity threshold, then determine this song audio as the candidate song audio for the target audio.
[0030] In one embodiment, the determining the melody similarity between the melody of the target audio and the melody of each song audio in the audio library includes:
[0031] Obtain the melody features of the target audio, and determine the cosine distance between the melody features of the target audio and the melody features of each song audio in the audio library;
[0032] Based on the cosine distance, determine the melody similarity between the melody of the target audio and the melody of each song audio in the audio library.
[0033] In a second aspect, the present application also provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0034] Obtain the candidate song audio obtained after performing song matching on the target audio;
[0035] Determine the beat number level corresponding to the beats in the target audio, and determine the beat number level corresponding to the beats in the candidate song audio;
[0036] Obtain the level difference between the beat number level of the target audio and the beat number level of the candidate song audio;
[0037] If the level difference is less than the preset difference threshold, then determine the candidate song audio as the original singer song audio of the target audio;
[0038] If the level difference is greater than or equal to the preset difference threshold, then determine that the recall of the original singer song audio of the target audio fails.
[0039] In a third aspect, the present application also provides a computer program product. The computer program product includes a computer program which, when executed by a processor, implements the following steps:
[0040] Obtain candidate song audio obtained by performing song matching on target audio;
[0041] Determine the number-of-beats level corresponding to the beats in the target audio, and determine the number-of-beats level corresponding to the beats in the candidate song audio;
[0042] Obtain the level difference between the number-of-beats level of the target audio and the number-of-beats level of the candidate song audio;
[0043] If the level difference is less than a preset difference threshold, determine the candidate song audio as the original song audio of the target audio;
[0044] If the level difference is greater than or equal to the preset difference threshold, determine that the recall of the original song audio of the target audio fails.
[0045] The above audio recognition method, computer device and computer program product can obtain candidate song audio obtained by performing song matching on target audio, determine the number-of-beats level corresponding to the beats in the target audio and the number-of-beats level corresponding to the beats in the candidate song audio, and further obtain the level difference between the number-of-beats level of the target audio and the number-of-beats level of the candidate song audio; if the level difference is less than the preset difference threshold, determine the candidate song audio as the original song audio of the target audio, and if the level difference is greater than or equal to the preset difference threshold, determine that the recall of the original song audio of the target audio fails. In the present application, by comparing the level difference between the number-of-beats levels of the target audio and the candidate song audio and comparing this level difference with the difference threshold, on the one hand, the degree of difference in rhythm between the audio content of the target audio and the audio content of the candidate song audio can be determined, so as to identify whether the target audio contains content unrelated to the candidate song audio and avoid misidentifying the candidate song audio as the original song audio of the target audio. On the other hand, it can avoid misfiltering caused by directly detecting the difference in the number of beats and improve the accuracy of the recall result of the original song audio. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 is a schematic flowchart of an audio recognition method in an embodiment;
[0047] Figure 2 is a schematic flowchart of the steps for determining the level difference of the number of beats in an embodiment;
[0048] Figure 3Schematic flowchart of steps for training a beat information recognition model in an embodiment;
[0049] Figure 4 Schematic flowchart of steps for obtaining candidate song audio in an embodiment;
[0050] Figure 5 Schematic flowchart of another audio recognition method in an embodiment;
[0051] Figure 6 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0052] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.
[0053] In one embodiment, as Figure 1 shown, an audio recognition method is provided. In this embodiment, it is exemplified that the method is applied to a server. It can be understood that the method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. Among them, the server can be implemented by an independent server or a server cluster composed of multiple servers; the terminal can be but is not limited to various personal computers, laptop computers, smart phones, tablet computers, etc.
[0054] In this embodiment, the method includes the following steps:
[0055] S101, obtain candidate song audio obtained after performing song matching on the target audio.
[0056] Among them, the target audio can be an audio to be recognized for cover singing and retrieve the corresponding original song audio. Exemplarily, the target audio can include a recorded audio uploaded by a user, such as the target audio can be read from an audio file or a video file pre-recorded by the user, or a real-time uploaded audio stream; of course, it can also include audio obtained by other means, such as a song audio downloaded from the network.
[0057] The candidate song audio can be an audio with a melody, such as a song containing lyrics, or a pure music without lyrics, such as an accompaniment.
[0058] In practical applications, song matching can be performed on the target audio. Song matching can refer to obtaining a song audio associated with the song content and the target audio, such as a song audio associated with the target audio in terms of lyrics or melody, so as to obtain candidate song audio.
[0059] S102. Determine the beat quantity level corresponding to the beats in the target audio, and determine the beat quantity level corresponding to the beats in the candidate song audio.
[0060] Among them, the beat can also be called the beat point, which can refer to the connection point between the previous beat and the next beat in the melody beat. In this embodiment, the beat quantity can be understood as the quantity corresponding to the beats in the melody.
[0061] For different beat quantities, hierarchical division can be performed in advance. Different beat quantities are divided into different intervals, and corresponding beat quantity levels are set for each interval to obtain multiple beat quantity levels. Exemplarily, the beat quantity level can be positively correlated with the beat quantity, that is, as the beat quantity increases, the beat quantity level also rises accordingly. For example, if the beat quantity is less than 5, the beat quantity level can be set to level 0; if the beat quantity is in the range of [5, 10), it is level 1; in the range of [10, 15), it is level 2; in the range of [15, 20), it is level 3; if the beat quantity is greater than 20, it can be set to level 4. Of course, in some other examples, the beat quantity level can also be negatively correlated with the beat quantity.
[0062] In this step, after obtaining the candidate song audio associated with the target audio, the beat quantity level corresponding to the beats in the target audio and the beat quantity level of the beats in the candidate song audio can be determined respectively. In some optional embodiments, the beat quantity level of the audio can be recognized by a pre-trained model, or, after the beats in the audio are recognized, according to the quantity interval to which the beats in the audio belong, the beat quantity level corresponding to this quantity interval can be used as the beat quantity level of the audio.
[0063] S103. Obtain the level difference between the beat quantity level of the target audio and the beat quantity level of the candidate song audio.
[0064] After determining the beat quantity levels of the target audio and the candidate song audio respectively, the beat quantity levels of the two can be compared to determine the level difference between the target audio and the candidate song audio in terms of the beat quantity level. For example, the level difference can be obtained based on the difference between the beat quantity level of the target audio and the beat quantity level of the candidate song audio.
[0065] S104. If the level difference is less than the preset difference threshold, determine the candidate song audio as the original singer song audio of the target audio.
[0066] S105. If the level difference is greater than or equal to the preset difference threshold, determine that the recall of the original singer song audio of the target audio fails.
[0067] As an example, the original song audio can be understood as the song audio that has not been adapted, such as the audio of the initial version of the song released publicly. In actual applications, users can perform one or more adaptation processes on the original song audio, such as adjusting the melody, lyrics, or music style, so as to obtain a new song audio. The new song audio can also be called the cover audio or adapted audio of the original song audio.
[0068] In specific implementation, for the case where the target audio contains human voices, there is information that, although it is human voice, is actually irrelevant to the song content. For example, for the target audio containing speech, if the speech content is similar to the lyrics content of part of a song, such as if the target audio involves a human voice reciting "Sing of our dear motherland, from now on towards prosperity and strength", then when performing song matching, there may be a situation where "Sing of the Motherland" is matched as the candidate song audio. However, the target audio may actually be just a recording of a poem recitation. The method of matching according to the lyrics content is likely to wrongly take the candidate song audio containing the same or similar lyrics content as the original song audio of the target audio containing irrelevant human voices.
[0069] Another example is that when performing song matching based on the pitch sequence (such as melody), although multiple pitches in the target audio can be matched with multiple pitches in the candidate song audio, the pitches in the target audio may be caused by noise and are not the melody of the song itself, that is, the target audio may not contain the melody of the candidate song audio or their melodies are not similar. In this case, it is also easy to wrongly take the candidate song audio as the original song audio of the target audio.
[0070] For the case where the target audio contains speech that is irrelevant to the song content, it can be understood that during the speech, the speaker mainly breaks sentences according to words and speaks, and usually does not have a sense of rhythm or has a very weak sense of rhythm. The lyrics in a song often match the rhythm of the song melody. For example, the position of word stress or the position of breathing during singing will match the beat positions of the song melody. Therefore, there will be a large difference in the density of beats between the speech that is irrelevant to the song and the density of beats when the user sings the lyrics. For the case where the target audio contains noise, due to the higher randomness of the noise, the beats in the audio will be sparser, and there will also be a large difference in the density of beats from the song audio with a melody.
[0071] In this regard, after obtaining the candidate song audio of the target audio, the present application can further obtain the level difference between the target audio and the candidate song audio in terms of the number of beat levels, and determine whether the level difference is less than a preset difference threshold. Since the number of beat levels can characterize the density of beats in the melody, if the level difference is less than the preset difference threshold, it can be determined that the number of beats in the target audio and the candidate song audio is similar. While the song content (such as lyrics or pitch sequence) matches, it is determined that the rhythm or rhyme of the two is similar, so that the candidate song audio can be determined as the original song audio of the target audio.
[0072] If the level difference is greater than or equal to the preset difference threshold, it can be determined that the difference in the density of beats between the target audio and the candidate song audio is large, and there is a significant difference in the rhythm and rhyme of the two. Then, the current candidate song audio can be filtered and excluded, and the candidate song audio is not used as the original song audio of the target audio, and it is determined that the recall of the original song audio of the target audio fails.
[0073] In some examples, a failure prompt for recalling the original song can be displayed to indicate that the original song audio of the target audio has not been obtained currently. The irrelevant vocals and noises in the target audio may be caused by factors such as a poor recording environment, accidental triggering by the user, or random recording. By returning a failure prompt for recalling the original song to the user, the user can be guided to record correctly and improve the recording quality of the target audio.
[0074] In addition, in some embodiments, although it is also possible to directly compare the difference in the number of beats between the target audio and the candidate song audio, this method has a low tolerance and is prone to misfiltering of the candidate song audio (for example, the candidate song audio is indeed the original song audio of the target audio, but due to incorrect lyrics singing rhythm during the recording of the target audio by the user, there is a difference in the number of beats). However, the present application compares the level difference between the target audio and the candidate song audio in terms of the number of beat levels, rather than directly comparing the difference in the number of beats between the target audio and the candidate song audio, which can avoid misfiltering caused by beat detection errors and improve the accuracy of recalling the original song audio.
[0075] In this embodiment, candidate song audio obtained after song matching for the target audio can be acquired, the beat quantity level corresponding to the beats in the target audio and the beat quantity level corresponding to the beats in the candidate song audio can be determined, and further, the level difference between the beat quantity level of the target audio and the beat quantity level of the candidate song audio can be obtained; if the level difference is less than a preset difference threshold, the candidate song audio is determined as the original song audio of the target audio, and if the level difference is greater than or equal to the preset difference threshold, it is determined that the recall of the original song audio of the target audio fails. In this application, by comparing the level difference between the beat quantity levels of the target audio and the candidate song audio and comparing this level difference with the difference threshold, on the one hand, the difference degree of the audio content rhythm between the target audio and the candidate song audio can be determined, so as to identify whether the target audio contains content unrelated to the candidate song audio and avoid wrongly taking the candidate song audio as the original song audio of the target audio. On the other hand, it can avoid the false filtering caused by directly detecting the beat quantity difference and improve the accuracy of the original song recall result.
[0076] In one embodiment, determining the beat quantity level corresponding to the beats in the target audio and determining the beat quantity level corresponding to the beats in the candidate song audio may include the following steps:
[0077] The target audio is divided into multiple target audio segments with a preset duration, and the beat quantity level corresponding to the beats in each target audio segment is determined; and the candidate song audio is divided into multiple candidate song audio segments with a preset duration, and the beat quantity level corresponding to the beats in each candidate song audio segment is determined.
[0078] In a specific implementation, the target audio and the candidate song audio can be divided into multiple audio segments according to a preset duration span, obtaining multiple target audio segments corresponding to the target audio and candidate song audio segments.
[0079] It can be understood that the beat quantity can change in different segments of the audio. In this step, the beat quantity level corresponding to the beats in each target audio segment and each candidate song audio segment can be determined.
[0080] For the target audio, based on the beat quantity levels of the multiple target audio segments respectively, the beat quantity level of the target audio at different playback progressions can be obtained; for the candidate song audio, based on the beat quantity levels of the multiple candidate song audio segments respectively, the beat quantity level of the candidate song audio segment at different playback progressions can be obtained.
[0081] In this embodiment, on the one hand, by dividing the target audio and the candidate song audio into multiple audio segments of a preset duration respectively, it is possible to compare the number of beat levels of beats within the same duration subsequently, improving the comparability of the number of beat levels. On the other hand, by determining the number of beat levels of each audio segment, the density of beats of the target audio and the candidate song audio in different time segments can be measured in detail, thereby effectively increasing the accuracy of the finally determined level difference.
[0082] Correspondingly, after obtaining the number of beat levels of each target audio segment and each candidate song audio segment, as Figure 2 shown, S103 obtaining the level difference between the number of beat levels of the target audio and the number of beat levels of the candidate song audio may include the following steps:
[0083] S201, determining the segment level difference between the number of beat levels of each target audio segment and the corresponding number of beat levels of each candidate song audio segment.
[0084] Specifically, for each target audio segment of the target audio, a corresponding candidate song audio segment can be determined in the candidate song audio. For example, the candidate song audio segment with a time progress matching a target audio segment can be used as the candidate song audio segment corresponding to this target audio segment.
[0085] Furthermore, after obtaining the number of beat levels of each target audio segment and the number of beat levels of each candidate song audio segment, for each target audio segment, the level difference between the number of beat levels of this target audio segment and the number of beat levels of the corresponding candidate song audio segment can be obtained as the segment-level number of beat level difference, and this level difference can be called the segment level difference.
[0086] S202, determining the level difference between the number of beat levels of the target audio and the candidate song audio according to multiple segment level differences.
[0087] After obtaining multiple segment level differences, the level difference of the beat count levels of the target audio and the candidate song audio can be determined by synthesizing the multiple segment level differences. Specifically, for example, the multiple segment level differences can be summed up, and the sum result can be used as the level difference of the beat count levels of the target audio and the candidate song audio. For example, for a 15s target audio and candidate song audio, they can be divided according to a preset duration of 3s, and 5 audio segments can be obtained respectively. If the beat count levels of the respective target audio segments of the target audio are 3, 0, 4, 2, 1 in sequence, and the beat count levels of the respective candidate song audio segments of the candidate song audio are 1, 0, 3, 1, 2 in sequence, then the level difference is 2 + 0 + 1 + 1 + 1 = 5. Of course, in some other embodiments, the largest segment level difference among the multiple segment level differences can also be used as the level difference between the target audio and the candidate song audio. When the largest segment level difference is greater than the preset difference threshold, the candidate song audio will not be recalled as the original song audio.
[0088] In this embodiment, the segment level differences between the target audio and the candidate song audio in different audio segments can be determined in a refined manner, the comparability of the beat count levels of the target audio and the candidate song audio can be improved, and the accuracy of the final obtained level difference result between the audios can be increased.
[0089] In one embodiment, S102 determining the beat count levels corresponding to the beats in the target audio and, determining the beat count levels corresponding to the beats in the candidate song audio may include the following steps:
[0090] Inputting the audio features corresponding to the target audio into the trained beat information recognition model to obtain the beat count level of the target audio output by the beat information recognition model; and, inputting the audio features corresponding to the candidate song audio into the beat information recognition model to obtain the beat count level of the candidate song audio output by the beat information recognition model.
[0091] Exemplarily, the audio features of the target audio and / or the candidate song audio may include at least one of the following: MFCC (Mel Frequency Cepstrum Coefficient) features, CQT features extracted based on the CQT algorithm (referring to a filter bank with center frequencies distributed exponentially, different filter bandwidths, but a constant Q ratio of the center frequency to the bandwidth), and HPCP (Harmonic Pitch Class Profile) features.
[0092] In specific implementation, the beat information recognition model can be pre-trained, and the beat information recognition model can be trained based on multiple audios with beat count levels.
[0093] After obtaining the target audio and the candidate song audio, for the target audio, the audio features of the target audio can be obtained and input into the beat information recognition model, and the model determines the beat number level of the target audio based on the input audio features of the target audio; correspondingly, for the candidate song audio, the audio features of the candidate song audio can be input into the beat information recognition model, and the beat number level output by the beat information recognition model is obtained. By inputting the audio features of the target audio and the candidate song audio into the beat information recognition model respectively, the beat number levels of each can be quickly obtained.
[0094] In one embodiment, as Figure 3 shown, the beat information recognition model can be trained through the following steps:
[0095] S301, Obtain a plurality of sample audios including speaker corpus audio and / or noise audio.
[0096] In practical applications, a plurality of sample audios for training the beat information recognition model can be obtained. The plurality of sample audios include not only song audio but also speaker corpus audio or noise audio. Among them, the speaker corpus audio contains speech or dialogue content unrelated to the song content, and the noise audio can contain disordered noise without melody rhythm.
[0097] S302, Based on the plurality of sample audios and the beat number level labels of each sample audio, perform supervised training on the beat information recognition model to be trained.
[0098] As an example, the beat information recognition model to be trained can be a neural network model, such as a resnet convolutional neural network.
[0099] After obtaining the plurality of sample audios, the beat number level corresponding to each sample audio can be pre-annotated, and the annotated beat number level is used as the beat number level label of the corresponding sample audio. Furthermore, the beat information recognition model can be supervised and trained using the plurality of sample audios and the beat number level labels of each sample audio. Specifically, for example, after inputting the audio features of the sample audio into the beat information recognition model to be trained, the difference value between the predicted beat number level output by the beat information recognition model and the preset beat number level label can be determined, and the model parameters of the beat information recognition model can be adjusted according to this difference value.
[0100] S303, When the training end condition is met, obtain the trained beat information recognition model.
[0101] When the training end condition is met, for example, when the number of model iterations reaches a preset number or the difference value between the predicted beat count level output by the model and the preset beat count level label is less than a threshold value, the current beat information recognition model can be regarded as a trained beat information recognition model.
[0102] In this embodiment, by using the speaker corpus audio or noise audio irrelevant to the song content as the sample audio to perform supervised training on the beat information recognition model, the model can accurately identify the beat count levels of the song audio or the human voice and noise irrelevant to the song.
[0103] In one embodiment, as Figure 4 shown, S101 to obtain the candidate song audio obtained after matching the target audio with a song may include the following steps:
[0104] S401, determine the lyrics of the target audio, and determine the lyric similarity between the lyrics of the target audio and the lyrics of each song audio in the audio library.
[0105] In a specific implementation, after obtaining the target audio, the audio features of the target audio can be obtained, such as MFCC features. Then, the audio features of the target audio can be input into the trained singing voice recognition model to obtain the lyric recognition result output by the singing voice recognition model, and the lyrics of the target audio can be obtained.
[0106] After obtaining the lyrics of the target audio, the recognized lyrics can be retrieved in the audio library to determine the lyric similarity between the lyrics of the target audio and the lyrics of each song audio in the audio library. Specifically, the edit distance between the lyrics of the target audio and the lyrics of each song audio in the audio library can be obtained, and the lyric similarity between the lyrics of the target audio and other lyrics can be determined through this edit distance. Among them, the edit distance refers to the minimum number of single-character edit operations (such as insertion, deletion, or replacement) required to convert one word <w1, w2> into another word w2 between two words, and the edit distance is negatively correlated with the lyric similarity.
[0107] S402, if there is a song audio with the largest lyric similarity and greater than the first similarity threshold, then determine this song audio as the candidate song audio of the target audio.
[0108] After obtaining multiple lyric similarities, the largest lyric similarity can be determined, and it can be judged whether the largest lyric similarity is also greater than the preset first similarity threshold; exemplarily, if the edit distance is used as the lyric similarity, it can be judged whether the edit distance is less than the preset edit distance threshold. If so, then this song audio can be determined as the candidate song audio of the target audio.
[0109] In this embodiment, after obtaining the target audio, candidate song audios associated with the lyrics content and the target audio can be obtained through lyrics matching with low matching difficulty and high accuracy. Subsequently, a fallback strategy for identifying the difference in the number of beat points levels can be combined to further identify the candidate song audios obtained by lyrics matching, ensuring the speed of finding the original singer's song while being able to eliminate the interference caused by irrelevant vocals in the target audio and ensuring the accuracy of the recall result.
[0110] In one embodiment, after determining the lyrics similarity between the lyrics information and the lyrics of multiple song audios in the audio library, the following steps may further be included:
[0111] If the lyrics similarity between the lyrics of the target audio and the lyrics of each song audio in the audio library is less than the first similarity threshold, then determine the melody similarity between the melody of the target audio and the melody of each song audio in the audio library; if there is a song audio with the largest melody similarity and greater than the second similarity threshold, then determine this song audio as the candidate song audio of the target audio.
[0112] After obtaining the lyrics similarity between the lyrics of the target audio and the lyrics of each song audio, if each lyrics similarity is less than the first similarity threshold, it can be determined that the confidence level of lyrics matching is low, and the lyrics of the target audio do not match the lyrics of each song audio. Furthermore, the melody similarity between the melody of the target audio and the melody of each song audio in the audio library can be obtained.
[0113] After obtaining multiple melody similarities, the largest melody similarity can be determined, and it can be judged whether the largest melody similarity is greater than the second similarity threshold. If so, this song audio can be determined as the candidate song audio of the target audio. If not, it can be determined that none of the song audios in the audio library can match the target audio, and a failure prompt for recalling the original singer's song can be returned.
[0114] In this embodiment, after the lyrics matching fails, the melody of the target audio can be further used for song matching. For the candidate song audios related to the melody, a fallback strategy for identifying the difference in the number of beat points levels can be combined to further identify the candidate song audios obtained by melody matching, eliminating the false recall caused by the noise interference in the target audio and ensuring the accuracy of the recall result.
[0115] In one embodiment, the determining the similarity between the melody of the target audio and the melody of each song audio in the audio library includes:
[0116] Obtain the melody features of the target audio, and determine the cosine distance between the melody features of the target audio and the melody features of each song audio in the audio library; based on the cosine distance, determine the similarity between the melody of the target audio and the melody of each song audio in the audio library.
[0117] In practical applications, the audio features of the target audio can be input into the melody feature extraction model to extract k-dimensional melody features (such as embedding features).
[0118] Then, the melody features of the target audio can be input into the melody feature retrieval library for retrieval. The retrieval library can store the melody features corresponding to each song audio in the audio library. During the retrieval process, the cosine distance between the melody features of the target audio and each melody feature in the retrieval library can be calculated, and based on the cosine distance, the melody similarity between the melody of the target audio and the melody of the song audio can be obtained. Among them, the cosine distance is negatively correlated with the melody similarity. By calculating the cosine distance between the melody features, the melody similarity between the target audio and each song audio in the audio library can be quickly obtained.
[0119] To enable those skilled in the art to better understand the above steps, the following provides an exemplary illustration of the embodiments of the present application through an example, but it should be understood that the embodiments of the present application are not limited thereto.
[0120] In specific implementation, the terminal can send a request for identifying the original singer song of the target audio to the server, as Figure 5 shown. In response to this request, the server can perform song matching on the target audio to obtain the candidate song audio of the target audio. During the song matching process, the server can obtain the audio features of the target audio. For example, the target audio can be divided into multiple target audio segments, and the MFCC features can be extracted in units of segments. Then, the audio features can be input into the singing voice recognition model to obtain the lyrics of the target audio recognized by the model, and the song audio with similar or identical lyrics can be retrieved in the audio library (i.e., Figure 5 the matching result 1 in). Then, when the confidence level is higher than the threshold, this song audio can be used as the candidate song audio. If the confidence level of the song audio obtained based on the lyrics matching is lower than the threshold, the extracted audio features can be input into the melody feature extraction model, and the melody features (such as embedding features) of each target audio segment of the target audio can be obtained based on the output result of the model, and the song audio with the same or similar melody can be retrieved based on this melody feature (i.e., Figure 5 the matching result 2 in). If the confidence level is higher than the threshold, this song audio can be used as the candidate song audio, otherwise, it can be determined that the recall of the original singer song of the target audio fails.
[0121] After obtaining the candidate song audio, it can be sliced and audio features can be extracted. Then, based on the audio features of the candidate song audio, the beat count level (which can also be called the beat level) of each candidate song audio segment of the candidate song audio can be determined. Moreover, for the target audio, the target audio can also be sliced and audio feature extraction processing can be performed to obtain the beat count level of each target audio segment. Then, based on the beat count level of each candidate song audio segment and the beat count level of each target audio segment, the level difference between the target audio and the candidate song audio can be determined.
[0122] When the level difference is less than the difference threshold, it can be determined that the confidence that the candidate song audio is the original singer song audio is higher than the threshold, and then the corresponding recall result can be obtained. Among them, the candidate song audio obtained based on lyric matching is recall result 1; if the candidate song audio based on lyric matching is not the original singer song audio, the candidate song audio obtained based on melody matching can be processed in the same way to obtain recall result 2. If the confidence of this candidate song audio is still lower than the threshold, the recall of the original singer song audio fails.
[0123] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0124] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 6As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store audio data of song audio. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. The computer program, when executed by the processor, implements an audio recognition method.
[0125] Those skilled in the art can understand that Figure 6 the structure shown in [the figure] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0126] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:
[0127] Obtain the candidate song audio obtained after performing song matching on the target audio;
[0128] Determine the number-of-beats level corresponding to the beats in the target audio, and determine the number-of-beats level corresponding to the beats in the candidate song audio;
[0129] Obtain the level difference between the number-of-beats level of the target audio and the number-of-beats level of the candidate song audio;
[0130] If the level difference is less than a preset difference threshold, determine the candidate song audio as the original song audio of the target audio;
[0131] If the level difference is greater than or equal to the preset difference threshold, determine that the recall of the original song audio of the target audio fails.
[0132] In one embodiment, when the processor executes the computer program, it also implements the steps in the above-mentioned other embodiments.
[0133] In one embodiment, a computer program product is provided, including a computer program which, when executed by a processor, implements the following steps:
[0134] Obtain candidate song audio obtained after performing song matching on target audio;
[0135] Determine the number of beat levels corresponding to the beats in the target audio, and determine the number of beat levels corresponding to the beats in the candidate song audio;
[0136] Obtain the level difference between the number of beat levels of the target audio and the number of beat levels of the candidate song audio;
[0137] If the level difference is less than a preset difference threshold, determine the candidate song audio as the original song audio of the target audio;
[0138] If the level difference is greater than or equal to the preset difference threshold, determine that the recall of the original song audio of the target audio fails.
[0139] In one embodiment, when the computer program is executed by a processor, it also implements the steps in the above-mentioned other embodiments.
[0140] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0141] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include Read-Only Memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0142] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0143] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. An audio recognition method, characterized in that, The method includes: Obtaining candidate song audios obtained by performing song matching on a target audio; Determining the beat count levels of multiple target audio segments in the target audio, and determining the beat count levels of multiple candidate song audio segments in the candidate song audio; Determining the segment level difference between the beat count level of each target audio segment and the corresponding beat count level of each candidate song audio segment; Determining the level difference in the beat count levels between the target audio and the candidate song audio based on the multiple segment level differences; If the level difference is less than a preset difference threshold, determining the candidate song audio as the original singer song audio of the target audio; If the level difference is greater than or equal to the preset difference threshold, determining that the recall of the original singer song audio of the target audio fails.
2. The method according to claim 1, characterized in that, The determining the beat count levels of multiple target audio segments in the target audio, and determining the beat count levels of multiple candidate song audio segments in the candidate song audio includes: Dividing the target audio into multiple target audio segments of a preset duration, and obtaining the beat count level corresponding to the beats in each target audio segment; and Dividing the candidate song audio into multiple candidate song audio segments of the preset duration, and obtaining the beat count level corresponding to the beats in each candidate song audio segment.
3. The method according to claim 1, wherein The determining the beat count level corresponding to the beats in the target audio, and determining the beat count level corresponding to the beats in the candidate song audio includes: Inputting the audio features corresponding to the target audio into a trained beat information recognition model to obtain the beat count level of the target audio output by the beat information recognition model; and Inputting the audio features corresponding to the candidate song audio into the beat information recognition model to obtain the beat count level of the candidate song audio output by the beat information recognition model.
4. The method according to claim 3, wherein The beat information recognition model is trained through the following steps: Obtaining multiple sample audios including speaker corpus audios and / or noise audios; Performing supervised training on the beat information recognition model to be trained based on the multiple sample audios and the beat count level label of each sample audio; When the training end condition is satisfied, obtaining the trained beat information recognition model.
5. The method according to any one of claims 1 to 4, characterized in that, The obtaining the candidate song audio obtained by performing song matching on a target audio includes: Determining the lyrics of the target audio, and determining the lyric similarity between the lyrics of the target audio and the lyrics of each song audio in the audio library; If there is a song audio with the maximum lyric similarity and greater than a first similarity threshold, determining the song audio as the candidate song audio of the target audio.
6. The method according to claim 5, characterized in that, After determining the lyric similarity between the lyrics and multiple song audios in the audio library, it further includes: If the lyric similarity between the lyrics of the target audio and the lyrics of each song audio in the audio library is less than the first similarity threshold, determining the melody similarity between the melody of the target audio and the melody of each song audio in the audio library; If there is a song audio with the maximum melody similarity that is greater than the second similarity threshold, then determine this song audio as the candidate song audio for the target audio.
7. The method according to claim 6, characterized in that, The determination of the melody similarity between the target audio and the melody of each song audio in the audio library includes: Obtain the melody features of the target audio and determine the cosine distance between the melody features of the target audio and the melody features of each song audio in the audio library; Based on the cosine distance, determine the melody similarity between the target audio and the melody of each song audio in the audio library.
8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 7.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method described in any one of claims 1 to 7.
Citation Information
Patent Citations
Cover version identification method and device and computer storage medium
CN111445923A