Audio recognition method and device, electronic device, and computer-readable storage medium

By segmenting the audio to be identified and matching the melody information, the copyright identification problem of cover songs and original songs is solved, and efficient audio similarity judgment is achieved.

CN115346515BActive Publication Date: 2025-09-16GUANGZHOU HUYA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210358586.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-06
Publication Date
2025-09-16
Estimated Expiration
2042-04-06

AI Technical Summary

Technical Problem

Existing technologies are unable to effectively identify the similarities between cover songs and original songs, making it difficult to determine copyright infringement.

Method used

By segmenting the audio to be recognized, extracting melody information, matching it with candidate audio segments in the audio database, and combining the time information to determine similar audio.

Benefits of technology

The accuracy and efficiency of audio recognition have been improved, and it can quickly identify the similarities between cover songs and original songs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115346515B_ABST
    Figure CN115346515B_ABST
Patent Text Reader

Abstract

The present application discloses an audio recognition method and device, an electronic device, and a computer-readable storage medium. The audio recognition method includes: obtaining the audio to be recognized; segmenting the audio to be recognized with a preset step size to obtain a number of audio segments to be recognized; extracting the melody information of each of the audio segments to be recognized; finding all candidate audio segments with the same melody information as each of the audio segments to be recognized from an audio database; and determining a target audio similar to the audio to be recognized based on the moment information of each of the audio segments to be recognized and the moment information of all the candidate audio segments corresponding to it. The above scheme can identify target audio similar to the audio to be recognized and can improve the recognition accuracy of similar audio.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of audio recognition, and in particular to an audio recognition method and device, an electronic device, and a computer-readable storage medium. Background Art

[0002] As people's living standards improve, they have more time and interest to enjoy music. Consequently, more and more people are covering their favorite songs and uploading them online to share. However, given the difficulties faced by creators, most songs are subject to copyright issues, prohibiting cover and commercial distribution without the song owner's permission. Due to the widespread distribution of music, it is difficult to determine copyright infringement solely through human review. Therefore, how to automatically identify cover songs and original songs to determine whether they infringe copyright has become a pressing issue. Summary of the Invention

[0003] The present application provides an audio recognition method and device, an electronic device, and a computer-readable storage medium, which can identify target audio similar to the audio to be recognized and improve the recognition accuracy of similar audio.

[0004] In a first aspect, the present application provides an audio recognition method, which includes: obtaining audio to be recognized; segmenting the audio to be recognized with a preset step size to obtain a plurality of audio segments to be recognized; extracting melody information of each of the audio segments to be recognized; finding all candidate audio segments with the same melody information as each of the audio segments to be recognized from an audio database; and determining a target audio similar to the audio to be recognized based on the moment information of each of the audio segments to be recognized and the moment information of all the candidate audio segments corresponding to it.

[0005] The step of obtaining the audio to be recognized includes performing background sound separation processing on the initial audio signal to be recognized to obtain the audio to be recognized that is derived from a human voice track of a speaking sound.

[0006] The step of segmenting the audio to be identified with a preset step size to obtain a plurality of audio segments to be identified includes: determining the preset step size based on the time length of a preset melody point and a preset hyperparameter of the melody point; and segmenting the audio to be identified from the starting moment according to the preset step size to obtain a plurality of audio segments to be identified.

[0007] The extracting of melody information of each audio segment to be identified includes: extracting a point with the largest pitch from the time length of each melody point by a landmark algorithm as the melody information of the melody point; and obtaining the melody information of the corresponding audio segment to be identified based on the melody information of all melody points in each audio segment to be identified.

[0008] The melody information includes pitch information; and finding all candidate audio segments with the same melody information as each of the audio segments to be identified from the audio database includes: for each of the audio segments to be identified, finding all candidate audio segments corresponding to the audio segment to be identified from the audio database; and the pitch information of all candidate audio segments corresponding to the audio segment to be identified and the pitch information of the audio segment to be identified meet a preset condition.

[0009] The preset condition includes that a difference between pitch information of all candidate audio segments corresponding to the audio segment to be identified and the pitch information of the audio segment to be identified is less than a preset threshold.

[0010] The method of determining the target audio similar to the audio to be identified based on the moment information of each audio segment to be identified and the moment information of all the candidate audio segments corresponding to the audio segment includes: obtaining the time difference between each audio segment to be identified and all the candidate audio segments corresponding to the audio segment according to the moment information of each audio segment to be identified and the moment information of all the candidate audio segments corresponding to the audio segment; and determining the target audio similar to the audio to be identified based on the time difference between each audio segment to be identified and all the candidate audio segments corresponding to the audio segment; wherein the target audio has at least one candidate audio segment corresponding to each audio segment to be identified, and among the time differences between each audio segment to be identified and all the candidate audio segments corresponding to the audio segment, there is at least one time difference that exists for all the audio segments to be identified.

[0011] In a second aspect, the present application provides an audio recognition device, comprising: an acquisition module for acquiring audio to be recognized; a segmentation module for segmenting the audio to be recognized with a preset step size to obtain a plurality of audio segments to be recognized; an extraction module for extracting melody information of each of the audio segments to be recognized; a comparison module for finding all candidate audio segments having the same melody information as each of the audio segments to be recognized from an audio database; and a determination module for determining a target audio similar to the audio to be recognized based on the moment information of each of the audio segments to be recognized and the moment information of all the candidate audio segments corresponding thereto.

[0012] A third aspect of the present application provides an electronic device, comprising a memory and a processor coupled to each other, wherein the processor is configured to execute program instructions stored in the memory to implement the audio recognition method in the first aspect.

[0013] In a fourth aspect, the present application provides a computer-readable storage medium having program instructions stored thereon. When the program instructions are executed by a processor, the audio recognition method in the first aspect is implemented.

[0014] In the above scheme, after obtaining the audio to be recognized, the present application segments the audio to be recognized with a preset step size to obtain a number of audio segments to be recognized. By extracting the melody information of each audio segment to be recognized, all candidate audio segments with the same melody information as the respective audio segments to be recognized can be found from the audio database. Therefore, based on the moment information of each audio segment to be recognized and the moment information of all corresponding candidate audio segments, the target audio similar to the audio to be recognized can be determined. By dividing the audio to be recognized into a number of audio segments to be recognized, finding all candidate audio segments with the same melody information as the respective audio segments to be recognized from the audio database, and considering that the candidate audio segments similar to the respective audio segments to be recognized are continuous sound sequences in the target audio, the target audio similar to the audio to be recognized can be identified with high recognition accuracy by referring to the moment information of each audio segment to be recognized and the moment information of all corresponding candidate audio segments. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 This is a flowchart of an embodiment of the audio recognition method of the present application;

[0016] Figure 2a and Figure 2b It is a schematic diagram of spectrum data for background sound separation processing of audio signals;

[0017] Figure 3 yes Figure 1 A flow chart of an embodiment of step S12;

[0018] Figure 4 yes Figure 1 A flow chart of an embodiment of step S13;

[0019] Figure 5 is a schematic diagram of the spectrum data of the pitch range of the audio segment to be identified;

[0020] Figure 6 yes Figure 1 A flow chart of an embodiment of step S15;

[0021] Figure 7a and Figure 7b It is a schematic diagram showing a comparison between the time information of an audio segment to be identified and the time information of all corresponding candidate audio segments in an application scenario;

[0022] Figure 8It is a schematic diagram showing a target audio that is similar to the audio to be recognized in an application scenario;

[0023] Figure 9 This is a schematic diagram of the framework of an embodiment of the audio recognition device of the present application;

[0024] Figure 10 This is a schematic diagram of the framework of an embodiment of the electronic device of the present application;

[0025] Figure 11 This is a schematic diagram of a framework of an embodiment of a computer-readable storage medium of the present application. DETAILED DESCRIPTION

[0026] The following describes the embodiments of the present application in detail with reference to the accompanying drawings.

[0027] In the following description, for the purpose of explanation rather than limitation, specific details such as specific system structures, interfaces, and technologies are provided to facilitate a thorough understanding of the present application.

[0028] The terms "system" and "network" are often used interchangeably in this document. The term "and / or" is simply a description of a relationship between related objects. Three possible relationships exist. For example, A and / or B can exist in three situations: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates an "or" relationship between the related objects. Furthermore, "more" in this document refers to two or more than two.

[0029] See also Figure 1 , Figure 1 This is a flow chart of an embodiment of the audio recognition method of the present application. In this embodiment, the audio recognition method may include the following steps:

[0030] Step S11: Acquire the audio to be recognized.

[0031] In the embodiment of the present application, the audio to be recognized can be all or part of a song performed by a user. Using the audio recognition method of the present application, target audio similar to the audio to be recognized can be identified from a pre-set audio database. The pre-set audio database includes a large number of candidate audios. The tracks in the audio database can be obtained from the authorization and network of various music companies, which is not specifically limited in the embodiment of the present application.

[0032] In one embodiment, the above step S11 may specifically include: performing background sound separation processing on the initial audio signal to be recognized to obtain the audio to be recognized that is derived from a human voice track of a speaking sound.

[0033] Please combine Figure 2a and Figure 2b ,in, Figure 2a is a schematic diagram of the spectrum data of the initial audio signal to be identified, Figure 2b It is a schematic diagram of the spectrum data of the audio to be identified obtained after the background sound separation processing is performed on the initial audio signal to be identified. It can be understood that when the initial audio signal to be identified is obtained, the initial audio signal to be identified can be the audio and video released by the information publisher such as the anchor or blogger. Since the initial audio signal to be identified includes a human voice track and a background sound track, wherein the human voice track is the audio signal of the speaking voice in the initial audio signal to be identified, and the background sound track is the audio signal of the background sound in the initial audio signal to be identified, the background sound can be a sound other than the speaking voice, background music, etc. Therefore, it is necessary to separate the human voice and the background sound of the initial audio signal to be identified. By performing background sound separation processing on the initial audio signal to be identified, the audio to be identified originating from the human voice track of the speaking voice is obtained, so as to better extract the melody information of the human voice.

[0034] Step S12: segmenting the audio to be recognized with a preset step size to obtain a plurality of audio segments to be recognized.

[0035] It is understandable that the audio to be identified is generally a longer audio segment. If a melody information of the entire audio to be identified is extracted, and the target audio with similar melody information is searched for based on the melody information of the entire audio, it may take a long time, and the target audio may not be found due to situations such as some measures in the song are different while other measures are similar. Therefore, the audio to be identified can be segmented according to a preset step size to obtain several audio segments to be identified, so as to extract the shorter melody information corresponding to the audio segments to be identified and perform subsequent melody comparison.

[0036] In one embodiment, the time length of the audio to be recognized can be 5 to 20 seconds, for example, 10 seconds or 13 seconds. It is understandable that when the audio to be recognized is longer, the time required for subsequent melody comparison will be longer, but the accuracy of identifying similar target audio will be improved. When the audio to be recognized is shorter, the time required for subsequent melody comparison will be shorter, but the accuracy of identifying similar target audio will be reduced. Therefore, setting the time length of the audio to be recognized within an appropriate range can achieve a relatively balanced speed and accuracy in identifying similar target audio.

[0037] Please combine Figure 3 , Figure 3 yes Figure 1 FIG. 1 is a flow chart of an embodiment of step S12. In this embodiment, step S12 may specifically include:

[0038] Step S121: determining the preset step length according to the preset time length of the melody point and the preset super parameters of the melody point.

[0039] Step S122: Segment the audio to be recognized from the starting time according to the preset step size to obtain a plurality of audio segments to be recognized.

[0040] Specifically, a melody point is used as the minimum audio segment for extracting melody information. The duration of a melody point is 50ms. Therefore, the melody information of all audio in the audio database is stored as a single melody point. However, considering that when searching the audio database for audio with similar melody information using a single melody point, too many similar melody points can be obtained, which is not representative, a melody point hyperparameter FAN_VALUE can be pre-set to define the length of the audio melody information stored in the audio database. When FAN_VALUE = 10, it means that the melody of the audio is stored for 500ms at a time. For example, FAN_VALUE = 3 means that the melody information of the audio is stored in the audio database for 150ms. Therefore, the audio to be recognized also needs to be searched in the audio database for audio with similar melody information according to the length of the audio segment. Therefore, the preset step size can be determined based on the time length of the preset melody point and the hyperparameters of the preset melody point, and then the audio to be identified is segmented from the starting time according to the preset step size to obtain several audio segments to be identified, so as to compare and search for audio segments with similar melody information to the audio segment to be identified in the audio database.

[0041] Step S13: extracting melody information of each of the audio segments to be identified.

[0042] It is understandable that after the audio to be recognized is segmented to obtain a plurality of audio segments to be recognized, the corresponding melody information can be extracted for each audio segment to be recognized.

[0043] Please combine Figure 4 , Figure 4 yes Figure 1 FIG. 1 is a flow chart of an embodiment of step S13. In this embodiment, step S13 may specifically include:

[0044] Step S131: extracting a point with the largest pitch from the duration of each melody point through a landmark algorithm as the melody information of the melody point.

[0045] Step S132: obtaining the melody information of the corresponding audio segment to be identified based on the melody information of all melody points in each of the audio segments to be identified.

[0046] Specifically, pitch refers to the various pitches of various sounds, i.e., the height of a sound. It is understood that pitch can be divided into four levels: C2-B2 (corresponding to a frequency of 65-123 Hz), C3-B3 (corresponding to a frequency of 131-247 Hz), C4-B4 (corresponding to a frequency of 262-494 Hz), and C5-B5 (corresponding to a frequency of 523-989 Hz). The pitch range for male voices is between C2 and B4, and the pitch range for female voices is between C3 and B5. For a melody point, the pitch value of the point with the highest pitch during the duration of the melody point can be used as its melody information. Therefore, a landmark algorithm can be used to extract the point with the highest pitch energy during the duration of each melody point as the melody information at the corresponding moment of the melody point. Then, based on the melody information of all melody points in each audio segment to be identified, the melody information of the corresponding audio segment to be identified is obtained. In other words, the melody information of each audio segment to be identified is composed of the melody information of all melody points contained therein.

[0047] Step S14: Find all candidate audio segments having the same melody information as the audio segments to be identified from the audio database.

[0048] It is understandable that after obtaining the melody information of all audio segments to be identified, for each audio segment to be identified, all candidate audio segments having the same melody information as the audio segment to be identified can be found from the audio database.

[0049] In one embodiment, the melody information includes pitch information, and the above-mentioned step S14 may specifically include: for each of the audio segments to be identified, finding all candidate audio segments corresponding to the audio segment to be identified from the audio database; wherein the pitch information of all candidate audio segments corresponding to the audio segment to be identified and the pitch information of the audio segment to be identified meet a preset condition.

[0050] Specifically, audio melody information includes audio pitch information. Therefore, finding candidate audio segments similar to the audio segment to be identified from the audio database primarily involves comparing the pitch, i.e., the frequency, of the audio segment to be identified with the candidate audio segments. For a particular audio segment to be identified, the pitch information of the segment to be identified is compared with the pitch information of all candidate audio segments in the audio database. If the pitch information of a candidate audio segment meets a predetermined condition with the pitch information of the audio segment to be identified, the candidate audio segment is considered similar to the audio segment to be identified. Thus, all candidate audio segments similar to the audio segment to be identified are found, i.e., all candidate audio segments corresponding to the audio segment to be identified are found from the audio database. For example, a hyperparameter FAN_VALUE = 3 can be used, meaning that each pitch comparison is performed for three melody points, each melody point being 50 ms, i.e., a total of approximately 150 ms of audio segments are compared. For example, if the audio segment to be identified is 10 seconds long, a total of 10,000 / 150 = 67 comparisons are required for the audio segment to be identified. In this embodiment, a hash algorithm can be used to compare the pitch information of the audio segment to be identified with the pitch information of the candidate audio segment, calculate the corresponding difference value, and regard the two audio segments to be identified and the candidate audio segment whose difference value is less than a preset range as similar audio segments.

[0051] In a preferred embodiment, the preset condition includes that a difference between pitch information of all candidate audio segments corresponding to the audio segment to be identified and the pitch information of the audio segment to be identified is less than a preset threshold.

[0052] It is understandable that when the audio to be identified is a cover audio and the target audio is the corresponding original audio, considering that the cover may have inaccurate pitch, the pitch of a single melody point will be offset within a certain range. For example, when singing a cover, the frequency of the pitch at the moment corresponding to a melody point is 220, then when the pitch of the cover is between 210 and 230, it can be considered that the pitch of the cover at this moment is 220. Figure 5 As shown, Figure 5 is a schematic diagram of the spectrum data of the pitch range of the audio segment to be identified. Figure 5 The area shown in the box represents the offset range of the pitch. Therefore, if the pitch of the melody point in a candidate audio segment is within the box, it can be said that the melody point in the candidate audio segment is different from the pitch of the melody point in the candidate audio segment. Figure 5 That is, when the difference between the pitch information of a candidate audio segment and the pitch information of a to-be-recognized audio segment is less than a preset threshold, it can be considered that the candidate audio segment and the to-be-recognized audio segment are similar.

[0053] Step S15: determining a target audio similar to the audio to be recognized based on the time information of each audio segment to be recognized and the time information of all the candidate audio segments corresponding thereto.

[0054] It is understandable that when the originally continuous melodies of do, re, mi, fa, and so in the audio to be identified hit the 5th, 20th, 30th, 50th, and 120th seconds of a candidate audio respectively, it is obvious that the audio to be identified is not similar to the candidate audio. In the scenario of matching cover songs with original songs, they are not the same song, but just happen to have some of the same audio fingerprint segments. Therefore, after finding all candidate audio segments with the same melody information as each audio segment to be identified, considering that the candidate audio segments similar to each audio segment to be identified are continuous sound sequences in the target audio, it is necessary to refer to the moment information of each audio segment to be identified and the moment information of all corresponding candidate audio segments, align each audio segment to be identified and all corresponding candidate audio segments according to the time in the audio, and find all candidate audio segments that meet the moment continuity of the audio segment to be identified, so that the target audio similar to the audio to be identified can be identified with high recognition accuracy.

[0055] Through the above steps, after obtaining the audio to be recognized, the audio to be recognized is segmented at a preset step size to obtain several audio segments to be recognized. By extracting the melody information of each audio segment to be recognized, all candidate audio segments with the same melody information as the respective audio segments to be recognized can be found from the audio database. Based on the time information of each audio segment to be recognized and the time information of all corresponding candidate audio segments, target audio similar to the audio to be recognized can be determined. By dividing the audio to be recognized into several audio segments to be recognized, finding all candidate audio segments with the same melody information as the respective audio segments to be recognized from the audio database, and considering that the candidate audio segments similar to the respective audio segments to be recognized are continuous sound sequences in the target audio, target audio similar to the audio to be recognized can be identified with high recognition accuracy by referring to the time information of each audio segment to be recognized and the time information of all corresponding candidate audio segments.

[0056] Please combine Figure 6 , Figure 6 yes Figure 1 FIG. 1 is a flow chart of an embodiment of step S15. In this embodiment, step S15 may specifically include:

[0057] Step S151: obtaining the time difference between each audio segment to be identified and all the corresponding candidate audio segments according to the time information of each audio segment to be identified and the time information of all the corresponding candidate audio segments.

[0058] Step S152: Determine a target audio similar to the audio to be identified based on the time difference between each of the audio segments to be identified and all of the candidate audio segments corresponding to it; wherein the target audio has at least one candidate audio segment corresponding to each of the audio segments to be identified, and among the time differences between each of the audio segments to be identified and all of the candidate audio segments corresponding to it, there is at least one time difference that exists in all of the audio segments to be identified.

[0059] It is understandable that in order to solve the problem that the candidate audio segments that are similar to the audio segments to be identified are continuous sound sequences in the target audio, a statistics can be done. Since the moment information of each audio segment to be identified and the moment information of all the candidate audio segments corresponding to it are known, the time difference between each audio segment to be identified and all the candidate audio segments corresponding to it can be calculated, and then based on the time difference between each audio segment to be identified and all the candidate audio segments corresponding to it, the target audio similar to the audio to be identified can be determined. Specifically, when the audio to be identified is similar to the target audio, each audio segment to be identified in the audio to be identified has at least one similar candidate segment in the target audio, and the time difference between each audio segment to be identified and at least one candidate audio segment corresponding to it is equal, that is, the time continuity of each audio segment to be identified in the audio to be identified is the same as the time continuity of each candidate audio segment in the target audio.

[0060] Please combine Figure 7a and Figure 7b , Figure 7a and Figure 7b This is a schematic diagram showing the comparison of the time information of the audio segment to be identified and the time information of all corresponding candidate audio segments in an application scenario. Figure 7a As shown, Figure 7a The candidate audio in is the target audio of the audio to be recognized. Specifically, the left side of the colon is the audio segments to be recognized, and the numbers in the brackets represent the time information of the audio segments to be recognized. The right side of the colon is the audio segments of a candidate audio in the audio database, and the numbers in the brackets represent the time information of the corresponding candidate audio segments. Furthermore, the audio to be recognized is a cover song, and the target audio is a cover song. Figure 7a As can be seen in the figure, when the cover song on the left jumps, the original song on the right also jumps synchronously. It should be noted that the error in the middle jump is only a few hundred milliseconds. It can be understood that the two people breathe at the same time, but the breathing time is different, and then the melody is right again, or a certain part in the middle is seriously out of tune. Figure 7b middle, Figure 7bThe candidate audio in is not the target audio of the audio to be recognized. Specifically, the left side of the colon is the audio segments to be recognized, and the numbers in the brackets represent the time information of the audio segments to be recognized. The right side of the colon is the candidate audio segments of a candidate audio in the audio database, and the numbers in the brackets represent the time information of the corresponding candidate audio segments. Furthermore, the audio to be recognized is a cover song, and the candidate audio is other songs. Figure 7a It can be seen that when the time of the cover song jumps, there is an obvious break in the other songs. The time of the cover song jumps from 46 to 53 with only 7 frames in between, which is equivalent to about 350ms. The closest time point that the other songs on the right can be aligned to is 2938, which jumps from 2885 to 2938, which is equivalent to 2.5s. This is obviously unreasonable. Assuming that the cover songs are sung continuously, the cover song corresponds to the other songs when singing the previous melody, but it takes 2.5s for the next melody to correspond to the other songs again. Therefore, it can be found that as the target audio of the audio to be identified, for each audio segment to be identified, it has one or more corresponding candidate audio segments, and in the time difference between each audio segment to be identified and all its corresponding candidate audio segments, there is a time difference that exists for each audio segment to be identified.

[0061] Please combine Figure 8 , Figure 8 It is a display schematic diagram of identifying a target audio similar to the audio to be identified in an application scenario. As shown in the figure, for a cover song, the above-mentioned audio recognition method is used to obtain the corresponding audio to be identified by removing the background sound, and the melody of the audio to be identified is compared with the candidate audio of the 34 original songs in the audio database after removing the background sound. Then, the confidence between the audio to be identified and each candidate audio can be obtained, that is, the similarity probability between the audio to be identified and each candidate audio. For example, the confidence between the audio to be identified and "Birch Forest" is 0.7. These confidences are arranged from high to low. It is roughly estimated that the confidence between the audio to be identified and the candidate audio exceeds 0.6, which can basically determine that the audio to be identified is a cover of the corresponding candidate audio. Finally, it is determined that "Birch Forest" is the target audio of the audio to be identified, and the confidence that "Birch Forest" is the original song corresponding to the cover song is 0.91, that is, the accuracy rate of "Birch Forest" as the original song corresponding to the cover song is 0.91. The audio recognition method of the present application has high recognition accuracy. In addition, query_time is the time to query the audio database, and align_time is the time to query the melody and make a judgment. These two times will increase with the increase in the number of audios. It can be found that the query time for 34 tracks in the current audio database is 2 seconds. The audio recognition method of this application has a high recognition speed.

[0062] See also Figure 9 , Figure 9 Schematic diagram of the framework of an embodiment of the audio recognition device of the present application. The audio recognition device 90 includes an acquisition module 900, a segmentation module 902, an extraction module 904, a comparison module 906, and a determination module 908. The acquisition module 900 is used to acquire the audio to be recognized; the segmentation module 902 is used to segment the audio to be recognized with a preset step size to obtain a plurality of audio segments to be recognized; the extraction module 904 is used to extract the melody information of each audio segment to be recognized; the comparison module 906 is used to find all candidate audio segments with the same melody information as each audio segment to be recognized from the audio database; the determination module 908 is used to determine the target audio similar to the audio to be recognized based on the moment information of each audio segment to be recognized and the moment information of all the candidate audio segments corresponding to it.

[0063] In the above scheme, after the acquisition module 900 acquires the audio to be recognized, the segmentation module 902 segments the audio to be recognized at a preset step size to obtain a number of audio segments to be recognized. The extraction module 904 extracts the melody information of each audio segment to be recognized. The comparison module 906 then identifies all candidate audio segments from the audio database that have the same melody information as the audio segments to be recognized. The determination module 908 then determines target audio similar to the audio to be recognized based on the time information of each audio segment to be recognized and the time information of all corresponding candidate audio segments. By dividing the audio to be recognized into a number of audio segments to be recognized, identifying all candidate audio segments from the audio database that have the same melody information as the audio segments to be recognized, and considering that the candidate audio segments similar to the audio segments to be recognized are continuous sound sequences in the target audio, target audio similar to the audio to be recognized can be identified with high recognition accuracy by referring to the time information of each audio segment to be recognized and the time information of all corresponding candidate audio segments.

[0064] In some embodiments, the acquisition module 900 executes the step of acquiring the audio to be recognized, including: performing background sound separation processing on the initial audio signal to be recognized to obtain the audio to be recognized that is derived from the human voice track of the speaking sound.

[0065] In some embodiments, the segmentation module 902 performs the step of segmenting the audio to be identified with a preset step size to obtain a plurality of audio segments to be identified, including: determining the preset step size based on the time length of the preset melody point and the hyperparameters of the preset melody point; and segmenting the audio to be identified from the starting time according to the preset step size to obtain a plurality of audio segments to be identified.

[0066] In some embodiments, the extraction module 904 performs the step of extracting the melody information of each audio segment to be identified, specifically including: extracting a point with the largest pitch in the time length of each melody point through a landmark algorithm as the melody information of the melody point; and obtaining the melody information of the corresponding audio segment to be identified based on the melody information of all melody points in each audio segment to be identified.

[0067] In some embodiments, the melody information includes pitch information; the comparison module 906 performs the step of finding all candidate audio segments with the same melody information as each of the audio segments to be identified from the audio database, including: for each of the audio segments to be identified, finding all candidate audio segments corresponding to the audio segment to be identified from the audio database; wherein the pitch information of all candidate audio segments corresponding to the audio segment to be identified and the pitch information of the audio segment to be identified meet a preset condition.

[0068] In some embodiments, the preset condition includes that a difference between pitch information of all candidate audio segments corresponding to the audio segment to be identified and the pitch information of the audio segment to be identified is less than a preset threshold.

[0069] In some embodiments, the determination module 908 performs the step of determining a target audio similar to the audio to be identified based on the moment information of each audio segment to be identified and the moment information of all the candidate audio segments corresponding to it, including: obtaining the time difference between each audio segment to be identified and all the candidate audio segments corresponding to it according to the moment information of each audio segment to be identified and the moment information of all the candidate audio segments corresponding to it; determining the target audio similar to the audio to be identified based on the time difference between each audio segment to be identified and all the candidate audio segments corresponding to it; wherein the target audio has at least one candidate audio segment corresponding to each audio segment to be identified, and among the time differences between each audio segment to be identified and all the candidate audio segments corresponding to it, there is at least one time difference that exists corresponding to all the audio segments to be identified.

[0070] See also Figure 10 , Figure 10 This is a schematic diagram of the framework of an embodiment of an electronic device of the present application. Electronic device 100 includes a memory 101 and a processor 102 coupled to each other. Processor 102 is configured to execute program instructions stored in memory 101 to implement the steps of any of the aforementioned audio recognition method embodiments. In a specific implementation scenario, electronic device 100 may include, but is not limited to, a microcomputer or a server.

[0071] Specifically, the processor 102 is used to control itself and the memory 101 to implement the steps of any of the above-mentioned audio recognition method embodiments. The processor 102 can also be called a CPU (Central Processing Unit). The processor 102 may be an integrated circuit chip with signal processing capabilities. The processor 102 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 102 can be implemented by an integrated circuit chip.

[0072] See also Figure 11 , Figure 11 The computer-readable storage medium 110 stores program instructions 1100 that can be executed by a processor, and the program instructions 1100 are used to implement the steps of any of the above-mentioned audio recognition method embodiments.

[0073] In the several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation methods described above are only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation. For example, units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.

[0074] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0075] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0076] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of each embodiment method of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

Claims

1. An audio recognition method, characterized in that: The audio recognition method comprises: Get the audio to be recognized; Segmenting the audio to be recognized with a preset step length to obtain a plurality of audio segments to be recognized; wherein the preset step length is determined based on a preset duration of a melody point and a preset hyperparameter of the melody point; Extracting melody information of each of the audio segments to be identified; Find all candidate audio segments with the same melody information as the audio segments to be identified from the audio database; Determining target audio similar to the audio to be recognized based on the time information of each audio segment to be recognized and the time information of all corresponding candidate audio segments; specifically comprising: According to the time information of each audio segment to be identified and the time information of all the candidate audio segments corresponding to it, the time difference between each audio segment to be identified and all the candidate audio segments corresponding to it is obtained.

2. The audio recognition method according to claim 1, wherein: The step of obtaining the audio to be recognized includes: The initial audio signal to be recognized is subjected to background sound separation processing to obtain the audio signal to be recognized which is derived from a human voice track of a speaking sound.

3. The audio recognition method according to claim 1, wherein: The step of segmenting the audio to be recognized with a preset step size to obtain a plurality of audio segments to be recognized includes: The audio to be recognized is segmented from the starting time of the audio according to the preset step size to obtain a plurality of audio segments to be recognized.

4. The audio recognition method according to claim 1, wherein: The extracting melody information of each of the to-be-identified audio segments includes: The landmark algorithm is used to extract the point with the largest pitch during the duration of each melody point as the melody information of the melody point. According to the melody information of all melody points in each of the audio segments to be identified, the melody information of the corresponding audio segment to be identified is obtained.

5. The audio recognition method according to claim 1, wherein: The melody information includes pitch information; The step of finding all candidate audio segments having the same melody information as the respective audio segments to be identified from the audio database includes: For each of the audio segments to be identified, all candidate audio segments corresponding to the audio segment to be identified are found from the audio database; wherein the pitch information of all the candidate audio segments corresponding to the audio segment to be identified and the pitch information of the audio segment to be identified meet a preset condition.

6. The audio recognition method according to claim 5, characterized in that The preset condition includes that a difference between pitch information of all candidate audio segments corresponding to the audio segment to be identified and the pitch information of the audio segment to be identified is less than a preset threshold.

7. The audio recognition method according to claim 1, wherein: The determining, based on the time information of each audio segment to be identified and the time information of all corresponding candidate audio segments, a target audio segment similar to the audio segment to be identified, includes: Based on the time difference between each of the audio segments to be identified and all of the corresponding candidate audio segments, a target audio segment similar to the audio segment to be identified is determined; wherein the target audio segment has at least one candidate audio segment corresponding to each of the audio segments to be identified, and among the time differences between each of the audio segments to be identified and all of the corresponding candidate audio segments, there is at least one time difference that exists in all of the audio segments to be identified.

8. An audio recognition device, characterized in that: The audio recognition device comprises: An acquisition module, used to acquire the audio to be recognized; A segmentation module, configured to segment the audio to be recognized with a preset step size to obtain a plurality of audio segments to be recognized; An extraction module, configured to extract melody information of each of the audio segments to be identified; A comparison module is used to find all candidate audio segments with the same melody information as the audio segments to be identified from the audio database; The determining module is configured to determine a target audio similar to the audio to be identified based on the time information of each audio segment to be identified and the time information of all the candidate audio segments corresponding thereto.

9. An electronic device, characterized in that: The system comprises a memory and a processor coupled to each other, wherein the processor is configured to execute program instructions stored in the memory to implement the audio recognition method according to any one of claims 1 to 7.

10. A computer-readable storage medium having program instructions stored thereon, characterized in that: When the program instructions are executed by a processor, the audio recognition method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Song melody information processing method and device

    CN107203571A

  • Intonation evaluation method based on Deep Conventional Neural Network DCNN and CTC algorithm

    CN110364184A