Audio Recognition Method, Device, Electronic Device, and Storage Medium
By selecting candidate audio with high similarity from the preset library and entering the detection model, the problem of users not being able to recognize melody or songs is solved, efficient audio recognition is achieved, and user experience is improved.
Patent Information
- Application Number
- CN202111109177.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-22
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-09-22
AI Technical Summary
When retrieving music, users often hear a melody or song but cannot recognize the song name or singer, which makes it impossible to query the corresponding song, reducing the user experience.
By obtaining the clip information of the audio to be identified, select candidate audio with high similarity from the preset library, and enter the trained detection model to obtain the target audio clip and target audio.
Use some clip information to identify matching target audio clips and target audio from the preset library, improving recognition efficiency and improving user experience.
Smart Images

Figure CN113889146B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data processing, and particularly to an audio recognition method, apparatus, electronic device, and storage medium. Background Art
[0002] Currently, the digital dissemination of music is becoming a popular trend, and users are accustomed to retrieving various rich and colorful music contents from the network. At present, when retrieving and querying music, users can use the song name or singer, etc. as retrieval conditions to obtain music. In practical applications, users often hear a melody or a song, such as watching short videos, mobile phone ringtones, etc., but they don't know information such as the song name or singer, resulting in being unable to query the corresponding song and reducing the user experience. Summary of the Invention
[0003] The present disclosure provides an audio recognition method, apparatus, electronic device, and storage medium to solve the deficiencies of the related art.
[0004] According to the first aspect of the embodiments of the present disclosure, an audio recognition method is provided, and the method includes:
[0005] Obtain query content; the query content includes segment information representing the audio to be recognized;
[0006] Select a preset number of candidate audios corresponding to the query content from a preset library; the candidate audios include candidate audio segments matching the segment information;
[0007] Input the candidate audio segments into a trained detection model to obtain target segment information including the segment information and the target audio where the target segment information is located.
[0008] In some embodiments, selecting a preset number of candidate audios corresponding to the query content from a preset library includes:
[0009] Determine the similarity between the morphemes of the segment information and the text information of each audio in the preset library;
[0010] Sort the audios in the preset library according to the similarity from large to small to obtain a sorting result;
[0011] Based on the sorting result, determine the preset number of audios with the front sorting positions as the candidate audios, and each candidate audio includes at least one audio segment matching the morphemes of the segment information;
[0012] Obtain the audio segment with the longest continuous matching morphemes from at least one audio segment of each candidate audio to obtain the candidate audio segment matching the segment information for each candidate audio.
[0013] In some embodiments, inputting the candidate audio segment into a trained detection model to obtain target segment information including the segment information and the target audio where the target segment information is located, includes:
[0014] Obtaining a vector to be detected corresponding to each candidate audio according to the segment information and the candidate audio segment;
[0015] Inputting the vector to be detected corresponding to each candidate audio into the detection model to obtain detection result data output by the detection model;
[0016] Obtaining target segment information including the segment information and the target audio where the target segment information is located according to the detection result data.
[0017] In some embodiments, obtaining a vector to be detected corresponding to each candidate audio according to the segment information and the candidate audio segment, includes:
[0018] Concatenating the segment information with the candidate audio segments of each candidate audio respectively to obtain a vector to be detected corresponding to each candidate audio;
[0019] Wherein, each vector to be detected at least includes a first identifier and a second identifier, the first identifier is used to identify the starting position of the vector to be detected, and the second identifier is used to identify the concatenation position and the ending position of the vector to be detected.
[0020] In some embodiments, the detection result data includes first probability data and second probability data representing the starting position and the ending position corresponding to each morpheme in the candidate audio segment respectively;
[0021] Obtaining target segment information including the segment information and the target audio where the target segment information is located according to the detection result data, includes:
[0022] In the case where the starting position is less than the ending position, determining a target audio segment from the candidate audio segment based on the product of the first probability data and the second probability data;
[0023] Taking the target audio segment as the target segment information identified from the query content and taking the audio where the target segment information is located as the target audio.
[0024] In some embodiments, determining a target audio segment from the candidate audio segment based on the product of the first probability data and the second probability data, includes: determining the starting morpheme at the starting position and the ending morpheme at the ending position when the product of the first probability data and the second probability data is the largest;
[0025] Determine that all morphemes between the starting morpheme and the ending morpheme constitute the target audio segment.
[0026] In some embodiments, the audio to be recognized is a song, and the segment information refers to some lyrics in the song.
[0027] According to a second aspect of the embodiments of the present disclosure, there is provided an audio recognition device, the device includes:
[0028] A query content acquisition module, configured to execute acquiring query content; the query content includes segment information characterizing the audio to be recognized;
[0029] A candidate audio acquisition module, configured to execute selecting a preset number of candidate audios corresponding to the query content from a preset library; the candidate audios include candidate audio segments matching the segment information;
[0030] A target audio acquisition module, configured to execute inputting the candidate audio segment into a trained detection model to obtain target segment information including the segment information and the target audio where the target segment information is located.
[0031] In some embodiments, the candidate audio acquisition module includes:
[0032] A similarity determination sub-module, configured to execute determining the similarity between the morphemes of the segment information and the text information of each audio in the preset library;
[0033] A sorting result acquisition sub-module, configured to execute sorting the audios in the preset library from large to small according to the similarity to obtain a sorting result;
[0034] An audio segment acquisition sub-module, configured to execute determining that the preset number of audios with a higher sorting position based on the sorting result are the candidate audios, and each candidate audio includes at least one audio segment matching the morphemes of the segment information;
[0035] A candidate segment acquisition sub-module, configured to execute obtaining the audio segment with the longest continuous matching morphemes from at least one audio segment of each candidate audio to obtain a candidate audio segment matching the segment information for each candidate audio.
[0036] In some embodiments, the target audio acquisition module includes:
[0037] A vector to be tested acquisition sub-module, configured to execute obtaining a vector to be detected corresponding to each candidate audio according to the segment information and the candidate audio segment;
[0038] A detection result acquisition sub-module, configured to input the vector to be detected corresponding to each of the candidate audios into the detection model, and obtain detection result data output by the detection model;
[0039] A target audio acquisition sub-module, configured to obtain target segment information including the segment information and a target audio where the target segment information is located according to the detection result data.
[0040] In some embodiments, the vector to be detected acquisition sub-module includes:
[0041] A vector to be detected acquisition unit, configured to splice the segment information with candidate audio segments of each candidate audio respectively, and obtain a vector to be detected corresponding to each candidate audio;
[0042] Wherein, each of the vectors to be detected includes at least a first identifier and a second identifier, the first identifier is used to identify the starting position of the vector to be detected, and the second identifier is used to identify the splicing position and the ending position of the vector to be detected.
[0043] In some embodiments, the detection result data includes first probability data and second probability data representing the starting position and the ending position corresponding to each morpheme in the candidate audio segment respectively;
[0044] The target audio acquisition sub-module includes:
[0045] A target segment acquisition unit, configured to determine a target audio segment from the candidate audio segments based on the product of the first probability data and the second probability data when the starting position is less than the ending position;
[0046] A target audio acquisition unit, configured to use the target audio segment as the target segment information identified from the query content and use the audio where the target segment information is located as the target audio.
[0047] In some embodiments, the target segment acquisition unit includes:
[0048] A morpheme determination sub-unit, configured to determine a starting morpheme at the starting position and an ending morpheme at the ending position when the product of the first probability data and the second probability data is the largest;
[0049] A segment determination sub-unit, configured to determine that all morphemes between the starting morpheme and the ending morpheme constitute the target audio segment.
[0050] In some embodiments, the audio to be recognized is a song, and the segment information refers to part of the lyrics in the song.
[0051] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including:
[0052] a processor;
[0053] a memory for storing a computer program executable by the processor;
[0054] wherein the processor is configured to execute the computer program in the memory to implement the method as described above.
[0055] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, which can implement the method as described above when the executable computer program in the storage medium is executed by a processor.
[0056] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:
[0057] As can be seen from the above embodiments, in the solution provided by the embodiments of the present disclosure, query content can be obtained; the query content includes segment information characterizing the audio to be recognized; then, a preset number of candidate audios corresponding to the query content are selected from a preset library; the candidate audios include candidate audio segments matching the segment information; thereafter, the candidate audio segments are input into a trained detection model to obtain target segment information including the segment information and a target audio where the target segment information is located. In this way, in this embodiment, the matching target audio segment and target audio can be identified from the preset library by using partial segment information, which is beneficial to improving the recognition efficiency and enhancing the user experience.
[0058] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.
[0060] Figure 1 is a flowchart of an audio recognition method shown according to an exemplary embodiment.
[0061] Figure 2 is a flowchart of obtaining candidate music segments shown according to an exemplary embodiment.
[0062] Figure 3 is a flowchart of obtaining detection result data shown according to an exemplary embodiment.
[0063] Figure 4 is a flowchart of an audio recognition method shown according to an exemplary embodiment.
[0064] Figure 5 It is a flowchart for obtaining a target music segment shown according to an exemplary embodiment.
[0065] Figure 6 It is a block diagram of an audio recognition device shown according to an exemplary embodiment.
[0066] Figure 7 It is a block diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners
[0067] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The exemplary embodiments described below do not represent all embodiments consistent with the present disclosure. On the contrary, they are only examples of devices consistent with some aspects of the present disclosure as detailed in the appended claims. It should be noted that, without conflict, the features in the following embodiments and implementation manners can be combined with each other.
[0068] Currently, the digital dissemination of music is becoming a popular trend, and users are accustomed to retrieving various rich and colorful music contents from the network. Currently, when retrieving and querying music, users can use the song name or singer, etc. as retrieval conditions to obtain music. In practical applications, users often hear a melody or a song, such as when watching short videos, mobile phone ringtones, etc., but they don't know information such as the song name or singer, resulting in being unable to query the corresponding song and reducing the user experience.
[0069] To solve the above technical problems, the embodiments of the present disclosure provide an audio recognition method, which can be applied to an electronic device. The electronic device can include, but is not limited to, mobile terminal devices such as tablet computers, smart phones, personal computers, smart TVs, desktop computers, large screens, etc. or fixed terminal devices, and can also be applied to a server or a server cluster.
[0070] Figure 1 It is a flowchart of an audio recognition method shown according to an exemplary embodiment. Refer to Figure 1 , an audio recognition method, including step 11 to step 13.
[0071] In step 11, obtain query content; the query content includes segment information representing the audio to be recognized.
[0072] In this embodiment, an application program can be installed on the electronic device, such as a built-in music program, or a third-party music program (such as QQ Music, Kuwo Music, etc.), or other application programs with music playback functions. The user can listen to audio or watch videos through the above application programs. When there is a need to query audio, the query content can be input into the electronic device. The above query content can include, but is not limited to, text, voice, and pictures, etc. In one example, during the music search process, the most commonly used query methods are text or voice.
[0073] For example, a search box can be provided in the electronic device or the application program, and the user can input the query content in the search box. In practical applications, the above query content can include at least one of the following: music name, singer, composer, or segment information of the audio to be recognized, etc. The above segment information can be part of the morphemes in the audio to be recognized, and the morpheme can be a single character, a word, a short phrase or a long sentence formed by multiple words. Considering the usage scenarios where the user hears a part of the audio segment and hopes to retrieve the original segment or the audio name of the audio, etc., the above query content can also include a part of the audio segment in the target audio, such as the text or semantics corresponding to the part of the audio segment. It can be understood that the above target audio refers to the audio to which the original audio segment containing the part of the audio segment in the query content belongs, that is, the audio that needs to be retrieved.
[0074] With the continuous development of Internet technology and the increasing demand of users for search convenience, the application of voice search is becoming more and more widespread. For example, search by inputting sentences such as "Which song is 'The sparrows outside the window are chattering on the telegraph pole'" or "Play 'The sparrows outside the window are chattering on the telegraph pole'" through voice. Another example is that it can also be searched by pictures. When receiving an audio segment, a screenshot can be taken and the above screenshot can be input into the search box for audio retrieval, or audio search can be performed by uploading the album cover of the singer, etc.
[0075] In step 12, select a preset number of candidate audios corresponding to the query content from the preset library; the candidate audios include candidate audio segments that match the segment information.
[0076] In this embodiment, the preset library can be pre-stored in the electronic device or in the cloud, and the preset library includes several audios, and each audio can be all or part of a piece of music; an index can be established for each audio, and the index can include audio auxiliary tag information, and the auxiliary tag information can include, but is not limited to, text information of the audio, audio name, album, year, type, songwriters, singers and other auxiliary tag information. The above index can be established manually or by automated means such as a computer. Therefore, the preset library in this embodiment can include indexes of several audios, thereby reducing the occupation of storage space and improving the reading and writing efficiency.
[0077] In this embodiment, the electronic device can select a preset number of candidate audios corresponding to the query content from a preset library. Refer to Figure 2 , which includes steps 21 to 24.
[0078] In step 21, the electronic device can determine the similarity between the morphemes of the segment information in the query content and the text information of each audio in the preset library. For example, the electronic device can segment the audio segments in the query content to obtain multiple morphemes. Then, the electronic device can obtain the similarity between each morpheme and the text information in each audio index. The BM25 algorithm can be used to calculate the similarity between the segment information and the text information in step 21. If the calculation process of the BM25 algorithm can refer to related technologies, it will not be elaborated here.
[0079] In step 22, the electronic device can sort the candidate audios in the preset library from largest to smallest according to the similarity to obtain a sorting result.
[0080] Considering that the number of indexes in the preset library is relatively large (such as more than tens of millions), it takes a long time to sort all the indexes; combined with only using a preset number of candidate audios in this disclosure, therefore, the tournament sorting algorithm can be used in this example. For example, to sort 25 audios and obtain 3 audios with larger sorting similarities while the remaining audios are not sorted, it includes:
[0081] The first step: Divide 25 audios into 5 groups, namely A1 - A5, B1 - B5, C1 - C5, D1 - D5, E1 - E5. Sort each group to obtain the audio with the largest similarity in each group. Assume that the sorting within each group is the same as the serial number of each audio, that is, the similarities in group A are in descending order as A1, A2, A3, A4, and A5. The second step: Sort the audios with the largest similarities in the 5 groups (A1, B1, C1, D1, and E1) for the 6th time to obtain the audio A1 with the largest similarity among the 25 audios. The third step: Among the remaining audios in all groups, only A2, A3, B1, B2, and C1 have the opportunity to participate in the sorting. These 5 audios A2, A3, B1, B2, and C1 are sorted for the 7th time to obtain the audios ranked second and third in similarity. The above first, second, and third places obtain the sorting from largest to smallest while the sorting of other audios is less than the third place and there is no ranking.
[0082] It should be noted that in this embodiment, by using the tournament sorting algorithm, the audios in the preset library can be sorted, that is, select a preset number (such as N) of candidate audios with larger similarities from the preset library, and the sorting of other audios can be set to a preset identifier (such as 0) to indicate that the sorting of this audio is after the Nth place, which can reduce the amount of data processed by the electronic device and improve the retrieval speed of the electronic device.
[0083] In step 23, the electronic device may determine a preset number of audios with a higher ranking position based on the ranking result, that is, select a preset number of audios starting from the audio with the highest similarity. Each audio includes at least one audio segment that matches the morpheme of the segment information in the query content. If the tournament sorting algorithm is used, the electronic device may directly obtain the audio with a sorting identifier as the candidate audio.
[0084] It should be noted that in the process of obtaining the similarity, the similarity is calculated based on each morpheme in the query content and the text information of each audio. Then, some audios in the preset library can be matched by at least one word in the query content, that is, each candidate audio includes at least one audio segment that matches the morpheme of the segment information in the query content.
[0085] In step 24, the electronic device may obtain the audio segment with the longest continuous matching morpheme from at least one audio segment of each candidate audio, and obtain the candidate audio segment that matches the segment information in the query content for each candidate audio. Or, in step 24, the electronic device may select one longest audio segment from each of the preset number of candidate audios, and finally obtain N audio segments.
[0086] In step 13, the candidate audio segment is input into the trained detection model, and the target segment information including the segment information and the target audio where the target segment information is located are obtained.
[0087] In this embodiment, a trained detection model may be stored in the electronic device, and the detection model may be a BERT model. The electronic device may use the above detection model to obtain the target audio and the target segment information. Refer to Figure 3 , including steps 31 to 33. In step 31, the electronic device may obtain a vector to be detected corresponding to each candidate audio according to the segment information of the query content and each candidate audio segment. In this step, the electronic device may splice the segment information and each candidate audio segment to obtain the vector to be detected. Each vector to be detected includes at least a first identifier and a second identifier. The first identifier is used to identify the start position of the vector to be detected, and the second identifier is used to identify the splicing position and the end position of the vector to be detected. For example, "[cls]query content[seg]lyrics of candidate audio segment[seg]" represents a vector to be detected, where [cls] represents the first identifier of the vector to be detected, and [seg] represents the second identifier of the vector to be detected.
[0088] In step 32, the electronic device may input the vector to be detected corresponding to each candidate audio into the detection model to obtain the detection result data output by the detection model. The electronic device may call the detection model, input the vector to be detected into the detection model in sequence, and obtain the detection result data output by the detection model.
[0089] In step 33, the electronic device may obtain the target segment information including the segment information and the target audio where the target segment information is located according to the detection result data.
[0090] In this step, the above detection result data includes first probability data indicating that each morpheme of the candidate audio segment in the vector to be detected is located at the starting position as the starting morpheme, second probability data indicating that each morpheme of the candidate audio segment in the vector to be detected is located at the ending position as the ending morpheme, and third probability data indicating whether there is a matching audio segment in the query content.
[0091] In this embodiment, the electronic device may compare the starting position and the ending position. When the starting position is less than the ending position, the electronic device may determine the target audio segment from the candidate audio segments based on the product of the first probability data and the second probability data. For example, the electronic device may obtain the product of the first probability data and the second probability data of each vector to be detected; and use the candidate audio segment with the largest product as the target audio segment, and use the target audio segment as the target segment information identified from the query content. It can be understood that after determining the target audio segment, the electronic device may further determine the auxiliary label information corresponding to the target audio segment, and determine the target audio according to the above auxiliary label information.
[0092] In one embodiment, after obtaining the above target audio or target audio segment, the electronic device may feedback the text information of the target audio or the target audio segment to the user to facilitate the user to read the above retrieval result. When the user hopes to listen to the target audio, the electronic device may obtain the corresponding audio data from the server through the application and cache it for local playback.
[0093] So far, in the solution provided by the embodiments of the present disclosure, the query content can be obtained; the query content includes segment information characterizing the audio to be recognized; then, a preset number of candidate audios corresponding to the query content are selected from the preset library; the candidate audios include candidate audio segments matching the segment information; afterwards, the candidate audio segments are input into the trained detection model to obtain the target segment information including the segment information and the target audio where the target segment information is located. In this way, in this embodiment, the matching target audio segment and target audio can be recognized from the preset library by using partial segment information, which is beneficial to improving the recognition efficiency and the user experience.
[0094] Taking music retrieval as an example, the working principle of an audio recognition method provided by the present disclosure will be described below in conjunction with Figure 4 and Figure 5 :
[0095] In this embodiment, the electronic device can decompose the recognition task of the segment information in the query content into three parts, namely: a retriever, a slicer, and a reader, which are specifically as follows:
[0096] Retriever: Obtain the index of each piece of music, where the index contains music auxiliary tag information. Then, based on the index, use the BM25 algorithm to calculate the similarity between the query content and the full text of each piece of music respectively, and obtain N candidate pieces of music; and recall the music auxiliary tag information of the N candidate pieces of music with the highest relevance (i.e., larger similarity) to the query content.
[0097] Slicer: Truncate and fill the full text of the music of the N candidate pieces of music with the highest relevance recalled by the index respectively to make each music segment form a complete music segment, and then cut out the music segment with the longest continuous matching words of each candidate piece of music. N candidate pieces of music obtain N candidate music segments. These N candidate music segments are respectively spliced with the query content and used as the input of the "reader" model later. The splicing method adopts Figure 4 in the form of "[cls] query content [seg] candidate music segment [seg]" to obtain the vector to be detected.
[0098] Reader: Adopt the BERT model and input the vectors to be detected obtained by the slicer into the BERT model in turn. The output of the BERT model includes two parts, namely whether it contains a music segment that meets the requirements (Has Answer Score) and the specific music segment (lyric span). See Figure 6 , for the Has Answer Score part, the vector representation of the [CLS] position of the BERT model passes through an additional fully connected layer and binary classification Softmax to obtain the third probability data (HAScore) of whether there is an answer. According to the third probability data and a preset probability threshold (which can take values from 0.85 to 0.95 and can be set), it is judged whether there is a music segment that meets the requirements. For the lyric span part, that is Figure 7In the Start / End Span part, after passing the candidate music segment vector through a fully connected layer and Softmax calculation, the probabilities Pstart (i.e., the first probability data) and Pend (i.e., the second probability data) of the start word and the end word are obtained for each segment unit (Token) as the answer. Then, a combination with the largest Pstart * Pend and the position start of the start word less than the position end of the end word is obtained, and the text between the start position start and the end position end at this time is used as the target music segment.
[0099] Based on the audio recognition method provided in the above embodiment, the present disclosure embodiment also provides an audio recognition device. Refer to Figure 6 , the device includes:
[0100] A query content acquisition module 61, configured to execute acquiring query content; the query content includes segment information representing the audio to be recognized;
[0101] A candidate music selection module 62, configured to execute selecting a preset number of candidate audios corresponding to the query content from a preset library; the candidate audios include candidate audio segments matching the segment information;
[0102] A target audio acquisition module 63, configured to execute inputting the candidate audio segments into a trained detection model to obtain target segment information including the segment information and the target audio where the target segment information is located.
[0103] In one embodiment, the candidate audio acquisition module includes:
[0104] A similarity determination sub-module, configured to execute determining the similarity between the morphemes of the segment information and the text information of each audio in the preset library;
[0105] A sorting result acquisition sub-module, configured to execute sorting the audios in the preset library from large to small according to the similarity to obtain a sorting result;
[0106] An audio segment acquisition sub-module, configured to execute determining that a preset number of audios with a higher sorting position based on the sorting result are the candidate audios, and each candidate audio includes at least one audio segment matching the morphemes of the segment information;
[0107] A candidate segment acquisition sub-module, configured to execute obtaining the audio segment with the longest continuously matching morphemes from at least one audio segment of each candidate audio to obtain the candidate audio segments matching the segment information for each candidate audio.
[0108] In one embodiment, the target audio acquisition module includes:
[0109] The sub-module for obtaining the vector to be detected is configured to execute the operation of obtaining the vector to be detected corresponding to each candidate audio according to the segment information and the candidate audio segments;
[0110] The sub-module for obtaining the detection result is configured to execute the operation of inputting the vector to be detected corresponding to each candidate audio into the detection model to obtain the detection result data output by the detection model;
[0111] The sub-module for obtaining the target audio is configured to execute the operation of obtaining the target segment information including the segment information and the target audio where the target segment information is located according to the detection result data.
[0112] In one embodiment, the sub-module for obtaining the vector to be detected includes:
[0113] The unit for obtaining the vector to be detected is configured to execute the operation of splicing the segment information with the candidate audio segments of each candidate audio respectively to obtain the vector to be detected corresponding to each candidate audio;
[0114] Wherein, each vector to be detected includes at least a first identifier and a second identifier, the first identifier is used to identify the starting position of the vector to be detected, and the second identifier is used to identify the splicing position and the ending position of the vector to be detected.
[0115] In one embodiment, the detection result data includes first probability data and second probability data representing the starting position and the ending position corresponding to each morpheme in the candidate audio segment respectively;
[0116] The sub-module for obtaining the target audio includes:
[0117] The unit for obtaining the target segment is configured to execute the operation of determining the target audio segment from the candidate audio segment based on the product of the first probability data and the second probability data when the starting position is less than the ending position;
[0118] The unit for obtaining the target audio is configured to execute the operation of taking the target audio segment as the target segment information identified from the query content and taking the audio where the target segment information is located as the target audio.
[0119] In one embodiment, the unit for obtaining the target segment includes:
[0120] The sub-unit for determining the morpheme is configured to execute the operation of determining the starting morpheme at the starting position and the ending morpheme at the ending position when the product of the first probability data and the second probability data is the largest;
[0121] A segment determination subunit is configured to determine that all morphemes between the starting morpheme and the ending morpheme constitute the target audio segment.
[0122] In one embodiment, the audio to be recognized is a song, and the segment information refers to partial lyrics in the song.
[0123] It should be noted that the devices and equipment shown in this embodiment match the content of the method embodiment. The content of the above method embodiment can be referred to and will not be elaborated here.
[0124] Figure 7 It is a block diagram of an electronic device shown according to an exemplary embodiment. For example, the electronic device 700 can be a smart phone, a computer, a digital broadcast terminal, a tablet device, a medical device, a fitness device, a personal digital assistant, etc. It can be understood that the above electronic device can be used as the first device or the second device.
[0125] Referring to Figure 7 , the electronic device 700 may include one or more of the following components: a processing component 702, a memory 704, a power supply component 706, a multimedia component 708, an audio component 710, an input / output (I / O) interface 712, a sensor component 714, a communication component 716, and an image acquisition component 718.
[0126] The processing component 702 generally controls the overall operation of the electronic device 700, such as operations associated with display, telephone call, data communication, camera operation, and recording operation. The processing component 702 may include one or more processors 720 to execute computer programs. In addition, the processing component 702 may include one or more modules to facilitate the interaction between the processing component 702 and other components. For example, the processing component 702 may include a multimedia module to facilitate the interaction between the multimedia component 708 and the processing component 702.
[0127] The memory 704 is configured to store various types of data to support the operation of the electronic device 700. Examples of these data include computer programs for any application or method operating on the electronic device 700, contact data, phone book data, messages, pictures, videos, etc. The memory 704 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.
[0128] The power supply component 706 provides power for various components of the electronic device 700. The power supply component 706 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 700. The power supply component 706 may include a power chip, and the controller can communicate with the power chip to control the power chip to turn on or off the switching device, so that the battery supplies power to the main board circuit or does not supply power.
[0129] The multimedia component 708 includes a screen that provides an output interface between the electronic device 700 and the target object. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input information from the target object. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation.
[0130] The audio component 710 is configured to output and / or input audio file information. For example, the audio component 710 includes a microphone (MIC), which is configured to receive external audio file information when the electronic device 700 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio file information can be further stored in the memory 704 or sent via the communication component 716. In some embodiments, the audio component 710 further includes a speaker for outputting audio file information.
[0131] The I / O interface 712 provides an interface between the processing component 702 and the peripheral interface module, and the peripheral interface module may be a keyboard, a click wheel, a button, etc.
[0132] The sensor component 714 includes one or more sensors for providing a status assessment of various aspects of the electronic device 700. For example, the sensor component 714 can detect the on / off state of the electronic device 700, the relative positioning of components, such as the display screen and the keypad of the electronic device 700. The sensor component 714 can also detect a change in the position of the electronic device 700 or a component, the presence or absence of contact between the target object and the electronic device 700, the orientation or acceleration / deceleration of the electronic device 700, and the temperature change of the electronic device 700. In this example, the sensor component 714 may include a magnetic sensor, a gyroscope, and a magnetic field sensor, and the magnetic field sensor includes at least one of the following: a Hall sensor, a thin film magnetoresistive sensor, and a magnetic fluid acceleration sensor.
[0133] The communication component 716 is configured to facilitate communication between the electronic device 700 and other devices in a wired or wireless manner. The electronic device 700 can access a communication standard-based wireless network, such as WiFi, 2G, 3G, 4G, 5G, or a combination thereof. In an exemplary embodiment, the communication component 716 receives broadcast information or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 716 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0134] In an exemplary embodiment, the electronic device 700 can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components.
[0135] In an exemplary embodiment, a non-transitory readable storage medium including an executable computer program is also provided, such as a memory 704 including instructions, and the executable computer program can be executed by a processor. The readable storage medium can be a ROM, Random Access Memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0136] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the disclosure herein. The present disclosure is intended to cover any variations, uses, or adaptations of the disclosure that follow the general principles of the disclosure and include known common knowledge or conventional technical means in the technical field not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0137] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. An audio recognition method, characterized in that, the method includes: obtaining query content; the query content includes segment information representing the audio to be recognized; selecting a preset number of candidate audios corresponding to the query content from a preset library; the candidate audios include candidate audio segments that match the segment information; the text information of the candidate audio segments matches the morphemes of the segment information; inputting the candidate audio segments into a trained detection model to obtain target segment information containing the segment information and the target audio where the target segment information is located; the detection result data output by the detection model includes first probability data and second probability data representing the starting position and ending position of each morpheme in the candidate audio segment respectively, and the detection result is used to obtain the target segment information and the target audio.
2. The method according to claim 1, characterized in that, selecting a preset number of candidate audios corresponding to the query content from a preset library includes: determining the similarity between the morphemes of the segment information and the text information of each audio in the preset library; sorting the audios in the preset library from large to small according to the similarity to obtain a sorting result; determining the preset number of audios with the top sorting positions as the candidate audios based on the sorting result, and each candidate audio includes at least one audio segment that matches the morphemes of the segment information; obtaining the audio segment with the longest continuous matching morphemes from at least one audio segment of each candidate audio to obtain the candidate audio segment that matches the segment information for each candidate audio.
3. The method according to claim 1, characterized in that, inputting the candidate audio segments into a trained detection model to obtain target segment information containing the segment information and the target audio where the target segment information is located includes: obtaining a detection vector corresponding to each candidate audio according to the segment information and the candidate audio segment; inputting the detection vector corresponding to each candidate audio into the detection model to obtain the detection result data output by the detection model; obtaining target segment information containing the segment information and the target audio where the target segment information is located according to the detection result data.
4. The method according to claim 3, characterized in that, obtaining a detection vector corresponding to each candidate audio according to the segment information and the candidate audio segment includes: concatenating the segment information with the candidate audio segments of each candidate audio respectively to obtain a detection vector corresponding to each candidate audio; wherein, each detection vector includes at least a first identifier and a second identifier, the first identifier is used to identify the starting position of the detection vector, and the second identifier is used to identify the concatenation position and the ending position of the detection vector.
5. The method according to claim 3, characterized in that, obtaining target segment information containing the segment information and the target audio where the target segment information is located according to the detection result data includes: When the starting position is less than the ending position, determine a target audio segment from the candidate audio segments based on the product of the first probability data and the second probability data; Use the target audio segment as the target segment information identified from the query content and use the audio where the target segment information is located as the target audio.
6. The method according to claim 5, wherein, Determining a target audio segment from the candidate audio segments based on the product of the first probability data and the second probability data includes: determining the starting morpheme at the starting position and the ending morpheme at the ending position when the product of the first probability data and the second probability data is the largest; Determine that all the morphemes between the starting morpheme and the ending morpheme constitute the target audio segment.
7. The method according to any one of claims 1 to 6, wherein, The audio to be recognized is a song, and the segment information refers to part of the lyrics in the song.
8. An audio recognition device, wherein, The device includes: A query content acquisition module configured to acquire query content; the query content includes segment information representing the audio to be recognized; A candidate audio acquisition module configured to select a preset number of candidate audios corresponding to the query content from a preset library; the candidate audios include candidate audio segments that match the segment information; the text information of the candidate audio segments matches the morphemes of the segment information; A target audio acquisition module configured to input the candidate audio segments into a trained detection model to obtain target segment information including the segment information and the target audio where the target segment information is located; the detection result data output by the detection model includes first probability data and second probability data representing the corresponding positions of each morpheme in the candidate audio segment at the starting position and the ending position respectively, and the detection result is used to obtain the target segment information and the target audio.
9. An electronic device, wherein, including: A processor; A memory for storing computer programs executable by the processor; wherein, the processor is configured to execute the computer program in the memory to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, wherein, When the executable computer program in the storage medium is executed by a processor, it can implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Humming type rhythm identification method based on hidden Markov model
CN101504834A
Query by humming method and system
CN104978962A