Audio recognition method, electronic device and computer readable storage medium
By adopting a similarity-based fuzzy matching method in audio recognition, the problem of misidentification when the fingerprint library does not store precise matching songs is solved, and the accuracy of audio recognition is improved.
Patent Information
- Application Number
- CN202310313241.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-27
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2043-03-27
AI Technical Summary
In the prior art, when the fingerprint library does not store the exact match song when it is recognized and the accuracy of audio recognition is reduced.
A fuzzy matching method based on similarity is used to find the first K matching songs with the highest similarity in the fingerprint library for each audio clip, and the target matching songs are determined based on the transfer probability and clip similarity of adjacent matching songs.
Even if the fingerprint library does not store precisely matched songs, this method can obtain multiple matching songs through fuzzy matching, improving the accuracy of audio recognition.
Smart Images

Figure CN116312496B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio processing, and particularly to an audio recognition method, an electronic device, and a computer-readable storage medium. Background Art
[0002] With the rapid development of audio processing technology, more and more users like to upload personalized processed audio to audio platforms. For example, users can splice multiple songs with natural melodies to create medley songs. Such behaviors make the quality of audio on audio platforms uneven, so it is necessary to identify medley songs through technical means.
[0003] Generally, the fingerprint exact matching method can be adopted to identify whether an audio is a medley song. Among them, if the audio is a non-medley ordinary song, multiple segments of the audio will all exactly match the same song. However, if the audio is a medley song, multiple segments of the audio can exactly match multiple different songs. Therefore, it is possible to determine whether the audio is a medley song by the number of songs matched. However, if there is no song in the fingerprint database that exactly matches a certain segment of the audio, it is possible to misidentify the audio from a medley song as a non-medley song, resulting in a low accuracy rate of audio recognition. Summary of the Invention
[0004] Embodiments of the present application provide an audio recognition method, an electronic device, and a computer-readable storage medium, which can effectively improve the accuracy rate of audio recognition.
[0005] In a first aspect, embodiments of the present application provide an audio recognition method, and the method includes:
[0006] Performing a slicing operation on a target audio to obtain multiple audio segments, and obtaining a melody fingerprint of each of the audio segments;
[0007] For each of the audio segments, searching in a fingerprint database for the top K pre-stored melody fingerprints with the highest similarity to the melody fingerprint of the audio segment and the matching songs corresponding to the K pre-stored melody fingerprints, to obtain K matching songs of the audio segment; K is a positive integer;
[0008] Based on a preset transition probability between adjacent matching songs and the similarity corresponding to the K matching songs of each audio segment, determining a target matching song for each audio segment from the K matching songs of each audio segment; wherein the adjacent matching songs refer to a pair of songs composed of a matching song corresponding to each of the adjacent audio segments, and the similarity corresponding to the matching song is the similarity between the pre-stored melody fingerprint corresponding to the matching song and the melody fingerprint of the audio segment corresponding to the matching song;
[0009] Determine the audio recognition result of the target audio according to the target matching songs of each of the audio segments, where the audio recognition result is used to indicate whether the target audio is a medley song.
[0010] Since the present application adopts a fuzzy matching method based on similarity, even if the fingerprint database does not store a song that exactly matches a certain audio segment, the present application can obtain K matching songs that are fuzzily matched to each audio segment. By retaining multiple matching songs for each audio segment, this method can improve the accuracy of determining the target matching song based on multiple matching songs, and further improve the accuracy of audio recognition.
[0011] In a possible implementation manner, the number of the multiple audio segments is S, and S is a positive integer greater than or equal to 3; determining the target matching song of each audio segment from the K matching songs of each audio segment based on the preset transition probability between adjacent matching songs and the similarity corresponding to the K matching songs of each audio segment includes:
[0012] Based on the preset transition probability between adjacent matching songs and the similarity corresponding to the K matching songs of each audio segment, construct multiple decoding paths and determine the path probability of each decoding path in the multiple decoding paths; where each decoding path includes S - 1 decoding sub-paths, and each decoding sub-path is used to indicate: a matching song of the m-th audio segment among the S audio segments, pointing to a matching song of the (m + 1)-th audio segment; the S audio segments are sorted in sequence according to their positions in the target audio, and m is a positive integer greater than or equal to 1 and less than S;
[0013] Take the decoding path with the maximum path probability among the multiple decoding paths as the target decoding path;
[0014] Take the matching songs of each audio segment indicated by each decoding sub-path in the target decoding path as the target matching songs of each audio segment.
[0015] Based on this method, multiple decoding paths can be constructed according to multiple decoding sub-paths first, and then the target matching song of each audio segment can be determined according to the decoding path with the maximum path probability among the multiple decoding paths.
[0016] In a possible implementation manner, constructing multiple decoding paths and determining the path probability of each decoding path in the multiple decoding paths based on the preset transition probability between adjacent matching songs and the similarity corresponding to the K matching songs of each audio segment includes:
[0017] Determine multiple decoding sub-paths at the first level and the sub-path probabilities of each decoding sub-path at the first level;
[0018] Set \(i\) to 2; based on the preset transition probability between adjacent matching songs, the similarities corresponding to \(K\) first to-be-processed matching songs, and the similarities corresponding to \(K\) second to-be-processed matching songs, determine multiple decoding sub-paths at the \(i\)th level and the sub-path probabilities of each decoding sub-path at the \(i\)th level; wherein, the first to-be-processed matching song is the matching song of the \(i\)th audio segment indicated by a decoding sub-path at the \((i - 1)\)th level; the second to-be-processed matching song is one of the \(K\) matching songs of the \((i + 1)\)th audio segment;
[0019] If \(i\) is less than \(S - 1\), perform an increment operation on \(i\), and return to execute the step of determining multiple decoding sub-paths at the \(i\)th level and the sub-path probabilities of each decoding sub-path at the \(i\)th level based on the preset transition probability between adjacent matching songs, the similarities corresponding to \(K\) first to-be-processed matching songs, and the similarities corresponding to \(K\) second to-be-processed matching songs;
[0020] If \(i\) is equal to \(S - 1\), set \(p\) to 1; for the \(p\)th matching song of the \(S\)th audio segment, obtain the decoding path corresponding to the \(p\)th matching song and the path probability of the decoding path according to the following steps: select one decoding sub-path from multiple decoding sub-paths at each level from the 1st level to the \((S - 1)\)th level, and form the decoding path corresponding to the \(p\)th matching song by the selected \(S - 1\) decoding sub-paths; the decoding sub-path at the \((S - 1)\)th level in the decoding path corresponding to the \(p\)th matching song points to the \(p\)th matching song of the \(S\)th audio segment; determine the path probability of the decoding path according to the sub-path probabilities of each decoding sub-path included in the decoding path; if \(p\) is less than \(K\), perform an increment operation on \(p\), and return to execute the step of obtaining the decoding path corresponding to the \(p\)th matching song of the \(S\)th audio segment and the path probability of the decoding path according to the following steps; if \(p\) is equal to \(K\), end the process.
[0021] Based on this method, the decoding sub-paths at the 2nd level to the \((S - 1)\)th level can be constructed in sequence first, and then the decoding path can be constructed according to the decoding sub-paths at the 1st level to the \((S - 1)\)th level.
[0022] In a possible implementation manner, the determining multiple decoding sub-paths at the \(i\)th level and the sub-path probabilities of each decoding sub-path at the \(i\)th level based on the preset transition probability between adjacent matching songs, the similarities corresponding to \(K\) first to-be-processed matching songs, and the similarities corresponding to \(K\) second to-be-processed matching songs includes:
[0023] Set \(q\) to 1. For the \(q\)-th second pending matching song among the \(K\) second pending matching songs, obtain the decoding sub-path corresponding to the \(q\)-th second pending matching song in the \(i\)-th level and the sub-path probability of the decoding sub-path according to the following steps:
[0024] According to the preset transition probability between adjacent matching songs, determine \(K\) pending transition probabilities from the \(K\) first pending matching songs to the \(q\)-th second pending matching song;
[0025] Based on the similarities corresponding to the \(K\) first pending matching songs, the \(K\) pending transition probabilities, and the similarity corresponding to the \(q\)-th second pending matching song, calculate the sub-path probabilities of each target decoding sub-path pointing to the \(q\)-th second pending matching song;
[0026] Take the target decoding sub-path with the largest sub-path probability as the decoding sub-path corresponding to the \(q\)-th second pending matching song in the \(i\)-th level;
[0027] If \(q\) is less than \(K\), perform an increment operation on \(q\), and return to execute the step of obtaining the decoding sub-path corresponding to the \(q\)-th second pending matching song in the \(i\)-th level and the sub-path probability of the decoding sub-path according to the following steps; if \(q\) is equal to \(K\), end the process.
[0028] Based on this method, when constructing the decoding sub-paths of each level from the 2nd level to the \((S - 1)\)-th level, multiple target decoding sub-paths can be constructed first according to the decoding sub-paths of the previous level, and then the path with the largest sub-path probability is selected from the multiple target decoding sub-paths as a decoding sub-path, so that each decoding sub-path of each level is a sub-path with a relatively large sub-path probability.
[0029] In a possible implementation manner, if the first pending matching song is the same as the \(q\)-th second pending matching song, the preset transition probability between adjacent matching songs is the first transition probability;
[0030] Alternatively, if the first pending matching song is different from the \(q\)-th second pending matching song, the preset transition probability between adjacent matching songs is the second transition probability; the first transition probability is greater than the second transition probability.
[0031] Based on this method, the preset transition probability between adjacent matching songs can be determined according to whether the first pending matching song is the same as the second pending matching song, and then the pending transition probability can be determined. The pending transition probability determined by this method is more accurate and can improve the accuracy of audio recognition.
[0032] In a possible implementation manner, determining the multiple decoding sub-paths of the first level and the sub-path probabilities of each decoding sub-path of the first level includes:
[0033] For each matching song of the second audio segment among the S audio segments, determine a decoding sub-path and the sub-path probability of the decoding sub-path according to the following steps:
[0034] Determine whether the matching song is the same as any matching song of the first audio segment among the S audio segments;
[0035] Determine the preset transition probability between the adjacent matching songs according to the judgment result, and determine the to-be-processed transition probability pointing from any matching song of the first audio segment to the matching song according to the preset transition probability between the adjacent matching songs;
[0036] Based on the K to-be-processed transition probabilities, the similarities corresponding to the K matching songs of the first audio segment, and the similarity corresponding to the matching song, determine the decoding sub-path corresponding to the matching song in the first level and the sub-path probability of the decoding sub-path.
[0037] Based on this method, the decoding sub-paths of the first level can be determined.
[0038] In a possible implementation manner, determining the audio recognition result of the target audio according to the target matching song of each audio segment includes:
[0039] If the target matching songs of each audio segment are the same, determine that the audio recognition result of the target audio is used to indicate that the target audio is not a medley song;
[0040] If there are at least two different target matching songs among the target matching songs of each audio segment, determine that the audio recognition result of the target audio is used to indicate that the target audio is a medley song.
[0041] Based on this method, it can be determined whether the audio recognition result indicates that the target audio is a medley song by judging whether the target matching songs of each audio segment are the same.
[0042] In a possible implementation manner, the method further includes:
[0043] When the audio recognition result of the target audio is used to indicate that the target audio is a medley song, determine the version identifier of the target matching song of each audio segment; the version identifier includes a first version identifier or a second version identifier, the first version identifier is used to indicate that the corresponding target matching song is an original song, and the second version identifier is used to indicate that the corresponding target matching song is an adapted song;
[0044] If the version identifier of the target matching song for each of the audio segments is the second version identifier, the audio recognition result of the target audio is further used to indicate that the target audio is a medley song of the adapted type;
[0045] If the version identifier of the target matching song for each of the audio segments includes the first version identifier and the second version identifier, the audio recognition result of the target audio is further used to indicate that the target audio is a medley song of the hybrid type based on the original singer and the adaptation.
[0046] Based on this method, the type of the target audio as a medley song can be further determined according to the version identifier of the target matching song.
[0047] In a possible implementation manner, the melody fingerprint is an embedding vector, and obtaining the melody fingerprint of each of the audio segments includes:
[0048] Extracting the frequency domain features of each of the audio segments;
[0049] Inputting the frequency domain features of each of the audio segments into a pre-trained embedding vector generation model to obtain the melody fingerprint of each of the audio segments.
[0050] Based on this method, the melody fingerprint of each audio segment can be obtained according to the pre-trained embedding vector generation model.
[0051] In a possible implementation manner, the method further includes:
[0052] Obtaining a training audio set; the training audio set includes original singer songs, at least one adapted song corresponding to the original singer songs, and at least one reference song, and the reference song is different from the original singer songs and the adapted songs;
[0053] Performing a slicing operation on each song in the training audio set to obtain N song segments of each song, where N is a positive integer, and the song content of the j-th song segment of the original singer song is the same as the song content of the j-th song segment of any of the adapted songs, and j is a positive integer less than or equal to N;
[0054] Extracting the frequency domain features of the N song segments of each song, and inputting the frequency domain features of the N song segments of each song into an initial embedding vector generation model to obtain the embedding vectors of the N song segments of each song;
[0055] Determining a first vector distance according to the embedding vectors of the N song segments of the original singer song and the embedding vectors of the N song segments of the adapted song;
[0056] Determine a second vector distance based on the embedding vectors of N song segments of the original song and the embedding vectors of N song segments of the reference song;
[0057] Taking reducing the first vector distance and increasing the second vector distance as the training objective, train the initial embedding vector generation model to obtain the pre-trained embedding vector generation model.
[0058] Based on the pre-trained embedding vector generation model obtained in this way, a higher-quality melody fingerprint can be generated for each audio segment. Furthermore, when performing fuzzy matching on the melody fingerprints of each audio segment, K more accurate matching songs can be obtained.
[0059] In a second aspect, an embodiment of the present application provides an audio recognition device, which includes:
[0060] A slicing module for performing slicing operations on the target audio to obtain multiple audio segments;
[0061] An acquisition module for acquiring the melody fingerprints of each of the audio segments;
[0062] A search module for, for each of the audio segments, searching in the fingerprint library for the top K pre-stored melody fingerprints with the highest similarity to the melody fingerprint of the audio segment and the matching songs corresponding to the K pre-stored melody fingerprints to obtain the K matching songs of the audio segment; K is a positive integer;
[0063] A determination module for determining the target matching song of each audio segment from the K matching songs of each audio segment based on the preset transition probability between the audio segments and the similarity corresponding to the K matching songs of each audio segment; wherein the similarity corresponding to the matching song is the similarity between the pre-stored melody fingerprint corresponding to the matching song and the melody fingerprint of the audio segment corresponding to the matching song; determine the audio recognition result of the target audio according to the target matching song of each audio segment, and the audio recognition result is used to indicate whether the target audio is a medley song.
[0064] In a third aspect, an embodiment of the present application provides an electronic device, which includes a memory and a processor; the memory is used to store a computer program, and the computer program includes program instructions; the processor is used to call the program instructions from the memory, so that the electronic device executes the method described in the first aspect above.
[0065] Fourthly, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. The computer program includes program instructions. When the program instructions are executed by a processor, the processor is caused to execute the method described in the first aspect above.
[0066] Fifthly, an embodiment of the present application provides a computer program product. The computer program product includes a computer program. The computer program includes program instructions, and the program instructions are executed by a processor to execute the method described in the first aspect above. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] To more clearly illustrate the technical solutions of the embodiments of the present application, the accompanying drawings required for the description of the embodiments will be briefly introduced below.
[0068] Figure 1 is a schematic diagram of an existing audio recognition process based on fingerprint exact matching;
[0069] Figure 2 is a schematic diagram of an audio recognition system provided by an embodiment of the present application;
[0070] Figure 3 is a schematic diagram of a process of an audio recognition method provided by an embodiment of the present application;
[0071] Figure 4 is a schematic diagram of an audio slice provided by an embodiment of the present application;
[0072] Figure 5 is a schematic diagram of a decoding path provided by an embodiment of the present application;
[0073] Figure 6 is a schematic diagram of a decoding sub-path provided by an embodiment of the present application;
[0074] Figure 7 is a schematic diagram of a training process of an embedded vector generation model provided by an embodiment of the present application;
[0075] Figure 8 is a schematic diagram of a training process based on vector distance provided by an embodiment of the present application;
[0076] Figure 9 is a schematic diagram of the structure of an audio recognition device provided by an embodiment of the present application;
[0077] Figure 10 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0078] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without making creative efforts shall fall within the protection scope of the present application.
[0079] In the description, claims and drawings of the present application, the terms "first", "second", etc. are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.
[0080] At present, the recognition of medley songs mainly adopts the method of fingerprint exact matching. Please refer to Figure 1 , Figure 1 which is a schematic diagram of an existing audio recognition process based on fingerprint exact matching. Among them, the audio recognition process includes the establishment stage of the hash code fingerprint library and the retrieval and matching stage of the audio.
[0081] In the establishment stage of the hash code fingerprint library, first, the training songs are preprocessed and Fourier-transformed to obtain the spectrograms of each segment of the training songs. Then, a peak point sequence is extracted from each segment's spectrogram. Each peak point sequence uniquely represents a segment, and each peak point included in each peak point sequence can be a point with a local maximum amplitude. Next, the peak point sequences of each segment are combined to generate multiple pairs of peak point combinations. Exemplarily, the peak point combination can be represented in the combination form of (t1, f1, t2, f2), where t1 and t2 represent the time corresponding to the peak point, and f1 and f2 represent the frequency corresponding to the peak point; for example, when (t1, f1, t2, f2) is (5, 440, 6.2, 700), it can represent that there is a peak point at 440 Hz at the 5th second and a peak point at 700 Hz at the 6.2nd second. Further, each peak point combination is compressed, and the compressed peak point combination is hash-coded to obtain the hash code corresponding to each segment; after obtaining the hash code corresponding to each segment, the hash code and the song identifier of each segment can be used as fingerprint information, and the fingerprint information is stored in the hash code fingerprint library.
[0082] In the retrieval and matching stage of audio, the song to be recognized is preprocessed, Fourier-transformed, a peak point sequence is extracted, peak point combinations are generated, compressed, and hash-coded in sequence to obtain the hash codes of each segment to be recognized of the song to be recognized. Then, it is retrieved in the hash code fingerprint database whether there are hash codes that exactly match the hash codes of each segment to be recognized, and song identifiers that match each segment to be recognized are obtained, that is, the songs that match each segment to be recognized are obtained. This exact match means that the hash code of the segment to be recognized is exactly the same as the hash code of a certain segment in the hash code fingerprint database. Finally, the recognition result of the song to be recognized can be determined according to the songs that match each segment to be recognized.
[0083] The above method can effectively obtain the recognition result of the song to be recognized. However, when the song identifier and the corresponding hash code of a certain segment are not pre-stored in the hash code fingerprint database, the above exact matching method cannot obtain the song that exactly matches a certain segment. In this case, a medley song may be recognized as a non-medley song, resulting in a low accuracy rate of audio recognition.
[0084] To improve the accuracy rate of audio recognition, an embodiment of the present application provides an audio recognition method, which can be applicable to Figure 2 the audio recognition system shown. As Figure 2 shown, the audio recognition system includes at least one server, such as server 20, server 21, and at least one client, such as client 22, client 23. Communication connections can be established between each server and each client through a network, and the network can be a wired network, a wireless network, etc. For any server in the audio recognition system, a fingerprint database can be pre-stored in the any server, and pre-stored melody fingerprints and identifiers of matching songs corresponding to the pre-stored melody fingerprints are stored in the fingerprint database.
[0085] The following takes the server 20 as an example to illustrate the process of executing the audio recognition method: First, after the server 20 obtains the target audio, it can slice the target audio to obtain multiple audio segments; then, the server 20 can obtain the melody fingerprints of each audio segment, and for each audio segment, look up the top K pre-stored melody fingerprints with the highest similarity to the melody fingerprint and the identifiers of the K matching songs corresponding to the top K pre-stored melody fingerprints in the fingerprint library. Through the identifiers of the K matching songs, the server 20 can determine the K matching songs that are vaguely matched to each audio segment; further, the server 20 performs decoding processing according to the similarity of the K matching songs corresponding to each audio segment (i.e., the similarity between the corresponding pre-stored melody fingerprint and the corresponding melody fingerprint), and the preset transition probability between the song pairs composed of a matching song corresponding to each adjacent audio segment, and determines the target matching song for each audio segment from the K matching songs of each audio segment; the server 20 determines the audio recognition result of the target audio according to the target matching song of each audio segment, and this audio recognition result can be used to indicate whether the target audio is a medley song.
[0086] Since the present application adopts a fuzzy matching method based on similarity, even if the fingerprint library does not store a song that exactly matches a certain audio segment, the present application can still obtain K matching songs that are vaguely matched to each audio segment. By retaining multiple matching songs for each audio segment, this method can improve the accuracy of determining the target matching song based on multiple matching songs, and thus improve the accuracy of audio recognition.
[0087] Optionally, the server 20 can further determine what type of medley song the target audio is according to the version identifier of the target matching song of each audio segment. For example, this version identifier can indicate that the target matching song is an original song and an adapted song. When the version identifiers of the target matching songs of all audio segments indicate that they are all adapted songs, the target audio is a medley song of the adapted type; when the version identifiers of the target matching songs of all audio segments indicate that they include both adapted songs and original songs, the target audio is a medley song of the mixed type of adapted and original.
[0088] Optionally, the above target audio can be stored in the server 20, or sent to the server 20 by any client, or sent to the server 20 by other servers. After the server 20 determines the audio recognition result of the target audio, it can perform operations such as deleting the song or taking the song offline on the target audio stored in itself according to the audio recognition result, or notify the client or other servers to perform operations such as deleting the song or taking the song offline.
[0089] Optionally, the audio recognition method proposed in this application can also be executed by any client, such as client 22. Among them, a fingerprint library can be pre-stored in client 22, or client 22 can download the fingerprint library from a server storing the fingerprint library in real time. Furthermore, when client 22 obtains the target audio, it can determine the audio recognition result of the target audio based on the fingerprint library and in the above manner.
[0090] It should be noted that the above server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms. The above client can be a terminal device, which can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, etc., but is not limited thereto.
[0091] The above briefly introduces the audio recognition system provided by the embodiments of this application. Next, the audio recognition method, audio recognition device, electronic device, computer-readable storage medium, etc. provided by the embodiments of this application will be described in detail respectively. Figures 3 to 10 Please refer to
[0092] Please refer to Figure 3 , Figure 3 is a schematic flowchart of an audio recognition method provided by an embodiment of this application. The method includes steps S301 to S304, and its execution subject can be a server or a client, or a chip in the server, or a chip in the client. Hereinafter, taking the server as the execution subject of the method as an example for description, the server can be the above Figure 2 The server 20 introduced in. Among them:
[0093] S301. The server performs a slicing operation on the target audio to obtain a plurality of audio segments, and obtains the melody fingerprint of each audio segment.
[0094] In the embodiments of the present application, the target audio may be an audio stored locally in the server or an audio obtained by the server from other devices, and this audio may be a song file. For example, the server may receive an audio sent by a client, where the audio is pre-stored in the client or is collected in real time by an audio collection device of the client; or for another example, the server may receive an audio sent by another server. By way of example, this other server may be a storage server corresponding to an audio platform, and the storage server may send the audio to the server for audio recognition before the audio goes online, or regularly send the audio to the server for audio recognition, etc. The embodiments of the present application do not limit the source of the target audio.
[0095] In a possible implementation manner, the server may first obtain a preset segment duration and a preset segment offset duration, and then perform a slicing operation on the target audio based on the preset segment duration and the preset segment offset duration to obtain a plurality of audio segments.
[0096] Among them, the preset segment duration is the duration of each audio segment, and the preset segment offset duration is the overlapping duration between two adjacent audio segments. For example, as Figure 4 shown, if the total duration of the target audio is 6 s, the preset segment duration is 2 s, and the preset segment offset duration is 1 s, then the server may slice the target audio into 5 audio segments. Among them, audio segment 1 is the segment of the target audio from the 0th s to the 2nd s, audio segment 2 is the segment of the target audio from the 1st s to the 3rd s, audio segment 3 is the segment of the target audio from the 2nd s to the 4th s, audio segment 4 is the segment of the target audio from the 3rd s to the 5th s, and audio segment 5 is the segment of the target audio from the 4th s to the 6th s. It can be Figure 4 seen that any two adjacent arranged audio segments have a part of overlapping information, and this overlapping information can ensure that after the server performs a slicing operation on the target audio, the information contained in the target audio will not be lost, and thus the complete information contained in the target audio can be fully utilized in subsequent processing.
[0097] In the embodiments of the present application, the melody fingerprint of each audio segment can be used to indicate the melody information contained in each audio segment, and this melody fingerprint can be expressed as a vector, etc., and the present application does not make any limitations in this regard.
[0098] In a possible implementation manner, the melody fingerprint is an embedding vector, and the manner in which the server obtains the melody fingerprint of each audio segment specifically includes: the server extracts the frequency domain features of each audio segment and inputs the frequency domain features of each audio segment into a pre-trained embedding vector generation model to obtain the melody fingerprint of each audio segment.
[0099] Among them, the frequency-domain features of each audio segment may include one or more of Mel Frequency Cepstrum Coefficient (MFCC) features, spectrogram features, features based on a filter bank (Fbank), Perceptual Linear Predictive (PLP) features, etc. The pre-trained embedding vector generation model can be a convolutional neural network model or a variant model of a convolutional neural network model, etc. For example, it can be a Deep Residual Network model (Resnet18). Resnet18 has added residual modules compared to a general convolutional neural network model, which can effectively avoid the problems of gradient disappearance and gradient explosion.
[0100] S302. For each audio segment, the server searches in the fingerprint database for the top K pre-stored melody fingerprints with the highest similarity to the melody fingerprint of the audio segment and the matching songs corresponding to the K pre-stored melody fingerprints, to obtain the K matching songs of the audio segment, where K is a positive integer.
[0101] Among them, the fingerprint database is a database storing the correspondence between pre-stored melody fingerprints and matching songs. The matching songs can be identified using information such as song names and performing artists.
[0102] After the server obtains the melody fingerprint of each audio segment, for each audio segment, it can calculate the similarity between the melody fingerprint of the audio segment and each pre-stored melody fingerprint in the fingerprint database. For example, when the melody fingerprint and the pre-stored melody fingerprint are embedding vectors, the similarity can be the cosine distance between the embedding vectors. When the cosine distance between the embedding vectors is smaller, the similarity between the melody fingerprint and the pre-stored melody fingerprint is higher; conversely, when the cosine distance between the embedding vectors is larger, the similarity between the melody fingerprint and the pre-stored melody fingerprint is lower. Then, the server sorts the similarities between the melody fingerprint and each pre-stored melody fingerprint to obtain the top K pre-stored melody fingerprints with the highest similarity to the melody fingerprint; based on the correspondence between the pre-stored melody fingerprints and the matching songs in the fingerprint database, the server can determine the K matching songs corresponding to the top K pre-stored melody fingerprints as the K matching songs of the audio segment. Based on this method, even if there is no song in the fingerprint database that exactly matches the audio segment, the server can obtain multiple matching songs that are vaguely matched to each audio segment.
[0103] S303. The server determines the target matching song for each audio segment from the K matching songs of each audio segment based on the preset transition probability between adjacent matching songs and the similarity corresponding to the K matching songs of each audio segment.
[0104] Among them, adjacent matching songs refer to song pairs composed of a matching song corresponding to each of adjacent audio segments, and the similarity corresponding to the matching song is the similarity between the pre-stored melody fingerprint corresponding to the matching song and the melody fingerprint of the audio segment corresponding to the matching song.
[0105] In a possible implementation manner, the number of multiple audio segments is S, and S is a positive integer greater than or equal to 3; the manner in which the server determines the target matching song for each audio segment based on the preset transition probability between adjacent matching songs and the similarity corresponding to the K matching songs of each audio segment specifically includes: constructing multiple decoding paths based on the preset transition probability between adjacent matching songs and the similarity corresponding to the K matching songs of each audio segment, and determining the path probabilities of the multiple decoding paths; taking the decoding path with the largest path probability among the multiple decoding paths as the target decoding path; and taking the matching songs of each audio segment indicated by each decoding sub-path in the target decoding path as the target matching song for each audio segment.
[0106] Among them, each decoding path includes S - 1 decoding sub-paths, and each decoding sub-path is used to indicate: a matching song of the m-th audio segment among the S audio segments, pointing to a matching song of the (m + 1)-th audio segment. The S audio segments are sorted in sequence according to their positions in the target audio, and m is a positive integer greater than or equal to 1 and less than S.
[0107] Exemplarily, as Figure 5 shown, the server divides the target audio into 5 audio segments, namely audio segment 1 to audio segment 5, and 5 matching songs are determined for each audio segment. For example, for audio segment 1, matching songs 1.1 to 1.5 are determined, for audio segment 2, matching songs 2.1 to 2.5 are determined, for audio segment 3, matching songs 3.1 to 3.5 are determined, for audio segment 4, matching songs 4.1 to 4.5 are determined, and for audio segment 5, matching songs 5.1 to 5.5 are determined. Exemplarily, a decoding path can be constructed by decoding sub-path 501, decoding sub-path 502, decoding sub-path 503, and decoding sub-path 504. Among them, decoding sub-path 501 points from matching song 1.1 to matching song 2.1, decoding sub-path 502 points from matching song 2.1 to matching song 3.2, decoding sub-path 503 points from matching song 3.2 to matching song 4.4, and decoding sub-path 504 points from matching song 4.4 to matching song 5.3.
[0108] In the embodiments of the present application, the server can first construct multiple decoding paths, and then determine the target matching song for each audio segment according to the multiple decoding paths. The following will first introduce in detail the manner in which the server constructs multiple decoding paths:
[0109] In a possible implementation, the server determines multiple decoding sub-paths at the first level and the sub-path probabilities of each decoding sub-path at the first level; sets i to 2; based on the preset transition probabilities between adjacent matching songs, the similarities corresponding to K first to-be-processed matching songs, and the similarities corresponding to K second to-be-processed matching songs, determines multiple decoding sub-paths at the i-th level and the sub-path probabilities of each decoding sub-path at the i-th level; if i is less than S - 1, performs an increment operation on i and returns to execute the step of "based on the preset transition probabilities between adjacent matching songs, the similarities corresponding to K first to-be-processed matching songs, and the similarities corresponding to K second to-be-processed matching songs, determines multiple decoding sub-paths at the i-th level and the sub-path probabilities of each decoding sub-path at the i-th level"; if i is equal to S - 1, sets p to 1; for the p-th matching song of the S-th audio segment, obtains the decoding path corresponding to the p-th matching song and the path probability of the decoding path according to the following steps: selects one decoding sub-path from multiple decoding sub-paths at each level from the first level to the S - 1 level, and the S - 1 selected decoding sub-paths form the decoding path corresponding to the p-th matching song; determines the path probability of the decoding path according to the sub-path probabilities of each decoding sub-path included in the decoding path; if p is less than K, performs an increment operation on p and returns to execute the step of "for the p-th matching song of the S-th audio segment, obtains the decoding path corresponding to the p-th matching song and the path probability of the decoding path according to the following steps"; if p is equal to K, ends the process.
[0110] Among them, any decoding sub-path at the i - 1 level is used to indicate: a matching song of the (i - 1)-th audio segment among the S audio segments, pointing to a matching song of the i-th audio segment. For example, for the Figure 5 decoding path mentioned above, the decoding sub-path 501 included in this decoding path can be understood as the decoding sub-path at the first level, the decoding sub-path 502 can be understood as the decoding sub-path at the second level, the decoding sub-path 503 can be understood as the decoding sub-path at the third level, and the decoding sub-path 504 can be understood as the decoding sub-path at the fourth level.
[0111] Among them, the first to-be-processed matching song is the matching song of the i-th audio segment indicated by a decoding sub-path at the i - 1 level; the second to-be-processed matching song is one of the K matching songs of the (i + 1)-th audio segment. For example, if Figure 5If the decoding sub-path 501 in [the above context] is a decoding sub-path of the first level, then when determining the decoding sub-paths of the second level (corresponding to i taking the value of 2), a first to-be-processed matching song can be the matching song 2.1 indicated by the decoding sub-path 501, and the second to-be-processed matching song can be any one of the matching songs 3.1 - 3.5. If the decoding sub-path 502 is a decoding sub-path of the second level, then when determining the decoding sub-paths of the third level (corresponding to i taking the value of 3), a first to-be-processed matching song can be the matching song 3.2 indicated by the decoding sub-path 502, and the second to-be-processed matching song can be any one of the matching songs 4.1 - 4.5, and so on.
[0112] In the embodiments of the present application, when the server constructs multiple decoding paths, it needs to first construct the decoding sub-paths of the first level to the (S - 1)th level in sequence (corresponding to the operation where i is less than S - 1), and then construct the decoding paths based on the decoding sub-paths of the first level to the (S - 1)th level (corresponding to the operation where i is equal to S - 1). And when the server constructs the decoding sub-paths of the second level to the (S - 1)th level, it needs to construct the decoding sub-paths of the next level according to the decoding sub-paths of the previous level.
[0113] The following first introduces the method for constructing the decoding sub-paths of the first level to the (S - 1)th level:
[0114] In a possible implementation manner, the way for the server to determine the multiple decoding sub-paths of the i-th level and the sub-path probabilities of each decoding sub-path of the i-th level based on the preset transition probabilities between adjacent matching songs, the similarities corresponding to the K first to-be-processed matching songs, and the similarities corresponding to the K second to-be-processed matching songs specifically includes: setting q to 1, and for the q-th second to-be-processed matching song among the K second to-be-processed matching songs, obtaining the decoding sub-path corresponding to the q-th second to-be-processed matching song in the i-th level and the sub-path probability of the decoding sub-path according to the following steps: determining K to-be-processed transition probabilities from the K first to-be-processed matching songs to the q-th second to-be-processed matching song according to the preset transition probabilities between adjacent matching songs; calculating the sub-path probabilities of each target decoding sub-path pointing to the q-th second to-be-processed matching song based on the similarities corresponding to the K first to-be-processed matching songs, the K to-be-processed transition probabilities, and the similarity corresponding to the q-th second to-be-processed matching song; taking the target decoding sub-path with the largest sub-path probability as the decoding sub-path corresponding to the q-th second to-be-processed matching song in the i-th level; if q is less than K, perform an increment operation on q, and return to execute the step of "obtaining the decoding sub-path corresponding to the q-th second to-be-processed matching song in the i-th level and the sub-path probability of the decoding sub-path according to the following steps"; if q is equal to K, end the process.
[0115] In the embodiments of the present application, when the server determines multiple decoding sub-paths at the i-th level: for each of the K second to-be-processed matching songs, a decoding sub-path pointing to the second to-be-processed matching song can be determined respectively. The decoding sub-path pointing to the second to-be-processed matching song is also called the decoding sub-path corresponding to the second to-be-processed matching song at the i-th level. The K decoding sub-paths pointing to the K second to-be-processed matching songs are the multiple decoding sub-paths at the i-th level. For example, in Figure 5 when determining multiple decoding sub-paths at the 2nd level, the decoding sub-path pointing to matching song 3.1, the decoding sub-path pointing to matching song 3.2,..., and the decoding sub-path pointing to matching song 3.5 are determined respectively; the 5 decoding sub-paths pointing to matching songs 3.1 to 3.5 are used as the multiple decoding sub-paths at the 2nd level. Among them, the decoding sub-path pointing to matching song 3.1 can also be called the decoding sub-path corresponding to matching song 3.1 at the 2nd level.
[0116] Specifically, when the server determines the decoding sub-path pointing to a second to-be-processed matching song (such as the q-th second to-be-processed matching song), it can first determine the K to-be-processed transition probabilities from the K first to-be-processed matching songs to this second to-be-processed matching song; then, according to the K to-be-processed transition probabilities, the similarities corresponding to the K first to-be-processed matching songs, and the similarity corresponding to this second to-be-processed matching song, determine the decoding sub-path pointing to a second to-be-processed matching song.
[0117] Optionally, the server can determine the K to-be-processed transition probabilities according to the following steps for determining a to-be-processed transition probability: determine the preset transition probability between adjacent matching songs corresponding to a song pair according to whether a first to-be-processed matching song is the same as the q-th second to-be-processed matching song, and this song pair is composed of this first to-be-processed matching song and the q-th second to-be-processed matching song; determine the to-be-processed transition probability from this first to-be-processed matching song to the q-th second to-be-processed matching song according to the preset transition probability between adjacent matching songs corresponding to this song pair.
[0118] In a possible implementation manner, if the first to-be-processed matching song is the same as the q-th second to-be-processed matching song, the preset transition probability between adjacent matching songs is the first transition probability; or, if the first to-be-processed matching song is different from the q-th second to-be-processed matching song, the preset transition probability between adjacent matching songs is the second transition probability; the first transition probability is greater than the second transition probability.
[0119] For example, if the first first song to be processed for matching is the same as the q-th second song to be processed for matching, the preset transition probability between adjacent matching songs corresponding to the song pair composed of the first first song to be processed for matching and the q-th second song to be processed for matching is the first transition probability; if the second first song to be processed for matching is different from the q-th second song to be processed for matching, the preset transition probability between adjacent matching songs corresponding to the song pair composed of the second first song to be processed for matching and the q-th second song to be processed for matching is the first transition probability.
[0120] In the embodiments of the present application, the server may directly determine the first transition probability or the second transition probability as the transition probability to be processed; alternatively, the server may first perform processing such as normalization on the first transition probability or the second transition probability, and then use the processed probability as the transition probability to be processed. The present application does not make any limitations in this regard.
[0121] It can be seen that when the first song to be processed for matching is the same as the second song to be processed for matching, the transition probability to be processed is greater; when the first song to be processed for matching is different from the second song to be processed for matching, the transition probability to be processed is smaller. Since in the real situation, the duration of each audio segment is short, it is generally more likely that adjacent audio segments are two segments of the same song, and less likely that they are two segments of different songs. Therefore, the transition probability to be processed determined in this way is more in line with the real situation. Based on this, the server can obtain a more accurate sub-path probability of the decoding sub-path, making each decoding sub-path determined at each level more accurate.
[0122] Next, taking the q-th second song to be processed for matching as the matching song 3.1 (corresponding to the value of 1 in the q area) as an example, a specific description will be given on how to determine the decoding sub-path corresponding to the q-th second song to be processed for matching in the i-th level according to the K transition probabilities to be processed, the similarities corresponding to the K first songs to be processed for matching, and the similarity corresponding to the q-th second song to be processed for matching:
[0123] Such as Figure 6As shown, when i takes the value of 2, if the decoding sub-paths at the first level include a total of 5 sub-paths, namely sub-path ① to sub-path ⑤, and the matching songs of the second audio segments indicated by sub-path ① to sub-path ⑤ are matching song 2.1, matching song 2.2, matching song 2.3, matching song 2.5, and matching song 2.4 respectively; then when the server determines the decoding sub-path corresponding to matching song 3.1 at the second level, the K first to-be-processed matching songs are matching song 2.1 to matching song 2.5. The server first determines 5 target decoding sub-paths based on matching song 3.1 and matching song 2.1 to matching song 2.5. These 5 target decoding sub-paths are, in sequence, sub-path ⑥ pointing from matching song 2.1 to matching song 3.1, sub-path ⑦ pointing from matching song 2.2 to matching song 3.1, sub-path ⑧ pointing from matching song 2.3 to matching song 3.1, sub-path ⑨ pointing from matching song 2.4 to matching song 3.1, and sub-path ⑩ pointing from matching song 2.5 to matching song 3.1. For sub-path ⑥ to sub-path ⑩, the server calculates the sub-path probability of the sub-path based on the respective similarities of the two matching songs in the sub-path and the to-be-processed transition probability related to these two matching songs. For example, for sub-path ⑥, the server can obtain the sub-path probability of sub-path ⑥ based on the similarity corresponding to matching song 2.1, the similarity corresponding to matching song 3.1, and the to-be-processed transition probability from matching song 2.1 to matching song 3.1. Exemplarily, the server can obtain the sub-path probability of sub-path ⑥ based on the following formula 1:
[0124] The sub-path probability of sub-path ⑥ = the similarity corresponding to matching song 2.1 * the to-be-processed transition probability from matching song 2.1 to matching song 3.1 + the similarity corresponding to matching song 3.1 (formula 1)
[0125] Furthermore, after obtaining the sub-path probabilities corresponding to sub-path ⑥ to sub-path ⑩ respectively, the server can select the sub-path with the maximum sub-path probability from sub-path ⑥ to sub-path ⑩ as the decoding sub-path corresponding to matching song 3.1 at the second level.
[0126] According to the above processing process, the server can also obtain the decoding sub-path corresponding to matching song 3.2 at the second level, the decoding sub-path corresponding to matching song 3.3 at the second level, the decoding sub-path corresponding to matching song 3.4 at the second level, and the decoding sub-path corresponding to matching song 3.5 at the second level.
[0127] Moreover, after the server determines the five decoding sub-paths at the second level as above, it can sequentially determine the five decoding sub-paths at the third level and the five decoding sub-paths at the fourth level, and so on according to the same process. And, for each level of matching songs, a decoding sub-path pointing to itself needs to be found, that is, 5 decoding sub-paths can be found for each level.
[0128] The following provides supplementary description on how the server determines the decoding sub-path of the first layer:
[0129] In a possible implementation manner, for each matching song of the second audio segment among the S audio segments, the server determines a decoding sub-path and the sub-path probability of the decoding sub-path according to the following steps: Determine whether the matching song is the same as any matching song of the first audio segment among the S audio segments; Determine the preset transition probability between adjacent matching songs according to the determination result, and determine the to-be-processed transition probability from any matching song of the first audio segment to the matching song of the second audio segment according to the preset transition probability between adjacent matching songs; Based on the K to-be-processed transition probabilities, the similarities corresponding to the K matching songs of the first audio segment, and the similarity corresponding to the matching song, determine the decoding sub-path corresponding to the matching song in the first layer and the sub-path probability of the decoding sub-path.
[0130] The following takes Figure 5 as an example to illustrate this possible implementation manner: Exemplarily, first, for the matching song 2.1 of audio segment 2, the server can determine the sub-path probability of the sub-path from each matching song in audio segment 1 to the matching song 2.1 according to the similarity of each matching song in audio segment 1, the to-be-processed transition probability from each matching song in audio segment 1 to the matching song 2.1, and the similarity corresponding to the matching song 2.1. Then, the server determines a sub-path pointing to the matching song 2.1 according to the sub-path probability of the sub-path from each matching song in audio segment 1 to the matching song 2.1, and takes this sub-path as the decoding sub-path corresponding to the matching song 2.1 in the first layer. Among them, the manner of determining the transition probability from each matching song to the matching song 2.1, and the manner of determining the sub-path probability of the sub-path from each matching song in audio segment 1 to the matching song 2.1 can refer to the above description and will not be elaborated here.
[0131] According to a similar process, the server can also determine the decoding sub-path corresponding to the matching song 2.2 in the first layer, the decoding sub-path corresponding to the matching song 2.3 in the first layer, the decoding sub-path corresponding to the matching song 2.4 in the first layer, and the decoding sub-path corresponding to the matching song 2.5 in the first layer.
[0132] Further, after the server determines the decoding sub-paths of the 1st to (S - 1)th levels according to the above method, multiple decoding paths can be constructed based on the decoding sub-paths of the 1st to (S - 1)th levels. Among them, the decoding sub-paths of adjacent levels in each decoding path are associated. The decoding sub-paths of adjacent levels being associated means that: the decoding sub-path of the next level is determined based on the decoding sub-path of the previous level, and the decoding sub-path of the previous level can be traced back from the decoding sub-path of the next level, or it can be understood that the matching song at the end point of the arrow in the decoding sub-path of the previous level is the same as the matching song at the starting point of the arrow in the decoding sub-path of the next level. For example, in Figure 6 if the sub-path ⑥ is the decoding sub-path of the 2nd level in the determined decoding path, then the decoding sub-path of the 1st level in this decoding path should be the sub-path ①.
[0133] In the embodiment of the present application, for each matching song of the Sth audio segment, the server can select a decoding sub-path from the multiple decoding sub-paths of the 1st to (S - 1)th levels, so as to obtain 1 decoding path corresponding to this matching song according to the selected decoding sub-path. For example, for the pth matching song of the Sth audio segment, first determine the decoding sub-path corresponding to the pth matching song in the (S - 1)th level, and then find the decoding sub-path of the (S - 2)th level associated with the decoding sub-path corresponding to the pth matching song in the (S - 1)th level; find the decoding sub-path of the (S - 3)th level associated with the decoding sub-path of the (S - 2)th level,..., until the decoding sub-path of the 1st level is found. The path formed by the S - 1 decoding sub-paths found is used as the decoding path corresponding to the pth matching song. According to a similar process, the server can finally determine the decoding paths corresponding to the 1st to Kth matching songs of the Sth audio segment, a total of K decoding paths.
[0134] Furthermore, the server determines the target decoding path according to the K decoding paths. The target decoding path is the decoding path with the highest path probability, and determines the target matching song of each audio segment from the K matching songs of each audio segment according to the target decoding path.
[0135] Among them, for any decoding path constructed by the server, the server can obtain the sub-path probabilities corresponding to the decoding sub-paths of each level included in any decoding path; then, based on the sub-path probabilities corresponding to the decoding sub-paths of each level, determine the path probability of any decoding path. For example, the server can determine the sum of the sub-path probabilities of the decoding sub-paths of each level included in the decoding path as the path probability of the decoding path.
[0136] Since the sub-path probability of the decoding sub-path is directly proportional to the similarity corresponding to the matching song, and the sub-path probability of the decoding sub-path is also directly proportional to the path probability of the decoding path to which it belongs, the path probability determined based on the sub-path probability can be used to indicate the probability that the matching song of each audio segment indicated by each decoding sub-path is the same as the song to which each audio segment actually belongs. When the path probability of the decoding path is larger, it indicates that the matching songs of each audio segment included in the decoding path are more likely to be the same as the songs to which each audio segment actually belongs. That is to say, in this application, the target decoding path with the largest path probability is the path that most conforms to the actual situation, and the target matching song obtained based on the target decoding path is the matching song that most conforms to the actual situation. For example, if Figure 5 the decoding path in it is the target decoding path, then the target matching songs of audio segments 1 to 5 are matching song 1.1, matching song 2.1, matching song 3.2, matching song 4.4, and matching song 5.3 respectively. Matching song 1.1, matching song 2.1, matching song 3.2, matching song 4.4, and matching song 5.3 are the matching songs that most conform to the actual situation of these 5 audio segments.
[0137] It should be noted that in specific implementation, the server can also construct multiple decoding paths and determine the target matching song of each audio segment based on decoding methods such as global traversal. This application does not make any limitations on this.
[0138] S304. The server determines the audio recognition result of the target audio according to the target matching song of each audio segment, and the audio recognition result is used to indicate whether the target audio is a medley song.
[0139] In a possible implementation manner, the server determines whether the target matching songs of each audio segment are all the same; if the target matching songs of each audio segment are all the same, it is determined that the audio recognition result of the target audio is used to indicate that the target audio is not a medley song; if there are at least two different target matching songs among the target matching songs of each audio segment, it is determined that the audio recognition result of the target audio is used to indicate that the target audio is a medley song.
[0140] Among them, when the server determines the target matching songs for each audio segment, the target matching songs can be identified by information such as the song name and the singer. Therefore, for example, the server can determine whether the target matching songs are the same by judging whether the information such as the song name and the singer of each target matching song is the same. When the information such as the song name and the singer of each target matching song is the same, it means that each audio segment of the target audio is from the same song, and the server can determine that the audio recognition result is used to indicate that the target audio is not a medley song. When there are different information in the information such as the song name and the singer of each target matching song, it means that each audio segment of the target audio is not from the same song, and the server can determine that the audio recognition result is used to indicate that the target audio is a medley song.
[0141] For example, the audio recognition result of the target audio can be indicated by "0" or "1", where "0" indicates that the target audio is not a medley song, and "1" indicates that the target audio is a medley song.
[0142] Furthermore, when the server determines that the audio recognition result is used to indicate that the target audio is a medley song, it can also determine the type of the medley song to which the target matching song belongs according to the version identifier of the target matching song. In a possible implementation manner, the server determines the version identifier of the target matching song for each audio segment; the version identifier includes a first version identifier or a second version identifier, the first version identifier is used to indicate that the corresponding target matching song is an original song, and the second version identifier is used to indicate that the corresponding target matching song is an adapted song; if the version identifier of the target matching song for each audio segment is the second version identifier, the audio recognition result of the target audio is also used to indicate that the target audio is a medley song of the adapted type; if the version identifiers of the target matching songs of multiple audio segments include the first version identifier and the second version identifier, the audio recognition result of the target audio is also used to indicate that the target audio is a medley song of the mixed type based on the original song and the adapted song.
[0143] For example, the server can determine the version identifier of each target matching song according to the information such as the song name and the singer of each target matching song, or the version identifier of each target matching song can also be stored in the fingerprint database. By determining the version identifier of each target matching song, the server can determine whether the target audio is a medley of multiple adapted songs or a medley of adapted songs and original songs.
[0144] In Figure 3In the described embodiments, the present application adopts a fuzzy matching method based on similarity. Therefore, even if the fingerprint database does not store a song that exactly matches a certain audio clip, the present application can obtain K matching songs that are fuzzily matched to each audio clip. By retaining multiple matching songs for each audio clip, this method can improve the accuracy of determining the target matching song based on multiple matching songs, and thus improve the accuracy of audio recognition.
[0145] The above describes the audio recognition method proposed in the embodiments of the present application. Next, the training process of the embedding vector generation model involved in this audio recognition method will be introduced: In a possible implementation manner, the server obtains a training audio set; the training audio set includes an original song, at least one adapted song corresponding to the original song, and at least one reference song, where the reference song is different from the original song and the adapted song; then, the server sequentially performs a slicing operation on each song in the training audio set to obtain N song segments of each song, where N is a positive integer, and the song content of the j-th song segment of the original song is the same as the song content of the j-th song segment of any of the adapted songs, and j is a positive integer less than or equal to N; the server extracts the frequency domain features of the N song segments of each song, inputs the frequency domain features of the N song segments of each song into the initial embedding vector generation model, and obtains the embedding vectors of the N song segments of each song; the server determines a first vector distance based on the embedding vectors of the N song segments of the original song and the embedding vectors of the N song segments of the adapted song; and the server determines a second vector distance based on the embedding vectors of the N song segments of the original song and the embedding vectors of the N song segments of the reference song; with the goal of reducing the first vector distance and increasing the second vector distance for training, the initial embedding vector generation model is trained to obtain a pre-trained embedding vector generation model.
[0146] Among them, the fact that the reference song is different from the original song and the adapted song means that the melody of the reference song is different from that of the original song and the adapted song. When the server performs a slicing operation on each song in the training audio set, this slicing operation matches the slicing operation in the above audio recognition method. For example, the preset segment duration and the preset segment offset duration used in the training process are the same as those in the above method embodiment.
[0147] Optionally, before performing the slicing operation, the server can also perform noise addition processing on each song in the training audio set. This noise addition processing can achieve data augmentation, making each song after noise addition more in line with the actual recording situation.
[0148] Next, the server can extract the frequency-domain features of each song segment obtained by the slicing operation, and input the frequency-domain features of each song segment into the initial embedding vector generation model to obtain the embedding vector of each song segment. Then, the server can determine the first vector distance based on the embedding vectors of the song segments of the original song and the embedding vectors of the song segments of the adapted song, and determine the second vector distance based on the embedding vectors of the song segments of the original song and the embedding vectors of the song segments of the reference song, determine the training objective (or called training loss) according to the first vector distance and the second vector distance, and obtain the embedding vector generation model according to the training objective.
[0149] For example, please refer to Figure 7 , Figure 7 which is a schematic diagram of a model training process provided by an embodiment of the present application. Among them, the server first performs noise addition and slicing processing on the original song, adapted song 1, adapted song 2, and reference song in sequence to obtain original segments 1 to original segment 10, adapted 1 segments 1 to adapted 1 segment 10, adapted 2 segments 1 to adapted 2 segment 10, and reference segments 1 to reference segment 10. Then the server extracts the MFCC features of each segment, and inputs the MFCC features of each extracted segment into the initial Resnet18 model to obtain the embedding features of each segment, determines the training loss according to the embedding features of each segment, and trains the initial Resnet18 model to obtain the Resnet18 model.
[0150] Exemplarily, if the embedding vector of the song segment of the original song is anchor, the embedding vector of the song segment of the adapted song is positive, and the embedding vector of the song segment of the reference song is negative, the training loss can be obtained by the following formula 2:
[0151] L = max(d(a,p) - d(a,n) + margin, 0) (Formula 2)
[0152] where d(a,p) represents the cosine distance between anchor and positive, that is, the first vector distance mentioned above, d(a,n) represents the cosine distance between anchor and negative, that is, the second vector distance mentioned above, and margin is the degree coefficient, which can take a value of 0.03, for example.
[0153] When the server trains the initial Resnet18 model based on L, it can follow Figure 8Train according to the training principle shown. Among them, the server can increase the value of d(a, n) in L as much as possible, so that the melody features of the original song and the reference song are more distant, and the server can reduce the value of d(a, p) in L as much as possible, so that the melody features of the original song and the adapted song are closer. When the value of d(a, n) is larger and the value of d(a, p) is smaller, the value of L is also smaller. If the value of L decreases to less than the preset threshold during the training process, the initial Resnet18 model is trained to completion to obtain the Resnet18 model.
[0154] It should be noted that the pre-trained embedding vector generation model obtained based on the above training method can first generate the embedding vectors stored in the fingerprint library in the above audio recognition method. Then, after each audio segment of the target audio is obtained, the embedding vector generation model generates an embedding vector for each audio segment.
[0155] Since the embedding vectors generated by the pre-trained embedding vector generation model have higher discrimination for songs with different melodies, and the vectors generated by the pre-trained embedding vector generation model have higher similarity for songs with similar melodies, higher-quality embedding vectors can be obtained based on the pre-trained embedding vector generation model. When the server performs fuzzy matching using higher-quality embedding vectors, more accurate matching songs can be obtained, thereby improving the accuracy of audio recognition.
[0156] Please refer to Figure 9 , Figure 9 FIG. is a schematic structural diagram of an audio recognition device provided by an embodiment of the present application. The audio recognition device includes a slicing module 901, an acquisition module 902, a search module 903, and a determination module 904.
[0157] Among them:
[0158] The slicing module 901 is used to perform slicing operations on the target audio to obtain multiple audio segments;
[0159] The acquisition module 902 is used to acquire the melody fingerprints of each of the audio segments;
[0160] The search module 903 is used to, for each of the audio segments, search in the fingerprint library for the top K pre-stored melody fingerprints with the highest similarity to the melody fingerprints of the audio segment and the matching songs corresponding to the K pre-stored melody fingerprints, to obtain the K matching songs of the audio segment; K is a positive integer;
[0161] A determination module 904, configured to determine, based on a preset transition probability between adjacent matching songs and the similarity corresponding to the K matching songs of each audio segment, a target matching song for each audio segment from the K matching songs of each audio segment; wherein the adjacent matching songs refer to song pairs composed of a matching song corresponding to each adjacent audio segment, and the similarity corresponding to the matching song is the similarity between the pre-stored melody fingerprint corresponding to the matching song and the melody fingerprint of the audio segment corresponding to the matching song; and determine an audio recognition result of the target audio according to the target matching song of each audio segment, where the audio recognition result is used to indicate whether the target audio is a medley song.
[0162] In a possible implementation manner, the number of the multiple audio segments is S, and S is a positive integer greater than or equal to 3; when determining, based on the preset transition probability between adjacent matching songs and the similarity corresponding to the K matching songs of each audio segment, a target matching song for each audio segment from the K matching songs of each audio segment, the determination module 904 is specifically configured to:
[0163] Based on the preset transition probability between the adjacent matching songs and the similarity corresponding to the K matching songs of each audio segment, construct multiple decoding paths and determine the path probability of each decoding path in the multiple decoding paths; wherein each decoding path includes S - 1 decoding sub-paths, and each decoding sub-path is used to indicate: a matching song of the m-th audio segment among the S audio segments, pointing to a matching song of the (m + 1)-th audio segment; the S audio segments are sorted in sequence according to their positions in the target audio, and m is a positive integer greater than or equal to 1 and less than S; take the decoding path with the maximum path probability in the multiple decoding paths as the target decoding path; and take the matching songs of each audio segment indicated by each decoding sub-path in the target decoding path as the target matching songs of each audio segment.
[0164] In a possible implementation manner, when constructing multiple decoding paths and determining the path probability of each decoding path in the multiple decoding paths based on the preset transition probability between the adjacent matching songs and the similarity corresponding to the K matching songs of each audio segment, the determination module 904 is specifically configured to:
[0165] Determine multiple decoding sub-paths at the first level and the sub-path probabilities of each decoding sub-path at the first level; set i to 2; based on the preset transition probabilities between adjacent matching songs, the similarities corresponding to K first to-be-processed matching songs, and the similarities corresponding to K second to-be-processed matching songs, determine multiple decoding sub-paths at the i-th level and the sub-path probabilities of each decoding sub-path at the i-th level; wherein, the first to-be-processed matching song is the matching song of the i-th audio segment indicated by a decoding sub-path at the i-1-th level; the second to-be-processed matching song is one of the K matching songs of the (i + 1)-th audio segment; if i is less than S - 1, perform an increment operation on i, and return to execute the step of determining multiple decoding sub-paths at the i-th level and the sub-path probabilities of each decoding sub-path at the i-th level based on the preset transition probabilities between adjacent matching songs, the similarities corresponding to K first to-be-processed matching songs, and the similarities corresponding to K second to-be-processed matching songs; if i is equal to S - 1, set p to 1; for the p-th matching song of the S-th audio segment, obtain the decoding path corresponding to the p-th matching song and the path probability of the decoding path according to the following steps: select one decoding sub-path from multiple decoding sub-paths at each level from the first level to the S - 1-th level, and form the decoding path corresponding to the p-th matching song by the selected S - 1 decoding sub-paths; the decoding sub-path at the S - 1-th level in the decoding path corresponding to the p-th matching song points to the p-th matching song of the S-th audio segment; determine the path probability of the decoding path according to the sub-path probabilities of each decoding sub-path included in the decoding path; if p is less than K, perform an increment operation on p, and return to execute the step of obtaining the decoding path corresponding to the p-th matching song of the S-th audio segment and the path probability of the decoding path according to the following steps; if p is equal to K, end the process.
[0166] In a possible implementation manner, when the determining module 904 determines multiple decoding sub-paths at the i-th level and the sub-path probabilities of each decoding sub-path at the i-th level based on the preset transition probabilities between adjacent matching songs, the similarities corresponding to K first to-be-processed matching songs, and the similarities corresponding to K second to-be-processed matching songs, it is specifically configured to:
[0167] Set q to 1. For the q-th second pending matching song among the K second pending matching songs, obtain the decoding sub-path corresponding to the q-th second pending matching song in the i-th level and the sub-path probability of the decoding sub-path according to the following steps: According to the preset transition probability between adjacent matching songs, determine K pending transition probabilities from the K first pending matching songs to the q-th second pending matching song; Based on the similarities corresponding to the K first pending matching songs, the K pending transition probabilities, and the similarity corresponding to the q-th second pending matching song, calculate the sub-path probabilities of each target decoding sub-path pointing to the q-th second pending matching song; Take the target decoding sub-path with the largest sub-path probability as the decoding sub-path corresponding to the q-th second pending matching song in the i-th level; If q is less than K, perform an increment operation on q, and return to execute the step of obtaining the decoding sub-path corresponding to the q-th second pending matching song in the i-th level and the sub-path probability of the decoding sub-path according to the following steps; If q is equal to K, end the process.
[0168] In a possible implementation manner, if the first pending matching song is the same as the q-th second pending matching song, the preset transition probability between adjacent matching songs is the first transition probability; or, if the first pending matching song is different from the q-th second pending matching song, the preset transition probability between adjacent matching songs is the second transition probability; The first transition probability is greater than the second transition probability.
[0169] In a possible implementation manner, when the determining module 904 determines the multiple decoding sub-paths in the first level and the sub-path probabilities of each decoding sub-path in the first level, it is specifically configured to:
[0170] For each matching song of the second audio segment among the S audio segments, determine a decoding sub-path and the sub-path probability of the decoding sub-path according to the following steps: Determine whether the matching song is the same as any matching song of the first audio segment among the S audio segments; Determine the preset transition probability between adjacent matching songs according to the judgment result, and determine the pending transition probability from any matching song of the first audio segment to the matching song according to the preset transition probability between adjacent matching songs; Based on the K pending transition probabilities, the similarities corresponding to the K matching songs of the first audio segment, and the similarity corresponding to the matching song, determine the decoding sub-path corresponding to the matching song in the first level and the sub-path probability of the decoding sub-path.
[0171] In a possible implementation manner, when determining the audio recognition result of the target audio according to the target matching songs of each audio segment, the determining module 904 is specifically configured to:
[0172] If the target matching songs of each audio segment are the same, determine that the audio recognition result of the target audio is used to indicate that the target audio is not a medley song; if there are at least two different target matching songs among the target matching songs of each audio segment, determine that the audio recognition result of the target audio is used to indicate that the target audio is a medley song.
[0173] In a possible implementation manner, the determining module 904 is further configured to:
[0174] When the audio recognition result of the target audio is used to indicate that the target audio is a medley song, determine the version identifier of the target matching song of each audio segment; the version identifier includes a first version identifier or a second version identifier, the first version identifier is used to indicate that the corresponding target matching song is an original song, and the second version identifier is used to indicate that the corresponding target matching song is an adapted song; if the version identifiers of the target matching songs of each audio segment are all the second version identifier, the audio recognition result of the target audio is further used to indicate that the target audio is an adapted medley song; if the version identifiers of the target matching songs of each audio segment include the first version identifier and the second version identifier, the audio recognition result of the target audio is further used to indicate that the target audio is a medley song of a mixed type based on the original song and the adaptation.
[0175] In a possible implementation manner, the melody fingerprint is an embedding vector. When the obtaining module 902 obtains the melody fingerprint of each audio segment, it is specifically configured to:
[0176] Extract the frequency domain features of each audio segment; input the frequency domain features of each audio segment into a pre-trained embedding vector generation model to obtain the melody fingerprint of each audio segment.
[0177] In a possible implementation manner, the audio recognition device further includes a training module, and the training module is used for:
[0178] Obtain a training audio set through the obtaining module 902; the training audio set includes an original song, at least one adapted song corresponding to the original song, and at least one reference song, and the reference song is different from the original song and the adapted song;
[0179] The slicing module 901 sequentially performs slicing operations on each song in the training audio set to obtain N song segments of each song, where N is a positive integer, and the song content of the j-th song segment of the original song is the same as the song content of the j-th song segment of any of the adapted songs, and j is a positive integer less than or equal to N;
[0180] The obtaining module 902 extracts the frequency domain features of the N song segments of each song, and inputs the frequency domain features of the N song segments of each song into the initial embedding vector generation model to obtain the embedding vectors of the N song segments of each song;
[0181] The determining module 904 determines a first vector distance according to the embedding vectors of the N song segments of the original song and the embedding vectors of the N song segments of the adapted song;
[0182] The determining module 904 determines a second vector distance according to the embedding vectors of the N song segments of the original song and the embedding vectors of the N song segments of the reference song;
[0183] Taking reducing the first vector distance and increasing the second vector distance as the training objective, the initial embedding vector generation model is trained to obtain the pre-trained embedding vector generation model.
[0184] It should be noted that the functions of the modules of the audio recognition device in the embodiments of the present application can be specifically implemented according to the methods in the above method embodiments, and the specific implementation process and beneficial effects can refer to the relevant descriptions of the above method embodiments, which will not be elaborated here.
[0185] Please refer to Figure 10 , Figure 10 which is a schematic structural diagram of an electronic device provided in the embodiments of the present application. The electronic device can be the server or the client in the above method embodiments, and the electronic device can include: one or more processors 1001 and a memory 1002. Optionally, the electronic device can further include a transceiver 1003. The above processors 1001, memory 1002 and transceiver 1003 can be connected through a bus 1004. The memory 1002 is used to store a computer program, and the computer program includes program instructions. The processor 1001 performs the following operations by running the program instructions stored in the memory 1002:
[0186] Performing a slicing operation on the target audio to obtain a plurality of audio segments, and obtaining the melody fingerprint of each of the audio segments;
[0187] For each of the audio segments, find the top K pre-stored melody fingerprints with the highest similarity to the melody fingerprint of the audio segment in the fingerprint database, as well as the matching songs corresponding to the K pre-stored melody fingerprints, to obtain the K matching songs of the audio segment; K is a positive integer;
[0188] Based on the preset transition probability between adjacent matching songs and the similarity corresponding to the K matching songs of each audio segment, determine the target matching song of each audio segment from the K matching songs of each audio segment; wherein the adjacent matching songs refer to a pair of songs composed of a matching song corresponding to each of the adjacent audio segments, and the similarity corresponding to the matching song is the similarity between the pre-stored melody fingerprint corresponding to the matching song and the melody fingerprint of the audio segment corresponding to the matching song;
[0189] According to the target matching song of each audio segment, determine the audio recognition result of the target audio, and the audio recognition result is used to indicate whether the target audio is a medley song.
[0190] In a possible implementation manner, the number of the multiple audio segments is S, and S is a positive integer greater than or equal to 3; when the processor 1001 determines the target matching song of each audio segment from the K matching songs of each audio segment based on the preset transition probability between adjacent matching songs and the similarity corresponding to the K matching songs of each audio segment, it is specifically configured to:
[0191] Based on the preset transition probability between adjacent matching songs and the similarity corresponding to the K matching songs of each audio segment, construct multiple decoding paths and determine the path probability of each decoding path in the multiple decoding paths; wherein each decoding path includes S - 1 decoding sub-paths, and each decoding sub-path is used to indicate: a matching song of the m-th audio segment among the S audio segments, pointing to a matching song of the (m + 1)-th audio segment; the S audio segments are sorted in sequence according to their positions in the target audio, and m is a positive integer greater than or equal to 1 and less than S; take the decoding path with the largest path probability in the multiple decoding paths as the target decoding path; and take the matching songs of each audio segment indicated by each decoding sub-path in the target decoding path as the target matching song of each audio segment.
[0192] In a possible implementation manner, when the processor 1001 constructs multiple decoding paths and determines the path probability of each decoding path in the multiple decoding paths based on the preset transition probability between adjacent matching songs and the similarity corresponding to the K matching songs of each audio segment, it is specifically configured to:
[0193] Determine multiple decoding sub-paths at the first level and the sub-path probabilities of each decoding sub-path at the first level; set i to 2; based on the preset transition probability between adjacent matching songs, the similarities corresponding to K first to-be-processed matching songs, and the similarities corresponding to K second to-be-processed matching songs, determine multiple decoding sub-paths at the i-th level and the sub-path probabilities of each decoding sub-path at the i-th level; wherein, the first to-be-processed matching song is the matching song of the i-th audio segment indicated by a decoding sub-path at the i-1 level; the second to-be-processed matching song is one of the K matching songs of the (i + 1)-th audio segment; if i is less than S - 1, perform an increment operation on i, and return to execute the step of determining multiple decoding sub-paths at the i-th level and the sub-path probabilities of each decoding sub-path at the i-th level based on the preset transition probability between adjacent matching songs, the similarities corresponding to K first to-be-processed matching songs, and the similarities corresponding to K second to-be-processed matching songs; if i is equal to S - 1, set p to 1; for the p-th matching song of the S-th audio segment, obtain the decoding path corresponding to the p-th matching song and the path probability of the decoding path according to the following steps: select one decoding sub-path from multiple decoding sub-paths at each level from the first level to the S - 1 level, and form the decoding path corresponding to the p-th matching song by the selected S - 1 decoding sub-paths; the decoding sub-path at the S - 1 level in the decoding path corresponding to the p-th matching song points to the p-th matching song of the S-th audio segment; determine the path probability of the decoding path according to the sub-path probabilities of each decoding sub-path included in the decoding path; if p is less than K, perform an increment operation on p, and return to execute the step of obtaining the decoding path corresponding to the p-th matching song of the S-th audio segment and the path probability of the decoding path according to the following steps; if p is equal to K, end the process.
[0194] In a possible implementation manner, when the processor 1001 determines multiple decoding sub-paths at the i-th level and the sub-path probabilities of each decoding sub-path at the i-th level based on the preset transition probability between adjacent matching songs, the similarities corresponding to K first to-be-processed matching songs, and the similarities corresponding to K second to-be-processed matching songs, it is specifically configured to:
[0195] Set q to 1. For the q-th second pending matching song among the K second pending matching songs, obtain the decoding sub-path corresponding to the q-th second pending matching song in the i-th level and the sub-path probability of the decoding sub-path according to the following steps: Determine K pending transition probabilities from the K first pending matching songs to the q-th second pending matching song according to the preset transition probability between adjacent matching songs; Based on the similarities corresponding to the K first pending matching songs, the K pending transition probabilities, and the similarity corresponding to the q-th second pending matching song, calculate the sub-path probabilities of the respective target decoding sub-paths pointing to the q-th second pending matching song; Take the target decoding sub-path with the largest sub-path probability as the decoding sub-path corresponding to the q-th second pending matching song in the i-th level; If q is less than K, perform an increment operation on q, and return to execute the step of obtaining the decoding sub-path corresponding to the q-th second pending matching song in the i-th level and the sub-path probability of the decoding sub-path according to the following steps; If q is equal to K, end the process.
[0196] In a possible implementation manner, if the first pending matching song is the same as the q-th second pending matching song, the preset transition probability between adjacent matching songs is the first transition probability; or, if the first pending matching song is different from the q-th second pending matching song, the preset transition probability between adjacent matching songs is the second transition probability; the first transition probability is greater than the second transition probability.
[0197] In a possible implementation manner, when the processor 1001 determines the multiple decoding sub-paths in the first level and the sub-path probabilities of the respective decoding sub-paths in the first level, it is specifically configured to:
[0198] For each matching song of the second audio segment among the S audio segments, determine a decoding sub-path and the sub-path probability of the decoding sub-path according to the following steps: Determine whether the matching song is the same as any matching song of the first audio segment among the S audio segments; Determine the preset transition probability between adjacent matching songs according to the judgment result, and determine the pending transition probability from any matching song of the first audio segment to the matching song according to the preset transition probability between adjacent matching songs; Based on the K pending transition probabilities, the similarities corresponding to the K matching songs of the first audio segment, and the similarity corresponding to the matching song, determine the decoding sub-path corresponding to the matching song in the first level and the sub-path probability of the decoding sub-path.
[0199] In a possible implementation, when determining the audio recognition result of the target audio by matching the target songs for each of the audio segments, the processor 1001 is specifically configured to:
[0200] If the target matching songs for each of the audio segments are the same, determine that the audio recognition result of the target audio is used to indicate that the target audio is not a medley song; if there are at least two different target matching songs among the target matching songs for each of the audio segments, determine that the audio recognition result of the target audio is used to indicate that the target audio is a medley song.
[0201] In a possible implementation, the processor 1001 is further configured to:
[0202] When the audio recognition result of the target audio is used to indicate that the target audio is a medley song, determine the version identifier of the target matching song for each of the audio segments; the version identifier includes a first version identifier or a second version identifier, the first version identifier is used to indicate that the corresponding target matching song is an original song, and the second version identifier is used to indicate that the corresponding target matching song is an adapted song; if the version identifiers of the target matching songs for each of the audio segments are all the second version identifier, the audio recognition result of the target audio is further used to indicate that the target audio is an adapted medley song; if the version identifiers of the target matching songs for each of the audio segments include the first version identifier and the second version identifier, the audio recognition result of the target audio is further used to indicate that the target audio is a mixed medley song based on the original and adapted songs.
[0203] In a possible implementation, the melody fingerprint is an embedding vector. When obtaining the melody fingerprint of each of the audio segments, the processor 1001 is specifically configured to:
[0204] Extract the frequency domain features of each of the audio segments; input the frequency domain features of each of the audio segments into a pre-trained embedding vector generation model to obtain the melody fingerprint of each of the audio segments.
[0205] In a possible implementation, the processor 1001 is further configured to:
[0206] Obtain a training audio set; the training audio set includes an original song, at least one adapted song corresponding to the original song, and at least one reference song, where the reference song is different from the original song and the adapted song; perform slicing operations on each song in the training audio set in sequence to obtain N song segments of each song, N is a positive integer, and the song content of the j-th song segment of the original song is the same as the song content of the j-th song segment of any of the adapted songs, j is a positive integer less than or equal to N; extract the frequency domain features of the N song segments of each song, and input the frequency domain features of the N song segments of each song into an initial embedding vector generation model to obtain the embedding vectors of the N song segments of each song; determine a first vector distance according to the embedding vectors of the N song segments of the original song and the embedding vectors of the N song segments of the adapted song; determine a second vector distance according to the embedding vectors of the N song segments of the original song and the embedding vectors of the N song segments of the reference song; use reducing the first vector distance and increasing the second vector distance as the training objective to train the initial embedding vector generation model to obtain the pre-trained embedding vector generation model.
[0207] It should be understood that in some feasible implementation manners, the above-mentioned processor 1001 may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc. The above-mentioned memory 1002 may include a read-only memory and a random access memory, and provide instructions and data to the processor 1001. A part of the memory 1002 may also include a non-volatile random access memory. For example, the memory 1002 may also store information about the device type.
[0208] In specific implementation, the above-mentioned electronic device may execute the implementation manners provided above through its built-in various functional modules Figures 3 - 8 The specific implementation process and beneficial effects can refer to the specific content of the above method embodiments, and will not be elaborated here.
[0209] The embodiments of the present application further provide a computer-readable storage medium, and the computer-readable storage medium stores a computer program executed by the aforementioned audio recognition device. The computer program includes program instructions. When the processor executes the above program instructions, it can execute the above Figures 3 - 8 content. Therefore, it will not be elaborated here. In addition, the beneficial effects of using the same method will not be elaborated either. For the technical details not disclosed in the embodiments of the computer-readable storage medium involved in the present application, please refer to the description of the method embodiments of the present application. As an example, the program instructions can be deployed on an electronic device, or executed on multiple electronic devices located at one place. Or, executed on multiple electronic devices distributed at multiple locations and interconnected through a communication network. The multiple electronic devices distributed at multiple locations and interconnected through a communication network can form a blockchain system.
[0210] According to one aspect of the present application, there is also provided a computer program product. The computer program product includes a computer program, and the computer program is stored in a computer-readable storage medium and includes program instructions. The processor of the electronic device reads the program instructions from the computer-readable storage medium, and the processor executes the program instructions, so that the electronic device can execute the above Figures 3 - 8 content. Therefore, it will not be elaborated here.
[0211] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The above program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. Among them, the above storage medium can be a magnetic disk, an optical disc, a read-only memory (ROM), or a random access memory (RAM), etc.
[0212] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or substitutions within the technical scope disclosed by the present invention. These modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. An audio recognition method, characterized in that: The method comprises: Slicing the target audio to obtain multiple audio segments, and obtaining a melody fingerprint of each of the audio segments; For each of the audio clips, searching in the fingerprint library for the first K pre-stored melody fingerprints with the highest similarity to the melody fingerprint of the audio clip and the matching songs corresponding to the K pre-stored melody fingerprints, to obtain K matching songs of the audio clip; K is a positive integer; Based on the preset transition probability between adjacent matching songs and the similarities corresponding to the K matching songs of each audio clip, a target matching song for each audio clip is determined from the K matching songs of each audio clip; wherein the adjacent matching songs refer to a song pair consisting of a matching song corresponding to each of the adjacent audio clips; if the two songs constituting the adjacent matching songs are the same, the preset transition probability between the adjacent matching songs is a first transition probability; if the two songs constituting the adjacent matching songs are different, the preset transition probability between the adjacent matching songs is a second transition probability, and the first transition probability is greater than the second transition probability; the similarities corresponding to the matching songs are the similarities between the pre-stored melody fingerprint corresponding to the matching songs and the melody fingerprint of the audio clip corresponding to the matching songs; According to the target matching song of each of the audio clips, an audio recognition result of the target audio is determined, and the audio recognition result is used to indicate whether the target audio is a medley song.
2. The method according to claim 1, characterized in that The number of the multiple audio clips is S, where S is a positive integer greater than or equal to 3; The step of determining a target matching song for each audio segment from the K matching songs of each audio segment based on a preset transition probability between adjacent matching songs and the similarities corresponding to the K matching songs of each audio segment comprises: Based on the preset transition probabilities between the adjacent matching songs and the similarities corresponding to the K matching songs of each of the audio clips, multiple decoding paths are constructed and the path probability of each decoding path in the multiple decoding paths is determined; wherein each decoding path includes S-1 decoding sub-paths, and each decoding sub-path is used to indicate: a matching song of the m-th audio clip among the S audio clips, pointing to a matching song of the m+1-th audio clip; the S audio clips are sorted in sequence according to their positions in the target audio, and m is a positive integer greater than or equal to 1 and less than S; Taking the decoding path with the highest path probability among the multiple decoding paths as the target decoding path; The matching song of each audio segment indicated by each decoding sub-path in the target decoding path is used as the target matching song of each audio segment.
3. The method according to claim 2, characterized in that The step of constructing multiple decoding paths and determining the path probability of each decoding path in the multiple decoding paths based on the preset transition probability between the adjacent matching songs and the similarity corresponding to the K matching songs of each of the audio clips includes: Determining a plurality of decoding subpaths of a first level and a subpath probability of each decoding subpath of the first level; i is set to 2; based on the preset transition probability between the adjacent matching songs, the similarities corresponding to the K first matching songs to be processed, and the similarities corresponding to the K second matching songs to be processed, multiple decoding sub-paths of the i-th level and the sub-path probabilities of each decoding sub-path of the i-th level are determined; wherein the first matching song to be processed is a matching song of the i-th audio clip indicated by a decoding sub-path of the i-1-th level; the second matching song to be processed is a matching song among the K matching songs of the i+1-th audio clip; If i is less than S-1, perform an operation of adding 1 to i, and return to the step of determining multiple decoding sub-paths of the i-th level and the sub-path probability of each decoding sub-path of the i-th level based on the preset transition probability between the adjacent matching songs, the similarities corresponding to the K first matching songs to be processed, and the similarities corresponding to the K second matching songs to be processed; If i is equal to S-1, set p to 1; for the p-th matching song of the S-th audio clip, obtain the decoding path corresponding to the p-th matching song and the path probability of the decoding path according to the following steps: select one decoding sub-path from multiple decoding sub-paths of each level from the 1st level to the S-1th level, and the selected S-1 decoding sub-paths constitute the decoding path corresponding to the p-th matching song; the decoding sub-path of the S-1th level in the decoding path corresponding to the p-th matching song points to the p-th matching song of the S-th audio clip; determine the path probability of the decoding path according to the sub-path probabilities of each decoding sub-path contained in the decoding path; if p is less than K, perform an addition operation on p, and return to the step of obtaining the decoding path corresponding to the p-th matching song and the path probability of the decoding path according to the following steps for the p-th matching song of the S-th audio clip; if p is equal to K, end the process.
4. The method according to claim 3, characterized in that The method of determining a plurality of decoding sub-paths of the i-th level and a sub-path probability of each decoding sub-path of the i-th level based on the preset transition probability between the adjacent matching songs, the similarities corresponding to the K first matching songs to be processed, and the similarities corresponding to the K second matching songs to be processed includes: Set q to 1, and for the qth second to-be-processed matching song among the K second to-be-processed matching songs, obtain the decoding subpath corresponding to the qth second to-be-processed matching song in the i-th level and the subpath probability of the decoding subpath according to the following steps: According to the preset transition probabilities between the adjacent matching songs, determining K to-be-processed transition probabilities from the K first to-be-processed matching songs to the qth second to-be-processed matching songs; Based on the similarities corresponding to the K first to-be-processed matching songs, the K to-be-processed transition probabilities, and the similarities corresponding to the qth second to-be-processed matching songs, calculating the subpath probabilities of each target decoding subpath pointing to the qth second to-be-processed matching song; The target decoding subpath with the largest subpath probability is used as the decoding subpath corresponding to the qth second matching song to be processed in the i-th level; If q is less than K, then add 1 to q and return to the step of obtaining the decoding sub-path corresponding to the qth second matching song to be processed in the i-th level and the sub-path probability of the decoding sub-path according to the following steps; if q is equal to K, then end the process.
5. The method according to claim 4, characterized in that If the first to-be-processed matching song is the same as the qth second to-be-processed matching song, the preset transition probability between the adjacent matching songs is a first transition probability; Alternatively, if the first matching song to be processed is different from the qth second matching song to be processed, the preset transition probability between the adjacent matching songs is the second transition probability; and the first transition probability is greater than the second transition probability.
6. The method according to claim 3, characterized in that: The determining of a plurality of decoding sub-paths of the first level and a sub-path probability of each decoding sub-path of the first level comprises: For each matching song of the second audio segment of the S audio segments, a decoding sub-path and a sub-path probability of the decoding sub-path are determined according to the following steps: Determining whether the matching song is the same as any matching song of the first audio segment among the S audio segments; Determine a preset transition probability between the adjacent matching songs according to the judgment result, and determine a to-be-processed transition probability from any matching song of the first audio segment to the matching song according to the preset transition probability between the adjacent matching songs; Based on the K to-be-processed transition probabilities, the similarities corresponding to the K matching songs of the first audio segment, and the similarities corresponding to the matching songs, the decoding sub-paths corresponding to the matching songs in the first level and the sub-path probabilities of the decoding sub-paths are determined.
7. The method according to any one of claims 1 to 6, characterized in that The step of determining the audio recognition result of the target audio according to the target matching song of each of the audio clips comprises: If the target matching songs of each of the audio clips are the same, determining the audio recognition result of the target audio is used to indicate that the target audio is not a medley song; If there are at least two target matching songs that are different among the target matching songs of each of the audio clips, then determining the audio recognition result of the target audio is used to indicate that the target audio is a medley song.
8. The method according to any one of claims 1 to 6, characterized in that The method further comprises: When the audio recognition result of the target audio is used to indicate that the target audio is a medley song, determine the version identifier of the target matching song of each of the audio clips; the version identifier includes a first version identifier or a second version identifier, the first version identifier is used to indicate that the corresponding target matching song is an original song, and the second version identifier is used to indicate that the corresponding target matching song is an adapted song; If the version identifier of the target matching song of each of the audio clips is the second version identifier, the audio recognition result of the target audio is also used to indicate that the target audio is an adapted medley song; If the version identifier of the target matching song of each of the audio clips includes the first version identifier and the second version identifier, the audio recognition result of the target audio is also used to indicate that the target audio is a mixed type medley song based on the original version and the adaptation.
9. The method according to any one of claims 1 to 6, characterized in that: The melody fingerprint is an embedded vector, and obtaining the melody fingerprint of each audio clip includes: Extracting frequency domain features of each of the audio clips; The frequency domain features of each of the audio clips are input into a pre-trained embedding vector generation model to obtain a melody fingerprint of each of the audio clips.
10. The method according to claim 9, characterized in that The method further comprises: Acquire a training audio set; the training audio set includes an original song, at least one adapted song corresponding to the original song, and at least one reference song, wherein the reference song is different from the original song and the adapted song; Performing a slicing operation on each song in the training audio set in turn to obtain N song segments of each song, where N is a positive integer, and the song content of the j-th song segment of the original song is the same as the song content of the j-th song segment of any of the adapted songs, and j is a positive integer less than or equal to N; Extracting frequency domain features of the N song segments of each song, and inputting the frequency domain features of the N song segments of each song into an initial embedding vector generation model to obtain embedding vectors of the N song segments of each song; Determine a first vector distance according to the embedding vectors of the N song segments of the original song and the embedding vectors of the N song segments of the adapted song; Determining a second vector distance according to the embedding vectors of the N song segments of the original song and the embedding vectors of the N song segments of the reference song; The initial embedding vector generation model is trained with the training goal of reducing the first vector distance and increasing the second vector distance to obtain the pre-trained embedding vector generation model.
11. An electronic device, characterized in that: The electronic device includes a memory and a processor; The memory is used to store a computer program, wherein the computer program includes program instructions; The processor is used to call the program instructions from the memory so that the electronic device executes the method according to any one of claims 1-10.
12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions. When the program instructions are executed by a processor, the processor executes the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Audio track identification method and device and readable storage medium
CN113468369A
Audio recognition method and device, electronic equipment and computer readable storage medium
CN115346515A