Audio recognition method and computer device
Through joint learning and measurement learning of the target audio recognition model, more accurate feature vectors are generated, which solves the problem of low recognition accuracy of cover songs and realizes more efficient cover song recognition and copyright management.
Patent Information
- Application Number
- CN202210719204.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-23
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2042-06-23
AI Technical Summary
In the prior art, the accuracy of cover song recognition is low and the robustness is insufficient, resulting in low accuracy of cover recognition, making it difficult to effectively manage massive music information and copyright management.
The target audio recognition model is adopted, and the parameters of the initial audio recognition model are adjusted through the first task module and the second task module of joint learning, and more accurate feature vectors to be identified are generated. The metric learning module is used to improve the similarity matching of the feature vectors, and the audio information in the music library is used for identification.
It improves the accuracy and efficiency of cover song recognition, and can more accurately find different versions of the same audio from the song library, enhancing the effects of song management and copyright protection.
Smart Images

Figure CN115101052B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an audio recognition method and computer equipment. Background Art
[0002] In recent years, with the rise of short videos and the emergence of a large amount of user-generated content (UGC) online, music recognition has become increasingly important when people encounter interesting music while watching short videos and other multimedia content and want to know the name of the music, the artist, and other information. Furthermore, the large number of audio and video works available online poses a significant challenge to song copyright management. Using music information retrieval (MIR) technologies, such as cover song recognition, to identify different versions of the same work, is crucial for song management and copyright management. Therefore, cover song identification (CSI) has become a new research hotspot. Summary of the Invention
[0003] This application provides an audio recognition method and computer device, which can effectively improve the accuracy and efficiency of audio recognition.
[0004] In a first aspect, the present application provides an audio recognition method, comprising:
[0005] Inputting the to-be-recognized spectrogram corresponding to the to-be-recognized audio segment into a target audio recognition model to obtain a to-be-recognized feature vector output by the target audio recognition model; wherein the target audio recognition model is obtained by adjusting model parameters of an initial audio recognition model using an adjustment parameter, the initial audio recognition model including a first task module and a second task module, and the adjustment parameter is determined based on a first loss parameter generated by the first task module and a second loss parameter generated by the second task module;
[0006] Determine a target feature vector from the music library that satisfies a preset condition with the feature vector to be identified;
[0007] The target audio pointed to by the target feature vector is determined as a recognition result of the audio segment to be recognized, and the recognition result indicates that the audio segment to be recognized and the target audio are different versions of the same audio.
[0008] It can be seen that the use of the target audio recognition model can obtain more accurate feature vectors to be identified, thereby improving the efficiency of querying matching target feature vectors from the music library and improving the accuracy of audio recognition.
[0009] In one implementation, the method further includes: inputting the training spectrogram into an initial audio recognition model to obtain a first training feature vector output by the first task module and a second training feature vector output by the second task module; determining a first loss parameter based on the first training feature vector; determining a second loss parameter based on the second training feature vector; and determining an adjustment parameter based on the first loss parameter and the second loss parameter, wherein the first training feature vector and the second training feature vector are different.
[0010] In one implementation, the method further includes: determining a predicted audio category label of the training spectrogram based on the first training feature vector; determining a prediction probability corresponding to the predicted audio category label; and determining a first loss parameter based on the prediction probability.
[0011] In one implementation, the training spectrum graph includes a first sample graph, a second sample graph, and a third sample graph, the first sample graph and the second sample graph have the same audio category label, and the first sample graph and the third sample graph have different audio category labels; the above method may further include: determining a first vector distance between a second training feature vector corresponding to the first sample graph and a second training feature vector corresponding to the second sample graph; determining a second vector distance between the second training feature vector corresponding to the first sample graph and the second training feature vector corresponding to the third sample graph; and determining a second loss parameter based on the first vector distance and the second vector distance.
[0012] It can be seen that by adjusting the model parameters of the initial audio recognition model using the adjustment parameters obtained by joint learning of the first task module and the second task module, the target audio recognition model can generate feature vectors to be recognized that are more accurate and robust, which is conducive to improving the accuracy and efficiency of audio recognition.
[0013] In one implementation, the recognition result also includes the audio category label of the audio segment to be recognized; the above method also includes: adding the correspondence between the audio category label of the audio segment to be recognized and the feature vector to be recognized to the music library.
[0014] In one implementation, the above method also includes: inputting the spectrum of the historical audio in the music library into the target audio recognition model to obtain the third eigenvector output by the target audio recognition model; the historical audio has an audio category label; and adding the correspondence between the audio category label of the historical audio and the third eigenvector to the music library.
[0015] In one implementation, the method further includes: calculating the similarity between the feature vector to be identified and a third feature vector in the music library; if the similarity between the feature vector to be identified and the third feature vector meets a preset condition, determining the third feature vector as the target feature vector.
[0016] In a second aspect, the present application provides an audio recognition method, comprising:
[0017] In response to a recognition request for a humming audio segment, obtaining a humming spectrogram of the humming audio segment;
[0018] Inputting the humming spectrogram into a target audio recognition model to generate a humming vector to be recognized; the target audio recognition model is obtained by adjusting model parameters of an initial audio recognition model using an adjustment parameter, the initial audio recognition model including a first task module and a second task module, the adjustment parameter being determined based on a first loss parameter generated by the first task module and a second loss parameter generated by the second task module;
[0019] Determine from the music library a similar feature vector that satisfies a preset condition with the humming vector to be identified, and determine the audio pointed to by the similar feature vector as the similar audio of the humming audio segment;
[0020] Output audio information of similar audio, where the audio information includes the audio name and singer name of the similar audio.
[0021] It can be seen that the humming vector to be identified obtained by the target audio recognition model can more accurately represent the humming audio segment, which is conducive to improving the efficiency of finding similar audio from the music library and improving the accuracy of humming recognition.
[0022] In a third aspect, the present application provides an audio recognition device.
[0023] In one possible design, the audio recognition device includes a processing unit and a retrieval unit.
[0024] a processing unit, configured to input a to-be-recognized spectrogram corresponding to the to-be-recognized audio segment into a target audio recognition model, and obtain a to-be-recognized feature vector output by the target audio recognition model; wherein the target audio recognition model is obtained by adjusting model parameters of the initial audio recognition model using an adjustment parameter, the adjustment parameter being determined based on a first loss parameter generated by the first task module and a second loss parameter generated by the second task module;
[0025] The retrieval unit is used to determine a target feature vector to be identified from the music library, whose feature vector meets preset conditions; determine the recognition result of the audio segment to be identified based on the target audio pointed to by the target feature vector, and the recognition result indicates that the audio segment to be identified and the target audio are different versions of the same audio.
[0026] In one implementation, the processing unit is further configured to: input the training spectrogram into the initial audio recognition model to obtain a first training feature vector output by the first task module and a second training feature vector output by the second task module; determine a first loss parameter based on the first training feature vector; determine a second loss parameter based on the second training feature vector; and determine an adjustment parameter based on the first loss parameter and the second loss parameter, wherein the first training feature vector and the second training feature vector are different.
[0027] In one implementation, the processing unit is further configured to: determine a predicted audio category label of the training spectrogram based on the first training feature vector; determine a prediction probability corresponding to the predicted audio category label; and determine a first loss parameter based on the prediction probability.
[0028] In one implementation, the training spectrum graph includes a first sample graph, a second sample graph, and a third sample graph, the first sample graph and the second sample graph have the same audio category label, and the first sample graph and the third sample graph have different audio category labels; the processing unit is further used to: determine a first vector distance between a second training feature vector corresponding to the first sample graph and a second training feature vector corresponding to the second sample graph; determine a second vector distance between the second training feature vector corresponding to the first sample graph and the second training feature vector corresponding to the third sample graph; and determine a second loss parameter based on the first vector distance and the second vector distance.
[0029] In one implementation, the recognition result also includes the audio category label of the audio segment to be recognized; the retrieval unit is specifically further used to: add the correspondence between the audio category label of the audio segment to be recognized and the feature vector to be recognized to the music library.
[0030] In one implementation, the processing unit is further used to: input the spectrum of the historical audio in the music library into the target audio recognition model to obtain a third eigenvector output by the target audio recognition model; the historical audio has an audio category label; and add the correspondence between the audio category label of the historical audio and the third eigenvector to the music library.
[0031] In one implementation, the retrieval unit is further configured to: calculate the similarity between the feature vector to be identified and a third feature vector in the music library; and determine the third feature vector as the target feature vector if the similarity between the feature vector to be identified and the third feature vector meets a preset condition.
[0032] In another possible design, the audio recognition device includes an acquisition unit, a processing unit, and a retrieval unit.
[0033] The acquiring unit is configured to: respond to a recognition request for a humming audio segment and acquire a humming spectrogram of the humming audio segment.
[0034] The processing unit is used to: input the humming spectrum into the target audio recognition model to generate a humming vector to be recognized; the target audio recognition model is obtained by adjusting the model parameters of the initial audio recognition model using the adjustment parameter, the initial audio recognition model includes a first task module and a second task module, and the adjustment parameter is determined based on the first loss parameter generated by the first task module and the second loss parameter generated by the second task module.
[0035] The retrieval unit is used to determine from the music library a similar feature vector that meets preset conditions with the humming vector to be identified, and determine the audio pointed to by the similar feature vector as the similar audio of the humming audio segment.
[0036] The processing unit is further used to: output audio information of similar audio, where the audio information includes the audio name and singer name of the similar audio.
[0037] In a fourth aspect, the present application provides a computer device comprising a processor, a network interface, and a storage device, wherein the processor, the network interface, and the storage device are interconnected. The network interface is controlled by the processor to transmit and receive data, and the storage device is used to store a computer program, which includes program instructions. The processor is configured to invoke the program instructions to implement the audio recognition method provided in the present application.
[0038] In a fifth aspect, the present application provides a computer-readable storage medium, which stores a computer program. The computer program includes program instructions, which, when executed by a processor, enable the processor to implement the audio recognition method provided by the present application.
[0039] In a sixth aspect, the present application provides a computer program product, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to implement the audio recognition method provided in the present application.
[0040] By adopting the present application, joint learning is achieved through the first task module and the second task module included in the target audio recognition model, so that the feature vector to be identified represents the audio more accurately, thereby improving the accuracy and efficiency of audio recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0042] Figure 1 A schematic diagram of a scenario of an audio recognition system provided in an embodiment of the present application;
[0043] Figure 2 A flowchart of an audio recognition method provided in an embodiment of the present application;
[0044] Figure 3 A schematic diagram of an audio recognition scenario provided in an embodiment of the present application;
[0045] Figure 4 A flowchart of a method for training an audio recognition model provided in an embodiment of the present application;
[0046] Figure 5 A schematic diagram of the structure of an audio recognition model provided in an embodiment of the present application;
[0047] Figure 6 A flowchart of an audio recognition method provided in an embodiment of the present application;
[0048] Figure 7 A schematic diagram of an audio recognition scenario provided in an embodiment of the present application;
[0049] Figure 8 A schematic diagram of the structure of an audio recognition device provided in an embodiment of the present application;
[0050] Figure 9 A schematic diagram of the structure of a computer device provided in this application. DETAILED DESCRIPTION
[0051] The following will be combined with the accompanying drawings to clearly and completely describe the technical solutions in this application. Obviously, the embodiments described are only some of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.
[0052] To facilitate understanding, the terms involved in this application are first explained.
[0053] 1. Machine learning (ML)
[0054] Machine learning studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. It is the core of artificial intelligence and the fundamental path to computer intelligence. Deep learning (DL) is a new research direction in machine learning. Deep learning studies the inherent patterns and representational hierarchies of sample data. The information gained during this learning process is highly helpful in interpreting data such as text, images, and sound. The ultimate goal of deep learning is to enable machines to acquire human-like analytical learning capabilities and recognize data such as text, images, and sound.
[0055] 2. Deep neutral networks (DNN)
[0056] Deep neural networks are the foundation of deep learning. The neural network layers within a DNN can be divided into three categories: input layer, hidden layer, and output layer. Generally speaking, the first layer is the input layer, the last layer is the output layer, and all layers in between are hidden layers, with each layer being fully connected. Backpropagation in a DNN refers to a method for calculating the gradients of neural network parameters. Generally speaking, backpropagation follows the chain rule from calculus. It calculates and stores the gradients of the objective function with respect to the intermediate variables and parameters of each neural network layer, sequentially from the output layer to the input layer. This allows the DNN's loss function to be derived. This loss function can be used to adjust network parameters and optimize the network.
[0057] 3. Vector Mapping (embedding)
[0058] Embedding is a distributed representation method that represents raw input data as a linear combination of features. This allows large, sparse vectors to be mapped into a low-dimensional space that preserves semantic relationships. Furthermore, the nature of embedding vectors ensures that objects corresponding to closely spaced vectors have similar meanings. For example, the embeddings for "weather" and "sunny" are closer, while the embeddings for "weather" and "table" are farther apart. Due to these characteristics, embeddings are widely used in deep learning. In audio recognition, using embeddings to represent audio can improve audio recognition performance.
[0059] A cover song is a song performed by someone else in a new style, including rewriting the lyrics and music. Cover song recognition, on the other hand, involves identifying songs with similar lyrics and music arrangements to the original. Its primary goal is to find different versions of the same source music within a vast amount of music data.
[0060] Currently, the probability that two audio clips are covers of each other is typically determined based on the audio's harmonic pitch class profile (HPCP) feature. However, HPCP features contain a large amount of interference information, resulting in low cover recognition accuracy. Furthermore, current deep learning solutions typically employ a single learning approach, which can easily lead to overfitting, making the learned song features less robust, affecting the generalization ability of the song representation, and ultimately resulting in insufficient accuracy in cover recognition.
[0061] Based on the above problems, this application provides an audio recognition method that can be used to identify cover songs. For example, but not limited to, the audio method provided in the embodiment of this application can be applied to Figure 1 The audio recognition system shown. Figure 1 This is a scene diagram of an audio recognition system. The audio recognition system may include but is not limited to: one or more terminals 120, one or more servers 110. For example, Figure 1 The figure shows a server and four terminals: smart phone, smart watch, car terminal and computer. The terminals and the server establish communication connection through wired network or wireless network and exchange data. Figure 1 The number and form of the devices shown are for illustrative purposes only and do not constitute a limitation on the embodiments of the present application.
[0062] In the embodiment of the present application, the terminal may include but is not limited to smart phones, tablet computers, laptops, desktop computers, smart speakers, smart watches, car terminals, smart home appliances, smart voice interaction devices and other smart devices.
[0063] Applied in this application, the terminal can be used as an audio recognition device, obtaining the audio segment to be recognized from the terminal or the server, inputting the spectrum graph to be recognized corresponding to the audio segment to be recognized into the target audio recognition model, and obtaining the feature vector to be recognized output by the target audio recognition model; the terminal can also determine the target feature vector that meets the preset conditions from the music library based on the feature vector to be recognized, and determine the recognition result of the audio segment to be recognized through the target feature vector.
[0064] In an embodiment of the present application, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, and big data and artificial intelligence platforms.
[0065] Applied in the embodiment of the present application, the server can serve as an audio recognition device to obtain the audio segment to be recognized from the terminal, or to obtain the audio segment to be recognized from the database in the server; the spectrum graph to be recognized corresponding to the audio segment to be recognized is input into the target audio recognition model to obtain the feature vector to be recognized output by the target audio recognition model; the server can also determine the target feature vector that meets the preset conditions from the music library based on the feature vector to be recognized, and determine the recognition result of the audio segment to be recognized through the target feature vector.
[0066] The audio segment to be identified, the spectrogram to be identified, and the feature vectors to be identified and target feature vectors generated during the audio recognition process can be stored in a cloud database. When executing the audio recognition method, the audio recognition device retrieves the aforementioned data from the cloud database. Alternatively, other data generated by the audio recognition method can also be stored in a blockchain. Blockchain is a novel application model for computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. It is essentially a decentralized database, a series of data blocks generated using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. Due to the aforementioned characteristics of blockchain, the data stored on the blockchain cannot be tampered with, ensuring data security.
[0067] It can be understood that in the specific implementation of this application, related data such as audio to be identified and spectrum graphs to be identified are involved. When the above embodiments of this application are applied to specific products or technologies, the relevant data must obtain the permission or consent of the relevant objects, and the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0068] Based on Figure 1 The audio recognition system shown in the figure, the embodiment of the present application proposes an audio recognition method, which can be applied to the audio recognition device. Figure 2 , Figure 2 This is a flow chart of an audio recognition method provided in an embodiment of the present application. Figure 2 As shown, the audio recognition method includes but is not limited to the following steps:
[0069] S201: Inputting the to-be-recognized spectrogram corresponding to the to-be-recognized audio segment into a target audio recognition model to obtain a to-be-recognized feature vector output by the target audio model.
[0070] The audio recognition device can obtain a complete song or a song segment from the terminal. If the complete song is obtained, the complete song can be sampled to obtain the audio segment to be recognized. If the song segment is obtained, the song segment can be used as the audio segment to be recognized. The audio segment to be recognized can be the original version of the song or a cover version of the song. For example, please refer to Figure 3 For the content shown in (a), the user can upload a song (the song can be a complete song or a song fragment) as the audio fragment to be identified through the audio recognition application (such as a music player) in a terminal such as a smartphone; or, the user can implement functions such as "humming recognition" and "song recognition by listening" through applications such as music players in a terminal such as a smartphone. Specifically, the user can hum a melody or song by himself, or record music played by other devices, which will be received by a terminal such as a smartphone and input into the target audio recognition model as the audio to be identified for audio recognition.
[0071] In one implementation, a Fourier transform can be performed on the audio segment to be identified, converting the audio segment from the time domain to the frequency domain to generate a spectrum of the audio segment to be identified. "Spectrum," short for frequency spectral density, is a frequency distribution curve. Complex oscillations are decomposed into harmonic oscillations of varying amplitudes and frequencies. A graph of the amplitudes of these harmonic oscillations arranged by frequency is called a spectrum. The spectrum shifts the study of signals from the time domain to the frequency domain, providing a more intuitive understanding. For a segment of audio, the audio is converted to the frequency domain to obtain a spectrum of size (T, F), where F is the frequency axis and T is the time axis. The spectrum can be viewed as a two-dimensional image, a spectrogram, which can be input into the target audio recognition model for processing.
[0072] The audio recognition device includes a target audio recognition model, and the target audio recognition model is obtained by adjusting the model parameters of the initial audio recognition model using adjustment parameters. The initial audio recognition model includes a first task module and a second task module. The input of the first task module is the same as the input of the second task module, and the execution order between the two modules is not limited. The first task module may execute the task before the second task module, or the second task module may execute the task before the first task module, or the first task module and the second task module may execute the task at the same time. This application does not limit this. In addition, the adjustment parameter is determined based on the first loss parameter generated by the first task module and the second loss parameter generated by the second task module, and is used to adjust the model parameters of the initial audio recognition model. The feature vector to be identified output by the target audio recognition model is an embedding vector, which contains the spectral information of the audio segment to be identified.
[0073] In one implementation, the first task module in the initial audio recognition model and the target audio recognition model can be a classification task module, which is used to predict the audio category label to which the spectrogram of the audio input to the model belongs. Specifically, when the spectrogram to be identified is input into the target audio recognition model, the first task module can classify the spectrogram to be identified, output a first embedding vector, and predict the audio category label to which the spectrogram to be identified belongs. The same audio category label can be used to label one or more spectrograms to be identified, and the spectrograms to be identified with the same audio category label can be different versions of the same audio. For example, song A sung by singer a and song B sung by singer b both have an audio category label of 2, indicating that song A and song B are different versions of the same audio.
[0074] In one implementation, the second task module may be a metric learning module that autonomously learns a metric distance function for the audio recognition task. By calculating the similarity between two spectrograms, the input spectrogram is classified into the audio category with the highest similarity. The second task module outputs a second embedding vector corresponding to the spectrogram to be recognized. If the audio category label of the spectrogram to be recognized is 2, the distance between the second embedding vector and the embedding vectors of song A and song B, both of which have the same audio category label of 2, are calculated in the target audio recognition model. The distance between the second embedding vector and the embedding vectors of songs with audio category labels different from 2 (e.g., 1, 3, 4, 5, etc.) in the target audio recognition model are also calculated. Based on the calculation results, under the constraints of the distance function, the distance between similar objects (embedding vectors with the same audio category label) is minimized, while the distance between dissimilar objects (embedding vectors with different audio category labels) is maximized, thereby improving the representation accuracy and robustness of the feature vector to be recognized output by the target audio recognition model.
[0075] In one implementation, the embedding vector output by the first task module can be used as the feature vector to be identified, or the embedding vector output by the second task module can be used as the feature vector to be identified, which is not limited in this application.
[0076] S202: Determine a target feature vector from the music library that satisfies a preset condition with the feature vector to be identified.
[0077] In this application, the music library is a database that records a large amount of audio information, including basic information of the audio, such as the name of the audio, the singer, etc. After the audio in the music library is processed by the target audio recognition model, the embedding vector and audio category label of the audio output by the target audio recognition model can be obtained, so that the music library also includes the embedding vector corresponding to the audio, the audio category label of the audio, and the correspondence between each embedding vector and the corresponding audio category label.
[0078] Optionally, the preset condition may mean that the attribute information corresponding to the feature vector to be identified is similar to or identical to the attribute information corresponding to the target feature vector, and the aforementioned attribute information may be an audio category label, the direction of the vector, etc.; for example, the preset condition may mean that the audio category label corresponding to the feature vector to be identified is the same as the audio category label corresponding to the target feature vector.
[0079] In one implementation, the similarity between the feature vector to be identified and the embedding vector corresponding to the audio in the music library can be calculated. If the similarity between the feature vector to be identified and the embedding vector corresponding to the audio in the music library meets a preset condition, the embedding vector is determined to be the target feature vector.
[0080] Alternatively, the preset condition may be a manually set similarity threshold range, or a similarity threshold range determined by artificial intelligence technology (such as machine learning); or alternatively, the preset condition may be that the target feature vector satisfies both the audio category label and the audio category label of the feature vector to be identified, and that the similarity between the target feature vector and the feature vector to be identified satisfies the preset condition. Artificial intelligence is a branch of computer science that utilizes the understood essence of intelligence to produce a new intelligent machine that can respond in a manner similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, and natural language processing.
[0081] Optionally, the cosine similarity between vectors can be used as the above-mentioned similarity, that is, the similarity between them is measured by measuring the cosine value of the angle between the vector in the music library and the vector to be identified. Among them, the value of cosine similarity is independent of the length of the vector, but only related to the pointing direction of the vector. The value range of cosine similarity is [-1,1], then the preset condition can be that the value of cosine similarity is in the value interval of [a,b], wherein a and b are integers between [-1,1], and b is greater than a, and the specific values of a and b depend on the actual application scenario, which is not limited in this application. For example, if the feature vector to be identified is the same as the vector in the music library, that is, the two vectors have the same direction, the value of cosine similarity is 1; if the feature vector to be identified is relatively similar to the vector in the music library, when the angle between the two vectors is 90°, the value of cosine similarity is 0; if the feature vector to be identified is completely different from the vector in the music library, that is, when the two vectors point in completely opposite directions, the value of cosine similarity is -1. Assume that the characteristic vector to be identified is vector A, a vector in the music library is vector B, and the angle is α. i With B i Represent the components of vector A and vector B respectively, then the cosine similarity between vector A and vector B can be calculated by the following formula, where n is the number of vector components:
[0082]
[0083] Alternatively, the method for calculating similarity can also be to calculate the Euclidean distance between the vector to be identified and the vector in the music library. If the value of the Euclidean distance obtained meets the preset conditions, the vector in the music library is determined to be the target feature vector. Among them, Euclidean distance, also known as Euclidean distance or Euclidean metric, is a commonly used distance definition. It is the true distance between two points in N (N is a positive integer) dimensional space. The Euclidean distance in two-dimensional and three-dimensional space is the distance between two points. In this application, the Euclidean distance between the vector to be identified and the vector in the music library is the distance in one-dimensional space.
[0084] S203: Determine the target audio pointed to by the target feature vector as the recognition result of the audio segment to be recognized, where the recognition result indicates that the audio segment to be recognized and the target audio are different versions of the same audio.
[0085] In this application, based on the target feature vector, the target audio corresponding to the target feature vector can be queried from the music library. Since the target feature vector and the feature vector to be identified meet the preset conditions, it can be determined that the target audio and the audio clip to be identified are different versions of the same audio. For example, if the audio clip to be identified is the original version of a song, the target audio can be a cover version of the song. Since there may be multiple cover versions of a song, there may also be multiple target audios to be determined. If the audio clip to be identified is a cover version of a song, the target audio can be the original version of the song.
[0086] Optionally, the recognition result may also include the audio category label of the audio segment to be identified. After obtaining the recognition result, the correspondence between the audio category label of the audio segment to be identified and the feature vector to be identified can be added to the music library, thereby increasing the feature vector in the music library, which is helpful for the identification of other audio to be identified.
[0087] Based on this, the application scenario of the audio recognition method provided in the embodiment of the present application can be to recognize songs by listening to them. For example, Figure 3 , Figure 3 A schematic diagram of an audio recognition scenario provided in an embodiment of the present application is shown as follows: Figure 3 As shown in (a) in the figure, the user uploads an audio clip through the terminal, and the terminal can upload the audio clip to the server. The server can process the audio clip to obtain the spectrum to be recognized, and input it into the target audio recognition model to obtain the recognition result of the audio clip. The recognition result may include two types: recognition success and recognition failure, such as Figure 3 As shown in (b), the recognition result is successful, and the terminal can display which song the audio is; Figure 3 As shown in (c), the recognition result is "no matching result", and the user can choose to re-recognize.
[0088] Optionally, in another application scenario, songs in the music library can be input into the target audio recognition model, and the cover versions of the songs can be clustered for managing the various versions of the songs.
[0089] Optionally, the audio recognition method provided in the embodiment of the present application can also be applied to song copyright management. For example, the spectrum of the producer's original song is input into the target audio recognition model to query whether there are any songs in the music library with a melody similar to the original song. If so, the original song can be modified accordingly to avoid copyright risks.
[0090] In an embodiment of the present application, a target audio recognition model for audio recognition of an audio to be recognized is obtained by adjusting the model parameters of an initial audio recognition model based on adjustment parameters, and the adjustment parameters are determined according to a first loss parameter generated by a first task module and a second loss parameter generated by a second task module. According to the audio recognition method provided by the present application, the audio to be recognized is input into the target audio recognition model, and a feature vector to be recognized output by the target audio model can be obtained. Through the feature vector to be recognized, a different version of the target audio that is the same audio as the audio to be recognized can be found from the music library. Using the method provided by the present application, joint learning is achieved through the first task module and the second task module included in the target audio recognition model, so that the feature vector to be recognized represents the audio more accurately, thereby improving the accuracy and efficiency of audio recognition.
[0091] See also Figure 4 , Figure 4 This is a flow chart of a method for training an audio recognition model provided in an embodiment of the present application. The audio recognition model training method can be used to train an initial audio recognition model to obtain a target audio recognition model, which is used in the audio recognition method in the above embodiment. Figure 4 As shown, the training method of the audio recognition model includes but is not limited to the following steps:
[0092] S401: Inputting the training spectrogram into the initial audio recognition model to obtain a first training feature vector output by the first task module and a second training feature vector output by the second task module.
[0093] In this application, the training spectrogram is a spectrogram generated after the audio sample is processed by Fourier transform, etc. The audio sample can be a complete song or a song fragment. The training spectrogram used to train the initial audio recognition model has been marked with real audio category labels. The aforementioned real audio category labels are audio category labels marked by relevant staff based on actual song information. For example, the audio category label of audio A and audio B is 1, and the audio category label of audio C is 2, indicating that audio A and audio B belong to the same group of audio and are different versions of the same audio, while audio C and audio A, audio C and audio B belong to different groups of audio and are different audios.
[0094] The audio recognition device includes an initial audio recognition model, which is a model built based on a deep neural convolutional network and includes a first task module and a second task module. For example, see Figure 5 , Figure 5 A structural diagram of an audio recognition model provided for an embodiment of the present application. Wherein, the input of the first task module is the same as the input of the second task module, the output of the first task module is the first training feature vector and the first predicted label of the first training feature vector, the output of the second task module is the second training feature vector and the second predicted label of the second training feature vector, and the first training feature vector and the second training feature vector are different embedding vectors. In addition, the first predicted label and the second prediction both include predicted audio category labels, and the predicted audio category labels may be the same as or different from the actual audio category labels, that is, there is an error. During the training process, it is necessary to use a loss function to adjust the model parameters so that the error remains within an acceptable range. In an embodiment of the present application, the first task module may execute the task before the second task module, or the second task module may execute the task before the first task module, or the first task module and the second task module may execute the task at the same time, and this application does not limit this.
[0095] Optionally, the deep neural convolutional network used in the initial audio recognition model can be an autoencoder (AE), a residual neural network (ResNet), a wide residual neural network (wide ResNet) and other networks. Among them, the autoencoder is a type of artificial neural network (ANNs) used in semi-supervised learning and unsupervised learning. Its function is to perform representation learning on the input information by taking the input information as the learning target; the residual block inside the residual neural network uses a jump connection, which alleviates the gradient vanishing problem caused by increasing the depth in the deep neural network. It is easy to optimize and can improve the accuracy by increasing the depth. The widened residual neural network improves the residual neural network from the perspective of increasing the network width, so that both performance and training speed are improved. According to the needs of the actual application scenario, a suitable deep neural convolutional network can be selected as the basis for constructing the initial audio recognition model, and this application does not limit this.
[0096] S402: Determine a first loss parameter according to the first training feature vector.
[0097] In actual training scenarios, each model training usually uses batch data. In this application, the input to the initial audio recognition model is a batch of training spectrograms, each of which has been annotated with the real audio category label. Assuming that the number of a batch is n, where n is a positive integer, a batch of training spectrograms (x1, x2, x3, ..., x i ,…,x n ) is input into the initial audio recognition model to obtain the first training feature vector (y1, y2, y3, ..., y i ,…,y n ) and the first predicted label, i.e., the predicted audio category label (l1, l2, l3, ..., l i ,…,l n), wherein there are p types of output predicted audio category labels, where p is a positive integer greater than or equal to 1 and less than or equal to the number of types of annotated real audio category labels. Exemplarily, if training spectrograms A and B belong to different versions of the same audio, the predicted audio category labels of training spectrograms A and B may be the same, such as both are 1. Based on the obtained first training feature vector, the prediction probability corresponding to each predicted audio category label can be determined by the activation function in the deep neural network; and the first loss parameter can be determined based on the loss function.
[0098] In one implementation, the first task module may be a classification task module, which is used to predict the audio category label to which the audio spectrogram of the input model belongs. Specifically, when the spectrogram to be identified is input into the initial audio recognition model, the first task module may classify the spectrogram to be identified, and output a first training feature vector and a predicted audio category label. The first training feature vector is an embedding vector, and the predicted audio category labels of the training spectrograms belonging to the same audio category are the same, and the predicted audio category labels of the training spectrograms belonging to different audio categories are different. For example, training spectrograms A and B are both versions of song a sung by different singers, and training spectrogram C is the original version of song b. Then, the predicted audio category labels of training spectrograms A and B are both 1, and the predicted audio category label of training spectrogram C is 2. Optionally, the log-likelihood loss function and the softmax activation function may be used to obtain the predicted probability, thereby obtaining the first loss parameter; the cross-entropy loss function and the sigmoid activation function may also be used to obtain the first loss parameter, which is not limited in this application.
[0099] S403: Determine a second loss parameter according to the second training feature vector.
[0100] In the present application, a batch of training spectrograms is input into the initial audio recognition model for training, and each training spectrogram is annotated with a real audio category label. Assuming that a batch of training spectrograms is input into the initial audio recognition model, a second training feature vector and a second predicted label are obtained for each training spectrogram in a batch output by the second task module, that is, the predicted audio category label of the training spectrogram. Among them, the second training feature vector is an embedding vector. According to the second training feature vector of a batch of training spectrograms, the second loss function can be determined according to a loss function, such as triplet loss, prototypical network loss, contrast loss, and other loss functions, which is not limited in the present application.
[0101] In one implementation, the second task module can be a metric learning module. When a batch of training spectrograms is input into the initial audio recognition model, the corresponding second training feature vector and second predicted label output by the second task module can be obtained. Optionally, each batch of training spectrograms includes one or more training spectrograms labeled with the same audio category label, and one or more training spectrograms labeled with different audio category labels.
[0102] Exemplarily, a triplet loss function is used to determine the second loss parameter. Assume that a batch of training spectra includes a first sample image, a second sample image, and a third sample image, the first sample image and the second sample image have the same audio category label, and the first sample image and the third sample image have different audio category labels. According to the definition of the triplet loss function, the first sample graph can be named as the fixed spectrogram a (anchor), the second sample graph can be named as the positive sample spectrogram p (positive), and the third sample graph can be named as the negative sample spectrogram n (negative). Spectrogram a and spectrogram p are a pair of positive samples, and spectrogram a and spectrogram n are a pair of negative samples. The positive sample pair indicates that spectrogram a and spectrogram p are different versions of the same audio, and the negative sample pair indicates that spectrogram a and spectrogram n are different audios. Correspondingly, if the distance function is used, the first vector distance d(a, p) of the second training feature vector of spectrogram a and the second training feature vector of spectrogram p are close, while the second vector distance d(a, n) of the second training feature vector of spectrogram a and the second training feature vector of spectrogram n are farther, that is, the following formula is satisfied:
[0103] ‖f(a)-f(p)‖ 2 =d(a,p)
[0104] ‖f(a)-f(n)‖ 2 =d(a,n)
[0105] ‖f(a)-f(p)‖ 2 ≤‖f(a)-f(n)‖ 2
[0106] Where f represents embedding, which is used to encode the spectrogram into Euclidean space; ‖f(a)-f(p)‖ 2 represents the Euclidean distance metric between spectrogram a and spectrogram p, ‖f(a)-f(n)‖ 2Represents the Euclidean distance metric between spectrogram a and spectrogram n. According to the above formula, taking the margin parameter β can widen the gap between the anchor and positive spectrogram pairs and the anchor and negative spectrogram pairs, thereby improving the accuracy and robustness of the representation of the feature vector to be identified output by the initial audio recognition model. The second loss parameter can be determined based on the first vector distance and the second vector distance. Based on this, the second loss function L(a, p, n) can be expressed as follows:
[0107] L(a,p,n)=max(‖f(a)-f(p)‖ 2 -‖f(a)-f(n)‖ 2 +β,0)
[0108] It should be noted that the present application does not limit the execution order between S402 and S403, that is, S402 and S403 can be executed simultaneously, S402 can be executed before S403, and S402 can also be executed after S403.
[0109] S404: Determine an adjustment parameter according to the first loss parameter and the second loss parameter, and adjust the model parameters of the initial audio recognition model according to the adjustment parameter to obtain a target audio recognition model.
[0110] Among them, the adjustment parameter can be obtained by adding the first loss parameter and the second loss parameter and calculating the average value. If the model parameters of the initial audio recognition model adjusted by the adjustment parameter make the loss function of the first task module and the loss function of the second task module converge, then the initial audio recognition model currently adjusted by the adjustment parameter can be determined as the target audio recognition model. The target audio recognition model obtained by joint learning of the first task module and the second task module can effectively avoid the overfitting problem that is easy to occur in a single learning task, thereby making the representation of the output feature vector (such as the first training feature vector and the second training feature vector) more accurate and robust, thereby improving the recognition accuracy and efficiency of the target audio recognition model.
[0111] S405: Input the spectrogram of the historical audio in the music library into the target audio recognition model to obtain the third eigenvector output by the target audio recognition model; add the correspondence between the audio category label of the historical audio and the third eigenvector to the music library.
[0112] In the present application, the music library is a database that records a large amount of audio information, including basic information of the audio, such as song title, singer and other information. The audio stored in the music library is called historical audio, and the spectrum graph corresponding to the historical audio in the music library is input into the target audio recognition model to obtain the third eigenvector of the historical audio. The third eigenvector can be the first eigenvector output by the first task module, or the second eigenvector output by the second task module. The obtained third eigenvector and the correspondence between the third eigenvector and the audio category label of the historical audio are added to the music library, so that the music library includes the basic information of the audio, the third eigenvector corresponding to the audio, the audio category label of the audio, and the correspondence between each third eigenvector and the corresponding audio category label. Based on this, for example, in the application scenario of cover song recognition, the spectrum graph to be recognized of the audio to be recognized can be input into the target audio recognition model to obtain the eigenvector to be recognized. According to the eigenvector to be recognized, the target eigenvector can be retrieved from the music library to obtain the recognition result of the audio to be recognized. Among them, the recognition result can be that the target feature vector that meets the preset conditions is found, and the audio in the music library pointed to by the target feature vector is a different version of the same audio as the audio to be recognized, or the recognition result is that the target feature vector that meets the preset conditions is not retrieved in the music library, that is, no audio that is a different version of the same audio as the audio to be recognized is found in the music library.
[0113] By using the audio recognition model training method provided in this application, the adjustment parameter can be obtained through the first loss parameter generated by the first task module in the initial audio recognition model and the second loss parameter generated by the second task module. The model parameters of the initial audio recognition model are adjusted using the adjustment parameter to obtain the target audio recognition model, thereby realizing the joint learning of the first task module and the second task module, making the output feature vector more accurate in representing the audio, thereby making the trained target audio recognition model have strong generalization ability, while improving the accuracy and efficiency of audio recognition.
[0114] See also Figure 6 , Figure 6 A flowchart of an audio recognition method provided in an embodiment of the present application is applicable to Figure 1 The audio recognition system shown in FIG. 1 can be applied to an audio recognition device, which can be a terminal or a server. Figure 6 As shown, the audio recognition method includes but is not limited to the following steps:
[0115] S601: In response to a recognition request for a humming audio segment, a humming spectrogram of the humming audio segment is obtained.
[0116] In this application, the user can initiate a humming recognition request to a terminal, which then performs humming recognition on the humming audio clip. Alternatively, the user can initiate a humming recognition request to a terminal equipped with a humming recognition service. The terminal can record the user's humming, generate a humming audio clip, and initiate a humming recognition request to a server serving as an audio recognition device. The humming recognition request includes the humming audio clip and may also include information such as the identity of the terminal initiating the request.
[0117] For example, Figure 7 As shown in (a) of FIG, a user can implement the humming recognition function through an application such as a music player in a terminal such as a smartphone. Specifically, the user can hum a melody or a song to the terminal's audio receiving device and initiate a humming recognition request to the terminal. The terminal can then act as an audio recognition device and perform audio recognition on the received humming audio clip.
[0118] The audio recognition device can perform Fourier transform processing on the received humming audio segment to obtain a humming spectrum diagram of the humming audio segment.
[0119] In one implementation, the humming recognition request may further include a humming spectrogram of the humming audio segment. When the audio recognition device receives the user's humming recognition request, it may process the humming audio segment and generate a humming spectrogram of the humming audio segment.
[0120] S602: Input the humming spectrum into the target audio recognition model to generate a humming vector to be recognized; the target audio recognition model is obtained by adjusting the model parameters of the initial audio recognition model using the adjustment parameters, the initial audio recognition model includes a first task module and a second task module, and the adjustment parameters are determined based on a first loss parameter generated by the first task module and a second loss parameter generated by the second task module.
[0121] The audio recognition device includes a target audio recognition model, which includes a first task module and a second task module. When a humming spectrogram is input into the target audio model, the first task module and the second task module will jointly learn the humming spectrogram to generate a humming vector that can accurately represent the humming audio segment.
[0122] It should be noted that, based on the same inventive concept, the technical details and principles of constructing the target audio recognition model can be found in S401-S405, which will not be repeated here for the sake of brevity.
[0123] S603: Determine from the music library a similar feature vector that meets a preset condition with the humming vector to be identified, and determine the audio pointed to by the similar feature vector as the similar audio of the humming audio segment.
[0124] The music library stores a large amount of audio information, including basic audio information such as the audio's name, singer, duration, and other information. After the audio in the music library is processed by the target audio recognition model, the feature vector and audio category label of each audio can be obtained. Therefore, the music library also includes the correspondence between the feature vector of each audio and the corresponding audio category label. In the music library, there can be one or more similar feature vectors that meet the preset conditions with the humming vector. For example, the audio clip hummed by the user is song A, and the music library can store multiple audios similar to the humming audio clip, such as the original version of song A, cover versions by different singers, etc. In this case, the feature vectors of these similar audios meet the preset conditions with the humming vector.
[0125] Optionally, the preset condition may be that the cosine similarity between the humming feature vector and the similarity feature vector satisfies a similarity condition, wherein the similarity condition may be a set similarity threshold range.
[0126] S604: Output the audio information of the similar audio, where the audio information includes the audio name and singer name of the similar audio.
[0127] In the present application, the audio information of the similar audio includes the song title and singer name of the similar audio. Optionally, the audio information of the similar audio may also include the degree of similarity between the similar audio and the humming audio clip, the audio file of the similar audio, or the playback link of the similar audio. There may be one or more similar audios to the humming audio clip. The object that initiates the humming recognition request may be a terminal equipped with services such as humming recognition, which may collect the audio clips hummed by the user, or may send a humming recognition request to a server serving as an audio recognition device. The audio recognition device may return the audio information of the similar audio to the terminal in response to the humming recognition request.
[0128] Optionally, the object initiating the humming recognition request may be a user, who initiates the humming recognition request to the terminal by humming a song clip. The terminal, as an audio recognition device, performs the recognition. When similar audio is found in the music library, the terminal may respond to the humming recognition request and display the result in the terminal. For example, see Figure 7 , Figure 7 A schematic diagram of a humming recognition scenario provided in an embodiment of the present application is shown in FIG. Figure 7 As shown in (a) in FIG, the terminal can receive the audio clip of the user humming and the humming recognition request, and perform recognition. Figure 7 As shown in (b) in the figure, the terminal can identify one or more audios similar to the humming audio clip, and display the similarity between each similar audio and the humming audio clip, as well as the audio information of the similar audio, such as the song name and singer, and the user can click "Play" to listen to it. Figure 7As shown in (c), if the song of the audio clip hummed by the user cannot be identified, "No matching result" can be displayed, waiting for the user's next recognition. According to the audio recognition method provided by the present application, the humming spectrum is input into the target audio recognition model, and the humming feature vector output by the target audio recognition model can be obtained; through the humming feature vector. One or more similar audios similar to the humming audio clip can be found in the music library. Using the method provided by the present application, the first task module and the second task module included in the target audio recognition model are used to achieve joint learning, which can more accurately characterize the humming audio clip, thereby improving the accuracy and efficiency of audio recognition.
[0129] See also Figure 8 , Figure 8 A schematic diagram of the structure of an audio recognition device provided in an embodiment of the present application.
[0130] In one possible design, the audio recognition device includes a processing unit 810 and a retrieval unit 820.
[0131] Processing unit 810 is configured to input a to-be-recognized spectrogram corresponding to the to-be-recognized audio segment into a target audio recognition model to obtain a to-be-recognized feature vector output by the target audio recognition model; wherein the target audio recognition model is obtained by adjusting model parameters of the initial audio recognition model using an adjustment parameter, the adjustment parameter being determined based on a first loss parameter generated by the first task module and a second loss parameter generated by the second task module;
[0132] The retrieval unit 820 is used to: determine a target feature vector from the music library that meets preset conditions with the feature vector to be identified; determine the target audio pointed to by the target feature vector as the recognition result of the audio segment to be identified, and the recognition result indicates that the audio segment to be identified and the target audio are different versions of the same audio.
[0133] In one implementation, the processing unit 810 is further configured to: input the training spectrogram into the initial audio recognition model to obtain a first training feature vector output by the first task module and a second training feature vector output by the second task module; determine a first loss parameter based on the first training feature vector; determine a second loss parameter based on the second training feature vector; and determine an adjustment parameter based on the first loss parameter and the second loss parameter, wherein the first training feature vector and the second training feature vector are different.
[0134] In one implementation, the processing unit 810 is further configured to: determine a predicted audio category label of the training spectrogram based on the first training feature vector; determine a prediction probability corresponding to the predicted audio category label; and determine a first loss parameter based on the prediction probability.
[0135] In one implementation, the training spectrum graph includes a first sample graph, a second sample graph, and a third sample graph, the first sample graph and the second sample graph have the same audio category label, and the first sample graph and the third sample graph have different audio category labels; the processing unit 810 is further used to: determine a first vector distance between a second training feature vector corresponding to the first sample graph and a second training feature vector corresponding to the second sample graph; determine a second vector distance between the second training feature vector corresponding to the first sample graph and the second training feature vector corresponding to the third sample graph; and determine a second loss parameter based on the first vector distance and the second vector distance.
[0136] In one implementation, the recognition result also includes the audio category label of the audio segment to be recognized; the retrieval unit 820 is further used to add the correspondence between the audio category label of the audio segment to be recognized and the feature vector to be recognized to the music library.
[0137] In one implementation, the processing unit 810 is further used to: input the spectrum of the historical audio in the music library into the target audio recognition model to obtain a third eigenvector output by the target audio recognition model; the historical audio has an audio category label; and add the correspondence between the audio category label of the historical audio and the third eigenvector to the music library.
[0138] In one implementation, the retrieval unit 820 is further configured to calculate the similarity between the feature vector to be identified and a third feature vector in the music library; and if the similarity between the feature vector to be identified and the third feature vector meets a preset condition, determine the third feature vector as the target feature vector.
[0139] In another possible design, the audio recognition device includes an acquisition unit 830, a processing unit 810, and a retrieval unit 820.
[0140] The acquiring unit 830 is configured to: respond to a recognition request for a humming audio segment and acquire a humming spectrogram of the humming audio segment.
[0141] Processing unit 810 is used to: input the humming spectrogram into the target audio recognition model to generate a humming vector to be recognized; the target audio recognition model is obtained by adjusting the model parameters of the initial audio recognition model using adjustment parameters, the initial audio recognition model includes a first task module and a second task module, and the adjustment parameters are determined based on a first loss parameter generated by the first task module and a second loss parameter generated by the second task module.
[0142] The retrieval unit 820 is configured to: determine from the music library a similar feature vector that satisfies a preset condition with the humming vector to be identified, and determine the audio pointed to by the similar feature vector as the similar audio of the humming audio segment;
[0143] The processing unit 810 is further configured to output audio information of the similar audio, where the audio information includes the audio name and singer name of the similar audio.
[0144] According to one embodiment of the present application, Figure 2 、 Figure 4 and Figure 6 The steps involved in the audio recognition method shown can be represented by Figure 8 The various units in the audio recognition device shown are executed. For example, Figure 2 Steps S201 and S202 shown in Figure 4 Steps S401, S402, S403, S404, and S405 shown in FIG. Figure 6 Steps S602 and S604 shown in FIG can be replaced by Figure 8 The processing unit 810 in the embodiment is used to execute, Figure 2 Steps S202, S203 and Figure 6 Step S603 shown can be performed by Figure 8 The retrieval unit 820 in the embodiment is used to perform the Figure 6 Step S601 in Figure 8 The acquisition unit 830 in is executed.
[0145] According to one embodiment of the present application, Figure 8 The various units in the audio recognition device shown can be individually or all combined into one or several units to constitute, or one (or some) of the units can be further divided into multiple functionally smaller sub-units to achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above-mentioned units are divided based on logical functions. In actual applications, the functions of one unit can also be implemented by multiple units, or the functions of multiple units can be implemented by one unit. In other embodiments of the present application, the audio recognition device may also include other units. In actual applications, these functions can also be implemented with the assistance of other units, and can be implemented by the collaboration of multiple units.
[0146] It can be understood that the functions of the various functional units of the audio recognition device described in the embodiments of the present application can be specifically implemented according to the methods in the above method embodiments. The specific implementation process can refer to the relevant description of the above method embodiments and will not be repeated here.
[0147] By adopting the audio recognition method provided in this application, joint learning is achieved through the first task module and the second task module included in the target audio recognition model, so that the feature vector to be recognized represents the audio more accurately, thereby improving the accuracy and efficiency of audio recognition.
[0148] See Figure 9 , is a schematic diagram of the structure of a computer device provided by this application. Figure 9 As shown, the computer device may include: a processor 910, a network interface 920 and a memory 930. The processor 910, the network interface 920 and the memory 930 may be connected via a bus or other means. The embodiment of the present application takes the bus connection as an example.
[0149] Among them, the processor 910 (or CPU (Central Processing Unit)) is the computing core and control core of the computer device. It can parse various instructions within the computer device and process various data of the computer device. For example, the CPU can be used to parse the power on and off instructions sent to the computer device and control the computer device to perform power on and off operations; for another example, the CPU can transmit various interactive data between the internal structures of the computer device, etc. The network interface 920 can optionally include a standard wired interface or a wireless interface (such as Wi-Fi, mobile communication interface, etc.), which is controlled by the processor 910 to send and receive data. The memory 930 (Memory) is a memory device in the computer device for storing programs and data. It is understandable that the memory 930 here can include both the built-in memory of the computer device and the extended memory supported by the computer device. The memory 930 provides storage space, which stores the operating system of the computer device, which may include but is not limited to: Android system, iOS system, Windows Phone system, etc., and this application is not limited to this.
[0150] In one implementation, the processor 910 performs the following operations by running the executable program code in the memory 930:
[0151] The spectrum graph to be identified corresponding to the audio segment to be identified is input into the target audio recognition model to obtain the feature vector to be identified output by the target audio model; a target feature vector that meets preset conditions with the feature vector to be identified is determined from the music library; the target audio pointed to by the target feature vector is determined as the recognition result of the audio segment to be identified, and the recognition result indicates that the audio segment to be identified and the target audio are different versions of the same audio.
[0152] Optionally, the processor 910 may further perform the following operations by running the executable program code in the memory 930: inputting the training spectrogram into the initial audio recognition model to obtain a first training feature vector output by the first task module and a second training feature vector output by the second task module; determining a first loss parameter based on the first training feature vector; determining a second loss parameter based on the second training feature vector; and determining an adjustment parameter based on the first loss parameter and the second loss parameter, wherein the first training feature vector and the second training feature vector are different.
[0153] Optionally, the processor 910 can also perform the following operations by running the executable program code in the memory 930: determining the predicted audio category label of the training spectrogram based on the first training feature vector; determining the predicted probability corresponding to the predicted audio category label; and determining the first loss parameter based on the predicted probability.
[0154] Optionally, the training spectrum graph includes a first sample graph, a second sample graph and a third sample graph, the audio category labels of the first sample graph and the second sample graph are the same, and the audio category labels of the first sample graph and the third sample graph are different; the processor 910 can also perform the following operations by running the executable program code in the memory 930: determine the first vector distance between the second training feature vector corresponding to the first sample graph and the second training feature vector corresponding to the second sample graph; determine the second vector distance between the second training feature vector corresponding to the first sample graph and the second training feature vector corresponding to the third sample graph; determine the second loss parameter based on the first vector distance and the second vector distance.
[0155] Optionally, the recognition result also includes the audio category label of the audio segment to be identified; the processor 910 can also perform the following operations by running the executable program code in the memory 930: adding the correspondence between the audio category label of the audio segment to be identified and the feature vector to be identified to the music library.
[0156] Optionally, the processor 910 can also perform the following operations by running the executable program code in the memory 930: input the spectrum of the historical audio in the music library into the target audio recognition model to obtain the third eigenvector output by the target audio recognition model; the historical audio has an audio category label; and the correspondence between the audio category label of the historical audio and the third eigenvector is added to the music library.
[0157] Optionally, the processor 910 can also perform the following operations by running the executable program code in the memory 930: calculating the similarity between the feature vector to be identified and the third feature vector in the music library; if the similarity between the feature vector to be identified and the third feature vector meets a preset condition, determining the third feature vector as the target feature vector.
[0158] In another implementation, the processor 910 may perform the following operations by running the executable program code in the memory 930:
[0159] In response to a request for identifying a humming audio clip, a humming spectrogram of the humming audio clip is obtained; the humming spectrogram is input into a target audio recognition model to generate a humming vector to be identified; the target audio recognition model is obtained by adjusting the model parameters of an initial audio recognition model using adjustment parameters, the initial audio recognition model includes a first task module and a second task module, and the adjustment parameters are determined based on a first loss parameter generated by the first task module and a second loss parameter generated by the second task module; a similar feature vector that meets preset conditions with the humming vector to be identified is determined from a music library, and the audio pointed to by the similar feature vector is determined as similar audio to the humming audio clip; audio information of the similar audio is output, and the audio information includes the audio name and singer name of the similar audio.
[0160] It should be understood that the computer device described in the embodiments of the present application can execute the above Figure 2 、 Figure 4 and Figure 6 The description of the above audio recognition method in the corresponding embodiment can also be performed Figure 8 The description of the audio recognition device in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of the same method will not be repeated here either.
[0161] The present application also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program executed by the aforementioned audio recognition device, and the computer program includes program instructions. When a processor executes the program instructions, the aforementioned audio recognition device can be executed. Figure 2 、 Figure 4 and Figure 6 The description of the audio recognition method in the corresponding embodiment will therefore not be repeated here. In addition, the description of the beneficial effects of using the same method will not be repeated here. For technical details not disclosed in the computer storage medium embodiment involved in this application, please refer to the description of the method embodiment of this application.
[0162] The present application provides a computer program product, which includes computer instructions, which are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the above-mentioned Figure 2 、 Figure 4 and Figure 6 The description of the above-mentioned audio recognition method in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of adopting the same method will not be repeated here. For technical details not disclosed in the computer-readable storage medium embodiment involved in this application, please refer to the description of the method embodiment of this application.
[0163] The terms "first", "second", etc. in the description, claims, and drawings of the embodiments of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, apparatus, product, or device comprising a series of steps or units is not limited to the listed steps or modules, but may optionally include steps or modules not listed, or may optionally include other step units inherent to these processes, methods, apparatuses, products, or devices.
[0164] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0165] The methods and related devices provided in the embodiments of the present application are described with reference to the method flow charts and / or structural diagrams provided in the embodiments of the present application. Specifically, each process and / or block in the method flow charts and / or structural diagrams, as well as the combination of processes and / or blocks in the flow charts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable computer device to generate a machine, so that the instructions executed by the processor of the computer or other programmable computer device generate instructions for implementing the steps in the process. Figure 1 Schematic diagram of one or more processes and / or structures Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable computer device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including an instruction device, which implements the function specified in the process. Figure 1 Schematic diagram of one or more processes and / or structures Figure 1 These computer program instructions can also be loaded onto a computer or other programmable computer device so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the process. Figure 1 The flow or flows and / or structures illustrate the steps of the functions specified in one block or multiple blocks.
[0166] The above disclosure is only a preferred embodiment of the present application, and certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.
Claims
1. An audio recognition method, characterized in that: The method comprises: Inputting the to-be-recognized spectrogram corresponding to the to-be-recognized audio segment into a target audio recognition model to obtain a to-be-recognized feature vector output by the target audio model; wherein the target audio recognition model is obtained by adjusting the model parameters of an initial audio recognition model using an adjustment parameter, the initial audio recognition model including a first task module and a second task module, the adjustment parameter being determined based on a first loss parameter generated by the first task module and a second loss parameter generated by the second task module; the first task module is used to predict the audio category label to which the input audio spectrogram belongs, and the second task module is used to autonomously learn a metric distance function for the audio recognition task; Determine a target feature vector from the music library that satisfies a preset condition with the feature vector to be identified; The target audio pointed to by the target feature vector is determined as a recognition result of the audio segment to be recognized, where the recognition result indicates that the audio segment to be recognized and the target audio are different versions of the same audio.
2. The method according to claim 1, characterized in that The method further comprises: Inputting the training spectrogram into the initial audio recognition model to obtain a first training feature vector output by the first task module and a second training feature vector output by the second task module; the first training feature vector is different from the second training feature vector; Determining the first loss parameter according to the first training feature vector; Determining the second loss parameter according to the second training feature vector; The adjustment parameter is determined according to the first loss parameter and the second loss parameter.
3. The method according to claim 2, characterized in that The determining the first loss parameter according to the first training feature vector includes: determining a predicted audio category label for the training spectrogram based on the first training feature vector; Determining a predicted probability corresponding to the predicted audio category label; The first loss parameter is determined according to the predicted probability.
4. The method according to claim 2, characterized in that The training spectrogram includes a first sample graph, a second sample graph, and a third sample graph, the first sample graph and the second sample graph have the same audio category label, and the first sample graph and the third sample graph have different audio category labels; The determining the second loss parameter according to the second training feature vector includes: Determining a first vector distance between a second training feature vector corresponding to the first sample image and a second training feature vector corresponding to the second sample image; Determining a second vector distance between a second training feature vector corresponding to the first sample image and a second training feature vector corresponding to the third sample image; The second loss parameter is determined according to the first vector distance and the second vector distance.
5. The method according to any one of claims 1 to 4, characterized in that The recognition result also includes an audio category label of the audio segment to be recognized; The method further comprises: The corresponding relationship between the audio category label of the audio segment to be identified and the feature vector to be identified is added to the music library.
6. The method according to any one of claims 1 to 4, characterized in that The method further comprises: Inputting a spectrogram of historical audio in the music library into the target audio recognition model to obtain a third feature vector output by the target audio recognition model; the historical audio has an audio category label; The correspondence between the audio category label of the historical audio and the third feature vector is added to the music library.
7. The method according to claim 6, characterized in that The step of determining a target feature vector from the music library that satisfies a preset condition with the feature vector to be identified includes: Calculating the similarity between the feature vector to be identified and a third feature vector in the music library; If the similarity between the feature vector to be identified and the third feature vector meets a preset condition, the third feature vector is determined to be the target feature vector.
8. An audio recognition method, characterized in that: The method comprises: In response to a recognition request for a humming audio segment, obtaining a humming spectrogram of the humming audio segment; Inputting the humming spectrogram into a target audio recognition model to generate a humming vector to be recognized; the target audio recognition model is obtained by adjusting model parameters of an initial audio recognition model using adjustment parameters, the initial audio recognition model including a first task module and a second task module, the adjustment parameters being determined based on a first loss parameter generated by the first task module and a second loss parameter generated by the second task module; the first task module is used to predict the audio category label to which the input audio spectrogram belongs, and the second task module is used to autonomously learn a metric distance function for the audio recognition task; Determine from the music library a similar feature vector that satisfies a preset condition with the humming vector to be identified, and determine the audio pointed to by the similar feature vector as the similar audio of the humming audio segment; Output the audio information of the similar audio, where the audio information includes the audio name and singer name of the similar audio.
9. A computer device, characterized in that: The computer device includes a processor, a network interface, and a storage device, wherein the processor, the network interface, and the storage device are interconnected, wherein the network interface is controlled by the processor to send and receive data, and the storage device is used to store a computer program, wherein the computer program includes program instructions, and the processor is configured to call the program instructions to execute the audio recognition method according to any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a computer program, and when the computer program is executed by a processor, it is used to implement the audio recognition method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Accoustic model learning apparatus, accoustic model learning method, and program
US20220122626A1
Audio recognition method and system, and device
WO2020156153A1