Audio recognition method, device, equipment and storage medium
By extracting spectrum feature and processing the audio detection model for the recognized audio, combining the cover recognition features in the audio library, calculating similarity and performing secondary recognition of note correlation, the problem of the existing technology being unable to identify unknown song versions is solved, and audio recognition with high accuracy and coverage is achieved.
Patent Information
- Application Number
- CN202110309254.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-23
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-03-23
AI Technical Summary
The prior art cannot accurately identify unknown song versions that are not stored in the music library, resulting in a reduced user experience of audio recognition.
Spectral feature extraction is performed by extracting the audio to be recognized, and cover recognition features are obtained using the audio detection model. Combining the cover recognition features in the audio library, the similarity is calculated to determine the candidate audio, and secondary recognition is performed based on the note correlation.
Accurate recognition of unknown song versions is achieved, the accuracy and coverage of audio recognition is improved, and the user experience is improved.
Smart Images

Figure CN115116472B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to an audio recognition method, device, equipment and storage medium. Background Art
[0002] With the popularity of various terminal applications and the massive growth of audio content, people's demand for audio recognition is increasing.
[0003] In the related technology, for audio recognition, the main audio retrieval technology currently used is based on fingerprint retrieval. This audio retrieval technology is generally used to identify which song in the music library is included in the audio to be identified. In other words, the identified song must be a song version that has been stored in the music library. However, for unknown song versions that are not stored in the music library, such as user singing, singer live performance, lyrics adaptation, and different singers performing the same song, they cannot be accurately identified by existing audio retrieval technology, thereby reducing the user's experience of audio recognition. Summary of the invention
[0004] The present disclosure provides an audio recognition method, device, equipment and storage medium to at least solve at least one problem in the related art, such as the inability to accurately identify unknown song versions and the reduction of the user's experience in audio recognition. The technical solution of the present disclosure is as follows:
[0005] According to a first aspect of an embodiment of the present disclosure, there is provided an audio recognition method, comprising:
[0006] Extract features of the audio to be recognized, and obtain frequency spectrum features of the audio to be recognized;
[0007] Extracting the frequency spectrum features through an audio detection model to obtain a first cover song recognition feature of the audio to be recognized; the audio detection model is used to recognize whether the audio is a cover song audio;
[0008] Acquire a second cover song recognition feature of each audio in the audio library, and determine the similarity between the audio to be recognized and each audio in the audio library according to the first cover song recognition feature and the second cover song recognition feature;
[0009] Determine candidate audio from the audio library according to the similarity determination result;
[0010] Based on the note correlation between the audio to be recognized and each of the candidate audios, a recognition result of the audio to be recognized is determined.
[0011] As an optional implementation manner, before the step of obtaining the second cover recognition feature of each audio in the audio library, the method further includes:
[0012] Acquire each audio in the audio library, perform feature extraction from each audio, and respectively acquire target spectrum features corresponding to each audio;
[0013] Processing each of the target frequency spectrum features through the audio detection model to obtain the cover feature of each audio in the audio library;
[0014] Based on the obtained cover features of each audio in the audio library, construct an audio feature search library;
[0015] Correspondingly, the step of obtaining the second cover recognition feature of each audio in the audio library includes:
[0016] The cover feature of each audio is obtained from the audio feature search library as the second cover recognition feature of each audio in the audio library.
[0017] As an optional implementation manner, the extracting features of the audio to be recognized and obtaining the frequency spectrum features of the audio to be recognized includes:
[0018] Perform overlapping frame processing on the audio data to obtain frame data corresponding to the audio to be recognized;
[0019] Performing Fourier transform on the frame data to obtain transformed data corresponding to the audio to be recognized;
[0020] The transformed data is logarithmically compressed to obtain frequency spectrum features corresponding to the audio to be recognized.
[0021] As an optional implementation manner, extracting the frequency spectrum feature through an audio detection model to obtain a first cover song recognition feature of the audio to be recognized includes:
[0022] Inputting the frequency spectrum feature into an audio detection model, and processing the frequency spectrum feature using a feature extraction module in the audio detection model to obtain a first cover song recognition feature of the audio to be recognized;
[0023] Among them, the feature extraction module includes at least one convolutional pooling unit connected in sequence, and a fully connected layer connected to the pooling layer in the last convolutional pooling unit, and each convolutional pooling unit includes at least one convolutional layer connected in sequence, and a pooling layer connected to the last convolutional layer.
[0024] As an optional implementation manner, determining the similarity between the audio to be identified and each audio in the audio library according to the first cover song identification feature and the second cover song identification feature includes:
[0025] Calculating a vector distance between the first cover song identification feature and each of the second cover song identification features;
[0026] Based on the calculated vector distance, determining the similarity between the audio to be identified and each audio in the audio library; the vector distance is inversely proportional to the similarity;
[0027] Accordingly, determining the candidate audio from the audio library according to the similarity calculation result includes:
[0028] Sort the similarities in descending order, and determine a preset number of audios with the highest order as candidate audios; or,
[0029] The audio corresponding to the similarity greater than the preset threshold is used as the candidate audio.
[0030] As an optional implementation manner, determining the recognition result of the audio to be recognized based on the note correlation between the audio to be recognized and each of the candidate audios includes:
[0031] Extracting note features from the audio to be recognized and each of the candidate audios respectively, and obtaining a first note feature and a second note feature correspondingly;
[0032] calculating a note correlation between the first note feature and each of the second note features;
[0033] If the note correlation is greater than a preset correlation threshold, determining that the corresponding audio to be identified and the candidate audio are the same audio version, and determining that the audio to be identified is a cover audio of the candidate audio of the same audio version;
[0034] If the note correlation is less than or equal to a preset correlation threshold, it is determined that the corresponding audio to be identified and the candidate audio are different audio versions, and the audio to be identified is determined to be a non-cover audio of the candidate audio of the same audio version.
[0035] As an optional implementation manner, extracting note features from the audio to be recognized and each of the candidate audios respectively to obtain a first note feature and a second note feature accordingly includes:
[0036] Separating a first background audio and a second background audio from the audio to be recognized and each of the candidate audios respectively;
[0037] Extract the fundamental frequencies of the first background audio and the second background audio respectively, and obtain a first fundamental frequency and a second fundamental frequency respectively;
[0038] According to the conversion relationship between the fundamental frequency and the note feature, a first note feature corresponding to the first fundamental frequency is determined, and a second note feature corresponding to the second fundamental frequency is determined.
[0039] As an optional implementation, the audio detection model is trained in the following manner:
[0040] Obtain a training set, the training set comprising audio sample pairs, each audio sample pair comprising a spectral feature of an original audio, a spectral feature of a cover audio corresponding to the original audio, and an actual label of whether each audio is a cover audio;
[0041] Calling an initial audio detection model, and extracting features of the audio sample pair through a feature extraction module in the initial audio detection model to obtain a first sample feature and a second sample feature respectively;
[0042] Performing cover detection on the first sample feature and the second sample feature by a classification module in the initial audio detection model, and determining a detection label corresponding to each audio in the audio sample pair;
[0043] Determining a loss function of an initial audio detection model based on the detection label and an actual label corresponding to each audio;
[0044] The initial video and audio detection model is trained according to the loss function to obtain a trained audio detection model.
[0045] According to a second aspect of an embodiment of the present disclosure, there is provided an audio recognition device, including:
[0046] A first feature extraction module is configured to perform feature extraction on the audio to be recognized, and obtain the frequency spectrum features of the audio to be recognized;
[0047] A second feature extraction module is configured to perform feature extraction on the frequency spectrum feature through an audio detection model to obtain a first cover song recognition feature of the audio to be recognized; the audio detection model is used to recognize whether the audio is a cover song audio;
[0048] a similarity determination module configured to obtain a second cover song recognition feature of each audio in the audio library, and determine the similarity between the audio to be recognized and each audio in the audio library according to the first cover song recognition feature and the second cover song recognition feature;
[0049] A screening module is configured to determine candidate audio from the audio library according to the similarity determination result;
[0050] The recognition module is configured to determine a recognition result of the audio to be recognized based on a note correlation between the audio to be recognized and each of the candidate audios.
[0051] As an optional implementation, the device further includes a search library construction module;
[0052] The search library construction module is configured to execute the acquisition of each audio in the audio library, extract features from each audio, and respectively obtain the target spectrum features corresponding to each audio; the third feature extraction module is configured to execute the processing of each target spectrum feature through the audio detection model to obtain the cover features of each audio in the audio library; based on the obtained cover features of each audio in the audio library, an audio feature search library is constructed.
[0053] Correspondingly, the similarity determination module includes an identification feature acquisition submodule; the identification feature acquisition submodule is configured to execute acquisition of the cover feature of each audio from the audio feature search library as the second cover recognition feature of each audio in the audio library.
[0054] As an optional implementation, the first feature extraction module includes:
[0055] The framing submodule is configured to perform overlapping framing processing on the audio data to obtain framing data corresponding to the audio to be recognized;
[0056] A transform submodule, configured to perform Fourier transform on the frame data to obtain transformed data corresponding to the audio to be recognized;
[0057] The compression submodule is configured to perform logarithmic compression processing on the transformed data to obtain the frequency spectrum features corresponding to the audio to be recognized.
[0058] As an optional implementation, the second feature extraction module is specifically configured to execute: inputting the frequency spectrum feature into the audio detection model, processing the frequency spectrum feature using the feature extraction module in the audio detection model, and obtaining a first cover recognition feature of the audio to be recognized;
[0059] Among them, the feature extraction module includes at least one convolutional pooling unit connected in sequence, and a fully connected layer connected to the pooling layer in the last convolutional pooling unit, and each convolutional pooling unit includes at least one convolutional layer connected in sequence, and a pooling layer connected to the last convolutional layer.
[0060] As an optional implementation, the similarity determination module includes:
[0061] A vector distance determination submodule, configured to calculate a vector distance between the first cover song identification feature and each of the second cover song identification features;
[0062] A similarity determination submodule is configured to determine the similarity between the audio to be identified and each audio in the audio library based on the calculated vector distance; the vector distance is inversely proportional to the similarity;
[0063] Accordingly, the screening module includes:
[0064] The first screening submodule is configured to sort the similarities in descending order, and determine a preset number of audios with the highest order as candidate audios; or
[0065] The second screening submodule is configured to select the audio corresponding to the similarity greater than a preset threshold as the candidate audio.
[0066] As an optional implementation, the identification module includes:
[0067] The note feature extraction submodule is configured to perform note feature extraction on the audio to be recognized and each of the candidate audios, respectively, to obtain a first note feature and a second note feature;
[0068] a correlation determination submodule, configured to calculate the note correlation of the first note feature and each of the second note features;
[0069] The first recognition submodule is configured to determine that the corresponding audio to be recognized and the candidate audio are the same audio version if the note correlation is greater than a preset correlation threshold, and determine that the audio to be recognized is a cover audio of the candidate audio of the same audio version;
[0070] The second identification submodule is configured to execute if the note correlation is less than or equal to a preset correlation threshold, determine that the corresponding audio to be identified and the candidate audio are different audio versions, and determine that the audio to be identified is a non-cover audio of the candidate audio of the same audio version.
[0071] As an optional implementation, the note feature extraction submodule includes:
[0072] A separation unit, configured to separate the first background audio and the second background audio from the audio to be recognized and each of the candidate audios respectively;
[0073] An extraction unit is configured to perform fundamental frequency extraction on the first background audio and the second background audio respectively, and obtain a first fundamental frequency and a second fundamental frequency accordingly;
[0074] The note feature determination unit is configured to determine the first note feature corresponding to the first fundamental frequency and the second note feature corresponding to the second fundamental frequency according to the conversion relationship between the fundamental frequency and the note feature.
[0075] As an optional implementation, the audio detection model is trained in the following manner:
[0076] Obtain a training set, the training set comprising audio sample pairs, each audio sample pair comprising a spectral feature of an original audio, a spectral feature of a cover audio corresponding to the original audio, and an actual label of whether each audio is a cover audio;
[0077] Calling an initial audio detection model, and extracting features of the audio sample pair through a feature extraction module in the initial audio detection model to obtain a first sample feature and a second sample feature respectively;
[0078] Performing cover detection on the first sample feature and the second sample feature by a classification module in the initial audio detection model, and determining a detection label corresponding to each audio in the audio sample pair;
[0079] Determining a loss function of an initial audio detection model based on the detection label and an actual label corresponding to each audio;
[0080] The initial video and audio detection model is trained according to the loss function to obtain a trained audio detection model.
[0081] According to a third aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the audio recognition method as described in any of the above embodiments.
[0082] According to a fourth aspect of an embodiment of the present disclosure, there is provided an electronic device, including:
[0083] processor;
[0084] a memory for storing instructions executable by the processor;
[0085] The processor is configured to execute the instructions to implement the audio recognition method as described in any of the above embodiments.
[0086] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided. The computer program product includes a computer program. When the computer program is executed by a processor, the audio recognition method provided in any one of the above-mentioned embodiments is implemented.
[0087] The technical solution provided by the embodiments of the present disclosure brings at least the following beneficial effects:
[0088] The disclosed embodiment extracts features from the audio to be identified to obtain the spectrum features of the audio to be identified; extracts features from the spectrum features through an audio detection model to obtain the first cover recognition feature of the audio to be identified; the audio detection model is used to identify whether the audio is a cover audio; obtains the second cover recognition feature of each audio in the audio library, and determines the similarity between the audio to be identified and each audio in the audio library based on the first cover recognition feature and the second cover recognition feature; determines the candidate audio from the audio library based on the similarity determination result; determines the recognition result of the audio to be identified based on the note correlation between the audio to be identified and each of the candidate audios. Thus, by first screening the candidate audio, and then performing secondary recognition on the audio to be identified based on the note correlation between the candidate audio and the audio to be identified, the audio of unknown audio versions can be fully recognized with high recognition accuracy, while also improving the coverage of audio recognition and the user's experience of audio recognition, thereby helping to improve the user's stickiness to the product.
[0089] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0090] The drawings herein are incorporated into the specification and constitute a part of the present disclosure, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute improper limitations on the present disclosure.
[0091] Figure 1 The invention is an architecture diagram of a system for applying an audio recognition method according to an exemplary embodiment.
[0092] Figure 2 The present invention is a flowchart of an audio recognition method according to an exemplary embodiment.
[0093] Figure 3 It is a partial flow chart of another audio recognition method according to an exemplary embodiment.
[0094] Figure 4 The present invention is a flowchart showing a step of determining a recognition result of audio to be recognized according to an exemplary embodiment.
[0095] Figure 5 It is a structural diagram of an audio detection model according to an exemplary embodiment.
[0096] Figure 6 The figure is a block diagram of an audio recognition device according to an exemplary embodiment.
[0097] Figure 7 It is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0098] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings.
[0099] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0100] Figure 1 is an architecture diagram of a system for applying an audio recognition method according to an exemplary embodiment. Figure 1 , the architecture diagram may include a terminal 10 and a server 20.
[0101] Among them, the terminal 10 can be but is not limited to a physical device such as a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart wearable device, a digital assistant, an augmented reality device, a virtual reality device, or one or more applications or applets running in the physical device.
[0102] The server 20 can provide background services such as an audio library for the terminal. As an example only, the server 20 can be, but is not limited to, an independent server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, and basic cloud computing services such as big data and artificial intelligence platforms. The terminal 10 and the server 20 can be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiments of the present disclosure.
[0103] The audio recognition method provided in the embodiments of the present disclosure may be executed by an audio recognition device, which may be integrated in an electronic device such as a client or a server in the form of hardware or software, or may be executed by a terminal or a server alone, or may be executed by a terminal and a server in collaboration.
[0104] First, the application scenarios involved in the embodiments of the present disclosure are introduced:
[0105] In an exemplary application scenario, for example, in a song identification scenario, the user can turn on the audio recognition function on the terminal, and the terminal can collect the songs to be identified on the spot, or the user can input the songs to be identified to the terminal by humming, uploading, etc.; then, the terminal executes the audio recognition method provided by the embodiment of the present disclosure on the songs to be identified, and determines the music version of the music library corresponding to the songs to be identified, or the identification result of whether the songs to be identified are cover songs.
[0106] In a video watching scenario, a user is interested in the music played in the video being watched (e.g., a short video, a film or TV work, etc.), and wants to know which song the music belongs to. The user can turn on the audio recognition function of the corresponding application on the terminal playing the video, collect the music of the video being played, and use the collected music as the audio to be recognized, and can send the audio to be recognized to the server, and the server performs the audio recognition method provided by the embodiment of the present disclosure on the audio to be recognized, and determines the music version of the music library corresponding to the audio to be recognized, or whether the audio to be recognized is a cover song.
[0107] It should be noted that the application scenarios of the embodiments of the present disclosure include but are not limited to the above application scenarios, and may also be applicable to other scenarios requiring audio recognition.
[0108] Figure 2 is a flow chart of an audio recognition method according to an exemplary embodiment. Figure 2 As shown, the audio recognition method can be applied to an electronic device, and the electronic device is taken as the terminal in the above implementation environment diagram as an example for explanation, including the following steps.
[0109] In step S201, feature extraction is performed on the audio to be recognized to obtain frequency spectrum features of the audio to be recognized.
[0110] Among them, the spectrum feature is data used to reflect the characteristic information of audio, and each audio has its own inherent spectrum feature.
[0111] Optionally, the terminal can obtain the audio to be recognized in response to the audio recognition instruction, and perform feature extraction on the obtained audio to be recognized, so as to extract the spectral features representing the audio characteristics from the audio to be recognized. The audio to be recognized can be at least one of the audio input by the user through humming, manual input, etc., the audio collected by the terminal through recording or real-time, and the audio obtained from the local storage library or other devices (such as the cloud or server, etc.). The audio to be recognized can be of the type of song, accompaniment, humming, video, etc., and the number can be one or more.
[0112] Optionally, after obtaining the audio to be identified, the audio to be identified may be preprocessed and uniformly converted into audio data after pulse code modulation (pcm), for example, it may be converted into formats including but not limited to 11K pcm, 16K pcm, etc. Next, feature extraction is performed on the converted audio data to obtain the spectrum features corresponding to the audio to be identified.
[0113] Among many spectral features, the constant Q transform feature (CQT) is a two-dimensional representation, which is data used to reflect the characteristic information of the notes and melodies of the audio. Here, the spectral feature including the constant Q transform feature is taken as an example to specifically illustrate the spectral feature extraction process. In an optional implementation, in the above step S201, the feature extraction of the audio to be identified to obtain the spectral features of the audio to be identified includes:
[0114] In step S2011, the audio data is subjected to overlapping frame processing to obtain frame data corresponding to the audio to be recognized;
[0115] In step S2012, Fourier transform is performed on the frame data to obtain transformed data corresponding to the audio to be recognized;
[0116] In step S2013, logarithmic compression processing is performed on the transformed data to obtain frequency spectrum features corresponding to the audio to be recognized.
[0117] Optionally, the audio data or the converted audio data is subjected to overlapping frame processing to obtain the frame data corresponding to the audio to be identified, that is, there is a fixed length of data overlap between adjacent frames obtained by framing, and the overlapping parts of the data can be averaged during the subsequent merging process. Next, the frame data is subjected to Fourier transform to obtain the transformed data corresponding to the audio to be identified. Afterwards, based on the Fourier transform, the transformed data obtained by the Fourier transform is subjected to logarithmic compression processing, so that the features subjected to logarithmic compression processing are more consistent with human ear perception, because the human ear's sensitivity to frequency perception roughly conforms to a logarithmic distribution, so as to better reflect the characteristics of audio notes and melodies.
[0118] In the above-mentioned embodiment, overlapping frame processing, Fourier transform and logarithmic compression processing are performed on the audio data in sequence to obtain the spectral features corresponding to the audio to be identified, so that the obtained spectral features are more in line with human ear perception and can better reflect the characteristics of audio notes and melodies, which is conducive to improving the accuracy of audio recognition.
[0119] In step S202, feature extraction is performed on the frequency spectrum feature through an audio detection model to obtain a first cover song recognition feature of the audio to be recognized; the audio detection model is used to recognize whether the audio is a cover song audio.
[0120] Optionally, the audio detection model can be a trained neural network model. As an example only, the audio detection model can be a convolutional neural network (CNN) model. Since the audio detection model is used to identify whether the audio is a cover song, that is, the audio detection model is a classification model for detecting whether the audio is a cover song. The learning goal of the audio detection model is to determine whether the audio is a cover song, so the features extracted in the model are cover song recognition features.
[0121] Optionally, after the frequency spectrum feature is acquired, the frequency spectrum feature may be input into an audio detection model, and the frequency spectrum feature may be calculated using the audio detection model to obtain a first cover song recognition feature corresponding to the frequency spectrum feature.
[0122] In step S203, the second cover song recognition feature of each audio in the audio library is obtained, and the similarity between the audio to be recognized and each audio in the audio library is determined according to the first cover song recognition feature and the second cover song recognition feature.
[0123] Optionally, the terminal may obtain the second cover recognition feature of each audio in the pre-stored audio library from the server, or the terminal may extract the second cover recognition feature of each audio in the audio library. The extraction step of the second cover recognition feature is similar to the extraction step of the first cover recognition feature, that is, feature extraction is performed on each audio in the audio library to obtain the corresponding spectrum feature, and then the extracted feature is processed by the audio detection model to obtain the second cover recognition feature of each audio in the audio library.
[0124] In another alternative, such as Figure 3 As shown, before the step of obtaining the second cover recognition feature of each audio in the audio library, the following steps may also be included:
[0125] In step S301, each audio in the audio library is obtained, and features are extracted from each audio to obtain target spectrum features corresponding to each audio;
[0126] In step S302, each of the target frequency spectrum features is processed by the audio detection model to obtain the cover feature of each audio in the audio library;
[0127] In step S303, an audio feature search library is constructed based on the acquired cover features of each audio in the audio library.
[0128] At this time, in the above step S203, the step of obtaining the second cover recognition feature of each audio in the audio library may include:
[0129] In step S304, the cover feature of each audio is obtained from the audio feature search library as the second cover recognition feature of each audio in the audio library.
[0130] Optionally, the terminal can perform feature extraction on each audio in the audio library to obtain the target spectrum features corresponding to each audio. Then, each target spectrum feature is processed by an audio detection model to obtain the cover features of each audio in the audio library. Then, the cover features of each audio in the audio library are obtained and stored correspondingly with the corresponding audio, an audio feature search library is constructed, and the audio feature search library can be stored in a local storage library or other device (such as a cloud or server, etc.), so as to obtain the pre-stored cover features of each audio from the audio feature search library when needed, as the second cover recognition feature of each audio in the audio library.
[0131] The above embodiment, by pre-building an audio feature search library, directly obtains the cover feature of each audio from the audio feature search library as the second cover recognition feature of each audio in the audio library when audio recognition is needed, thereby reducing the time spent on cover feature recognition processing for each audio in the audio library separately, improving audio recognition efficiency and user experience of audio recognition, and helping to improve user stickiness to the product.
[0132] In an optional implementation manner, in the above step S203, determining the similarity between the audio to be identified and each audio in the audio library according to the first cover song identification feature and the second cover song identification feature includes:
[0133] In step S2031, calculating the vector distance between the first cover song identification feature and each of the second cover song identification features;
[0134] In step S2032, based on the calculated vector distance, the similarity between the audio to be identified and each audio in the audio library is determined; the vector distance is inversely proportional to the similarity.
[0135] Optionally, the first cover song recognition feature and each second cover song recognition feature may be represented by vectors, and then the vector distance between the first cover song recognition feature represented by the vector and each second cover song recognition feature represented by the vector is calculated, and the vector distance may include but is not limited to cosine distance and Euclidean distance. Then, based on converting the calculated vector distance into the similarity between the audio to be recognized and each audio in the audio library. The smaller the calculated vector distance, the greater the similarity between the two, and vice versa.
[0136] In step S204, candidate audio is determined from the audio library according to the similarity determination result.
[0137] Optionally, determining the candidate audio from the audio library according to the similarity calculation result may include: sorting the similarities in descending order, and determining a preset number of audios with the highest order as the candidate audio. Alternatively, determining the audio corresponding to the similarity greater than a preset threshold as the candidate audio.
[0138] Specifically, according to the order of similarity, a preset number k of audios ranked at the top, that is, the TOP k audios with the greatest similarity, can be determined as candidate audios. As an example only, the preset number k can be, but is not limited to, any value between 20 and 100, such as 30, 50, etc. Alternatively, the determined similarities can be compared with the size of a preset threshold, and the audio corresponding to the similarity greater than the preset threshold can be determined as the candidate audio. As an example only, the specific value of the preset threshold can be set according to the actual situation, and the present disclosure does not make specific limitations on this.
[0139] The above embodiment calculates the vector distance between the first cover recognition feature and each second cover recognition feature; then, based on the calculated vector distance, determines the similarity between the audio to be recognized and each audio in the audio library, which is conducive to preliminary screening of each audio in the audio library, quickly screening out candidate audio with high similarity, and excluding other audio with low similarity, which not only greatly reduces the amount of calculation for subsequently determining the recognition result of the audio to be recognized, and reduces the time consumption of audio recognition, but also improves the accuracy of subsequent audio recognition to a certain extent.
[0140] In step S205, a recognition result of the audio to be recognized is determined based on the note correlation between the audio to be recognized and each of the candidate audios.
[0141] In an optional implementation, in the above step S205, if Figure 4 As shown, the step of determining the recognition result of the audio to be recognized based on the note correlation between the audio to be recognized and each of the candidate audios includes:
[0142] In step S401, note features are extracted from the audio to be recognized and each of the candidate audios, respectively, to obtain a first note feature and a second note feature accordingly.
[0143] Optionally, the note feature may be a MIDI (Musical Instrument Digital Interface) feature. MIDI is a digital storage format for music, and is also the music score that is best understood by computers. It can accurately tell the music player the playing time, pitch, timbre, duration, and other information of each note.
[0144] The note features extracted here refer to the note density, average note pitch, note pitch variance, overall pitch, and time features reflecting the order in which notes appear, which can characterize rich audio melodies and styles.
[0145] In an optional implementation, taking the note feature as a MIMD feature as an example, extracting the note features of the audio to be recognized and each of the candidate audios respectively, and correspondingly obtaining the first note feature and the second note feature may include:
[0146] In step S4011, a first background audio and a second background audio are separated from the audio to be recognized and each of the candidate audios respectively;
[0147] In step S4012, fundamental frequencies are extracted from the first background audio and the second background audio respectively, to obtain a first fundamental frequency and a second fundamental frequency respectively;
[0148] In step S4013, according to the conversion relationship between the fundamental frequency and the note feature, the first note feature corresponding to the first fundamental frequency is determined, and the second note feature corresponding to the second fundamental frequency is determined.
[0149] Optionally, in order to reduce the interference of human voice, human voice can be separated from the audio to be recognized and each candidate audio, that is, the first background audio and the second background audio are separated accordingly. Then, the fundamental frequency is extracted from the separated first background audio and the second background audio, respectively, to obtain the first fundamental frequency and the second fundamental frequency accordingly. Based on the conversion relationship between the fundamental frequency and the note feature, the first note feature corresponding to the first fundamental frequency is determined, and the second note feature corresponding to the second fundamental frequency is determined. Among them, the expression of the conversion relationship between the fundamental frequency and the note feature can be implemented by the following formula:
[0150]
[0151] Among them, p represents the MIDI feature corresponding to the note, f represents the frequency corresponding to the note, and k1, k2 and k3 are constants.
[0152] In the above embodiment, the human voice separation and fundamental frequency extraction are performed on the audio to be recognized and each candidate audio respectively, and the corresponding note features are determined according to the conversion relationship between the fundamental frequency and the note features, so that the acquired note features better reflect the characteristics of the audio, which is conducive to improving the accuracy of subsequent audio recognition.
[0153] In step S402, the note correlation between the first note feature and each of the second note features is calculated.
[0154] Optionally, the first note feature and each second note feature may be vectorized first, and then a DTW (Dynamic Time Warping) algorithm may be used to respectively calculate the vector distance between the vectorized first note feature and each second note feature, thereby obtaining the corresponding note correlation. The note correlation is inversely proportional to the vector distance.
[0155] In step S403, if the note correlation is greater than a preset correlation threshold, it is determined that the corresponding audio to be identified and the candidate audio are of the same audio version, and the audio to be identified is determined to be a cover version of the candidate audio of the same audio version;
[0156] In step S404, if the note correlation is less than or equal to a preset correlation threshold, it is determined that the corresponding audio to be identified and the candidate audio are different audio versions, and the audio to be identified is determined to be a non-cover audio of the candidate audio of the same audio version.
[0157] Optionally, if the calculated note correlation is greater than a preset correlation threshold, it means that the melody of the audio to be identified is close to that of one of the candidate audios, and it can be determined that the two audios are the same audio version. Next, fingerprint recognition can be performed on the audio to be identified and the original audio of the same audio version. If the fingerprint recognition result shows that the audio to be identified is not the original audio, the audio to be identified is determined to be a cover audio; otherwise, the audio to be identified is determined to be the original audio. If the calculated note correlation is less than or equal to the preset correlation threshold, it is determined that the corresponding audio to be identified and the candidate audio are different audio versions, and the audio to be identified is determined to be a non-cover audio of the candidate audio of the same audio version.
[0158] The preset relevance threshold may be set according to actual conditions, and the present disclosure does not make any specific limitation on this.
[0159] In the above embodiment, note features are extracted from the audio to be identified and each candidate audio respectively, and the note correlation between the extracted first note feature and each second note feature is calculated; then, according to the relationship between the calculated note correlation and the preset correlation threshold, it is determined whether the audio to be identified is a unified audio version of the corresponding candidate audio. If the determination result is yes, the audio to be identified is determined to be a cover audio, otherwise it is determined to be a non-cover audio. In this way, various types of audio to be identified can be fully identified, the coverage of audio identification is improved, and it is beneficial to improve the user experience of audio identification and the stickiness of product use.
[0160] In an optional implementation, in the above step S202, extracting the frequency spectrum feature by using an audio detection model to obtain the first cover song recognition feature of the audio to be recognized may include:
[0161] The frequency spectrum feature is input into the audio detection model, and the frequency spectrum feature is processed by a feature extraction module in the audio detection model to obtain a first cover song recognition feature of the audio to be recognized.
[0162] Optionally, taking the audio detection model as a convolutional neural network (CNN) model as an example, the feature extraction module in the audio detection model may include at least one convolutional pooling unit connected in sequence, and a fully connected layer connected to the pooling layer in the last convolutional pooling unit, and each convolutional pooling unit includes at least one convolutional layer connected in sequence, and a pooling layer connected to the last convolutional layer. Specifically, Figure 5 As shown, the feature extraction module may include 5 convolutional pooling units and a fully connected layer, wherein the first 4 convolutional pooling units of the 5 convolutional pooling units include 2 convolutional layers and 1 maximum pooling layer connected in sequence, and the 5th convolutional pooling unit includes 2 convolutional layers and 1 adaptive maximum pooling layer connected in sequence, and the output of the adaptive maximum pooling layer is connected to the fully connected layer. Figure 5 As shown, the spectral features are input into the audio detection model, and are calculated in sequence through 5 convolutional pooling units. The calculation result of each convolutional pooling unit is used as the input of the next convolutional pooling unit, and the calculation result of the last convolutional pooling unit is used as the input of the first fully connected layer. The output of the first fully connected layer is the cover recognition feature, that is, the first cover recognition feature of the audio to be recognized is obtained.
[0163] In an optional implementation manner, the audio detection model is trained in the following manner:
[0164] Obtain a training set, the training set comprising audio sample pairs, each audio sample pair comprising a spectral feature of an original audio, a spectral feature of a cover audio corresponding to the original audio, and an actual label of whether each audio is a cover audio;
[0165] Calling an initial audio detection model, and extracting features of the audio sample pair through a feature extraction module in the initial audio detection model to obtain a first sample feature and a second sample feature respectively;
[0166] Performing cover detection on the first sample feature and the second sample feature by a classification module in the initial audio detection model, and determining a detection label corresponding to each audio in the audio sample pair;
[0167] Determining a loss function of an initial audio detection model based on the detection label and an actual label corresponding to each audio;
[0168] The initial video and audio detection model is trained according to the loss function to obtain a trained audio detection model.
[0169] As an example only, the loss function of the initial audio detection model may be a cross entropy loss function. The initial audio detection model is back-propagated according to the loss function, and the network parameters are continuously optimized until the training end condition is met to obtain a trained audio detection model. The training end condition may include but is not limited to minimizing the loss function, reaching a preset number of training times, etc.
[0170] The training process of the above-mentioned audio detection model is similar to that of the above-mentioned method embodiment. Only audio sample pairs are used in the training process. Each audio sample pair includes the spectral features of the original audio, the spectral features of the cover audio corresponding to the original audio, and the actual label of whether each audio is a cover audio. The specific model training process will not be repeated here.
[0171] In the above embodiment, the spectral features of the audio to be identified are obtained by extracting features from the audio to be identified; the first cover recognition feature of the audio to be identified is obtained by extracting features from the spectral features through an audio detection model; the audio detection model is used to identify whether the audio is a cover audio; the second cover recognition feature of each audio in the audio library is obtained, and the similarity between the audio to be identified and each audio in the audio library is determined based on the first cover recognition feature and the second cover recognition feature; the candidate audio is determined from the audio library based on the similarity determination result; the recognition result of the audio to be identified is determined based on the note correlation between the audio to be identified and each of the candidate audios. Thus, by first screening the candidate audio, and then performing secondary recognition on the audio to be identified based on the note correlation between the candidate audio and the audio to be identified, the audio of unknown audio versions can be fully recognized with high recognition accuracy, while also improving the coverage of audio recognition and the user's experience of audio recognition, thereby helping to improve the user's stickiness to the product.
[0172] Figure 6 FIG. 1 is a block diagram of an audio recognition device according to an exemplary embodiment. Figure 6 , the device is applied to electronic equipment, including:
[0173] A first feature extraction module 610 is configured to perform feature extraction on the audio to be recognized, and obtain the frequency spectrum features of the audio to be recognized;
[0174] The second feature extraction module 620 is configured to perform feature extraction on the frequency spectrum feature through an audio detection model to obtain a first cover song recognition feature of the audio to be recognized; the audio detection model is used to recognize whether the audio is a cover song audio;
[0175] The similarity determination module 630 is configured to obtain a second cover song recognition feature of each audio in the audio library, and determine the similarity between the audio to be recognized and each audio in the audio library according to the first cover song recognition feature and the second cover song recognition feature;
[0176] A screening module 640 is configured to determine candidate audio from the audio library according to the similarity determination result;
[0177] The recognition module 650 is configured to determine the recognition result of the audio to be recognized based on the note correlation between the audio to be recognized and each of the candidate audios.
[0178] As an optional implementation, the device further includes a search library construction module;
[0179] The search library construction module is configured to execute the acquisition of each audio in the audio library, extract features from each audio, and respectively obtain the target spectrum features corresponding to each audio; the third feature extraction module is configured to execute the processing of each target spectrum feature through the audio detection model to obtain the cover features of each audio in the audio library; based on the obtained cover features of each audio in the audio library, an audio feature search library is constructed.
[0180] Correspondingly, the similarity determination module includes an identification feature acquisition submodule; the identification feature acquisition submodule is configured to execute acquisition of the cover feature of each audio from the audio feature search library as the second cover recognition feature of each audio in the audio library.
[0181] As an optional implementation, the first feature extraction module includes:
[0182] The framing submodule is configured to perform overlapping framing processing on the audio data to obtain framing data corresponding to the audio to be recognized;
[0183] A transform submodule, configured to perform Fourier transform on the frame data to obtain transformed data corresponding to the audio to be recognized;
[0184] The compression submodule is configured to perform logarithmic compression processing on the transformed data to obtain the frequency spectrum features corresponding to the audio to be recognized.
[0185] As an optional implementation, the second feature extraction module is specifically configured to execute: inputting the frequency spectrum feature into the audio detection model, processing the frequency spectrum feature using the feature extraction module in the audio detection model, and obtaining a first cover recognition feature of the audio to be recognized;
[0186] Among them, the feature extraction module includes at least one convolutional pooling unit connected in sequence, and a fully connected layer connected to the pooling layer in the last convolutional pooling unit, and each convolutional pooling unit includes at least one convolutional layer connected in sequence, and a pooling layer connected to the last convolutional layer.
[0187] As an optional implementation, the similarity determination module includes:
[0188] A vector distance determination submodule, configured to calculate a vector distance between the first cover song identification feature and each of the second cover song identification features;
[0189] A similarity determination submodule is configured to determine the similarity between the audio to be identified and each audio in the audio library based on the calculated vector distance; the vector distance is inversely proportional to the similarity;
[0190] Accordingly, the screening module includes:
[0191] The first screening submodule is configured to sort the similarities in descending order, and determine a preset number of audios with the highest order as candidate audios; or
[0192] The second screening submodule is configured to select the audio corresponding to the similarity greater than a preset threshold as the candidate audio.
[0193] As an optional implementation, the identification module includes:
[0194] The note feature extraction submodule is configured to perform note feature extraction on the audio to be recognized and each of the candidate audios, respectively, to obtain a first note feature and a second note feature;
[0195] a correlation determination submodule, configured to calculate the note correlation of the first note feature and each of the second note features;
[0196] The first recognition submodule is configured to determine that the corresponding audio to be recognized and the candidate audio are the same audio version if the note correlation is greater than a preset correlation threshold, and determine that the audio to be recognized is a cover audio of the candidate audio of the same audio version;
[0197] The second identification submodule is configured to execute if the note correlation is less than or equal to a preset correlation threshold, determine that the corresponding audio to be identified and the candidate audio are different audio versions, and determine that the audio to be identified is a non-cover audio of the candidate audio of the same audio version.
[0198] As an optional implementation, the note feature extraction submodule includes:
[0199] A separation unit, configured to separate the first background audio and the second background audio from the audio to be recognized and each of the candidate audios respectively;
[0200] An extraction unit is configured to perform fundamental frequency extraction on the first background audio and the second background audio respectively, and obtain a first fundamental frequency and a second fundamental frequency accordingly;
[0201] The note feature determination unit is configured to determine the first note feature corresponding to the first fundamental frequency and the second note feature corresponding to the second fundamental frequency according to the conversion relationship between the fundamental frequency and the note feature.
[0202] As an optional implementation, the audio detection model is trained in the following manner:
[0203] Obtain a training set, the training set comprising audio sample pairs, each audio sample pair comprising a spectral feature of an original audio, a spectral feature of a cover audio corresponding to the original audio, and an actual label of whether each audio is a cover audio;
[0204] Calling an initial audio detection model, and extracting features of the audio sample pair through a feature extraction module in the initial audio detection model to obtain a first sample feature and a second sample feature respectively;
[0205] Performing cover detection on the first sample feature and the second sample feature by a classification module in the initial audio detection model, and determining a detection label corresponding to each audio in the audio sample pair;
[0206] Determining a loss function of an initial audio detection model based on the detection label and an actual label corresponding to each audio;
[0207] The initial video and audio detection model is trained according to the loss function to obtain a trained audio detection model.
[0208] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0209] In an exemplary embodiment, an electronic device is also provided, which includes a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to implement the steps of any audio recognition method in the above embodiments when executing the instructions stored in the memory.
[0210] The electronic device may be a terminal, a server or a similar computing device. For example, the electronic device is a terminal. Figure 7 is a block diagram of an electronic device for audio recognition according to an exemplary embodiment, specifically:
[0211] The terminal may include components such as an RF (Radio Frequency) circuit 1110, a memory 1120 including one or more computer-readable storage media, an input unit 1130, a display unit 1140, a sensor 1150, an audio circuit 1160, a WiFi (wireless fidelity) module 1170, a processor 1180 including one or more processing cores, and a power supply 1190. Those skilled in the art will appreciate that Figure 7 The terminal structure shown in the figure does not constitute a limitation on the terminal, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0212] The RF circuit 1110 can be used for receiving and sending signals during information transmission or calls. In particular, after receiving the downlink information of the base station, it is handed over to one or more processors 1180 for processing; in addition, the data related to the uplink is sent to the base station. Generally, the RF circuit 1110 includes but is not limited to an antenna, at least one amplifier, a tuner, one or more oscillators, a user identity module (SIM) card, a transceiver, a coupler, an LNA (Low Noise Amplifier), a duplexer, etc. In addition, the RF circuit 1110 can also communicate with the network and other terminals through wireless communication. The wireless communication can use any communication standard or protocol, including but not limited to GSM (Global System of Mobile communication), GPRS (General Packet Radio Service), CDMA (Code Division Multiple Access), WCDMA (Wideband Code Division Multiple Access), LTE (Long Term Evolution), email, SMS (Short Messaging Service), etc.
[0213] The memory 1120 can be used to store software programs and modules, and the processor 1180 executes various functional applications and data processing by running the software programs and modules stored in the memory 1120. The memory 1120 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, application programs required for functions, etc.; the data storage area may store data created according to the use of the terminal, etc. In addition, the memory 1120 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. Accordingly, the memory 1120 may also include a memory controller to provide the processor 1180 and the input unit 1130 with access to the memory 1120.
[0214] The input unit 1130 can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control. Specifically, the input unit 1130 may include a touch-sensitive surface 1131 and other input devices 1132. The touch-sensitive surface 1131, also known as a touch display screen or touch pad, can collect user touch operations on or near it (such as operations performed by users using fingers, styluses, or any other suitable objects or accessories on or near the touch-sensitive surface 1131), and drive corresponding connection devices according to a pre-set program. Optionally, the touch-sensitive surface 1131 may include a touch detection device and a touch controller. Among them, the touch detection device detects the user's touch orientation, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into touch point coordinates, and then sends it to the processor 1180, and can receive and execute commands sent by the processor 1180. In addition, the touch-sensitive surface 1131 may be implemented in various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch-sensitive surface 1131, the input unit 1130 may further include other input devices 1132. Specifically, the other input devices 1132 may include, but are not limited to, one or more of a physical keyboard, a function key (such as a volume control key, a switch key, etc.), a trackball, a mouse, a joystick, and the like.
[0215] The display unit 1140 can be used to display information input by the user or information provided to the user and various graphical user interfaces of the terminal, which can be composed of graphics, text, icons, videos and any combination thereof. The display unit 1140 may include a display panel 1141. Optionally, the display panel 1141 may be configured in the form of LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode), etc. Further, the touch-sensitive surface 1131 may cover the display panel 1141. When the touch-sensitive surface 1131 detects a touch operation on or near it, it is transmitted to the processor 1180 to determine the type of touch event, and then the processor 1180 provides corresponding visual output on the display panel 1141 according to the type of touch event. Among them, the touch-sensitive surface 1131 and the display panel 1141 can be two independent components to realize input and output functions, but in some embodiments, the touch-sensitive surface 1131 and the display panel 1141 can also be integrated to realize input and output functions.
[0216] The terminal may also include at least one sensor 1150, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor, wherein the ambient light sensor may adjust the brightness of the display panel 1141 according to the brightness of the ambient light, and the proximity sensor may turn off the display panel 1141 and / or the backlight when the terminal is moved to the ear. As a type of motion sensor, the gravity acceleration sensor can detect the magnitude of acceleration in each direction (generally three axes), and can detect the magnitude and direction of gravity when stationary. It can be used for applications that identify the terminal posture (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors such as gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc. that can be configured in the terminal, they will not be repeated here.
[0217] The audio circuit 1160, the speaker 1161, and the microphone 1162 can provide an audio interface between the user and the terminal. The audio circuit 1160 can transmit the received audio data to the speaker 1161 after converting the received audio data into an electrical signal, which is converted into a sound signal for output; on the other hand, the microphone 1162 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1160 and converted into audio data, and then the audio data is processed by the output processor 1180 and sent to, for example, another terminal through the RF circuit 1110, or the audio data is output to the memory 1120 for further processing. The audio circuit 1160 may also include an earplug jack to provide communication between an external headset and the terminal.
[0218] WiFi is a short-range wireless transmission technology. The terminal can help users send and receive emails, browse web pages, and access streaming media through the WiFi module 1170, which provides users with wireless broadband Internet access. Figure 7 A WiFi module 1170 is shown, but it is understandable that it is not an essential component of the terminal and can be omitted as needed without changing the essence of the invention.
[0219] The processor 1180 is the control center of the terminal, and uses various interfaces and lines to connect various parts of the entire terminal. By running or executing software programs and / or modules stored in the memory 1120, and calling data stored in the memory 1120, the processor 1180 performs various functions of the terminal and processes data, thereby monitoring the terminal as a whole. Optionally, the processor 1180 may include one or more processing cores; preferably, the processor 1180 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 1180.
[0220] The terminal also includes a power supply 1190 (such as a battery) for supplying power to each component. Preferably, the power supply can be logically connected to the processor 1180 through a power management system, so that the power management system can manage charging, discharging, and power consumption management. The power supply 1190 can also include any components such as one or more DC or AC power supplies, recharging systems, power failure detection circuits, power converters or inverters, and power status indicators.
[0221] Although not shown, the terminal may also include a camera, a Bluetooth module, etc., which will not be described in detail herein. Specifically in this embodiment, the terminal also includes a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by one or more processors. The one or more programs include instructions for executing the virtual resource claiming method provided in the above method embodiment.
[0222] In an exemplary embodiment, a computer storage medium is also provided. When instructions in the computer storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the steps of the method provided in any one of the above embodiments.
[0223] In an exemplary embodiment, a computer program product is also provided, the computer program product comprising a computer program / instruction, and the computer program / instruction is executed by a processor to implement the method provided in any of the above embodiments. Optionally, the computer program is stored in a computer-readable storage medium. The processor of the electronic device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the electronic device performs the method provided in any of the above embodiments.
[0224] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0225] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.
[0226] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. An audio recognition method, It is characterized in that include: Perform feature extraction on the audio to be recognized, and obtain the frequency spectrum features of the audio to be recognized; the audio to be recognized includes audio of an unknown audio version; Extracting the frequency spectrum features through an audio detection model to obtain a first cover song recognition feature of the audio to be recognized; the audio detection model is used to recognize whether the audio is a cover song audio; Performing feature search from an audio feature search library to obtain a second cover song recognition feature of each audio in the audio library, and determining the similarity between the audio to be recognized and each audio in the audio library according to the first cover song recognition feature and the second cover song recognition feature; Determine candidate audio from the audio library according to the similarity determination result; Based on the note correlation between the audio to be identified and each of the candidate audios, determining the identification result of the audio to be identified, including: if the note correlation between the audio to be identified and the corresponding candidate audio is greater than a preset correlation threshold, determining that the audio to be identified and the corresponding candidate audio are the same audio version; performing fingerprint identification on the audio to be identified and the original audio of the same audio version, if the fingerprint identification result indicates that the audio to be identified is not the original audio, determining that the audio to be identified is a cover audio of the original audio; Among them, the note correlation is obtained by calculating the first note feature of the audio to be identified and the second note feature of each of the candidate audios; the first note feature and the second note feature are respectively obtained by performing vocal separation and fundamental frequency extraction on the audio to be identified and each of the candidate audios, and are determined based on the conversion relationship between the fundamental frequency and the note feature.
2. The audio recognition method according to claim 1, It is characterized in that Before the step of obtaining the second cover recognition feature of each audio in the audio library, the method further includes: Acquire each audio in the audio library, perform feature extraction from each audio, and respectively acquire target spectrum features corresponding to each audio; Processing each of the target frequency spectrum features through the audio detection model to obtain the cover feature of each audio in the audio library; Based on the obtained cover features of each audio in the audio library, construct an audio feature search library; Correspondingly, the step of obtaining the second cover recognition feature of each audio in the audio library includes: The cover feature of each audio is obtained from the audio feature search library as the second cover recognition feature of each audio in the audio library.
3. The audio recognition method according to claim 1, It is characterized in that The extracting features of the audio to be recognized and obtaining the frequency spectrum features of the audio to be recognized includes: Perform overlapping frame processing on the audio data to obtain frame data corresponding to the audio to be recognized; Performing Fourier transform on the frame data to obtain transformed data corresponding to the audio to be recognized; The transformed data is logarithmically compressed to obtain frequency spectrum features corresponding to the audio to be recognized.
4. The audio recognition method according to any one of claims 1 to 3, It is characterized in that The extracting the frequency spectrum feature by using an audio detection model to obtain a first cover song recognition feature of the audio to be recognized includes: Inputting the frequency spectrum feature into an audio detection model, and processing the frequency spectrum feature using a feature extraction module in the audio detection model to obtain a first cover song recognition feature of the audio to be recognized; Among them, the feature extraction module includes at least one convolutional pooling unit connected in sequence, and a fully connected layer connected to the pooling layer in the last convolutional pooling unit, and each convolutional pooling unit includes at least one convolutional layer connected in sequence, and a pooling layer connected to the last convolutional layer.
5. The audio recognition method according to any one of claims 1 to 3, It is characterized in that The determining, according to the first cover song recognition feature and the second cover song recognition feature, the similarity between the audio to be recognized and each audio in the audio library comprises: Calculating a vector distance between the first cover song identification feature and each of the second cover song identification features; Based on the calculated vector distance, determining the similarity between the audio to be identified and each audio in the audio library; the vector distance is inversely proportional to the similarity; Accordingly, determining the candidate audio from the audio library according to the similarity calculation result includes: Sort the similarities in descending order, and determine a preset number of audios with the highest order as candidate audios; or, The audio corresponding to the similarity greater than the preset threshold is used as the candidate audio.
6. The audio recognition method according to any one of claims 1 to 3, It is characterized in that The step of determining the recognition result of the audio to be recognized based on the note correlation between the audio to be recognized and each of the candidate audios comprises: If the note correlation is greater than a preset correlation threshold, determining that the corresponding audio to be identified and the candidate audio are the same audio version, and determining that the audio to be identified is a cover audio of the candidate audio of the same audio version; If the note correlation is less than or equal to a preset correlation threshold, it is determined that the corresponding audio to be identified and the candidate audio are different audio versions, and the audio to be identified is determined to be a non-cover audio of the candidate audio of the same audio version.
7. The audio recognition method according to claim 6, It is characterized in that The extracting note features of the audio to be recognized and each of the candidate audios respectively to obtain a first note feature and a second note feature accordingly comprises: Separating a first background audio and a second background audio from the audio to be recognized and each of the candidate audios respectively; Extract the fundamental frequencies of the first background audio and the second background audio respectively, and obtain a first fundamental frequency and a second fundamental frequency respectively; According to the conversion relationship between the fundamental frequency and the note feature, a first note feature corresponding to the first fundamental frequency is determined, and a second note feature corresponding to the second fundamental frequency is determined.
8. The audio recognition method according to any one of claims 1 to 3, It is characterized in that The audio detection model is trained in the following way: Obtain a training set, the training set comprising audio sample pairs, each audio sample pair comprising a spectral feature of an original audio, a spectral feature of a cover audio corresponding to the original audio, and an actual label of whether each audio is a cover audio; Calling an initial audio detection model, and extracting features of the audio sample pair through a feature extraction module in the initial audio detection model to obtain a first sample feature and a second sample feature respectively; Performing cover detection on the first sample feature and the second sample feature by a classification module in the initial audio detection model, and determining a detection label corresponding to each audio in the audio sample pair; Determining a loss function of an initial audio detection model based on the detection label and an actual label corresponding to each audio; The initial video and audio detection model is trained according to the loss function to obtain a trained audio detection model.
9. An audio recognition device, It is characterized in that include: A first feature extraction module is configured to perform feature extraction on the audio to be recognized, and obtain the spectrum features of the audio to be recognized; the audio to be recognized includes audio of an unknown audio version; A second feature extraction module is configured to perform feature extraction on the frequency spectrum feature through an audio detection model to obtain a first cover song recognition feature of the audio to be recognized; the audio detection model is used to recognize whether the audio is a cover song audio; A similarity determination module is configured to perform feature search from an audio feature search library, obtain a second cover song recognition feature of each audio in the audio library, and determine the similarity between the audio to be recognized and each audio in the audio library according to the first cover song recognition feature and the second cover song recognition feature; A screening module is configured to determine candidate audio from the audio library according to the similarity determination result; The recognition module is configured to determine the recognition result of the audio to be recognized based on the note correlation between the audio to be recognized and each of the candidate audios, including: if the note correlation between the audio to be recognized and the corresponding candidate audio is greater than a preset correlation threshold, determining that the audio to be recognized and the corresponding candidate audio are the same audio version; performing fingerprint recognition on the audio to be recognized and the original audio of the same audio version, and if the fingerprint recognition result indicates that the audio to be recognized is not the original audio, determining that the audio to be recognized is a cover audio of the original audio; Among them, the note correlation is obtained by calculating the first note feature of the audio to be identified and the second note feature of each of the candidate audios; the first note feature and the second note feature are respectively obtained by performing vocal separation and fundamental frequency extraction on the audio to be identified and each of the candidate audios, and are determined based on the conversion relationship between the fundamental frequency and the note feature.
10. The audio recognition device according to claim 9, It is characterized in that The device also includes a search library construction module; the search library construction module is configured to execute the acquisition of each audio in the audio library, extract features from each audio, and respectively obtain target spectrum features corresponding to each audio; and is configured to execute the processing of each target spectrum feature through the audio detection model to obtain the cover feature of each audio in the audio library; based on the obtained cover feature of each audio in the audio library, construct an audio feature search library; Correspondingly, the similarity determination module includes an identification feature acquisition submodule; the identification feature acquisition submodule is configured to execute acquisition of the cover feature of each audio from the audio feature search library as the second cover recognition feature of each audio in the audio library.
11. The audio recognition device according to claim 9, It is characterized in that The first feature extraction module comprises: The framing submodule is configured to perform overlapping framing processing on the audio data to obtain framing data corresponding to the audio to be recognized; A transform submodule, configured to perform Fourier transform on the frame data to obtain transformed data corresponding to the audio to be recognized; The compression submodule is configured to perform logarithmic compression processing on the transformed data to obtain the frequency spectrum features corresponding to the audio to be recognized.
12. The audio recognition device according to any one of claims 9 to 11, It is characterized in that The second feature extraction module is specifically configured to execute: inputting the frequency spectrum feature into the audio detection model, processing the frequency spectrum feature using the feature extraction module in the audio detection model, and obtaining the first cover recognition feature of the audio to be recognized; Among them, the feature extraction module includes at least one convolutional pooling unit connected in sequence, and a fully connected layer connected to the pooling layer in the last convolutional pooling unit, and each convolutional pooling unit includes at least one convolutional layer connected in sequence, and a pooling layer connected to the last convolutional layer.
13. The audio recognition device according to any one of claims 9 to 11, It is characterized in that The similarity determination module comprises: A vector distance determination submodule, configured to calculate a vector distance between the first cover song identification feature and each of the second cover song identification features; A similarity determination submodule is configured to determine the similarity between the audio to be identified and each audio in the audio library based on the calculated vector distance; the vector distance is inversely proportional to the similarity; Accordingly, the screening module includes: The first screening submodule is configured to sort the similarities in descending order, and determine a preset number of audios with the highest order as candidate audios; or The second screening submodule is configured to select the audio corresponding to the similarity greater than a preset threshold as the candidate audio.
14. The audio recognition device according to any one of claims 9 to 11, It is characterized in that The identification module comprises: The note feature extraction submodule is configured to perform note feature extraction on the audio to be recognized and each of the candidate audios, respectively, to obtain a first note feature and a second note feature; a correlation determination submodule, configured to calculate the note correlation of the first note feature and each of the second note features; The first recognition submodule is configured to determine that the corresponding audio to be recognized and the candidate audio are the same audio version if the note correlation is greater than a preset correlation threshold, and determine that the audio to be recognized is a cover audio of the candidate audio of the same audio version; The second identification submodule is configured to execute if the note correlation is less than or equal to a preset correlation threshold, determine that the corresponding audio to be identified and the candidate audio are different audio versions, and determine that the audio to be identified is a non-cover audio of the candidate audio of the same audio version.
15. The audio recognition device according to claim 14, It is characterized in that The note feature extraction submodule includes: A separation unit, configured to separate the first background audio and the second background audio from the audio to be recognized and each of the candidate audios respectively; An extraction unit is configured to perform fundamental frequency extraction on the first background audio and the second background audio respectively, and obtain a first fundamental frequency and a second fundamental frequency accordingly; The note feature determination unit is configured to determine the first note feature corresponding to the first fundamental frequency and the second note feature corresponding to the second fundamental frequency according to the conversion relationship between the fundamental frequency and the note feature.
16. The audio recognition device according to any one of claims 9 to 11, It is characterized in that The audio detection model is trained in the following way: Obtain a training set, the training set comprising audio sample pairs, each audio sample pair comprising a spectral feature of an original audio, a spectral feature of a cover audio corresponding to the original audio, and an actual label of whether each audio is a cover audio; Calling an initial audio detection model, and extracting features of the audio sample pair through a feature extraction module in the initial audio detection model to obtain a first sample feature and a second sample feature respectively; Performing cover detection on the first sample feature and the second sample feature by a classification module in the initial audio detection model, and determining a detection label corresponding to each audio in the audio sample pair; Determining a loss function of an initial audio detection model based on the detection label and an actual label corresponding to each audio; The initial video and audio detection model is trained according to the loss function to obtain a trained audio detection model.
17. An electronic device, It is characterized in that include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the audio recognition method according to any one of claims 1 to 8. 18 . A computer-readable storage medium, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the audio recognition method according to claim 1 .
Citation Information
Patent Citations
Music similarity processing method
CN101552000A
Cover version identification method and device and computer storage medium
CN111445923A