Training of Audio Recognition Model, Audio Recognition Method, Device and Computer Equipment
By training an audio recognition model to jointly process song lyrics and melody features, the model improves song recognition accuracy by integrating lyrical and melodic information, addressing the limitations of traditional song recognition methods.
Patent Information
- Application Number
- CN202210866381.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-22
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-07-22
AI Technical Summary
Traditional song recognition technology fails to effectively combine lyrics comprehension and melody characteristics, resulting in low recognition accuracy, especially in the recognition of cover songs, which is prone to loss of prediction information.
By training the singing recognition model and melody recognition model, iteratively train the sample song audio in combination with lyrics text, obtain predicted phoneme sequences and melody vector representations, and fuse phoneme sequences and melody vectors to generate song recognition results.
It improves the accuracy of song recognition, can extract the semantic and melody information of the song more comprehensively, reduces the loss of information during prediction, and is suitable for the recognition of cover songs and the classification of same song groups.
Smart Images

Figure CN115240656B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of audio processing, and in particular, to a method for training an audio recognition model, an audio recognition method, an apparatus, and a computer device. Background Art
[0002] Cover song recognition technology is an important branch of the song recognition technology. With the rapid development of various short video platforms, many high-quality music works have stood out from the crowd, and many of these works are cover songs. The high-quality cover songs are widely spread and have become an important part of the music works loved by the public.
[0003] The song recognition technology in the traditional technology does not fully consider the understanding of the lyrics therein. It only extracts and classifies the melody features of the song to obtain the recognition result based on melody recognition, and through the lyrics search technology, matches the lyrics similar to the recognition result in the lyrics library to obtain the song recognition result based on lyrics recognition. Since the above two recognition tasks are independent and not mutually related, it is easy to cause the situation of prediction errors due to the loss of prediction information during recognition, which is not conducive to improving the recognition accuracy of song audio. Summary of the Invention
[0004] Based on this, in view of the above technical problems, it is necessary to provide a method for training an audio recognition model, an audio recognition method, an apparatus, a computer device, a computer-readable storage medium, and a computer program product that can improve the audio recognition accuracy.
[0005] In a first aspect, the present application provides a method for training an audio recognition model, where the audio recognition model includes a singing voice recognition model and a melody recognition model, and the method includes:
[0006] Obtain training sample data; the training sample data includes each sample song audio and the lyric text corresponding to the sample song audio;
[0007] Input the human voice audio in the sample song audio into the singing voice recognition model to obtain a predicted phoneme sequence, and input the sample song audio into the melody recognition model to obtain a melody vector representation corresponding to the sample song audio;
[0008] Determine the actual phoneme sequence of the sample song audio based on the lyrics text, and iteratively train the singing voice recognition model according to the difference between the actual phoneme sequence and the predicted phoneme sequence. Also, iteratively train the melody recognition model according to the difference between the melody vector representation and the prototype vector representation until the two iterative trainings meet the preset training end conditions to obtain a trained audio recognition model. Among them, different prototype vector representations are used to represent different melody categories. The trained audio recognition model is used to output the song recognition result corresponding to the song audio to be recognized.
[0009] In one embodiment, the iteratively training the melody recognition model according to the difference between the melody vector representation and the prototype vector representation includes:
[0010] Cluster each of the melody vector representations to obtain at least one category of post-clustering vector representations.
[0011] Based on the distance relationship between the post-clustering vector representations of the same category, determine the prototype vector representation corresponding to each category. Among them, the distance between the prototype vector representation and each post-clustering vector representation of the category corresponding to the prototype vector representation satisfies a preset condition.
[0012] Iteratively train the melody recognition model based on the distance between the melody vector representation and the prototype vector representation corresponding to each category.
[0013] In one embodiment, the inputting the human voice audio in the sample song audio into the singing voice recognition model to obtain a predicted phoneme sequence includes:
[0014] Input the sample song audio into a pre-trained voice separation model to obtain the human voice audio in the sample song audio.
[0015] Extract the first spectral feature from the human voice audio and input the first spectral feature into the singing voice recognition model to obtain the predicted phoneme sequence corresponding to the human voice audio.
[0016] In one embodiment, the singing voice recognition model includes a first convolutional layer, a first fully connected layer, and a first classification layer. The inputting the first spectral feature into the singing voice recognition model to obtain the predicted phoneme sequence corresponding to the human voice audio includes:
[0017] Input the first spectral feature into the first convolutional layer so that the first convolutional layer extracts the first spectral convolutional feature corresponding to the first spectral feature.
[0018] Input the first spectral convolution feature into the first fully connected layer, so that the first fully connected layer transforms the dimension type of the first spectral convolution feature from a spatial feature to a temporal feature, obtaining a spectral convolution feature with reduced dimensions;
[0019] Input the spectral convolution feature with reduced dimensions into the first classification layer, so that the first classification layer performs classification processing on the spectral convolution feature with reduced dimensions, obtaining a predicted phoneme sequence corresponding to the human voice audio.
[0020] In one embodiment, the melody recognition model includes a second convolutional layer and a second fully connected layer. The step of inputting the sample song audio into the melody recognition model to obtain a melody vector representation corresponding to the sample song audio includes:
[0021] Extract a second spectral feature from the sample song audio, and input the second spectral feature into the second convolutional layer, so that the second convolutional layer extracts a second spectral convolution feature corresponding to the second spectral feature;
[0022] Input the second spectral convolution feature into the second fully connected layer, so that the second fully connected layer transforms the dimension type of the second spectral convolution feature from a spatial feature to a temporal feature, obtaining a melody vector representation corresponding to the sample song audio.
[0023] In a second aspect, the present application further provides an audio recognition method, and the method includes:
[0024] Obtain a trained audio recognition model; the trained audio recognition model is trained according to the training method as described above; the trained audio recognition model includes a trained singing voice recognition model and a trained melody recognition model;
[0025] Input the human voice audio in the song audio to be recognized into the trained singing voice recognition model, obtaining a predicted phoneme sequence corresponding to the song audio to be recognized, and input the song audio to be recognized into the trained melody recognition model, obtaining a melody vector representation corresponding to the song audio to be recognized;
[0026] Output a song recognition result corresponding to the song audio to be recognized according to the predicted phoneme sequence and the melody vector representation.
[0027] In one embodiment, the step of outputting a song recognition result corresponding to the song audio to be recognized according to the predicted phoneme sequence and the melody vector representation includes:
[0028] Obtain a phoneme sequence vector representation corresponding to the predicted phoneme sequence;
[0029] Fuse the phoneme sequence vector representation and the melody vector representation to obtain a fused vector representation;
[0030] Generate a song recognition result corresponding to the to-be-recognized song audio according to the fused vector representation.
[0031] In one embodiment, the fusing the phoneme sequence vector representation and the melody vector representation to obtain a fused vector representation includes:
[0032] Obtain the phoneme weight corresponding to the phoneme sequence vector representation and the melody weight corresponding to the melody vector representation respectively;
[0033] Perform weighted summation on the phoneme sequence vector representation and the melody vector representation according to the phoneme weight and the melody weight to obtain the fused vector representation;
[0034] In one embodiment, the fusing the phoneme sequence vector representation and the melody vector representation to obtain a fused vector representation includes:
[0035] Concatenate the phoneme sequence vector representation and the melody vector representation to obtain a concatenated vector representation;
[0036] Use the concatenated vector representation as the fused vector representation.
[0037] In one embodiment, the generating a song recognition result corresponding to the to-be-recognized song audio according to the fused vector representation includes:
[0038] Obtain the candidate song vector representations corresponding to the candidate song audios;
[0039] Determine the similarity between the song vector representation corresponding to the to-be-recognized song audio and each of the candidate song vector representations;
[0040] Use the audio identification information of the candidate song audio corresponding to the candidate song vector representation with the highest similarity as the song recognition result corresponding to the to-be-recognized song audio.
[0041] In a third aspect, the present application also provides a computer device. The computer device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the steps of the above method are implemented.
[0042] In a fourth aspect, the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, and when the computer program is executed by the processor, the steps of the above method are implemented.
[0043] Fifth aspect, the present application also provides a computer program product. The computer program product includes a computer program which, when executed by a processor, implements the steps of the above method.
[0044] For the training of the above audio recognition model, audio recognition method, device, computer device, storage medium and computer program product, training sample data is obtained; the training sample data includes each sample song audio and the corresponding lyric text; the human voice audio in the sample song audio is input into the singing voice recognition model to obtain a predicted phoneme sequence, and, the sample song audio is input into the melody recognition model to obtain a melody vector representation corresponding to the sample song audio; the actual phoneme sequence of the sample song audio is determined based on the lyric text, and the singing voice recognition model is iteratively trained according to the difference between the actual phoneme sequence and the predicted phoneme sequence, and, the melody recognition model is iteratively trained according to the difference between the melody vector representation and the prototype vector representation until the preset training end condition is met to obtain a trained audio recognition model; wherein, the prototype vector representation is determined by clustering each melody vector representation and based on the distance relationship between the clustered melody vector representations; different prototype vector representations are used to represent different melody categories; the trained audio recognition model is used to output the song recognition result corresponding to the song audio to be recognized; thus, by jointly optimizing the singing voice recognition model and the melody recognition model using the same sample song audio, the melody recognition model in the trained audio recognition model can effectively learn the melody vector representation for accurately representing the melody features of the song audio, and the singing voice recognition model in the trained audio recognition model can effectively learn the phoneme sequence for accurately representing the semantic features of the song lyrics of the song audio, and the song information corresponding to the song audio is recognized by using the fusion vector representation between the melody vector representation and the phoneme sequence corresponding to the song audio, so that the audio recognition model can not only extract the semantic information in the song, but also extract the melody information in the song, prevent information loss during prediction, more fully explore the similar information in the song, and thus achieve a better song recognition effect. Description of the Drawings
[0045] Figure 1 It is a schematic flow chart of a method for training an audio recognition model in an embodiment;
[0046] Figure 2 It is a schematic diagram of a phoneme posterior probability matrix in an embodiment;
[0047] Figure 3 It is a schematic diagram of the vector representation after clustering in an embodiment;
[0048] Figure 4 It is a schematic flow chart of a method for training an audio recognition model in an embodiment;
[0049] Figure 5 It is an application environment diagram of an audio recognition method in an embodiment;
[0050] Figure 6 It is a schematic flowchart of an audio recognition method in an embodiment;
[0051] Figure 7 It is a schematic flowchart of an audio recognition method in an embodiment;
[0052] Figure 8 It is a model block diagram of a training method for an audio recognition model in an embodiment;
[0053] Figure 9 It is an internal structure diagram of a computer device in an embodiment. Specific implementation manners
[0054] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0055] In an embodiment, as Figure 1 shown, a training method for an audio recognition model is provided. The audio recognition model includes a singing voice recognition model and a melody recognition model. Taking the application of this method to an electronic device as an example, the method includes the following steps:
[0056] Step S102, obtain training sample data; the training sample data includes each sample song audio and the corresponding lyrics text of each sample song audio.
[0057] Among them, the training sample data may refer to the sample data used to train the audio recognition model.
[0058] In practical applications, the training sample data may include multiple sample song audios. The sample song audio may be the music audio corresponding to various songs. Among them, each sample song audio has a corresponding lyrics text.
[0059] In specific implementation, during the process of training the audio recognition model by the electronic device, the electronic device may obtain the music audio corresponding to each song and the lyrics text of each song as the training sample data for the audio recognition model.
[0060] Step S104, input the human voice audio in the sample song audio into the singing voice recognition model to obtain a predicted phoneme sequence, and input the sample song audio into the melody recognition model to obtain the melody vector representation corresponding to the sample song audio.
[0061] In a specific implementation, after the electronic device obtains each sample song audio for training the audio recognition model, the electronic device can input the human voice audio in the sample song audio into the singing voice recognition model to obtain a predicted phoneme sequence.
[0062] Specifically, the electronic device can input the sample song audio into a pre-trained music human voice separation model (for example, the Spleeter separation network) so that the pre-trained music human voice separation model separates the human voice audio from the sample song audio. Then, the electronic device inputs the spectral features corresponding to the human voice audio in the sample song audio into the singing voice recognition model so that the singing voice recognition model outputs a predicted phoneme sequence corresponding to the human voice audio.
[0063] Among them, the predicted phoneme sequence can refer to the prediction result of the pronunciation phonemes corresponding to each pronunciation time period in the human voice audio.
[0064] The electronic device can input the sample song audio into the melody recognition model to obtain a melody vector representation corresponding to the sample song audio. Specifically, the electronic device can input the spectral features corresponding to the accompaniment audio in the sample song audio into the melody recognition model so that the melody recognition model outputs a melody vector representation corresponding to the sample song audio.
[0065] Step S106, determine the actual phoneme sequence of the sample song audio based on the lyric text, and iteratively train the singing voice recognition model according to the difference between the actual phoneme sequence and the predicted phoneme sequence, and iteratively train the melody recognition model according to the difference between the melody vector representation and the prototype vector representation until the preset training end condition is met to obtain a trained audio recognition model for generating a song vector representation corresponding to the song audio to be recognized.
[0066] Among them, the prototype vector representation is determined by clustering each melody vector representation and based on the distance relationship between the clustered melody vector representations; different prototype vector representations are used to represent different melody categories.
[0067] In a specific implementation, in the training stage, after the electronic device obtains the predicted phoneme sequence corresponding to the human voice audio in the sample song audio, the electronic device can determine the actual phoneme sequence of the sample song audio according to the lyric text, and use the actual phoneme sequence to represent the annotation result of the pronunciation phonemes corresponding to each pronunciation time period in the human voice audio as the classification training target of the singing voice recognition model. The electronic device can use the actual phoneme sequence as a supervision signal for training the singing voice recognition model, that is, iteratively train the singing voice recognition model according to the difference between the actual phoneme sequence and the predicted phoneme sequence.
[0068] Specifically, the electronic device can adopt a preset loss function (e.g., CTC (Connectionist Temporal Classification) loss function), determine the loss function value of the singing recognition model by using the actual phoneme sequence and the predicted phoneme sequence, and update the network parameters (e.g., weights and biases) in the singing recognition model by using the loss function value until the trained singing recognition model meets the preset training end condition, and obtain the trained singing recognition model.
[0069] For example, taking the lyric text "Hello" as an example, the actual phoneme sequence corresponding to this lyric text can be expressed as n i h a o, and the predicted phoneme sequence output by the singing recognition model (i.e., the phoneme posterior probability matrix in the training stage, as Figure 2 shown) is used to represent the probability that the singing recognition model calculates that each frame belongs to the entire phoneme set. The electronic device uses the actual phoneme sequence as the supervision signal for training the singing recognition model, so that the trained singing recognition model can find the best decoding path corresponding to the phonemes contained in the input human voice audio, and then obtain the target phoneme sequence that can effectively represent the phoneme information in the human voice audio.
[0070] After the electronic device obtains the melody vector representation corresponding to the sample song audio, it clusters the melody vector representations corresponding to each sample song audio, and determines the prototype vector representation based on the distance relationship between the clustered melody vector representations; among them, different prototype vector representations have different melody types.
[0071] The electronic device can perform iterative training on the melody recognition model according to the difference (e.g., vector distance) between the melody vector representation and the prototype vector representation until the preset training target is met to obtain the trained melody recognition model, so that the trained melody recognition model can learn the target melody vector representations corresponding to various sample song audios.
[0072] In practical applications, the training target can be that the closer the melody vector representations of the same category are to the category center, the better, and the farther the melody vector representations that do not belong to the current category are from the category center, the better. It should be noted that the training process of the melody recognition model will be introduced in detail later, and will not be elaborated here for the time being.
[0073] The electronic device can splice the output ends of the two branch modules of the trained singing recognition model and the trained melody recognition model with a fully connected convergence layer to obtain the trained audio recognition model, and the trained audio recognition model can be used to output the song recognition result corresponding to the song audio to be recognized.
[0074] Specifically, the electronic device can input the human voice audio in the song audio to be recognized (such as cover audio) into the singing voice recognition model in the trained audio recognition model to generate a predicted phoneme sequence of the song audio to be recognized, and input the song audio to be recognized into the melody recognition model in the trained audio recognition model to generate a melody vector representation corresponding to the song audio to be recognized; then, the electronic device fuses the vector representation corresponding to the predicted phoneme sequence (phoneme sequence embedding, phoneme sequence representation) and the melody vector representation corresponding to the song audio to be recognized (melody embedding, melody representation), and based on the fused vector representation, outputs the song recognition result corresponding to the song audio to be recognized.
[0075] Taking the scene of identifying songs by listening as an example, since there are a large number of cover version songs in the song library and a large number of new cover songs are added every day, and the fingerprint library for identifying songs by listening cannot cover all of them, using the audio recognition model trained by the above audio recognition model training method for cover recognition can make up for this part of the gap and improve the recall rate of identifying songs by listening.
[0076] In addition, the audio recognition model trained by the above audio recognition model training method is also used to generate song groups of the same song, which can effectively organize the song library information and efficiently classify songs of the same song. There are a large number of user cover works that have not been tagged with songs. Combining the technology of identifying songs by listening and cover recognition can tag these user works, which is convenient for subsequent work analysis.
[0077] In the training method of the above-mentioned audio recognition model, training sample data is obtained; the training sample data includes each sample song audio and the corresponding lyrics text; the human voice audio in the sample song audio is input into the singing recognition model to obtain a predicted phoneme sequence, and the sample song audio is input into the melody recognition model to obtain a melody vector representation corresponding to the sample song audio; the actual phoneme sequence of the sample song audio is determined based on the lyrics text, and the singing recognition model is iteratively trained according to the difference between the actual phoneme sequence and the predicted phoneme sequence, and the melody recognition model is iteratively trained according to the difference between the melody vector representation and the prototype vector representation, until the preset training end condition is met to obtain a trained audio recognition model; wherein the prototype vector representation is obtained by clustering each melody vector representation and determining it based on the distance relationship between each melody vector representation after clustering; different prototype vector representations are used to represent different Melody category; the trained audio recognition model is used to output the song recognition result corresponding to the song audio to be recognized; in this way, by using the same sample song audio to jointly optimize the singing recognition model and the melody recognition model, the melody recognition model in the trained audio recognition model can effectively learn the melody vector representation used to accurately characterize the melody features of the song audio, and the singing recognition model in the trained audio recognition model can effectively learn the phoneme sequence used to accurately characterize the semantic features of the lyrics of the song audio, and identify the song information corresponding to the song audio by utilizing the melody vector representation corresponding to the song audio and the fusion vector representation between the phoneme sequence, so that the audio recognition model can extract both the semantic information and the melody information in the song, prevent information loss during prediction, and more fully mine similar information in the song, thereby achieving better song recognition effect.
[0078] In another embodiment, the melody recognition model is iteratively trained based on the difference between the melody vector representation and the prototype vector representation, including: clustering each melody vector representation to obtain a clustered vector representation of at least one category; determining the prototype vector representation corresponding to each category based on the distance relationship between the clustered vector representations of the same category; wherein the distance between the prototype vector representation and the clustered vector representation within the corresponding category satisfies a preset condition; and iteratively training the melody recognition model based on the distance between the melody vector representation and the prototype vector representation corresponding to each category.
[0079] In a specific implementation, during the process of iteratively training the melody recognition model based on the difference between the melody vector representation and the prototype vector representation, the electronic device can cluster each melody vector representation to obtain at least one category of post-clustering vector representations (samples); then, based on the distance relationship between the post-clustering vector representations of the same category, determine the prototype vector representation corresponding to each category. Specifically, for the post-clustering vector representation of any category, the electronic device can calculate the mean of the post-clustering vector representation of this category and determine the prototype vector representation (i.e., the category center) corresponding to this category according to the mean of the post-clustering vector representation.
[0080] The electronic device can iteratively train the melody recognition model according to the difference (such as vector distance) between the melody vector representation and the prototype vector representation until the preset training objective is met to obtain a trained melody recognition model, so that the trained melody recognition model can learn the target melody vector representations corresponding to various sample song audios.
[0081] In practical applications, the training objective can be that the closer the melody vector representations of the same category are to the category center, the better, and the farther the melody vector representations that do not belong to the current category are from the category center, the better.
[0082] For example, please refer to Figure 3 , Figure 3 Exemplarily shows a schematic diagram of the post-clustering vector representations of three categories. Among them, the three centers of the circles represent the three category centers, which can be determined by the mean of the samples embedding (post-clustering vector representations) within the same category.
[0083] Among them, the category center can be expressed as:
[0084]
[0085] Among them, x j,k represents the k-th sample within category j, and c j is the center of sample category j, that is, the prototype center.
[0086] The electronic device obtains the Euclidean distance d k from each sample x j to the category center c k,j and uses the Euclidean distance between each melody vector representation and the prototype vector representation to characterize the difference between each melody vector representation and the prototype vector representation.
[0087] Among them, the Euclidean distance d k from sample x j to the category center c k,j can be expressed as:
[0088] d k,j= ||x k -c j || 2 , x k ∈Q;
[0089] In order to make the melody vector representation output by the trained melody recognition model close to the prototype vector representation of the same melody category in the vector space and far from the prototype vector representations of different melody categories in the vector space, the following loss function can be adopted:
[0090]
[0091] wherein,
[0092]
[0093] In the technical solution of this embodiment, by clustering each melody vector representation, at least one category of clustered vector representations is obtained, and based on the distance relationship between the clustered vector representations of the same category, the prototype vector representation corresponding to each category is determined; wherein, the distance between the prototype vector representation and the clustered vector representation within the corresponding category satisfies a preset condition; then, based on the distance between the melody vector representation and the prototype vector representations corresponding to each category, the melody recognition model is iteratively trained, so that the melody vector representation output by the trained melody recognition model is close to the prototype vector representation of the same melody category in the vector space and at the same time far from the prototype vector representations corresponding to other melody categories in this vector space, thereby enabling the trained melody recognition model to effectively learn the melody vector representation for characterizing the melody features of each sample song.
[0094] In another embodiment, the human voice audio in the sample song audio is input into the singing voice recognition model to obtain a predicted phoneme sequence, including: inputting the sample song audio into a pre-trained human voice separation model to obtain the human voice audio in the sample song audio; extracting the first spectral features in the human voice audio and inputting the first spectral features into the singing voice recognition model to obtain the predicted phoneme sequence corresponding to the human voice audio.
[0095] In a specific implementation, when the electronic device inputs the human voice audio in the sample song audio into the singing voice recognition model to obtain a predicted phoneme sequence, the electronic device can input the sample song audio into a pre-trained music human voice separation model (for example, the Spleeter separation network) so that the pre-trained music human voice separation model separates the human voice audio from the sample song audio.
[0096] After the electronic device separates the human voice audio from the sample song audio, the electronic device can extract the Mel Frequency Cepstral Coefficients (MFCCs) corresponding to the human voice audio, and further extract the first spectral feature (i.e., acoustic feature) corresponding to the human voice audio. The electronic device inputs the extracted first spectral feature into the singing voice recognition model to obtain the predicted phoneme sequence corresponding to the human voice audio.
[0097] In another embodiment, the singing voice recognition model includes a first convolutional layer, a first fully connected layer, and a first classification layer. Inputting the first spectral feature into the singing voice recognition model to obtain the predicted phoneme sequence corresponding to the human voice audio includes: inputting the first spectral feature into the first convolutional layer so that the first convolutional layer extracts the first spectral convolutional feature corresponding to the first spectral feature; inputting the first spectral convolutional feature into the first fully connected layer so that the first fully connected layer transforms the dimension type of the first spectral convolutional feature from a spatial feature to a temporal feature, obtaining the reduced-dimensional spectral convolutional feature; and inputting the reduced-dimensional spectral convolutional feature into the first classification layer so that the first classification layer performs classification processing on the reduced-dimensional spectral convolutional feature to obtain the predicted phoneme sequence corresponding to the human voice audio.
[0098] Among them, the first convolutional layer can be constructed by a convolutional neural network.
[0099] The technical solution of this embodiment is to input the first spectral feature into the first convolutional layer so that the first convolutional layer extracts the first spectral convolutional feature corresponding to the first spectral feature; input the first spectral convolutional feature into the first fully connected layer so that the first fully connected layer transforms the dimension type of the first spectral convolutional feature from a spatial feature to a temporal feature, obtaining the reduced-dimensional spectral convolutional feature; then, input the reduced-dimensional spectral convolutional feature into the first classification layer so that the first classification layer performs classification processing on the reduced-dimensional spectral convolutional feature to obtain the predicted phoneme sequence corresponding to the human voice audio.
[0100] In another embodiment, the melody recognition model includes a second convolutional layer and a second fully connected layer. Inputting the sample song audio into the melody recognition model to obtain the melody vector representation corresponding to the sample song audio includes: extracting the second spectral feature from the sample song audio and inputting the second spectral feature into the second convolutional layer so that the second convolutional layer extracts the second spectral convolutional feature corresponding to the second spectral feature; inputting the second spectral convolutional feature into the second fully connected layer so that the second fully connected layer transforms the dimension type of the second spectral convolutional feature from a spatial feature to a temporal feature, obtaining the melody vector representation corresponding to the sample song audio.
[0101] Among them, the second convolutional layer can be constructed by a convolutional neural network.
[0102] In the technical solution of this embodiment, by extracting the second spectral feature in the sample song audio and inputting the second spectral feature into the second convolutional layer, the second convolutional layer can extract the second spectral convolutional feature corresponding to the second spectral feature; inputting the second spectral convolutional feature into the second fully connected layer, so that the second fully connected layer can transform the dimension type of the second spectral convolutional feature from a spatial feature to a temporal feature, and obtain the melody vector representation corresponding to the sample song audio; in this way, it is possible to efficiently and accurately extract the melody vector representation for characterizing the melody information of the sample song audio. At the same time, transforming the dimension type of the second spectral convolutional feature from a spatial feature to a temporal feature makes the melody vector representation have the same dimension as the predicted phoneme sequence corresponding to the human voice audio, which is convenient for subsequent processing using the melody vector representation and the predicted phoneme sequence recognition of the human voice audio, and effectively reduces the amount of data processing of the recognition model in the subsequent process.
[0103] In another embodiment, as Figure 4 shown, a training method for an audio recognition model is provided. The audio recognition model includes a singing voice recognition model and a melody recognition model. The singing voice recognition model includes a first convolutional layer, a first fully connected layer, and a first classification layer. The melody recognition model includes a second convolutional layer and a second fully connected layer, and includes the following steps:
[0104] Step S410, obtaining training sample data; the training sample data includes each sample song audio and the corresponding lyric text.
[0105] Step S420, inputting the sample song audio into a pre-trained voice separation model to obtain the human voice audio in the sample song audio.
[0106] Step S430, extracting the first spectral feature in the human voice audio and inputting the first spectral feature into the first convolutional layer, so that the first convolutional layer extracts the first spectral convolutional feature corresponding to the first spectral feature.
[0107] Step S440, inputting the first spectral convolutional feature into the first fully connected layer, so that the first fully connected layer transforms the dimension type of the first spectral convolutional feature from a spatial feature to a temporal feature, and obtains the dimension-reduced spectral convolutional feature.
[0108] Step S450, inputting the dimension-reduced spectral convolutional feature into the first classification layer, so that the first classification layer performs classification processing on the dimension-reduced spectral convolutional feature to obtain the predicted phoneme sequence corresponding to the human voice audio.
[0109] Step S460, extracting the second spectral feature in the sample song audio and inputting the second spectral feature into the second convolutional layer, so that the second convolutional layer extracts the second spectral convolutional feature corresponding to the second spectral feature.
[0110] Step S470: Input the second spectral convolution feature into the second fully connected layer, so that the second fully connected layer transforms the dimension type of the second spectral convolution feature from spatial feature to temporal feature, and obtain the melody vector representation corresponding to the sample song audio.
[0111] Step S480: Determine the actual phoneme sequence of the sample song audio based on the lyric text, and iteratively train the singing voice recognition model according to the difference between the actual phoneme sequence and the predicted phoneme sequence, and iteratively train the melody recognition model according to the difference between the melody vector representation and the prototype vector representation, until the preset training end condition is met to obtain the trained audio recognition model.
[0112] It should be noted that the specific limitations of the above steps can refer to the specific limitations of a training method for an audio recognition model in the above text.
[0113] The audio recognition method provided by the embodiments of this application can be applied to an Figure 5 application environment as shown. Among them, the terminal 502 communicates with the electronic device 504 through the network. The data storage system can store the data that the server 504 needs to process. The data storage system can be integrated on the electronic device 504, or can be placed in the cloud or other network servers. In practical applications, the song recognition function of the terminal 502 can record the received sound signal in real time to generate the music audio to be recognized. Then, the terminal 502 sends the music audio to be recognized to the electronic device 504. The electronic device 504 obtains the trained audio recognition model; the trained audio recognition model is trained according to the training method as described above; the trained audio recognition model includes the trained singing voice recognition model and the trained melody recognition model; the electronic device 504 inputs the human voice audio in the song audio to be recognized into the trained singing voice recognition model to obtain the predicted phoneme sequence corresponding to the song audio to be recognized, and inputs the song audio to be recognized into the trained melody recognition model to obtain the melody vector representation corresponding to the song audio to be recognized; the electronic device 504 outputs the song recognition result corresponding to the song audio to be recognized according to the predicted phoneme sequence and the melody vector representation. Among them, the terminal 502 can be but is not limited to various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The electronic device 504 can be implemented by an independent server or a server cluster composed of multiple servers.
[0114] In one embodiment, as Figure 6 shown, an audio recognition method is provided. Taking the method applied to the Figure 5 electronic device in it as an example, it includes the following steps:
[0115] Step S602, obtain a trained audio recognition model.
[0116] Step S604, input the human voice audio in the song audio to be recognized into the trained singing voice recognition model to obtain the predicted phoneme sequence corresponding to the song audio to be recognized, and input the song audio to be recognized into the trained melody recognition model to obtain the melody vector representation corresponding to the song audio to be recognized.
[0117] In a specific implementation, the electronic device obtains a trained audio recognition model, and the trained audio recognition model includes a trained singing voice recognition model and a trained melody recognition model. It should be noted that the training method adopted can refer to the specific limitations on the training method of the audio recognition model above, and will not be elaborated here.
[0118] Specifically, the electronic device can input the human voice audio in the song audio to be recognized (such as, cover audio) into the singing voice recognition model in the trained audio recognition model to generate the predicted phoneme sequence of the song audio to be recognized, and input the song audio to be recognized into the melody recognition model in the trained audio recognition model to generate the melody vector representation corresponding to the song audio to be recognized.
[0119] Step S606, output the song recognition result corresponding to the song audio to be recognized according to the predicted phoneme sequence and the melody vector representation.
[0120] In a specific implementation, the electronic device can fuse the predicted phoneme sequence and the melody vector representation to obtain a fused vector representation; the electronic device can determine the song recognition result corresponding to the song audio to be recognized based on the fused vector representation. Specifically, the electronic device can input the fused vector representation into a pre-trained song recognition network, and classify the fused vector through the pre-trained song recognition network, and determine the song recognition result corresponding to the song audio to be recognized based on the classification result.
[0121] Of course, the electronic device can also adopt a method of performing similarity matching between the fused vector and the song vector representations corresponding to each candidate song to determine the song recognition result corresponding to the song audio to be recognized.
[0122] In the above audio recognition method, by obtaining a trained audio recognition model and inputting the human voice audio in the song audio to be recognized into the trained singing voice recognition model, a predicted phoneme sequence corresponding to the song audio to be recognized is obtained, and, by inputting the song audio to be recognized into the trained melody recognition model, a melody vector representation corresponding to the song audio to be recognized is obtained; then, according to the predicted phoneme sequence and the melody vector representation, a song recognition result corresponding to the song audio to be recognized is output, so that the audio recognition model can not only understand the semantic information in the song, but also understand the melody information in the song, more fully explore the similar information in the song, and thus achieve a better recognition effect.
[0123] In another embodiment, outputting a song recognition result corresponding to the song audio to be recognized according to the predicted phoneme sequence and the melody vector representation includes: obtaining a phoneme sequence vector representation corresponding to the predicted phoneme sequence; fusing the phoneme sequence vector representation and the melody vector representation to obtain a fused vector representation; and generating a song recognition result corresponding to the song audio to be recognized according to the fused vector representation.
[0124] In a specific implementation, the electronic device fuses the vector representation (phoneme sequence embedding, phoneme sequence characterization) corresponding to the predicted phoneme sequence and the melody vector representation (melody embedding, melody characterization) corresponding to the song audio to be recognized through a fully connected layer in the trained singing voice recognition model, and uses the fused feature vector as the song vector representation corresponding to the song audio to be recognized.
[0125] Wherein, in the process of the electronic device fusing the phoneme sequence vector representation and the melody vector representation to obtain a fused vector representation, it can respectively obtain a phoneme weight corresponding to the phoneme sequence vector representation and a melody weight corresponding to the melody vector representation; and perform weighted summation on the phoneme sequence vector representation and the melody vector representation according to the phoneme weight and the melody weight to obtain a fused vector representation.
[0126] Of course, in the process of the electronic device fusing the phoneme sequence vector representation and the melody vector representation to obtain a fused vector representation, it can also splice the phoneme sequence vector representation and the melody vector representation to obtain a spliced vector representation; and use the spliced vector representation as the fused vector representation.
[0127] Then, the electronic device can generate a song recognition result corresponding to the song audio to be recognized according to the fused vector representation.
[0128] Specifically, the electronic device can respectively calculate the similarity between the song vector representation corresponding to the song audio to be recognized and each candidate song vector representation, and use the song corresponding to the candidate song vector representation with the highest similarity as the song recognition result corresponding to the song audio to be recognized.
[0129] In the technical solution of this embodiment, by obtaining the phoneme sequence vector representation corresponding to the predicted phoneme sequence; fusing the phoneme sequence vector representation and the melody vector representation to obtain the fused vector representation, and generating the song recognition result corresponding to the song audio to be recognized according to the fused vector representation; it can be ensured that the audio recognition model can not only understand the semantic information in the song, but also understand the melody information in the song, more fully explore the similar information in the song, and thus achieve a better recognition effect.
[0130] In another embodiment, generating the song recognition result corresponding to the song audio to be recognized according to the fused vector representation includes: obtaining the candidate song vector representations corresponding to the candidate song audios; determining the similarity between the song vector representation corresponding to the song audio to be recognized and each of the candidate song vector representations; and using the audio identification information of the candidate song audio corresponding to the candidate song vector representation with the highest similarity as the song recognition result corresponding to the song audio to be recognized.
[0131] Among them, the audio identification information may refer to the song title.
[0132] In specific implementation, after the electronic device determines the song vector representation corresponding to the song audio to be recognized, the electronic device obtains the candidate song vector representations corresponding to the candidate song audios. It should be noted that the electronic device can use the method for determining the song vector representation corresponding to the song audio to be recognized in the above text to obtain the candidate song vector representations corresponding to the candidate song audios, which will not be elaborated here. Then, the electronic device calculates the similarity (such as cosine similarity, vector distance, etc.) between the song vector representation corresponding to the song audio to be recognized and each of the candidate song vector representations respectively. Then, the electronic device uses the audio identification information of the candidate song audio corresponding to the candidate song vector representation with the highest similarity as the song recognition result for the song audio to be recognized.
[0133] For example, if the similarity between the song vector representation corresponding to the cover audio and the song vector representation of song a is 80%, and the similarity between the song vector representation corresponding to the cover audio and the song vector representation of song b is 60%, then the song title of song a is used as the song recognition result for this cover audio.
[0134] The technical solution of this embodiment is to obtain the candidate song vector representations corresponding to each of the candidate song audios, and determine the similarity between the song vector representation corresponding to the song audio to be recognized and each of the candidate song vector representations; by using the audio identification information of the candidate song audio corresponding to the candidate song vector representation with the highest similarity as the song recognition result of the candidate song audio, thereby improving the song recognition accuracy of the candidate song audio and facilitating subsequent tagging operations on the song audio to be recognized using the audio identification information.
[0135] In another embodiment, as Figure 7 shown, an audio recognition method is provided. Taking the terminal in Figure 5 as an example, the method includes the following steps:
[0136] Step S710: Obtain a trained audio recognition model; the trained audio recognition model includes a trained singing voice recognition model and a trained melody recognition model.
[0137] Step S720: Input the human voice audio in the song audio to be recognized into the trained singing voice recognition model to obtain the predicted phoneme sequence corresponding to the song audio to be recognized, and input the song audio to be recognized into the trained melody recognition model to obtain the melody vector representation corresponding to the song audio to be recognized.
[0138] Step S730: Obtain the phoneme sequence vector representation corresponding to the predicted phoneme sequence.
[0139] Step S740: Fuse the phoneme sequence vector representation and the melody vector representation to obtain the fused vector representation.
[0140] Step S750: Obtain the candidate song vector representations corresponding to each candidate song audio.
[0141] Step S760: Determine the similarity between the song vector representation corresponding to the song audio to be recognized and each candidate song vector representation.
[0142] Step S770: Use the audio identification information of the candidate song audio corresponding to the candidate song vector representation with the highest similarity as the song recognition result corresponding to the song audio to be recognized.
[0143] It should be noted that the specific limitations of the above steps can be referred to the specific limitations of an audio recognition method above.
[0144] For the convenience of those skilled in the art, Figure 8 exemplarily provides a model block diagram of a training method for an audio recognition model. As Figure 8As shown, during the process of training an audio recognition model, the electronic device can obtain the music audio corresponding to each song and the lyric text of each song as the training sample data for the audio recognition model.
[0145] Then, the electronic device can extract the first spectral features of the human voice audio in the sample song audio and input the first spectral features into the singing voice recognition model. Specifically, the electronic device can input the first spectral features into the first convolutional network 810 in the singing voice recognition model to obtain the first spectral convolutional features; then, the electronic device inputs the first spectral convolutional features into the first fully connected layer 820 in the singing voice recognition model to obtain the spectral convolutional features after dimensionality reduction; then, the electronic device inputs the spectral convolutional features after dimensionality reduction into the classification layer 830 (Softmax layer) to obtain the predicted phoneme sequence corresponding to the human voice audio. Then, the electronic device can use a preset loss function (such as, the CTC (Connectionist temporal classification) loss function) to determine the loss function value of the singing voice recognition model using the actual phoneme sequence and the predicted phoneme sequence, and use the loss function value to update the network parameters (such as, weights and biases) in the singing voice recognition model until the trained singing voice recognition model meets the preset training end condition to obtain the trained singing voice recognition model.
[0146] In addition, the electronic device extracts the second spectral features in the sample song audio and inputs the second spectral features into the second convolutional network 840 to obtain the second spectral convolutional features; then, the electronic device inputs the second spectral convolutional features into the second fully connected layer 850 to obtain the melody vector representation (embedding) corresponding to the sample song audio. Then, the electronic device can perform iterative training on the melody recognition model according to the difference (such as, vector distance) between the melody vector representation and the prototype vector representation until the preset training objective is met to obtain the trained melody recognition model, so that the trained melody recognition model can learn the target melody vector representations corresponding to various sample song audios.
[0147] The electronic device can input the human voice audio in the song audio to be recognized into the trained singing voice recognition model to obtain the predicted phoneme sequence corresponding to the song audio to be recognized, and input the song audio to be recognized into the trained melody recognition model to obtain the melody vector representation corresponding to the song audio to be recognized. Then, through the softmax layer 860, using the predicted phoneme sequence and the melody vector representation corresponding to the song audio to be recognized, the song vector representation corresponding to the song audio to be recognized is determined. And by calculating the similarity (such as cosine similarity, vector distance, etc.) between the song vector representation corresponding to the song audio to be recognized and each candidate song vector representation. Then, the audio identification information of the candidate song audio corresponding to the candidate song vector representation with the highest similarity is used as the song recognition result for the song audio to be recognized.
[0148] In another embodiment, the singing voice recognition model includes a first convolutional layer, a first fully connected layer, and a first classification layer. Inputting the first spectral feature into the singing voice recognition model to obtain the predicted phoneme sequence corresponding to the human voice audio includes: inputting the first spectral feature into the first convolutional layer so that the first convolutional layer extracts the first spectral convolutional feature corresponding to the first spectral feature; inputting the first spectral convolutional feature into the first fully connected layer so that the first fully connected layer transforms the dimension type of the first spectral convolutional feature from a spatial feature to a temporal feature to obtain the dimension-reduced spectral convolutional feature; inputting the dimension-reduced spectral convolutional feature into the first classification layer so that the first classification layer performs classification processing on the dimension-reduced spectral convolutional feature to obtain the predicted phoneme sequence corresponding to the human voice audio.
[0149] When the electronic device inputs the sample song audio into the melody recognition model to obtain the melody vector representation corresponding to the sample song audio, the electronic device can extract the Mel Frequency Cepstral Coefficients (MFCC) corresponding to the sample song audio, and further extract the second spectral feature (i.e., acoustic feature) corresponding to the sample song audio.
[0150] Then, the electronic device inputs the second spectral feature into the second convolutional layer so that the second convolutional layer extracts the second spectral convolutional feature corresponding to the second spectral feature. The electronic device inputs the second spectral convolutional feature into the second fully connected layer so that the second fully connected layer transforms the dimension type of the second spectral convolutional feature from a spatial feature to a temporal feature to obtain the melody vector representation corresponding to the sample song audio.
[0151] The electronic device can use a preset loss function (e.g., the CTC (Connectionist Temporal Classification) loss function), determine the loss function value of the singing recognition model using the actual phoneme sequence and the predicted phoneme sequence, and update the network parameters (e.g., weights and biases) in the singing recognition model using the loss function value until the trained singing recognition model meets the preset training end condition, thus obtaining a trained singing recognition model. During the process of iteratively training the melody recognition model according to the difference between the melody vector representation and the prototype vector representation, the electronic device can cluster each melody vector representation to obtain at least one category of post-clustering vector representations (samples); then, based on the distance relationship between the post-clustering vector representations of the same category, determine the prototype vector representation corresponding to each category. Specifically, for the post-clustering vector representation of any category, the electronic device can determine the category center corresponding to the any category, that is, the prototype vector representation, according to the mean value of the post-clustering vector representation of the any category.
[0152] The electronic device can iteratively train the melody recognition model according to the difference (e.g., vector distance) between the melody vector representation and the prototype vector representation until the preset training objective is met to obtain a trained melody recognition model, so that the trained melody recognition model can learn the target melody vector representations corresponding to various sample song audios.
[0153] The electronic device can input the human voice audio in the song audio to be recognized (e.g., cover audio) into the singing recognition model in the trained audio recognition model to generate a predicted phoneme sequence of the song audio to be recognized, and input the song audio to be recognized into the melody recognition model in the trained audio recognition model to generate a melody vector representation corresponding to the song audio to be recognized; then, the electronic device fuses the vector representation corresponding to the predicted phoneme sequence (phoneme sequence embedding, phoneme sequence characterization) and the melody vector representation corresponding to the song audio to be recognized (melody embedding, melody characterization), and outputs the song recognition result corresponding to the song audio to be recognized based on the fused vector representation.
[0154] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless specifically stated herein, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least some of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least some of the steps or stages in other steps or other steps.
[0155] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structural diagram may be as Figure 9 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements an audio recognition method.
[0156] Those skilled in the art can understand that Figure 9 the structure shown in
[0157] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different component layout.
[0158] In one embodiment, a computer program product is provided, including a computer program, which when executed by a processor, causes the processor to execute the steps of training an audio recognition model and an audio recognition method as described above. The steps of training an audio recognition model and an audio recognition method here can be the steps in the training of an audio recognition model and an audio recognition method in each of the above embodiments.
[0159] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0160] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include Read-Only Memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0161] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0162] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A training method for an audio recognition model, characterized in that, The audio recognition model includes a singing voice recognition model and a melody recognition model, and the method includes: Obtaining training sample data; the training sample data includes each sample song audio and the lyric text corresponding to the sample song audio; Inputting the human voice audio in the sample song audio into the singing voice recognition model to obtain a predicted phoneme sequence, and inputting the sample song audio into the melody recognition model to obtain a melody vector representation corresponding to the sample song audio; Determining the actual phoneme sequence of the sample song audio based on the lyric text, and iteratively training the singing voice recognition model according to the difference between the actual phoneme sequence and the predicted phoneme sequence, and iteratively training the melody recognition model according to the difference between the melody vector representation and the prototype vector representation until the two iterative trainings meet the preset training end condition to obtain a trained audio recognition model; wherein, different prototype vector representations are used to represent different melody categories; the trained audio recognition model is used to output a song recognition result corresponding to the song audio to be recognized.
2. The method according to claim 1, wherein The iteratively training the melody recognition model according to the difference between the melody vector representation and the prototype vector representation includes: Clustering each of the melody vector representations to obtain at least one category of clustered vector representations; Determining the prototype vector representation corresponding to each category based on the distance relationship between the clustered vector representations of the same category; wherein, the distance between the prototype vector representation and each of the clustered vector representations of the category corresponding to the prototype vector representation meets a preset condition; Iteratively training the melody recognition model based on the distance between the melody vector representation and the prototype vector representations corresponding to each category.
3. The method according to claim 1, wherein The inputting the human voice audio in the sample song audio into the singing voice recognition model to obtain a predicted phoneme sequence includes: Inputting the sample song audio into a pre-trained voice separation model to obtain the human voice audio in the sample song audio; Extracting a first spectral feature from the human voice audio and inputting the first spectral feature into the singing voice recognition model to obtain a predicted phoneme sequence corresponding to the human voice audio.
4. The method according to claim 3, wherein The singing voice recognition model includes a first convolutional layer, a first fully connected layer, and a first classification layer. The inputting the first spectral feature into the singing voice recognition model to obtain a predicted phoneme sequence corresponding to the human voice audio includes: Inputting the first spectral feature into the first convolutional layer so that the first convolutional layer extracts a first spectral convolutional feature corresponding to the first spectral feature; Inputting the first spectral convolutional feature into the first fully connected layer so that the first fully connected layer transforms the dimension type of the first spectral convolutional feature from a spatial feature to a temporal feature to obtain a dimension-reduced spectral convolutional feature; Inputting the dimension-reduced spectral convolutional feature into the first classification layer so that the first classification layer performs a classification process on the dimension-reduced spectral convolutional feature to obtain a predicted phoneme sequence corresponding to the human voice audio.
5. The method according to claim 1, characterized in that, The melody recognition model includes a second convolutional layer and a second fully-connected layer. Inputting the sample song audio into the melody recognition model to obtain the melody vector representation corresponding to the sample song audio includes: Extracting second spectral features from the sample song audio and inputting the second spectral features into the second convolutional layer, so that the second convolutional layer extracts second spectral convolutional features corresponding to the second spectral features; Inputting the second spectral convolutional features into the second fully-connected layer, so that the second fully-connected layer transforms the dimension type of the second spectral convolutional features from spatial features to temporal features, and obtains the melody vector representation corresponding to the sample song audio.
6. An audio recognition method, characterized in that, The method includes: Obtaining a trained audio recognition model; the trained audio recognition model is trained according to the training method described in any one of claims 1 to 5; the trained audio recognition model includes a trained singing voice recognition model and a trained melody recognition model; Inputting the human voice audio in the song audio to be recognized into the trained singing voice recognition model to obtain a predicted phoneme sequence corresponding to the song audio to be recognized, and inputting the song audio to be recognized into the trained melody recognition model to obtain a melody vector representation corresponding to the song audio to be recognized; Outputting a song recognition result corresponding to the song audio to be recognized according to the predicted phoneme sequence and the melody vector representation.
7. The method according to claim 6, characterized in that, The outputting a song recognition result corresponding to the song audio to be recognized according to the predicted phoneme sequence and the melody vector representation includes: Obtaining a phoneme sequence vector representation corresponding to the predicted phoneme sequence; Fusing the phoneme sequence vector representation and the melody vector representation to obtain a fused vector representation; Generating a song recognition result corresponding to the song audio to be recognized according to the fused vector representation.
8. The method according to claim 7, characterized in that, The fusing the phoneme sequence vector representation and the melody vector representation to obtain a fused vector representation includes: Respectively obtaining a phoneme weight corresponding to the phoneme sequence vector representation and a melody weight corresponding to the melody vector representation; Performing weighted summation on the phoneme sequence vector representation and the melody vector representation according to the phoneme weight and the melody weight to obtain the fused vector representation.
9. The method according to claim 7, wherein The fusing the phoneme sequence vector representation and the melody vector representation to obtain a fused vector representation includes: Concatenating the phoneme sequence vector representation and the melody vector representation to obtain a concatenated vector representation; Taking the concatenated vector representation as the fused vector representation.
10. The method according to claim 7, wherein The generating a song recognition result corresponding to the song audio to be recognized according to the fused vector representation includes: Obtaining a candidate song vector representation corresponding to each candidate song audio; Determining the similarity between the song vector representation corresponding to the song audio to be recognized and each candidate song vector representation; Taking the candidate song audio corresponding to the candidate song vector representation with the highest similarity as the song recognition result corresponding to the song audio to be recognized.
11. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 10 are implemented.
12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 10 are implemented.
Citation Information
Patent Citations
Lyrics generation method and device
CN106547789A
Method and device for extracting main melody track in audio data, terminal, and storage medium
CN108831423A