Audio melody recognition model training method, audio processing method and related equipment

By extracting and splicing the spectrum peak eigenvectors and irrelevant eigenvectors of audio data, training neural network models solves the problem of interference in the prior art audio processing accuracy, and achieving higher melody recognition accuracy.

CN115691511BActive Publication Date: 2025-05-09TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211350427.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-31
Publication Date
2025-05-09
Estimated Expiration
2042-10-31

AI Technical Summary

Technical Problem

When identifying whether a song is a cover, existing audio processing technology is interfered with information that is not related to the audio processing task, such as semantic information, speaker information, song style information, song emotional information, etc., which affects the accuracy.

Method used

By obtaining the sample data set, the spectrum peak eigenvector and irrelevant eigenvectors of each set of audio data are extracted, and stitched and processed to obtain the target eigenvector. These feature vectors and song annotation data are input into the neural network model for training to obtain an audio melody recognition model.

Benefits of technology

By removing information that is not related to the audio melody, the recognition accuracy of the audio melody recognition model is improved, and it is possible to more accurately determine whether the audio to be identified is a cover song.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115691511B_ABST
    Figure CN115691511B_ABST
Patent Text Reader

Abstract

This application discloses a training method, audio processing method, and related equipment for an audio melody recognition model. The method includes: acquiring a sample dataset comprising multiple sets of audio data, each set including three categories: original songs, cover versions of original songs, and other songs besides original and cover versions, with each category having its own annotation data; extracting the spectral peak feature vector and irrelevant feature vector from each category of audio data, concatenating the spectral peak feature vector and irrelevant feature vector to obtain the target feature vector for each category of song data; and inputting the target feature vector and song annotation data into a neural network model for training to obtain the audio melody recognition model. This method can improve the reliability and accuracy of the audio melody recognition model during training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and in particular to a training method for an audio melody recognition model, an audio processing method and related equipment. Background Art

[0002] With the development of audio processing technology, the application of audio processing technology is becoming more and more extensive. Among them, audio processing technology can be applied to song cover recognition. Most of the current audio processing technologies extract the spectral features of the audio signal to determine whether it is a cover song. However, since these spectral features contain some information that is irrelevant to the task of audio processing, such as semantic information, speaker information, song style information, song emotional information, etc., directly using the spectral features to identify whether a song is a cover song will be interfered by information that is irrelevant to the task of audio processing, affecting the accuracy of audio processing. Therefore, it is very important to improve the accuracy of audio processing. Summary of the invention

[0003] The embodiments of the present application provide an audio melody recognition model training method, an audio processing method and related equipment, which can improve the accuracy of audio melody recognition.

[0004] In a first aspect, an embodiment of the present application provides a method for training an audio melody recognition model, comprising:

[0005] Acquire a sample data set, the sample data set comprising multiple groups of audio data, each group of audio data comprising three types of song data, namely, original song data, cover song data of the original song data, and other song data other than the original song data and the cover song data, and each type of song data has respective song annotation data;

[0006] Extracting the spectrum peak feature vector and irrelevant feature vector of each type of song data from each group of audio data, and concatenating the spectrum peak feature vector and the irrelevant feature vector of each type of song data to obtain the target feature vector of each type of song data in each group of audio data; the irrelevant feature vector is a feature vector irrelevant to the melody of the audio data;

[0007] The target feature vector and song annotation data of each type of song data in each group of audio data are input into a neural network model for training to obtain an audio melody recognition model.

[0008] In a second aspect, an embodiment of the present application provides an audio processing method, comprising:

[0009] Acquire audio to be recognized, and extract spectrum peak feature vectors and irrelevant feature vectors from the audio to be recognized, wherein the irrelevant feature vectors are feature vectors irrelevant to the melody of the audio to be recognized;

[0010] The spectrum peak feature vector and the irrelevant feature vector are concatenated to obtain the feature vector to be identified of the audio to be identified; the feature vector to be identified is input into the audio melody recognition model as described in the first aspect to obtain the melody feature vector of the audio to be identified;

[0011] If the minimum distance between the melody feature vector of the audio to be identified and the melody feature vectors of each original song data in the specified database is less than or equal to a preset threshold, it is determined that the audio to be identified is a cover song.

[0012] In a third aspect, an embodiment of the present application provides a computer device, the device comprising: a processor and a memory, the processor being configured to execute:

[0013] Acquire a sample data set, the sample data set comprising multiple groups of audio data, each group of audio data comprising three types of song data, namely, original song data, cover song data of the original song data, and other song data other than the original song data and the cover song data, and each type of song data has respective song annotation data;

[0014] Extracting the spectrum peak feature vector and irrelevant feature vector of each type of song data from each group of audio data, and concatenating the spectrum peak feature vector and the irrelevant feature vector of each type of song data to obtain the target feature vector of each type of song data in each group of audio data; the irrelevant feature vector is a feature vector irrelevant to the melody of the audio data;

[0015] The target feature vector and song annotation data of each type of song data in each group of audio data are input into a neural network model for training to obtain an audio melody recognition model.

[0016] In a fourth aspect, an embodiment of the present application provides another computer device, the device comprising: a processor and a memory, the processor being configured to execute:

[0017] Acquire audio to be recognized, and extract spectrum peak feature vectors and irrelevant feature vectors from the audio to be recognized, wherein the irrelevant feature vectors are feature vectors irrelevant to the melody of the audio to be recognized;

[0018] The spectrum peak feature vector and the irrelevant feature vector are concatenated to obtain the feature vector to be identified of the audio to be identified; the feature vector to be identified is input into an audio melody recognition model to obtain the melody feature vector of the audio to be identified;

[0019] If the minimum distance between the melody feature vector of the audio to be identified and the melody feature vectors of each original song data in the specified database is less than or equal to a preset threshold, it is determined that the audio to be identified is a cover song.

[0020] In a fifth aspect, an embodiment of the present application further provides a computer-readable storage medium, in which program instructions are stored. When the program instructions are executed, they are used to implement the method described in the first aspect or the second aspect above.

[0021] In the embodiment of the present application, a sample data set is obtained, which includes multiple groups of audio data, each group of audio data includes three types of song data, namely, original song data, cover song data of the original song data, and other song data other than the original song data and the cover song data, and each type of song data has its own song annotation data, and the spectrum peak feature vector and irrelevant feature vector of each type of song data are extracted from each group of audio data, and the irrelevant feature vector is a feature vector irrelevant to the melody of the audio data; the spectrum peak feature vector and irrelevant feature vector of each type of song data are spliced ​​to obtain the target feature vector of each type of song data in each group of audio data; the target feature vector and song annotation data of each type of song data in each group of audio data are input into the neural network model for training, and the audio melody recognition model is obtained. In this way, a large amount of information irrelevant to the spectrum melody is removed, and the accuracy of the audio melody recognition model in recognizing the melody is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0023] Figure 1 It is a flowchart of an audio melody recognition model training method;

[0024] Figure 2 is a flowchart of another audio melody recognition model training method;

[0025] Figure 3 This is an example of a spectrum graph;

[0026] Figure 4 This is an example graph of a peak point;

[0027] Figure 5 It is a flowchart of another method for training an audio melody recognition model;

[0028] Figure 6 It is a schematic diagram of the structure of a deep neural network;

[0029] Figure 7 It is a schematic diagram for calculating the loss function value;

[0030] Figure 8 It is a flowchart of another method for training an audio melody recognition model;

[0031] Fig. 9 is a flowchart of an audio processing method;

[0032] Fig.10 It is a structural schematic diagram of a training device for an audio melody recognition model;

[0033] Fig.11 It is a structural diagram of an audio processing device. DETAILED DESCRIPTION

[0034] The following will be combined with the drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0035] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.

[0036] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, large image processing technology, operation / interaction systems, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0037] The key technologies of speech technology include automatic speech recognition technology (ASR), text-to-speech technology (TTS) and voiceprint recognition technology. Enabling computers to listen, see, speak and feel is the future development direction of human-computer interaction, among which speech has become one of the most promising human-computer interaction methods in the future.

[0038] Natural language processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between people and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field will involve natural language, that is, the language people use in daily life, so it is closely related to the study of linguistics. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies.

[0039] Machine Learning (ML) is a multi-disciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning / deep learning usually includes artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.

[0040] Based on the machine learning and other technologies mentioned in the above-mentioned artificial intelligence technology, the present application proposes a training method, an audio processing method and related equipment for an audio melody recognition model. The audio melody recognition model is obtained by removing data irrelevant to the audio melody through training, thereby improving the accuracy of model recognition, which helps to improve the accuracy of the model in recognizing audio melodies.

[0041] The training method of the audio melody recognition model provided in the embodiment of the present application can be applied to a training device for an audio melody recognition model, and the training device for the audio melody recognition model can be set in a computer device. In some embodiments, the computer device can include but is not limited to smart phones, tablet computers, laptop computers, desktop computers, vehicle-mounted smart terminals, smart watches and other smart terminal devices. In some embodiments, the computer device includes one or more databases, and the databases can be used to store audio data.

[0042] The audio processing method provided in the embodiments of the present application can be applied to an audio processing device, which can be set in a computer device. In some embodiments, the computer device can include but is not limited to smart phones, tablet computers, laptop computers, desktop computers, vehicle-mounted smart terminals, smart watches and other smart terminal devices. In some embodiments, the computer device includes one or more databases, which can be used to store audio data.

[0043] In some embodiments, the training method of the audio melody recognition model and the audio processing method provided in the embodiments of the present application can be applied to the scenario of song cover recognition: for example, judging whether a song is a cover song based on the melody feature vector obtained by song processing, etc. Of course, the above application scenarios are only examples, and in other embodiments, the audio processing of the embodiments of the present application can be applied to any scenario associated with audio processing.

[0044] The following is a schematic illustration of the audio melody recognition model training method and the audio processing method provided in the embodiments of the present application in conjunction with the accompanying drawings.

[0045] For details, please see Figure 1 , Figure 1 1 is a flow chart of a method for training an audio melody recognition model provided in an embodiment of the present application. The method for training an audio melody recognition model in an embodiment of the present application can be performed by a training device for an audio melody recognition model, wherein the training device for the audio melody recognition model is arranged in a terminal or a computer device, wherein the specific explanation of the terminal or the computer device is as above. Specifically, the method in an embodiment of the present application includes the following steps.

[0046] S101: Obtain a sample data set, which includes multiple groups of audio data, each group of audio data includes three types of song data: original song data, cover song data of the original song data, and other song data except the original song data and the cover song data, and each type of song data has its own song annotation data.

[0047] In an embodiment of the present application, the cover song data may include cover song data of the original song data, wherein the cover song data of the original song data may include one or more different versions of cover song data corresponding to the original song data. In some embodiments, the song annotation data may include a first tag for indicating the original song data, a second tag for indicating the cover song data of the original song, and a third tag for indicating other song data; in other embodiments, the song annotation data may also include a song group identifier for indicating an audio group, such as song identifier 1 for indicating audio group 1, and audio group 1 may include an original song data, a cover song data of the original song data, and other song data except the original song data and the cover song data. In some embodiments, the audio data may include but is not limited to an audio signal.

[0048] S102: extracting the spectrum peak feature vector and irrelevant feature vector of each type of song data from each group of audio data, and concatenating the spectrum peak feature vector and irrelevant feature vector of each type of song data to obtain the target feature vector of each type of song data in each group of audio data.

[0049] In an embodiment of the present application, the spectrum peak feature vector may include the melody information of the audio data; in some embodiments, the irrelevant feature vector is a feature vector irrelevant to the melody of the audio data, and the irrelevant feature vector may include user information such as speaker information, audio semantic information, and audio emotional information. By introducing irrelevant feature vectors such as speaker information in the audio melody recognition model training process, sufficient speaker information prior is provided, so that in the model training process, it is possible to learn how to remove information irrelevant to the audio melody such as speaker information, which helps to make the melody features extracted by the audio melody recognition model irrelevant to the singer, further ensuring the accuracy of the melody features.

[0050] In one embodiment, when a computer device extracts the spectrum peak feature vector and irrelevant feature vector of each type of song data from each group of audio data, it can extract the peak point sequence of each type of song data from each group of audio data, and normalize the peak point sequence of each type of song data to obtain the spectrum peak feature vector of each type of song data; and extract the Mel spectrum feature of each type of song data from each group of audio data, and determine the irrelevant feature vector of each type of song data based on the Mel spectrum feature of each type of song data.

[0051] S103: Input the target feature vector and song annotation data of each type of song data in each group of audio data into the neural network model for training to obtain an audio melody recognition model.

[0052] In an embodiment of the present application, when the computer device inputs the target feature vector and song annotation data of each type of song data in each group of audio data into the neural network model for training to obtain an audio melody recognition model, the target feature vector and song annotation data of each type of song data in each group of audio data can be input into the neural network model to obtain a target loss function value; the model parameters are adjusted according to the target loss function value, and the target feature vector and song annotation data are input into the neural network model with the adjusted model parameters for retraining; when the target loss function value obtained by retraining is less than the function threshold, it is determined that the audio melody recognition model is obtained.

[0053] This application uses spectral peak point sequences and irrelevant feature vectors to train an audio melody recognition model, thereby directly removing a large amount of information irrelevant to the melody information. Even for songs with a strong cover or adaptation, the melody feature vector can be accurately extracted to determine whether it is a cover song based on the melody feature vector.

[0054] In one embodiment, when the computer device inputs the target feature vector and song annotation data of each type of song data in each group of audio data into the neural network model for training to obtain the audio melody recognition model, the original song data and the song annotation data corresponding to the original song data in each group of audio data can be input into the audio melody recognition model to obtain the melody feature vector of the original song data; the melody feature vector of the original song data is stored in a designated database. In this way, it is helpful to query the original melody feature vector similar to the audio to be recognized from the designated database during subsequent retrieval and recognition, and further determine whether the audio to be recognized is a cover song.

[0055] The embodiment of the present application obtains a sample data set, which includes multiple groups of audio data, each group of audio data includes three types of song data, namely original song data, cover song data of the original song data, and other song data except the original song data and the cover song data, and each type of song data has its own song annotation data, extracts the spectrum peak feature vector and irrelevant feature vector of each type of song data from each group of audio data, and obtains an audio melody recognition model through training based on the spectrum peak feature vector and irrelevant feature vector of each type of song data, thereby removing a large amount of information irrelevant to the spectrum melody, which helps to improve the accuracy of the audio melody recognition model in recognizing melody.

[0056] See also Figure 2 , Figure 2It is a flowchart of another method for training an audio melody recognition model provided by an embodiment of the present application. The method for training an audio melody recognition model of the embodiment of the present application can be executed by a training device for an audio melody recognition model, wherein the training device for the audio melody recognition model is arranged in a terminal or a computer device, wherein the specific explanation of the terminal or the computer device is as above. Specifically, the embodiment of the present application mainly describes the target feature vector extraction process of audio data, which specifically includes the following steps.

[0057] S201: Obtain a sample data set, which includes multiple groups of audio data, each group of audio data includes three types of song data: original song data, cover song data of the original song data, and other song data except the original song data and the cover song data, and each type of song data has its own song annotation data.

[0058] S202: extracting a peak point sequence of each type of song data from each group of audio data, and normalizing the peak point sequence of each type of song data to obtain a spectrum peak feature vector of each type of song data, and extracting Mel spectrum features of each type of song data from each group of audio data, and determining an irrelevant feature vector of each type of song data based on the Mel spectrum features of each type of song data.

[0059] In one embodiment, when the computer device extracts the peak point sequence of each type of song data from each group of audio data, it can transform each type of song data in each group of audio data to obtain a frequency spectrum of each type of song data in each group of audio data; extract one or more peak points from the frequency spectrum of each type of song data, and determine the one or more peak points as the peak point sequence of each type of song data.

[0060] Furthermore, when the computer device performs transformation processing on each type of song data in each set of audio data to obtain a frequency spectrum diagram of each type of song data in each set of audio data, it can perform Fourier transformation processing on the audio signal of each type of song data in each set of audio data to obtain a frequency spectrum diagram of each type of song data in each set of audio data, such as Figure 3 As shown, Figure 3 It is a schematic diagram of a spectrum diagram.

[0061] Furthermore, when determining that one or more peak points are a peak point sequence, the computer device may take the energy values ​​of one or more peak points in the spectrum graph to obtain a peak point sequence, such as Figure 4 As shown, Figure 4 It is a schematic diagram of a peak point, where x is used to represent the peak point.

[0062] In one embodiment, when a computer device normalizes a peak point sequence to obtain a spectrum peak feature vector, it can normalize the peak point sequence to obtain a normalized sequence, and calculate the mean and variance of the normalized sequence; based on the normalized sequence, the mean and variance, the spectrum peak feature vector of each type of song data is calculated.

[0063] Furthermore, when the computer device calculates the spectrum peak feature vector based on the normalized sequence, mean and variance, it can use the normalized sequence to subtract the mean and then divide it by the variance to obtain the spectrum peak feature vector of each type of song data. For example, assuming the peak point sequence x = (x1, x2, ..., xn), then the peak point sequence is normalized and the mean is subtracted. Divide by the variance σ to get x', and the calculation formula is as shown in the following formula (1):

[0064]

[0065] The present application calculates the spectrum peak feature vector to help the audio melody recognition model quickly learn the melody features in the audio data.

[0066] In one embodiment, when determining the irrelevant feature vector based on the mel spectrum feature, the computer device may input the mel spectrum feature into a pre-trained user recognition model to obtain a user feature vector; and input the mel spectrum feature into a pre-trained audio emotion category recognition model to obtain an audio emotion category feature vector; and determine that the user feature vector and the audio emotion category feature vector are irrelevant feature vectors. For example, the user feature vector may be a feature vector of a singer who sings a song.

[0067] Since the user and emotion categories of audio data are irrelevant information, whether two songs are the same song is mainly determined by the melody. Therefore, this application helps the audio melody recognition model learn to remove information in the audio data that is irrelevant to the melody feature vector by determining irrelevant feature vectors, thereby improving the accuracy of the feature vectors, so that the audio melody recognition model can more accurately identify the melody feature vectors.

[0068] S203: Concatenate the spectrum peak feature vectors and irrelevant feature vectors of each type of song data to obtain a target feature vector of each type of song data in each group of audio data.

[0069] In the embodiment of the present application, when the computer device performs splicing processing on the spectrum peak feature vector and the irrelevant feature vector of each type of song data to obtain the target feature vector of each type of song data in each group of audio data, the spectrum peak feature vector and the irrelevant feature vector can be summed to obtain the target feature vector of each type of song data in each group of audio data. In other embodiments, other processing methods can be used to splice the spectrum peak feature vector and the irrelevant feature vector of each type of song data, which is not specifically limited in the present application.

[0070] S204: Input the target feature vector and song annotation data of each type of song data in each group of audio data into the neural network model for training to obtain an audio melody recognition model.

[0071] In an embodiment of the present application, the computer device can input the target feature vector and song annotation data of each type of song data in each group of audio data into a neural network model for training to obtain an audio melody recognition model. In some embodiments, the neural network model can include but is not limited to a deep neural network model, a convolutional neural network model, etc.

[0072] In one embodiment, the computer device can add a first label to the original song data in each group of audio data, add a second label to the cover song data in each group of audio data, and add a third label to other song data in each group of audio data, and input the original song data with the first label, the cover song data with the second label, and the other song data with the third label into the feature extraction module of the neural network model to obtain the spectrum peak feature vector and irrelevant feature vector of each type of song data, and concatenate the spectrum peak feature vector and irrelevant feature vector of each type of song data to obtain the target feature vector of each type of song data in each group of audio data; further input the target feature vector into the prediction module of the neural network model for training to obtain an audio melody recognition model.

[0073] The embodiment of the present application extracts a peak point sequence from each type of song data in each group of audio data in a sample data set, and normalizes the peak point sequence to obtain a spectrum peak feature vector, which helps the audio melody recognition model to quickly learn the melody features in the audio data. At the same time, Mel spectrum features are extracted from each group of audio data, and irrelevant feature vectors are determined based on the Mel spectrum features, which helps the audio melody recognition model to further identify the melody feature vector more accurately.

[0074] See also Figure 5 , Figure 5It is a flowchart of another method for training an audio melody recognition model provided in an embodiment of the present application. The method for training an audio melody recognition model in an embodiment of the present application can be executed by a training device for an audio melody recognition model, wherein the training device for the audio melody recognition model is arranged in a terminal or a computer device, wherein the specific explanation of the terminal or the computer device is as above. Specifically, the embodiment of the present application mainly describes the training process of the audio melody recognition model, which specifically includes the following steps.

[0075] S501: Obtain a sample data set, which includes multiple groups of audio data, each group of audio data includes three types of song data: original song data, cover song data of the original song data, and other song data except the original song data and the cover song data, and each type of song data has its own song annotation data.

[0076] In some embodiments, the song annotation data includes a first label for indicating original song data, a second label for indicating cover song data, and a third label for indicating other song data. The first label is used to indicate that the category of the audio data is original song data, including but not limited to any one or more of text, numbers, letters, etc.; the second label is used to indicate that the category of the audio data is cover song data, including but not limited to any one or more of text, numbers, letters, etc.; the third label is used to indicate that the category of the audio data is other song data, including but not limited to any one or more of text, numbers, letters, etc., wherein the first label, the second label and the third label are different.

[0077] S502: extracting the spectrum peak feature vector and irrelevant feature vector of each type of song data from each group of audio data, and concatenating the spectrum peak feature vector and irrelevant feature vector of each type of song data to obtain the target feature vector of each type of song data in each group of audio data.

[0078] S503: Input the target feature vector and song annotation data of each type of song data in each group of audio data into the neural network model to obtain the target loss function value.

[0079] In some embodiments, the present application may adopt a deep neural network model resnet, such as the resnet may be an 18-layer structure, such as Figure 6 As shown, Figure 6 It is a structural diagram of a deep neural network. The target feature vector of the audio data is input into the 18-layer resnet to obtain the output result, which is the melody feature vector.

[0080] In one embodiment, when a computer device inputs the target feature vector and song annotation data of each type of song data in each group of audio data into a neural network model to obtain a target loss function value, the target feature vector, first label, second label and third label of each type of song data in each group of audio data can be input into the neural network model to obtain the first feature vector of the original song data, the second feature vector of the cover song data and the third feature vector of the other song data of each type of song data in each group of audio data; and the target loss function value is determined based on the first feature vector, the second feature vector and the third feature vector.

[0081] In one embodiment, when the computer device inputs the target feature vector, first label, second label and third label of each type of song data in each group of audio data into the neural network model, it can add a first label to the original song data in each group of audio data, add a second label to the cover song data in each group of audio data, and add a third label to other song data in each group of audio data, and input the original song data with the first label, the cover song data with the second label and the other song data with the third label into the neural network model.

[0082] In one embodiment, when a computer device determines a target loss function value based on a first eigenvector, a second eigenvector, and a third eigenvector, it can calculate a first distance between original song data and cover song data in each group of audio data based on the first eigenvector and the second eigenvector; calculate a second distance between original song data and other song data in each group of audio data based on the first eigenvector and the third eigenvector; and determine the target loss function value based on the first distance and the second distance of each group of audio data.

[0083] Specific can Figure 7 A schematic diagram of calculating the loss function is provided as an example. Figure 7 As shown, this application uses a ternary loss function to calculate the target loss function value, where Anchor is the embedding feature of the original song data, Positive is the embedding feature of the cover song data corresponding to the original song data, and Negative is the embedding feature of other song data. The target loss function value L is defined as shown in the following formula (2):

[0084]

[0085] Where i is a song group, the value of i is 1, 2, ..., n, d(a,p) is the cosine distance between anchor and positive, d(a,n) is the cosine distance between anchor and negative, and marg in is a preset adjustable degree coefficient.

[0086] S504: Retrain the neural network model according to the target loss function value to obtain an audio melody recognition model.

[0087] In an embodiment of the present application, the computer device can adjust the model parameters according to the target loss function value, and input the target feature vector and song annotation data into the neural network model with adjusted model parameters for retraining; when the target loss function value obtained by retraining is less than the function threshold, it is determined that the audio melody recognition model is obtained.

[0088] The embodiment of the present application trains the audio melody recognition model through the target loss function value, which helps to make the feature vectors of the original song data and the cover song data closer, and at the same time makes the feature vectors of the original song data and other cover song data more distant, which helps to improve the recognition accuracy of the audio melody recognition model.

[0089] See also Figure 8 , Figure 8 FIG. 1 is a flow chart of another method for training an audio melody recognition model provided in an embodiment of the present application. Figure 8 As shown, the present application performs Fourier transform on the audio data to extract the spectrum peak sequence and the Mel spectrum feature, obtains the spectrum peak feature vector by normalizing the spectrum peak sequence, and obtains the user feature vector by inputting the Mel spectrum feature into the user recognition model. Further, the spectrum peak feature vector and the user feature vector are feature spliced, and the spliced ​​target feature vector is input into the 18-layer deep neural network resnet18 to obtain the output result, which is a melody feature vector, so that the song of the same song group corresponding to the melody feature vector can be found according to the melody feature vector. In some embodiments, the melody embedding feature of the original song data, i.e., the melody feature vector, can be extracted in the layer before the output layer, and a library is built, which helps to calculate and find the embedding features of the audio to be recognized, such as the audio to be recognized, and the embedding features of the original song data closest to the library in the recognition stage, thereby obtaining the embedding features of the melody.

[0090] See also Fig. 9 , Fig. 91 is a flowchart of an audio processing method provided in an embodiment of the present application. The audio processing method in the embodiment of the present application can be performed by an audio processing device, wherein the audio processing device is arranged in a terminal or a computer device, wherein the specific explanation of the terminal or the computer device is as above. Specifically, the embodiment of the present application mainly describes the audio recognition process, which specifically includes the following steps.

[0091] S901: Acquire the audio to be recognized.

[0092] S902: extracting spectrum peak feature vectors and irrelevant feature vectors from the audio to be recognized, where the irrelevant feature vectors are feature vectors irrelevant to the melody of the audio to be recognized.

[0093] S903: Concatenate the spectrum peak feature vector and the irrelevant feature vector to obtain a feature vector to be identified of the audio to be identified.

[0094] S904: Input the feature vector to be recognized into the audio melody recognition model to obtain the melody feature vector of the audio to be recognized.

[0095] S905: Determine whether the audio to be identified is a cover song based on the melody feature vector.

[0096] In an embodiment of the present application, a computer device can input a feature vector to be identified into an audio melody recognition model to obtain an audio feature vector of the audio to be identified; calculate the distance between the audio feature vector and each original singer's melody feature vector stored in a specified database; determine the original singer's melody feature vector in the specified database with the smallest distance to the audio feature vector as the melody feature vector of the audio to be identified; if the minimum distance between the melody feature vector of the audio to be identified and the original singer's melody feature vector in the specified database is less than or equal to a preset threshold, it can be determined that the audio to be identified is a cover song.

[0097] The embodiment of the present application helps to more accurately obtain the melody feature vector of the audio to be identified by inputting the audio to be identified into a pre-trained audio melody recognition model, so as to determine whether the audio to be identified is a cover song based on the melody feature vector, thereby improving the accuracy of cover recognition.

[0098] See also Fig.10 , Fig.10 1001 is a schematic diagram of a structure of a training device for an audio melody recognition model provided in an embodiment of the present application. Specifically, the device includes: a memory 1001 and a processor 1002.

[0099] In one embodiment, the device further includes a data interface 1003, and the data interface 1003 is used to transfer data information between the computer device and other devices.

[0100] The memory 1001 may include a volatile memory; the memory 1001 may also include a non-volatile memory; the memory 1001 may also include a combination of the above-mentioned types of memories. The processor 1002 may be a central processing unit (CPU). The processor 1002 may further include a hardware chip. The above-mentioned hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD) or a combination thereof. The above-mentioned PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA) or any combination thereof.

[0101] The memory 1001 is used to store programs, and the processor 1002 can call the program stored in the memory 1001 to perform the following steps:

[0102] Acquire a sample data set, the sample data set comprising multiple groups of audio data, each group of audio data comprising three types of song data, namely, original song data, cover song data of the original song data, and other song data other than the original song data and the cover song data, and each type of song data has respective song annotation data;

[0103] Extracting the spectrum peak feature vector and irrelevant feature vector of each type of song data from each group of audio data, and concatenating the spectrum peak feature vector and the irrelevant feature vector of each type of song data to obtain the target feature vector of each type of song data in each group of audio data; the irrelevant feature vector is a feature vector irrelevant to the melody of the audio data;

[0104] The target feature vector and song annotation data of each type of song data in each group of audio data are input into a neural network model for training to obtain an audio melody recognition model.

[0105] Further, when the processor 1002 extracts the spectrum peak feature vector and the irrelevant feature vector of each type of song data from each group of audio data, it is specifically used to:

[0106] Extracting the peak point sequence of each type of song data from each group of audio data, normalizing the peak point sequence of each type of song data to obtain the spectrum peak feature vector of each type of song data; and,

[0107] The Mel frequency spectrum features of each type of song data are extracted from each group of audio data, and the irrelevant feature vectors of each type of song data are determined according to the Mel frequency spectrum features of each type of song data.

[0108] Further, when the processor 1002 determines the irrelevant feature vector of each type of song data according to the Mel frequency spectrum feature of each type of song data, it is specifically used to:

[0109] Inputting the Mel frequency spectrum features of each type of song data into a pre-trained user recognition model to obtain a user feature vector of each type of song data;

[0110] Inputting the Mel frequency spectrum features of each type of song data into a pre-trained audio emotion category recognition model to obtain an audio emotion category feature vector of each type of song data;

[0111] Determine the user feature vector of each category of song data and the audio emotion category feature vector of each category of song data as irrelevant feature vectors of each category of song data.

[0112] Further, when the processor 1002 extracts the peak point sequence of each type of song data from each group of audio data, it is specifically used to:

[0113] Performing transformation processing on each type of song data in each group of audio data to obtain a frequency spectrum diagram of each type of song data in each group of audio data;

[0114] One or more peak points are extracted from the frequency spectrum of each type of song data, and the one or more peak points are determined to be a peak point sequence of each type of song data.

[0115] Further, when the processor 1002 performs normalization processing on the peak point sequence of each type of song data to obtain the spectrum peak feature vector of each type of song data, it is specifically used to:

[0116] Normalizing the peak point sequence of each type of song data to obtain a normalized sequence, and calculating the mean and variance of the normalized sequence;

[0117] The spectrum peak feature vector of each type of song data is calculated based on the normalized sequence, mean and variance.

[0118] Furthermore, the song annotation data includes a first label for indicating the original song data, a second label for indicating the cover song data, and a third label for indicating the other song data; the processor 1002 inputs the target feature vector and song annotation data of each type of song data in each group of audio data into the neural network model for training, and when the audio melody recognition model is obtained, it is used to:

[0119] Input the target feature vector and the first label of the original song data, the target feature vector and the second label of the cover song data, and the target feature vector and the third label of the other song data in each group of audio data into the neural network model to obtain the first feature vector of the original song data, the second feature vector of the cover song data, and the third feature vector of the other song data in each group of audio data;

[0120] A target loss function value is determined according to the first eigenvector, the second eigenvector, and the third eigenvector, and the neural network model is trained according to the target loss function value to obtain the audio melody recognition model.

[0121] Further, when the processor 1002 determines the target loss function value according to the first eigenvector, the second eigenvector, and the third eigenvector, it is specifically used to:

[0122] Calculate a first distance between the original song data and the cover song data in each group of audio data according to the first feature vector and the second feature vector;

[0123] Calculate the second distance between the original song data and the other song data in each group of audio data according to the first feature vector and the third feature vector;

[0124] The target loss function value is determined according to the first distance and the second distance of each group of audio data.

[0125] Further, the processor 1002 is further configured to:

[0126] Inputting the original song data and the song annotation data corresponding to the original song data in each group of audio data into the audio melody recognition model to obtain a melody feature vector of the original song data;

[0127] The melody feature vector of the original song data is stored in a designated database.

[0128] The embodiment of the present application extracts spectral peak feature vectors and irrelevant feature vectors from each group of audio data, and trains an audio melody recognition model based on the spectral peak feature vectors and irrelevant feature vectors, thereby removing a large amount of information irrelevant to the spectral melody, which helps to improve the accuracy of the audio melody recognition model in recognizing melody.

[0129] See also Fig.11 , Fig.11 11 is a schematic diagram of the structure of an audio processing device provided in an embodiment of the present application. Specifically, the device includes: a memory 1101 and a processor 1102.

[0130] In one embodiment, the device further includes a data interface 1103, and the data interface 1103 is used to transfer data information between the computer device and other devices.

[0131] The memory 1101 may include a volatile memory; the memory 1101 may also include a non-volatile memory; the memory 1101 may also include a combination of the above-mentioned types of memories. The processor 1102 may be a central processing unit (CPU). The processor 1102 may further include a hardware chip. The above-mentioned hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD) or a combination thereof. The above-mentioned PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA) or any combination thereof.

[0132] The memory 1101 is used to store programs, and the processor 1102 can call the program stored in the memory 1101 to perform the following steps:

[0133] Acquire audio to be recognized, and extract spectrum peak feature vectors and irrelevant feature vectors from the audio to be recognized, wherein the irrelevant feature vectors are feature vectors irrelevant to the melody of the audio to be recognized;

[0134] The spectrum peak feature vector and the irrelevant feature vector are concatenated to obtain the feature vector to be identified of the audio to be identified; the feature vector to be identified is input into an audio melody recognition model to obtain the melody feature vector of the audio to be identified;

[0135] If the minimum distance between the melody feature vector of the audio to be identified and the melody feature vectors of each original song data in the specified database is less than or equal to a preset threshold, it is determined that the audio to be identified is a cover song.

[0136] The embodiment of the present application helps to more accurately obtain the melody feature vector of the audio to be identified by inputting the audio to be identified into a pre-trained audio melody recognition model, so as to determine whether the audio to be identified is a cover song based on the melody feature vector, thereby improving the accuracy of cover recognition.

[0137] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the method described in the embodiment corresponding to the present application, and can also implement the device of the embodiment corresponding to the present application, which will not be described in detail here.

[0138] The computer-readable storage medium may be an internal storage unit of the device described in any of the foregoing embodiments, such as a hard disk or memory of the device. The computer-readable storage medium may also be an external storage device of the device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the device. Furthermore, the computer-readable storage medium may also include both an internal storage unit of the device and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by the terminal. The computer-readable storage medium may also be used to temporarily store data that has been output or is to be output.

[0139] A person skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.

[0140] The above disclosure is only part of the embodiments of the present application, which certainly cannot be used to limit the scope of rights of the present application. Ordinary technicians in this field can understand that all or part of the processes of implementing the above embodiments and making equivalent changes according to the claims of this application are still within the scope of the invention.

Claims

1. A training method for an audio melody recognition model, characterized in that: include: Acquire a sample data set, the sample data set comprising multiple groups of audio data, each group of audio data comprising three types of song data, namely, original song data, cover song data of the original song data, and other song data other than the original song data and the cover song data, and each type of song data has respective song annotation data; the song annotation data comprises a first label for indicating the original song data, a second label for indicating the cover song data, and a third label for indicating the other song data; Extracting the spectrum peak feature vector and irrelevant feature vector of each type of song data from each group of audio data, and concatenating the spectrum peak feature vector and the irrelevant feature vector of each type of song data to obtain the target feature vector of each type of song data in each group of audio data; the irrelevant feature vector is a feature vector irrelevant to the melody of the audio data; The target feature vector and song annotation data of each type of song data in each group of audio data are input into a neural network model for training to obtain an audio melody recognition model; wherein, the target feature vector and the first label of the original song data, the target feature vector and the second label of the cover song data, and the target feature vector and the third label of the other song data in each group of audio data are input into the neural network model for training.

2. The method according to claim 1, characterized in that: The step of extracting the spectrum peak feature vector and irrelevant feature vector of each type of song data from each group of audio data comprises: Extracting the peak point sequence of each type of song data from each group of audio data, normalizing the peak point sequence of each type of song data to obtain the spectrum peak feature vector of each type of song data; and, The Mel frequency spectrum features of each type of song data are extracted from each group of audio data, and the irrelevant feature vectors of each type of song data are determined according to the Mel frequency spectrum features of each type of song data.

3. The method according to claim 2, characterized in that The step of determining the irrelevant feature vector of each type of song data according to the Mel frequency spectrum feature of each type of song data comprises: Inputting the Mel frequency spectrum features of each type of song data into a pre-trained user recognition model to obtain a user feature vector of each type of song data; Inputting the Mel frequency spectrum features of each type of song data into a pre-trained audio emotion category recognition model to obtain an audio emotion category feature vector of each type of song data; Determine the user feature vector of each category of song data and the audio emotion category feature vector of each category of song data as irrelevant feature vectors of each category of song data.

4. The method according to claim 2, characterized in that: The step of normalizing the peak point sequence of each type of song data to obtain the spectrum peak feature vector of each type of song data includes: Normalizing the peak point sequence of each type of song data to obtain a normalized sequence, and calculating the mean and variance of the normalized sequence; The spectrum peak feature vector of each type of song data is calculated based on the normalized sequence, mean and variance.

5. The method according to claim 1, characterized in that The target feature vector and song annotation data of each type of song data in each group of audio data are input into the neural network model for training to obtain the audio melody recognition model, including: Input the target feature vector and the first label of the original song data, the target feature vector and the second label of the cover song data, and the target feature vector and the third label of the other song data in each group of audio data into the neural network model to obtain the first feature vector of the original song data, the second feature vector of the cover song data, and the third feature vector of the other song data in each group of audio data; A target loss function value is determined according to the first eigenvector, the second eigenvector, and the third eigenvector, and the neural network model is trained according to the target loss function value to obtain the audio melody recognition model.

6. The method according to claim 5, characterized in that The determining the target loss function value according to the first eigenvector, the second eigenvector, and the third eigenvector includes: Calculate a first distance between the original song data and the cover song data in each group of audio data according to the first feature vector and the second feature vector; Calculate the second distance between the original song data and the other song data in each group of audio data according to the first feature vector and the third feature vector; The target loss function value is determined according to the first distance and the second distance of each group of audio data.

7. The method according to claim 1, characterized in that Also includes: Inputting the original song data and the song annotation data corresponding to the original song data in each group of audio data into the audio melody recognition model to obtain a melody feature vector of the original song data; The melody feature vector of the original song data is stored in a designated database.

8. An audio processing method, characterized in that: include: Acquire audio to be recognized, and extract spectrum peak feature vectors and irrelevant feature vectors from the audio to be recognized, wherein the irrelevant feature vectors are feature vectors irrelevant to the melody of the audio to be recognized; The spectrum peak feature vector and the irrelevant feature vector are concatenated to obtain the feature vector to be identified of the audio to be identified; the feature vector to be identified is input into the audio melody recognition model according to any one of claims 1 to 7 to obtain the melody feature vector of the audio to be identified; If the minimum distance between the melody feature vector of the audio to be identified and the melody feature vectors of each original song data in the specified database is less than or equal to a preset threshold, it is determined that the audio to be identified is a cover song.

9. A computer device, characterized in that: The method comprises a processor, an input device, an output device and a memory, wherein the processor, the input device, the output device and the memory are interconnected, wherein the memory is used to store a computer program, the computer program comprises a program, and the processor is configured to call the program to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program instructions, which, when executed, are used to implement the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Audio recognition model training method and device, audio recognition method and device and computer equipment

    CN115240656A