Audio representation model training method, audio processing method and related equipment
By introducing pitch features that match the note scale during the training process of the audio representation model, the problem of inaccurate features of MAE when processing music is solved, and the performance of downstream tasks is improved.
Patent Information
- Application Number
- CN202510211732.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-30
AI Technical Summary
The audio characterization model trained on masked autoencoder (MAE) is not accurate enough when processing music, resulting in poor performance in downstream audio processing tasks.
By introducing pitch features that match the note scale extracted from the training music, training losses are calculated and model parameters are adjusted to improve the performance of the audio characterization model in music processing.
After introducing pitch features that match the note scale, the trained audio representation model is able to output features that perform well in downstream tasks, improving the accuracy of music processing.
Smart Images

Figure CN120071966A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of music processing, and in particular, to a training method for an audio representation model, an audio processing method, and related devices. Background Art
[0002] An audio representation model is a type of model used to extract, encode, and represent useful features from audio signals. The purpose of these models is to convert the original audio signal into a more easily processed and more efficient feature or vector representation, facilitating subsequent downstream tasks such as audio classification, speech recognition, sentiment analysis, sound source separation, etc.
[0003] Masked autoencoders (MAE) can be used as a general audio representation model. During the training process, MAE masks the audio data for prediction and uses the loss obtained after processing the aforementioned audio data by the neural audio codec EnCodec as the prediction target.
[0004] However, in practical applications, it is found that the MAE trained based on the above method outputs inaccurate features when processing audio inputs of the music type, resulting in poor performance in downstream audio processing tasks. Summary of the Invention
[0005] The embodiments of the present application provide a training method for an audio representation model, an audio processing method, and related devices, which are used to improve the performance of the features output by the audio representation model in downstream tasks.
[0006] The first aspect of the embodiments of the present application provides a training method for an audio representation model, including:
[0007] Inputting training music into an initial audio representation model to obtain predicted features of the training music output by the initial audio representation model;
[0008] Inputting the training music into a pre-trained audio quantization network and using the output of the audio quantization network as the quantization features of the training music;
[0009] Extracting pitch features that conform to the note scale from the training music;
[0010] Calculating the training loss of the training music based on the predicted features, the quantization features, and the pitch features;
[0011] Adjusting the model parameters based on the training loss of the training music until a trained audio representation model is obtained, and the trained audio representation model is used to extract deep features of music.
[0012] In a specific implementation, calculating the training loss of the training music based on the prediction feature, the quantization feature, and the pitch feature includes:
[0013] Calculating a first loss of the training music based on the distance between the prediction feature and the quantization feature;
[0014] Calculating a second loss of the training music based on the distance between the prediction feature and the pitch feature;
[0015] Determining the training loss based on the first loss, a preset loss weight corresponding to the first loss, the second loss, and a preset loss weight corresponding to the second loss.
[0016] In a specific implementation, the training music includes multiple audio frames, at least one of the multiple audio frames is a masked frame, the prediction feature of the training music includes the prediction feature of each audio frame, the pitch feature of the training music includes the pitch feature of each audio frame, and calculating the second loss of the training music based on the distance between the prediction feature and the pitch feature includes:
[0017] For the pitch feature of each masked frame, calculating the pitch loss of the masked frame based on the pitch feature of the masked frame and the prediction feature of the masked frame;
[0018] Calculating the second loss of the training music based on the pitch loss of each masked frame.
[0019] In a specific implementation, the training music includes multiple audio frames, at least one of the multiple audio frames is a masked frame and at least one is a visible frame, the prediction feature of the training music includes the prediction feature of each audio frame, the pitch feature of the training music includes the pitch feature of each audio frame, and calculating the second loss of the training music based on the distance between the prediction feature and the pitch feature includes:
[0020] For the pitch feature of each audio frame, calculating the pitch loss of the audio frame based on the pitch feature of the audio frame and the prediction feature of the audio frame;
[0021] Determining a preset masked weight corresponding to the masked frame and a preset visible weight corresponding to the visible frame, where the preset masked weight is greater than the preset visible weight;
[0022] Calculating the second loss of the training music based on the pitch loss of each masked frame, the preset masked weight, the pitch loss of each visible frame, and the preset visible weight.
[0023] In a specific implementation manner, extracting pitch features that conform to the note scale from the training music includes:
[0024] Obtain the feature extraction conditions of the training music, where the feature extraction conditions include the sampling rate, sampling step, audio padding strategy, the number of frequency points per audio frame, the number of semitones included in each octave, and the lowest sampling frequency;
[0025] Input the feature extraction conditions and the training music into a CQT feature extractor to obtain the CQT features of each audio frame in the training music, and use the CQT features of each audio frame in the training music as the pitch features of the training music.
[0026] In a specific implementation manner, extracting pitch features that conform to the note scale from the training music includes:
[0027] Obtain the feature extraction conditions of the training music, where the feature extraction conditions include the sampling rate, sampling step, audio padding strategy, the number of frequency points per audio frame, the number of semitones included in each octave, and the lowest sampling frequency;
[0028] Input the feature extraction conditions and the training music into a CQT feature extractor to obtain the CQT features of each audio frame in the training music;
[0029] Based on the CQT features of each audio frame in the training music, calculate the pitch contour features of each audio frame in the training music, and use the pitch contour features of each audio frame in the training music as the pitch features of the training music.
[0030] In a specific implementation manner, the training music includes multiple music frames, and the initial audio representation model includes a feature extractor, an initial masking encoder, and an initial masking decoder. Inputting the training music into the initial audio representation model to obtain the predicted features of the training music output by the initial audio representation model includes:
[0031] Input the multiple audio frames into the feature extractor to obtain the initial features of each audio frame;
[0032] Determine at least one masking frame from the multiple audio frames, and determine the audio frames other than the masking frame among the multiple audio frames as visible frames;
[0033] Input the initial features of each visible frame into the masking encoder to obtain the deep features of each visible frame output by the masking encoder;
[0034] Input the frame index of each of the masked frames and the deep features of each of the visible frames into the masked decoder to obtain the predicted features of the multiple audio frames, where the predicted feature of each of the masked frames is a preset masking feature;
[0035] Adjusting the model parameters based on the training loss of the training music until a trained audio representation model is obtained, including:
[0036] Adjust the parameters of the initial masked encoder and the parameters of the initial masked decoder based on the training loss;
[0037] Use the feature extractor and the masked decoder with adjusted parameters as the trained audio representation model.
[0038] A second aspect of the embodiments of the present application provides an audio processing method, including:
[0039] Obtain a trained audio representation model, the trained audio representation model is trained based on the training method of the audio representation model in the first aspect, and the trained audio representation model includes a feature extractor and a masked encoder with adjusted parameters;
[0040] Input the target music into the feature extractor to obtain the initial features of the target music, input the initial features into the masked encoder with adjusted parameters, and use the features output by the trained masked encoder as the deep features of the target music;
[0041] Based on the deep features of the target music, perform a processing task for the target music.
[0042] A third aspect of the embodiments of the present application provides a computer device, including:
[0043] A central processing unit, a memory, and an input / output interface;
[0044] The memory is a transient storage memory or a persistent storage memory;
[0045] The central processing unit is configured to communicate with the memory and execute the instruction operations in the memory to execute the training method of the audio representation model in the first aspect or the audio processing method in the second aspect.
[0046] A fourth aspect of the embodiments of the present application provides a computer program product containing instructions, when the computer program product runs on a computer, it causes the computer to execute the training method of the audio representation model in the first aspect or the audio processing method in the second aspect.
[0047] A fifth aspect of the embodiments of the present application provides a computer storage medium, in which instructions are stored. When the instructions are executed on a computer, the computer is caused to execute the training method of the audio characterization model described in the first aspect or the audio processing method described in the second aspect.
[0048] As can be seen from the above technical solutions, the embodiments of the present application have the following advantages: By analyzing different types of audio such as music, white noise, and conversations, it is found that whether it is pure music, dry voice, or a song, an important characteristic that music differs from other types of audio is the musical score, that is, each piece of music has its corresponding musical score. And a musical score records the occurrence time, duration, and sequence of different notes in the music. And notes are specific identifiers for different pitches. Based on this, the embodiments of the present application consider introducing attention to pitch in the training loss. Specifically, in order to better extract the features that distinguish music from other types of audio, the embodiments of the present application introduce pitch features that conform to the note scale extracted from the training music in the step of calculating the training loss. And it is found in actual applications that after introducing the pitch features that conform to the note scale, the trained audio characterization model can output features that perform well in downstream tasks even when processing audio inputs of this type of music. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 is a schematic flowchart of a training method for an audio characterization model disclosed in an embodiment of the present application;
[0050] Figure 2 is a schematic flowchart of an audio processing method disclosed in an embodiment of the present application;
[0051] Figure 3 is another schematic flowchart of a training method for an audio characterization model disclosed in an embodiment of the present application;
[0052] Figure 4 is a schematic structural diagram of a computer device disclosed in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0054] The embodiments of the present application provide a training method for an audio characterization model, an audio processing method, and related devices, which are used to improve the performance of the features output by the audio characterization model in downstream tasks.
[0055] Please refer to Figure 1 , an embodiment of the present application provides a training method for an audio representation model, which can be deployed in software on computer devices such as servers and terminals, including:
[0056] 101. Input the training music into the initial audio representation model to obtain the predicted features of the training music output by the audio representation model.
[0057] In order to better guide the audio representation model to learn the features beneficial to downstream task execution in music, an embodiment of the present application uses audio such as music as the training samples of the audio representation model, that is, the training music referred to in the embodiment of the present application. Generally, the training music adopted in the embodiment of the present application can be audio such as pure music, songs or dry vocals played or sung based on a specific musical score.
[0058] After inputting the training into the initial audio representation model, the initial audio representation model will extract audio features from the training music based on the initial network parameters as the predicted features of the training music. Among them, the initial audio representation model in the embodiment of the present application can have the same network structure as models such as Wav2Vec or MAE, which is not limited here.
[0059] 102. Input the training music into the pre-trained audio quantization network, and use the output of the audio quantization network as the quantization features of the training music.
[0060] In the training method of the embodiment of the present application, the quantization features output by the audio quantization network are used as the learning target of the initial audio representation model. Therefore, in addition to inputting the training music into the initial audio representation model, it is also necessary to input the training music into the pre-trained audio quantization network to obtain the quantization features for the audio representation model to learn.
[0061] Among them, the audio quantization network in the embodiment of the present application can be a neural audio codec such as EnCodec or SoundStream that realizes audio feature extraction through quantization technology.
[0062] 103. Extract pitch features that conform to the note scale from the training music.
[0063] In the process of using the audio representation model trained by the existing model training method to extract features from music-type audio and using the extracted features for downstream tasks, it is found that the output of the audio representation model performs mediocrely in downstream tasks. For example, in automatic music transcription, which is a downstream task for converting audio into symbolic sheet music (MIDI or staff notation), the existing audio feature extraction does not particularly focus on the pitch information in the audio, resulting in incorrect note recording due to pitch detection errors; or in music style classification, which is a downstream task for classifying music styles based on music features (especially melody, harmony, mode), such as classical, jazz, pop, etc. Some music styles rely on unique pitch patterns, and the existing audio feature extraction does not particularly focus on the pitch information in the audio, so incorrect style classification will occur due to pitch detection errors.
[0064] Therefore, the embodiments of the present application focused on analyzing different types of audio and found that the important characteristic that differentiates music-type audio from other types of audio, whether it is pure music, dry voice or song, is the sheet music, that is, each piece of music is played or sung based on a specific sheet music.
[0065] In order to consider introducing the attention and learning of the sheet music corresponding to the music in the model training process in the embodiments of the present application, the embodiments of the present application further analyzed the characteristics that reflect the music characteristics in the sheet music. It can be understood that a sheet music records the occurrence time, duration and sequence of different notes in the music. And notes are specific identifiers for different pitches. Based on this, the embodiments of the present application consider introducing the learning of pitch features in the model training method of the embodiments of the present application.
[0066] It should be noted that there are many features that can reflect pitch, such as Mel spectrogram, linear spectrogram, and constant-Q transform (CQT) features, etc. Among them, different features reflect the pitch features in the audio from different scales or graduations. For example, the Mel spectrogram is the pitch feature that conforms to the Mel scale, the linear spectrogram is the pitch feature that conforms to the frequency scale, and the CQT feature is the pitch feature that conforms to the note scale. Among them, the Mel scale is designed based on the human ear's perception of pitch and corresponds to the logarithmic transformation of frequency. It is approximately linearly distributed in the low-frequency range and logarithmically distributed in the high-frequency range; the frequency scale is distributed linearly, that is, the bandwidth of each frequency band is equal. This scale directly maps the physical frequency of the signal (such as hertz, Hz) without considering the non-linear perception of frequency by the human ear; the note scale is divided according to notes (or scales), usually based on the equal temperament, dividing an octave into twelve semitones. The CQT feature uses this scale and distributes frequencies at logarithmic intervals of notes.
[0067] The pitch features referred to in the embodiments of the present application conform to the Mel scale, the frequency scale, or the musical note scale. In fact, it refers to the frequency distribution manner of the extracted pitch features. Since different scales correspond to different frequency distribution manners, mapping the audio spectrum to the feature spaces of different scales can reflect the characteristics of the audio under this specific frequency distribution manner, rather than a single value.
[0068] In summary, the embodiments of the present application introduce the step of extracting pitch features that conform to the musical note scale from the training music in the model training method, and use the extracted pitch features that conform to the musical note scale as the learning objective. Specifically, the pitch features that conform to the musical note scale include, but are not limited to, the CQT features and the pitch contour features further extracted based on the CQT features, that is, the chroma features.
[0069] 104. Calculate the training loss of the training music based on the prediction features, quantization features, and pitch features.
[0070] The embodiments of the present application use the quantization features and pitch features to process the prediction features output by the initial audio representation model. And, based on the foregoing embodiments, it can be known that the learning objective of the audio representation model in the embodiments of the present application is the quantization features and pitch features extracted in the foregoing steps. That is to say, the training objective of the embodiments of the present application is to make the prediction features output by the audio representation model as close as possible to the corresponding quantization features and pitch features.
[0071] Specifically, the training loss of the embodiments of the present application can be calculated based on the distance between the prediction features and the quantization features, and the distance between the prediction features and the pitch features, and the training loss is positively correlated with the foregoing two distances, that is, the greater the distance, the greater the training loss. The above calculation method of the training loss can measure the learning degree, or the closeness degree, of the audio representation model to the quantization features and pitch features reflected in the prediction features output by the audio representation model.
[0072] 105. Adjust the model parameters based on the training loss of the training music until an audio representation model that has been trained is obtained. The trained audio representation model is used to extract the deep features of the music.
[0073] The embodiments of the present application can adjust the network parameters of the initial audio representation model according to the training loss calculated in step 104. Until after a certain round of training is completed, it is found based on the output of the audio representation model that the current audio representation model meets the preset convergence condition, then the current audio representation model can be used as the trained audio representation model and applied to extract the audio features in the music.
[0074] Among them, the convergence conditions that the current audio representation model needs to meet include, but are not limited to: the training loss changes less than a preset loss threshold in multiple consecutive rounds, the key metrics (such as accuracy, F1 score, AUC, BLEU, etc.) change less than a preset metric threshold in multiple rounds, the parameter changes of the audio representation model in multiple rounds are less than a preset change threshold, and any conditions for determining that the audio representation model reaches the expected stable state and no longer improves significantly.
[0075] In the embodiments of the present application, in order to better extract the features of music that are different from other types of audio, in the step of calculating the training loss in the embodiments of the present application, pitch features that conform to the note scale are extracted from the training music. And it is found in actual applications that after introducing the pitch features that conform to the note scale, the trained audio representation model can output features that perform well in downstream tasks even when processing audio inputs of this type of music.
[0076] Based on the foregoing embodiments, step 104 in the embodiments of the present application can be specifically implemented through the following steps: calculate the first loss of the training music based on the distance between the predicted features and the quantization features; calculate the second loss of the training music based on the distance between the predicted features and the pitch features; determine the training loss based on the first loss, the preset loss weight corresponding to the first loss, the second loss, and the preset loss weight corresponding to the second loss.
[0077] Specifically, the embodiments of the present application independently calculate the second loss reflected by the distance between the predicted features and the pitch features, and independently calculate the first loss reflected by the distance between the predicted features and the quantization features. Finally, in a weighted summation manner, the product of the first loss and the corresponding preset loss weight and the product of the second loss and the corresponding preset loss weight are calculated to balance the proportions of the audio representation model learning the pitch features and the quantization features respectively. Among them, the preset loss weight corresponding to the first loss and the preset loss weight corresponding to the second loss can be configured as needed and are not limited here.
[0078] Further, in some specific implementation manners, the training music usually includes multiple audio frames, and the pitch feature of the training music includes the pitch feature of each audio frame, and the prediction feature of the training music includes the prediction feature of each audio frame. Additionally, the present application adopts a masking training mechanism, that is to say, there is at least one masking frame among the multiple audio frames included in the training music. That is to say, the audio representation model needs to learn how to recover and extract features from the masking frame according to the context of the masking frame during the training process. In this way, during the calculation of the second loss, it can be considered to calculate the second loss only using the masking frames, because the masking frames contain more information provided by the audio representation model than the visible frames. Therefore, the model parameters are adjusted based on the training loss calculated only based on the masking frames, without calculating the training loss based on each audio frame, which can improve the overall training efficiency of the model while ensuring the accuracy of the model.
[0079] Specifically, for each masking frame, the distance or error between its pitch feature and its prediction feature is used as its pitch loss. Then, the sum of the pitch losses of each masking frame or the average value of the sum of the pitch losses of each masking frame is used as the second loss, which is not specifically limited here.
[0080] Furthermore, in addition to the masking frames among the multiple audio frames included in the training music, there are also visible frames that are not masked. In order to make full use of the prediction features of each obtained audio frame, in the process of calculating the second loss in the embodiments of the present application, it is also considered to learn how to accurately extract features from the audio frames, such as introducing the pitch loss of the visible frames. The embodiments of the present application are similar to the foregoing embodiments, the difference being that in the embodiments of the present application, for each audio frame, the distance or error between its pitch feature and its prediction feature is used as its pitch loss. Among them, each audio frame includes visible frames and masking frames. However, considering that the masking frames contain more information provided by the audio representation model, therefore, the embodiments of the present application also additionally introduce a preset masking weight to weight the pitch loss of the masking frames, and introduce a preset visible weight to weight the pitch loss of the visible frames. Finally, the final second loss is calculated according to the product of the preset weight corresponding to each audio frame (the preset visible weight corresponding to the visible frame and the preset masking weight corresponding to the masking frame) and the pitch loss of the audio frame.
[0081] Specifically, the calculation of the second loss in the embodiments of the present application can be calculated by any error-based loss such as mean square error loss, mean absolute error loss, smooth mean absolute error loss, etc., which is not specifically limited here.
[0082] Based on the foregoing embodiments, in some specific implementation manners, if the CQT feature is used as the pitch feature of the training music, the foregoing step 103 may be implemented with reference to the following manner: Obtain the feature extraction conditions of the training music, where the feature extraction conditions include the sampling rate, the sampling step, the audio padding strategy, the number of frequency points in each audio frame, the number of semitones included in each octave, and the lowest sampling frequency; input the feature extraction conditions and the training music into a CQT feature extractor to obtain the CQT features of each audio frame in the training music, and use the CQT features of each audio frame in the training music as the pitch features of the training music. That is, use the CQT features of each audio frame in the training music as the pitch features of each audio frame in the training music.
[0083] Specifically, the CQT feature extractor adopted in the embodiments of the present application may be an audio processing tool that can directly perform CQT feature extraction, such as nnaudio, Librosa, etc., and the embodiments of the present application are not limited thereto. Among them, in order to ensure that the frequency distribution, time resolution, and data padding method of the extracted CQT features meet the requirements, the embodiments of the present application need to provide the CQT feature extractor with feature extraction conditions including, but not limited to, the sampling rate, the sampling step, the audio padding strategy, the number of frequency points in each audio frame, the number of semitones included in each octave, and the lowest sampling frequency.
[0084] In other specific implementation manners, in order to ensure that the extracted pitch features conform to the note scale, the embodiments of the present application may also use the pitch contour feature further processed based on the CQT feature, that is, the chroma feature, as the pitch feature of the embodiments of the present application. In the embodiments of the present application, the foregoing step 103 may be implemented with reference to the following manner: Obtain the feature extraction conditions of the training music, where the feature extraction conditions include the sampling rate, the sampling step, the audio padding strategy, the number of frequency points in each audio frame, the number of semitones included in each octave, and the lowest sampling frequency; input the feature extraction conditions and the training music into a CQT feature extractor to obtain the CQT features of each audio frame in the training music; calculate the pitch contour features of each audio frame in the training music based on the CQT features of each audio frame in the training music, and use the pitch contour features of each audio frame in the training music as the pitch features of the training music. Among them, the audio padding strategy includes whether to center and the padding method.
[0085] Specifically, the pitch contour in the embodiments of this application is the chroma feature. The chroma feature mainly focuses on note information and is usually aggregated over 12 pitches. Therefore, the calculation of the Chroma feature based on the CQT feature is usually performed on the CQT feature of each frame. In other words, the Chroma feature of each audio frame can be obtained by aggregating the spectral information of the corresponding CQT feature in different octaves. That is, the embodiments of this application can group the spectral energy of the CQT by pitch and aggregate the spectral information on these pitches.
[0086] It can be understood that the CQT feature extractor adopted in the embodiments of this application is similar to the CQT feature extractor in the foregoing embodiments. Similar to the CQT feature extraction, in the embodiments of this application, the step of calculating the pitch contour feature of each audio frame in the training music based on the CQT feature of each audio frame in the training music can also be implemented by audio processing tools such as the foregoing nnaudio and Librosa that can perform chroma feature extraction based on the CQT feature, which is not limited herein.
[0087] On the basis of the foregoing embodiments, the specific structure of the audio representation model in the embodiments of this application is briefly described. First, the training music in the embodiments of this application includes multiple music frames, and the initial audio representation model includes a feature extractor, an initial masking encoder, and an initial masking decoder. On this basis, step 101 above can be specifically implemented in the following manner: input multiple audio frames into the feature extractor to obtain the initial features of each audio frame; determine at least one masking frame from the multiple audio frames, and determine the audio frames other than the masking frames in the multiple audio frames as visible frames; input the initial features of each visible frame into the masking encoder to obtain the deep features of each visible frame output by the masking encoder; input the frame index of each masking frame and the deep features of each visible frame into the masking decoder to obtain the predicted features of the multiple audio frames, where the predicted feature of each masking frame is a preset masking feature; on this basis, step 105 above can be specifically implemented in the following manner: adjust the parameters of the initial masking encoder and the parameters of the initial masking decoder based on the training loss; use the feature extractor and the masking decoder with adjusted parameters as the trained audio representation model.
[0088] Specifically, in the forward propagation process, the input training music needs to pass through the feature extractor, the initial masking encoder, and the initial masking decoder in the audio representation model in sequence. During the model training process, among the various structures in the audio representation model, only the parameters in the initial masking encoder and the initial masking decoder need to be adjusted or learned. Among them, the feature extractor can be pre-trained to output initial features (or surface features), and the initial masking encoder needs to extract deep features from audio frames. In particular, since the embodiment of the present application adopts a masking training mechanism, at least one masking frame needs to be determined from multiple audio frames. And subsequently, the audio features in the masking frame are restored and extracted through the masking decoder and context information. Finally, it should be noted that the masking training mechanism is only used during the model training process, that is, there is no need to perform masking processing on the input audio during the model application process. During the application process of the audio representation model, only the feature extractor and the masking decoder with adjusted parameters are used as the trained audio representation model, that is, the input target music will only pass through the feature extractor and the masking decoder with adjusted parameters.
[0089] The foregoing described multiple embodiments of the training method of the audio representation model of the present application. The application of the audio representation model obtained by training the foregoing embodiments will be described below. Please refer to Figure 2 , the embodiment of the present application provides an audio processing method, including the following steps:
[0090] 201. Obtain a trained audio representation model, where the trained audio representation model is trained based on the training method of the audio representation model in any of the foregoing embodiments, and the trained audio representation model includes a feature extractor and a masking encoder with adjusted parameters.
[0091] As described in step 105 above, the trained audio representation model is used to output the deep features of music. Then, in order to obtain the deep features of the target music, the embodiment of the present application first needs to obtain a trained audio representation model trained based on the training method of the audio representation model in any of the foregoing embodiments. Based on the foregoing embodiments, although it is necessary to adjust the parameters of multiple structures in the audio representation model during the training process, before finally using the audio representation model, only the feature extractor and the masking encoder with adjusted parameters are used as the trained audio representation model.
[0092] 202. Input the target music into the feature extractor to obtain the initial features of the target music, input the initial features into the masking encoder with adjusted parameters, and use the features output by the trained masking encoder as the deep features of the target music.
[0093] After obtaining the trained audio representation model in step 201, the embodiments of the present application can input the target music into the feature extractor to obtain the initial features of the target music, and input the initial features into the masked encoder with adjusted parameters to obtain the deep features output by it. It can be understood that since the masked encoder acts as a deep feature extractor during the model training process, therefore, in the model application, the features output by the trained masked encoder can be used as the deep features of the target music. These deep features are also the audio features of the trained music output by the audio representation model.
[0094] 203. Based on the deep features of the target music, perform a processing task for the target music.
[0095] Regardless of what processing task is performed on the target music, first, the audio features of the target music need to be extracted. In the embodiments of the present application, the deep features of the target music are output through the trained audio representation model. Based on the foregoing embodiments, it can be known that the deep features obtained in the above manner can be used as the input of the downstream task and have good performance in the downstream task.
[0096] To better implement the training method and audio processing method of the audio representation model in the embodiments of the present application, the embodiments of the present application provide a training architecture as Figure 3 shown. In the embodiments of the present application, MAE is used as the audio representation model, EnCodec is used as the audio quantization network, and CQT features are used as the pitch features that conform to the note scale.
[0097] The existing training objective of MAE is to calculate the classification loss with the output of the EnCodec quantizer, but the training objective has the following disadvantages: 1. Encoding loss: During the encoding process of the EnCodec Encoder, in order to achieve a high compression ratio, some details of the original audio signal will be lost, especially those that are not important to the human auditory system, but may be key information for some downstream tasks. 2. Quantization error: The quantization process of the EnCodec quantizer will introduce errors because continuous audio features are quantized into embedded representations in the codebook and then represented by the corresponding indices in the codebook, and it is impossible to completely capture all the nuances. In addition, the size of the codebook is limited and it is impossible to accurately represent all audio features. 3. Insufficient generalization: During the training process, EnCodec is pre-trained and frozen, and its generalization ability to data is limited. As the amount of training data of MAE increases, the EnCodec quantizer is not sufficient to output sufficiently robust labels, thus affecting the generalization ability of the model.
[0098] MAE is a large general audio representation model, which consists of a feature extractor, a feature transformer, a masked autoencoder MAE, and EnCodec. In the embodiments of this application, after adding the CQT feature reconstruction training objective, during the training process of MAE, the original audio signal will be input into three branches respectively. The first branch is EnCodec, which is used to extract the quantization features quantized by the EnCodec quantizer as the label of the first loss; the second branch is the CQT feature extractor, which is used to extract the CQT features as pitch features as the label of the second loss; the third branch is MAE, which is used to extract the predicted features. In addition, the calculation of the first loss and the second loss is also required.
[0099] First, the feature extractor extracts features from the input original audio signal to obtain the mel spectrogram and inputs it into the feature transformer. The feature transformer transforms the feature dimension of the mel spectrogram to make it consistent with the feature dimension of MAE, so that it can be input into MAE for the next operation. MAE consists of an MAE encoder and an MAE decoder. The MAE encoder first adds position encoding to the input feature vector, then randomly selects a frame of the input feature vector according to a certain ratio and uses it as the masked frame, that is, directly discards the selected feature vector, and then inputs the un-discarded feature vectors into the MAE encoder to extract deep and fine features, and finally inputs them into the MAE decoder. The MAE decoder first restores the masked part of the input feature vector, that is, inserts learnable masked tokens at the masked positions, then adds position encoding to the restored feature vector, and finally inputs it into the MAE decoder. The output of the MAE decoder is the output of MAE.
[0100] Secondly, EnCodec consists of an EnCodec encoder and an EnCodec quantizer. It should be noted that during the training process of EnCodec, the EnCodec decoder is also required to participate. However, in the application scenario of EnCodec, what this application embodiment needs is the quantized features output by the EnCodec quantizer. Therefore, the EnCodec decoder is not required. In this application embodiment, the EnCodec encoder encodes the input original audio signal to obtain high-level fine features, that is, a series of feature vectors, and finally inputs them into the EnCodec quantizer. The main structure of the EnCodec quantizer is residual vector quantization (RVQ), which has n_q codebooks, and each codebook has n_bins vectors (i.e., embedding representations). RVQ first maps the input feature vectors (output by the EnCodec encoder) to the closest embedding representation in the first codebook to obtain the corresponding index, then calculates the residual between the embedding representation and the feature vector, and then uses the calculated residual as the input of the next codebook, repeating the above process to obtain a new residual. This process is repeated for each codebook until all codebooks are processed. The output of RVQ is the index corresponding to the embedding representation of each codebook. Therefore, EnCodec plays the role of encoding, quantizing, and compressing the original audio signal into a series of integer indexes, and during the training process of MAE, EnCodec is pre-trained and frozen.
[0101] The training method of MAE is clustering prediction. The label of each frame of feature vectors comes from the output of the EnCodec quantizer, that is, n_q indexes. During training, if the feature dimensions output by EnCodec and MAE are different, the predicted features output by MAE need to be transformed through a fully connected layer before they can be used to calculate the first loss together with the n_q indexes (i.e., the quantized features corresponding to each audio frame) corresponding to each audio frame.
[0102] Then, the calculation principle and calculation method of the CQT features in this application embodiment are described.
[0103] The frequency axis of the CQT features is distributed according to a logarithmic scale. The ratio of the frequency bandwidth of the filter to its center frequency is a constant Q. Since the frequency corresponding to the standard pitch of audio doubles every octave, the frequency axis distribution of the CQT features is consistent with the human auditory system's perception of frequency. After determining the number of semitones b included in each octave, the calculation formula of Q is as follows:
[0104]
[0105] Among them, the number of semitones b included in each octave can be set as needed. The larger the value of b, the more frequency details can be extracted.
[0106] After the value of Q is determined, the calculation formula for the window size corresponding to each frequency point is as follows, where k cq represents the frequency point of each CQT feature, and s represents the sampling rate.
[0107]
[0108] where, N kcq represents the window size corresponding to each frequency point, represents the frequency value of the frequency point of each CQT feature; ceil in the formula is a mathematical function representing rounding up (ceiling function), and the function is to round up a real number to the nearest integer.
[0109] The complete calculation formula for CQT features is as follows:
[0110]
[0111] where, X cq represents the feature value of the frequency point of each CQT feature; x[n] is the amplitude value of the nth sampling point of the audio signal; i represents the phase information of the audio signal.
[0112] The CQT feature describes the spectral information of an audio segment changing with time, where the horizontal axis represents time and the vertical axis represents frequency distribution. In practical applications, the embodiments of this application use the third-party Python library nnAudio to extract CQT features.
[0113] The MAE in the embodiments of this application downsamples the original audio signal by 320 times in the time dimension and upsamples it by 1024 times in the feature dimension. That is, when inputting a 1-second 24kHz mono audio, it outputs a 75-frame 1024-dimensional feature vector. To add the training objective of CQT feature reconstruction, it is necessary to make the number of frames of the CQT feature consistent with the number of frames output by EnCodecMAE, or in other words, make the features output by the MAE decoder and the CQT features extracted by nnAudio have the same dimension. The following parameters can be set when extracting CQT features to achieve this:
[0114] Sampling rate = 24000
[0115] Sampling step = 320
[0116] Is centered = Yes
[0117] Padding method = Constant padding
[0118] Among them, whether centered = yes indicates that the signal will be aligned in the time dimension. Specifically, by padding zeros at both ends of the signal, it is ensured that each time frame in the frequency domain transformation result is aligned with the center point of the audio signal, so that the frequency components correspond more precisely to the time information of the signal. When whether centered = yes, the audio signal also needs to be padded according to the indication of the padding method. It affects the padding effect of the signal boundary and the edge processing method of the signal to reduce distortion during the frequency transformation process. Among them, optional padding methods include but are not limited to reflection padding, constant padding, or copy padding, etc.
[0119] In addition, the following parameters also need to be set so that the number of frequency points of each audio frame (i.e., the number of frequency points of a CQT feature corresponding to each audio frame) n_bins is 336, the number of semitones included in each octave bins_per_octave is 48, and the lowest sampling frequency is the Hertz value corresponding to the C1 note:
[0120] n_bins = 336
[0121] bins_per_octave = 48
[0122] fmin = C1
[0123] After the above settings, the process of CQT feature extraction is to input a 1-second 24kHz mono audio, and output 75 frames of 336-dimensional CQT features, which is consistent with the output of MAE in the time dimension, and the feature dimension contains spectral information spanning 7 octaves.
[0124] Finally, it is necessary to calculate the CQT reconstruction loss based on both the masked part and the unmasked part in the predicted features output by MAE. The overall calculation formula of the CQT reconstruction loss is as follows:
[0125]
[0126] Among them, n bins represents the number of frequency points in each frame of CQT features, M represents the masked part, α represents the preset masking weight, β represents the preset visible weight, |M| represents the number of masked frames, T f represents the total number of audio frames in the training music, δ is an adjustable hyperparameter (which can be configured as needed), taking δ = 0.9. After δ is determined, α and β are also determined. represents the pitch feature of the k cq th frequency point of the tth frame. represents the predicted feature of the k cq th frequency point of the tth frame. represents the k cqFor the loss of each frequency point, the mean square error loss is adopted, and the second loss of the model is calculated by taking the square of the error between the pitch feature and the predicted feature.
[0127] In the application of MAE, the input target music only needs to pass through the primary feature extractor and the feature converter, then add the positional encoding and enter the MAE encoder. Finally, the features output by the MAE encoder are the deep features of the input music. It can be seen that there is no need for masking operation during the application process of MAE, nor is it necessary to calculate the first loss and the second loss.
[0128] In practical applications, the large audio representation model provides a powerful foundation for key application fields such as recommendation systems, search optimization, audio understanding, and speech synthesis by extracting and understanding the complex features of audio signals, greatly improving the performance of these tasks. The large audio representation model significantly enhances the performance of audio understanding tasks, including but not limited to speech recognition, speech emotion recognition, speaker recognition, and audio event detection. By capturing the nuances and context information of speech, the large audio representation model provides a richer and more accurate audio feature representation, thus significantly improving the accuracy and robustness of the audio understanding system. The large audio representation model has brought significant improvements to the field of speech synthesis, including improving the naturalness of the synthesized speech, achieving style and emotion transfer, optimizing the acoustic model, enhancing data diversity, improving the joint performance of speech recognition and synthesis, enhancing the personalization and context awareness of the system, and enhancing the expressiveness and accuracy of the speech synthesis system.
[0129] The technical solutions provided by the embodiments of this application will bring beneficial effects in the following aspects:
[0130] 1. Improve the spectral resolution ability: The extraction process of the CQT feature is constant-Q filtering, enabling the audio representation model to have better resolution ability in the high-frequency part, which helps to capture the details of the audio signal, especially the parts that are less important to the human auditory system, and is beneficial for the application of the model in a wider range of tasks and scenarios.
[0131] 2. Multi-sampling rate compatibility ability: The representation of the CQT feature enables the model to consistently process a wide range of audio signals from low frequency to high frequency, and can capture the key acoustic features regardless of the original sampling rate, while EnCodec is trained at a fixed sampling rate and does not have multi-sampling rate compatibility ability.
[0132] 3. Enhance the model generalization ability: The CQT feature has the general processing ability for different types of audio, has a good mathematical definition, and does not require training. As the amount of training data of MAE increases, it can output sufficiently robust features, thereby enhancing the generalization ability of the model.
[0133] Figure 4It is a schematic structural diagram of a computer device provided by an embodiment of the present application. The computer device 400 may include one or more central processing units (CPUs) 401 and a memory 405, and one or more application programs or data are stored in the memory 405.
[0134] Among them, the memory 405 may be volatile storage or persistent storage. The programs stored in the memory 405 may include one or more modules, and each module may include a series of instruction operations on the computer device. Further, the central processing unit 401 may be configured to communicate with the memory 405 and execute a series of instruction operations in the memory 405 on the computer device 400.
[0135] The computer device 400 may further include one or more power supplies 402, one or more wired or wireless network interfaces 403, one or more input / output interfaces 404, and / or one or more operating systems, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0136] The central processing unit 401 may execute the operations performed by the computer device in the foregoing Figures 1 to 3 illustrated embodiment, and details are not described herein again.
[0137] It should be noted that although the steps in the flowcharts involved in the embodiments are drawn in sequence according to the arrows, unless otherwise clearly stated in this article, the execution of these steps is not strictly limited in order, and these steps may be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but may be executed at different times. The execution order of these steps or stages is not necessarily sequential, but may be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.
[0138] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and details are not described herein again.
[0139] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0140] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0141] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0142] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, read-only memory), random access memories (RAM, random access memory), magnetic disks, or optical discs and other various media that can store program codes.
[0143] The embodiments of the present application also provide a computer program product containing instructions. When the computer program product runs on a computer, it enables the computer to execute the training method or audio processing method of the audio characterization model as described above.
Claims
1. A method for training an audio representation model, characterized in that: include: Inputting the training music into an initial audio representation model to obtain prediction features of the training music output by the initial audio representation model; Inputting the training music into a pre-trained audio quantization network, and using the output of the audio quantization network as the quantization feature of the training music; Extracting pitch features that conform to the note scale from the training music; Calculating the training loss of the training music based on the prediction feature, the quantization feature and the pitch feature; The model parameters are adjusted based on the training loss of the training music until a trained audio representation model is obtained, and the trained audio representation model is used to extract deep features of the music.
2. The method for training an audio representation model according to claim 1, characterized in that: The calculating the training loss of the training music based on the prediction feature, the quantization feature and the pitch feature includes: Calculating a first loss of the training music based on the predicted feature and the distance between the quantized feature; Calculating a second loss of the training music based on the predicted feature and the distance between the pitch feature; The training loss is determined based on the first loss, a preset loss weight corresponding to the first loss, the second loss, and the preset loss weight corresponding to the second loss.
3. The method for training an audio representation model according to claim 2, characterized in that: The training music includes a plurality of audio frames, the plurality of audio frames include at least one masked frame, the prediction feature of the training music includes the prediction feature of each of the audio frames, the pitch feature of the training music includes the pitch feature of each of the audio frames, and the second loss of the training music is calculated based on the prediction feature and the distance between the pitch features, including: For each pitch feature of the masked frame, based on the pitch feature of the masked frame and the predicted feature of the masked frame, calculating the pitch loss of the masked frame; Based on the pitch loss of each of the masked frames, a second loss of the training music is calculated.
4. The method for training an audio representation model according to claim 2, characterized in that: The training music includes a plurality of audio frames, the plurality of audio frames include at least one masked frame and at least one visible frame, the prediction feature of the training music includes the prediction feature of each of the audio frames, the pitch feature of the training music includes the pitch feature of each of the audio frames, and the second loss of the training music is calculated based on the distance between the prediction feature and the pitch feature, including: For each pitch feature of the audio frame, based on the pitch feature of the audio frame and the predicted feature of the audio frame, calculating the pitch loss of the audio frame; Determine a preset masking weight corresponding to the masked frame and a preset visible weight corresponding to the visible frame, wherein the preset masking weight is greater than the preset visible weight; A second loss of the training music is calculated based on the pitch loss of each of the masked frames, the preset masking weight, the pitch loss of each of the visible frames, and the preset visible weight.
5. The method for training an audio representation model according to claim 1, characterized in that: The step of extracting pitch features that conform to the note scale from the training music comprises: Acquire feature extraction conditions of the training music, wherein the feature extraction conditions include sampling rate, sampling step, audio filling strategy, the number of frequency points of each audio frame, the number of semitones contained in each octave, and the minimum sampling frequency; The feature extraction conditions and the training music are input into a CQT feature extractor to obtain the CQT features of each audio frame in the training music, and the CQT features of each audio frame in the training music are used as pitch features of the training music.
6. The method for training an audio representation model according to claim 1, characterized in that: The step of extracting pitch features that conform to the note scale from the training music comprises: Acquire feature extraction conditions of the training music, wherein the feature extraction conditions include sampling rate, sampling step, audio filling strategy, the number of frequency points of each audio frame, the number of semitones contained in each octave, and the minimum sampling frequency; Input the feature extraction condition and the training music into a CQT feature extractor to obtain a CQT feature of each audio frame in the training music; Based on the CQT feature of each audio frame in the training music, the pitch contour feature of each audio frame in the training music is calculated, and the pitch contour feature of each audio frame in the training music is used as the pitch feature of the training music.
7. The method for training an audio representation model according to any one of claims 1 to 6, characterized in that: The training music includes a plurality of music frames, the initial audio representation model includes a feature extractor, an initial masking encoder and an initial masking decoder, and the step of inputting the training music into the initial audio representation model to obtain the predicted features of the training music output by the initial audio representation model includes: Inputting the plurality of audio frames into the feature extractor to obtain an initial feature of each of the audio frames; Determine at least one masked frame from the plurality of audio frames, and determine audio frames other than the masked frame from the plurality of audio frames as visible frames; Input the initial features of each visible frame into the mask encoder to obtain the deep features of each visible frame output by the mask encoder; Inputting the frame index of each masked frame and the deep features of each visible frame into the masked decoder to obtain prediction features of the multiple audio frames, wherein the prediction features of each masked frame are preset masked features; The adjusting the model parameters based on the training loss of the training music until a trained audio representation model is obtained includes: Adjusting parameters of the initial masked encoder and parameters of the initial masked decoder based on the training loss; The feature extractor and the masked decoder with parameter adjustment are used as a trained audio representation model.
8. An audio processing method, characterized in that: include: Acquire a trained audio representation model, wherein the trained audio representation model is trained based on the training method for an audio representation model according to any one of claims 1 to 7, and the trained audio representation model includes a feature extractor and a masking encoder that completes parameter adjustment; Input the target music into the feature extractor to obtain the initial features of the target music, input the initial features into the masked encoder after parameter adjustment, and use the features output by the trained masked encoder as the deep features of the target music; Based on the deep features of the target music, a processing task for the target music is performed.
9. A computer device, characterized in that: include: CPU, memory and input / output interface; The memory is a short-term storage memory or a persistent storage memory; The central processing unit is configured to communicate with the memory and execute instruction operations in the memory to perform the training method of the audio representation model described in any one of claims 1 to 7 or the audio processing method described in claim 8.
10. A computer program product comprising instructions, characterized in that When the computer program product is run on a computer, the computer is enabled to execute the audio representation model training method according to any one of claims 1 to 7 or the audio processing method according to claim 8.
11. A computer storage medium, characterized in that: The computer storage medium stores instructions, and when the instructions are executed on a computer, the computer executes the training method for the audio representation model according to any one of claims 1 to 7 or the audio processing method according to claim 8.