Music Audio Restoration Method and System Based on DCT-DDPM
Through the DCT-DDPM-based music audio repair method, combined with audio and music score information, the problem of the inability to restore long gap audio in the existing technology is solved, and high-quality and robust audio repair effects are achieved, which is suitable for lightweight repair of music audio.
Patent Information
- Application Number
- CN202310105130.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-07
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2043-02-07
AI Technical Summary
The prior art cannot effectively restore the original clip information of long gap music audio, resulting in a sudden feeling among listeners, and the existing methods can only repair clips of fixed lengths, which are poorly robust and have poor quality of generated clips.
The music audio repair method based on DCT-DDPM is used to obtain the audio and music score, convert it into the Mel spectrogram using a short-time Fourier transform and a Mel filter, train it after adding Mask, combine the DCT-DDPM model of Unet structure to perform audio repair, and use a neural network or PGHI phase reconstruction algorithm to generate audio waveforms.
It realizes high-quality audio repair that is context-compatible, can repair gaps of any length, reduce abruptness, improves the robustness of the model and the authenticity of generated clips, and is suitable for lightweight environments.
Smart Images

Figure CN116072134B_ABST
Abstract
Description
Technical Field
[0001] A music audio repair method and system based on DCT-DDPM, which is used for music audio repair and belongs to the fields of speech processing and deep learning. Background Technique
[0002] During the transmission and storage of audio signals, parts of them are often damaged. For example, packet loss occurs during long-distance transmission of audio signals. When storing with media such as optical discs, if a part of the media is damaged, it will also cause local damage to the audio signal. And a study on recovering information from damaged audio segments (also called gaps) based on semantic information of audio context is called audio repair.
[0003] In the field of music audio repair, when the length of the gap where the music audio signal is damaged does not exceed 100 ms, the purpose of audio repair is to accurately restore the original signal according to context features. Existing studies in this regard include models based on sparsity, probabilistic non-negative matrix factorization, or neural networks. However, for long gaps of several hundred or even several thousand milliseconds (greater than 100 ms), it is unrealistic to accurately recover the lost information of the audio signal without additional information.
[0004] Therefore, the primary purpose of current related work for repairing long gaps is to generate segments that are compatible with the context semantics and insert them into the gaps so that people will not or will have as little a sense of abruptness as possible when hearing the audio. An existing study attempts to insert appropriate candidate segments into the gaps by using the information already existing in the audio signal. This method does not require collecting data for training, but its limitations are also great because it cannot generate new information, and the existing information may not be suitable for insertion in many cases. Another piece of work uses a generative adversarial network (GAN) containing a multi-scale context encoder to generate audio segments in the frequency domain. This method generates very poor-quality audio segments when the amount of training data is small, and at the same time, it can only repair gaps of fixed lengths, such as 480 ms and 960 ms, and it is difficult to be applied in practice. In addition, the existing audio repair work currently focuses on unconditional repair, that is, it can only generate new segments and cannot restore the information of the original segments.
[0005] In summary, the existing audio repair methods have the following technical problems:
[0006] 1. The problem of being able to only perform unconditional modification and unable to restore the information of the original segment, which will cause a great sense of discomfort when the listener hears familiar music;
[0007] 2. Traditional machine learning methods can only repair segments less than 100 ms, while existing deep learning methods can only repair segments of fixed lengths when repairing segments longer than 100 ms, and need to be retrained for segments of different lengths;
[0008] 3. The segments generated by existing machine learning and deep learning methods have poor quality, with a large gap in melody and clarity compared to the context segments.
[0009] 4. Existing machine learning and deep learning methods have poor robustness. When performing cross-dataset repair, even for the same instrument, if there are differences in the audio content features, the repair quality will be greatly reduced. Summary of the Invention
[0010] Aiming at the problems in the above research, the purpose of the present invention is to provide a music audio repair method and system based on DCT-DDPM, which solves the problem that the prior art can only perform unconditional modification and cannot restore the original segment information.
[0011] To achieve the above purpose, the present invention adopts the following technical solutions:
[0012] A music audio repair method based on DCT-DDPM includes the following steps:
[0013] Step 1: Obtain the audio of the music and the corresponding musical score of the audio, where the musical score is a MIDI file;
[0014] Step 2: Use the short-time Fourier transform and Mel filter banks to convert the audio into a Mel spectrogram, and evenly divide the Mel spectrogram. Convert the musical score into the pianoroll format, and after conversion, perform division. Align the divided Mel spectrogram of the audio with the musical score in the pianoroll format.
[0015] Step 3: Replace the random range of the Mel spectrogram with Gaussian noise as a Mask to obtain a Mel spectrogram with a Mask, where Mask represents occlusion;
[0016] Step 4: In the case of unconditional repair, splice the Mel spectrogram with a Mask and the Mel spectrogram obtained in Step 2 using a splicing function in the channel dimension and then send it to DCT-DDPM for training. Or in the case of conditional repair, extract the features of the musical score in the pianoroll format and then splice them with the Mel spectrogram with a Mask using a splicing function in the channel dimension and send it to DCT-DDPM for training. Finally, obtain the corresponding trained DCT-DDPM, where DCT-DDPM is a discrete cosine transform denoising diffusion probability model;
[0017] Step 5: After the DCT-DDPM training is completed, for the audio with gaps, after converting the audio to be repaired into a mel spectrogram to be repaired, in the case of unconditional repair, first randomly initialize a Gaussian noise with the same shape as the mel spectrogram to be repaired, and then concatenate the mel spectrogram to be repaired and the Gaussian noise in the channel dimension and send them into the DCT-DDPM to obtain a complete mel spectrogram. Or in the case of conditional repair, after using the score encoder to extract the score features of the score in Pianoroll format corresponding to the audio to be repaired, concatenate them with the masked mel spectrogram in the channel dimension and send them into the DCT-DDPM to obtain a complete mel spectrogram;
[0018] Step 6: Use a neural network vocoder or the PGHI phase reconstruction algorithm to invert the obtained complete mel spectrogram into an audio waveform.
[0019] Furthermore, in the above Step 1, the duration of the audio of the music is at least 0.5 hours, and the sampling rates of the obtained audio are all resampled to 24 kHz.
[0020] Furthermore, in the above Step 2, the parameters of the short-time Fourier transform are: win_size = 50 ms, fft_size = 50 ms, hop_size = 12.5 ms. Here, STFT represents the short-time Fourier transform, win_size represents the window size of the window function in STFT, fft_size represents how long the local data is used for Fourier transform, and hop_size represents the overlapping range of the window function when taking different windows. The values of hop_size, fft_size, and win_size result in the audio length corresponding to the mel spectrogram Figure 1 frame being 12.5 ms;
[0021] The segmentation length of the segmented mel spectrogram is 512 frames;
[0022] The score in Pianoroll format is segmented in the same way as the mel spectrogram, and the segmented score in Pianoroll format has the same shape as the segmented mel spectrogram. Among them, the segmented Pianoroll format file includes pretty-midi.
[0023] Furthermore, in the above Step 3, the value range of the random range is [0, 128], and the corresponding audio length is from 0 ms to 1600 ms.
[0024] Further, in the step 4, the structure of DCT-DDPM is a Unet structure, including an encoder and a decoder with symmetric structures. The encoder consists of an input convolution and 3 residual modules in sequence. The decoder consists of 3 residual modules and an output convolution in sequence. The last residual module of the encoder is connected to the first residual module of the decoder through another residual convolution module. The output of the i-th residual module of the encoder is skip-connected to the input of the (4 - i)-th residual module of the decoder, where 3 ≥ i ≥ 1. The skip connection means that the output of the i-th residual module of the encoder will be concatenated with the input of the (4 - i)-th residual module of the decoder in the channel dimension and then sent into the residual module of the decoder;
[0025] Each residual module in the encoder consists of a residual convolution module and a resampling layer connected in sequence, where the resampling layer in the encoder is downsampling. Each residual module in the decoder consists of a resampling layer and a residual convolution module connected in sequence, where the resampling layer in the decoder is upsampling;
[0026] There is a separate convolution at the first layer of the encoder and the last layer of the decoder, namely the input convolution and the output convolution, which are respectively used for: in the case of unconditional repair, the convolution at the first layer of the encoder changes the number of channels of the feature obtained by concatenating the mel spectrogram and the mel spectrogram with Mask in the channel dimension from 2 to 64, and the convolution at the last layer of the decoder changes the number of channels of the output data from 64 to 1; in the case of conditional repair, the score in pianoroll format is first input into the score encoder to obtain score features, then the convolution at the first layer of the encoder changes the number of channels of the feature obtained by concatenating the score features output by the score encoder and the mel spectrogram with Mask in the channel dimension from 2 to 64, and the convolution at the last layer of the decoder changes the number of channels of the output data from 64 to 1.
[0027] Each residual convolution module includes two parallel convolution modules. The two convolution modules respectively receive the input of the corresponding feature map and time embedding. The convolution module receiving the feature map input consists of a regularization function, an activation function, and a 1×3 convolution connected in sequence that receive the feature map input; while the convolution module receiving the time embedding input consists of an activation function, a Linear, and a 1×3 convolution connected in sequence that receive the time embedding input. The convolution of the two convolution modules in the encoder will double the channel dimension of the input feature, and in the decoder, it will halve the channel dimension of the input feature. Time embedding represents the time step embedding. When DDPM is training, it will add noise to the data according to the range, and the set value range is between [1, 1000];
[0028] The input feature map and the time embedding are respectively passed through the corresponding convolutional modules and then added together, and are successively input into the regularization function, the activation function, and the DCTS module to obtain the output feature map. Finally, the feature map output by the DCTS module is added to the feature map input to the residual convolution module as the final output feature map;
[0029] The DCTS module includes a DCT module, a 1×1 convolution, an activation function, and an iDCT connected in sequence. In the residual convolution, the stride of all convolutional kernels is 1, that is, the stride of the 1×1 convolution is 1. The DCT module represents the discrete cosine transform module, and the iDCT represents the inverse discrete cosine transform module;
[0030] The structure of the score encoder includes a 1×1 convolution, a Transformer layer, and a 1×1 convolution connected in sequence. The first convolution transforms the channel dimension of the score in the pianoroll format from 1 to 128, which is the input convolution, and the second convolution transforms the channel dimension of the features output by the Transformer from 128 to 1, which is the output convolution.
[0031] Furthermore, in step 6, if the obtained complete mel spectrogram is greater than 10 hours, the neural network vocoder is used, otherwise the PGHI phase reconstruction algorithm is used to invert the obtained complete mel spectrogram into a waveform;
[0032] The neural network vocoder is HifiGAN.
[0033] A music audio repair system based on DCT-DDPM
[0034] Acquisition module: Acquire the audio of the music and the score corresponding to the audio, where the score is a MIDI file;
[0035] Processing module: Use the short-time Fourier transform and the mel filter bank to convert the audio into a mel spectrogram, and evenly split the mel spectrogram. Convert the score into the pianoroll format, and after conversion, perform splitting, and align the mel spectrogram of the split audio with the score in the pianoroll format;
[0036] Mask module: Replace a random range of the mel spectrogram with Gaussian noise as the Mask to obtain a mel spectrogram with Mask, where Mask represents occlusion;
[0037] Training module: In the case of unconditional repair, the masked Mel spectrogram and the Mel spectrogram obtained in step 2 are concatenated in the channel dimension using the concatenation function (torch.cat) and then fed into DCT-DDPM for training. Or in the case of conditional repair, the musical score features in Pianoroll format are extracted and then concatenated with the masked Mel spectrogram in the channel dimension using the concatenation function (torch.cat) and then fed into DCT-DDPM for training. Finally, the corresponding trained DCT-DDPM is obtained, where DCT-DDPM is the discrete cosine transform denoising diffusion probabilistic model;
[0038] Repair module: When DCT-DDPM is trained, for the audio with gaps, after converting the audio to be repaired into a Mel spectrogram to be repaired, in the case of unconditional repair, first randomly initialize a Gaussian noise with the same shape as the Mel spectrogram to be repaired, and then concatenate the Mel spectrogram to be repaired and the Gaussian noise in the channel dimension and feed them into DCT-DDPM to obtain a complete Mel spectrogram. Or in the case of conditional repair, after using the musical score encoder to extract the musical score features of the Pianoroll format musical score corresponding to the audio to be repaired, then concatenate them with the masked Mel spectrogram in the channel dimension and feed them into DCT-DDPM to obtain a complete Mel spectrogram;
[0039] Audio waveform processing module: Use a neural network vocoder or the PGHI phase reconstruction algorithm to invert the obtained complete Mel spectrogram into an audio waveform.
[0040] Furthermore, in the acquisition module, the duration of the audio of the music is at least 0.5 hours, and the sampling rates of the acquired audio are all resampled to 24 kHz;
[0041] In the processing module, the parameters of the short-time Fourier transform are: win_size = 50 ms, fft_size = 50 ms, hop_size = 12.5 ms, where STFT represents the short-time Fourier transform, win_size represents the window size of the window function in STFT, fft_size represents how long the local data is used for Fourier transform, and hop_size represents the overlapping range of the window function when taking different windows. The values of hop_size, fft_size, and win_size result in the audio length corresponding to the Mel spectrogram frame being 12.5 ms; Figure 1 The audio length corresponding to the Mel spectrogram frame is 12.5 ms;
[0042] The segmentation length of the segmented Mel spectrogram is 512 frames;
[0043] The music score in pianoroll format is segmented in the same way as the mel spectrogram, and the shape of the segmented music score in pianoroll format is the same as that of the segmented mel spectrogram. Among them, the file for segmenting the pianoroll format includes pretty-midi;
[0044] In the Mask module, the value range of the random range is [0, 128], and the corresponding audio length is from 0 ms to 1600 ms.
[0045] Furthermore, in the training module, the structure of DCT-DDPM is a Unet structure, including an encoder and a decoder with symmetric structures. The encoder is composed of an input convolution and 3 residual modules in sequence, and the decoder is composed of 3 residual modules and an output convolution in sequence. The last residual module of the encoder and the first residual module of the decoder are connected by another residual convolution module. The output of the i-th residual module of the encoder and the input of the (4 - i)-th residual module of the decoder are skip-connected, where 3 ≥ i ≥ 1. Skip connection means that the output of the i-th residual module of the encoder will be concatenated with the input of the (4 - i)-th residual module of the decoder in the channel dimension and then sent into the residual module of the decoder;
[0046] Each residual module in the encoder is composed of a residual convolution module and a resampling layer connected in sequence. Among them, the resampling layer in the encoder is downsampling. Each residual module in the decoder is composed of a resampling layer and a residual convolution module connected in sequence. Among them, the resampling layer in the decoder is upsampling;
[0047] There is a separate convolution in the first layer of the encoder and the last layer of the decoder, namely the input convolution and the output convolution, which are respectively used for: in the case of unconditional restoration, the convolution in the first layer of the encoder changes the number of channels of the feature obtained by concatenating the mel spectrogram and the mel spectrogram with Mask in the channel dimension from 2 to 64, and the convolution in the last layer of the decoder changes the number of channels of the output data from 64 to 1; in the case of conditional restoration, the music score in pianoroll format is first input into the music score encoder to obtain music score features, and then the convolution in the first layer of the encoder changes the number of channels of the feature obtained by concatenating the music score features output by the music score encoder and the mel spectrogram with Mask in the channel dimension from 2 to 64, and the convolution in the last layer of the decoder changes the number of channels of the output data from 64 to 1.
[0048] Each residual convolution module includes two parallel convolution modules. The two convolution modules respectively receive the input of the corresponding feature map and time embedding. The convolution module that receives the feature map input consists of a regularization function, an activation function, and a 1×3 convolution connected in sequence to receive the feature map input; while the convolution module that receives the time embedding input consists of an activation function, a Linear, and a 1×3 convolution connected in sequence to receive the time embedding input. The convolutions of the two convolution modules in the encoder will double the channel dimension of the input features, and in the decoder, will halve the channel dimension of the input features. Time embedding represents the time step embedding. When training DDPM, noise is added to the data according to a range, and the set value range is between [1, 1000];
[0049] The input feature map and time embedding are added after passing through the corresponding convolution modules respectively, and are then input into the regularization function, activation function, and DCTS module in sequence to obtain the output feature map. Finally, the feature map output by the DCTS module is added to the feature map input to the residual convolution module as the final output feature map;
[0050] The DCTS module includes a DCT module, a 1×1 convolution, an activation function, and an iDCT connected in sequence. In the residual convolution, the stride of all convolution kernels is 1, that is, the stride of the 1×1 convolution is 1. iDCT represents the inverse discrete cosine transform module;
[0051] The structure of the music score encoder includes a 1×1 convolution, a Transformer layer, and a 1×1 convolution connected in sequence. The first convolution changes the channel dimension of the music score in pianoroll format from 1 to 128, which is the input convolution, and the second convolution changes the channel dimension of the features output by the Transformer from 128 to 1, which is the output convolution.
[0052] Further, in the audio waveform processing module, if the obtained complete Mel spectrogram is greater than 10 hours, a neural network vocoder is used, otherwise the PGHI phase reconstruction algorithm is used to invert the obtained complete Mel spectrogram into a waveform;
[0053] The neural network vocoder is HifiGAN.
[0054] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0055] First, the present invention uses DCT-DDPM to repair music audio. Compared with the GAN used in current related work, the training is more stable, the generated audio segments are more realistic and natural, can be compatible with the signal context, and are conducive to deployment in an environment that requires a lightweight model;
[0056] II. In combination with Mel-spectrum features, the present invention proposes a model DCT-DDPM based on the fast Fourier transform structure, reducing the model parameters of DDPM to 10M, with much reduced modeling ability, and being able to ensure the quality of restoration. That is, due to the addition of the DCTS (DCT Structure) module, which uses DCT and iDCT modules that do not require training, when the feature map is input, the global receptive field can be obtained through transformation, and at the same time, the frequency information of the Mel-spectrum image can be captured. DCTS can well ensure the generation quality of the model without increasing parameters; in addition, as a frequency-domain transformation method, the output obtained after DCT transformation is still a real value, which makes it easier for deep learning frameworks such as pytorch to process.
[0057] III. When processing audio data, the present invention adds random masks to enhance the robustness of the model (DCT-DDPM), enabling the model to repair music audio gaps of random lengths after training. Compared with the current work that can only repair gaps of fixed lengths, the application value has been greatly improved;
[0058] IV. The present invention can fuse the score features corresponding to the audio, generate segments similar to the information of the original audio segment, and can greatly reduce the sense of abruptness when people hear the repaired segment, especially when repairing well-known music;
[0059] V. The present invention integrates the FFTS and Unet structures, enabling a lightweight model to have good restoration quality while increasing the inference speed of the model;
[0060] VI. When repairing gaps, the present invention has no requirements on the position of the gaps and can handle gaps at the beginning or end of the segment, having a wider application scenario compared with previous work. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 is a schematic flow diagram of the present invention;
[0062] Figure 2 is the specific process of the present invention, including unconditional restoration and conditional restoration;
[0063] Figure 3 is a schematic diagram of the audio segment of the music in the present invention and its corresponding pianoroll matrix;
[0064] Figure 4 is the model structure of DCT-DDPM in the present invention;
[0065] Figure 5This is a rendering generated by the present invention, wherein the figure shows an example of a mel-spectrogram restoration of music with a duration of 6.4 seconds. The white strip portion represents the restoration area, and the length of the restoration area is 1.6 seconds. DETAILED DESCRIPTION
[0066] The present invention will be further described below in conjunction with the accompanying drawings and specific implementation methods.
[0067] The invention provides a music audio restoration method based on DCT-DDPM, which solves the problem that the restoration quality of existing research is poor and can only repair fixed-length fragments, and can restore the information of the original fragment according to the music score.
[0068] The main process of the present invention includes: 1) obtaining music audio data; 2) transforming audio data into frequency domain; 3) processing to obtain Mel-spectrogram with Mask; 4) training DCT-DDPM; 5) repairing audio based on the trained DCT-DDPM; 6) transforming the repaired Mel-spectrogram into time domain. The specific implementation steps are as follows:
[0069] 1. Acquisition of Music Audio Data
[0070] Obtain music audio data, such as single-instrument performance audio such as piano and violin, or multi-instrument performance audio such as string quartet. The required audio data duration is at least 0.5 hours. The acquired audio sampling rate is resampled to 24kHZ. If conditions permit, the corresponding music score (MIDI file) of the audio clip must also be obtained.
[0071] 2. Transform audio data into frequency domain
[0072] Use STFT (Short Time Fourier Transform) and Mel filter to convert the audio into Mel spectrum, and divide the Mel spectrum into certain lengths on average. The detailed parameters of STFT are: hop_size = 12.5ms, fft_size = 40ms, win_size = 12.5ms, so the Mel spectrum obtained is Figure 1 The audio length corresponding to the frame is 12.5ms. The segmentation length when segmenting the Mel spectrum is 512 frames. The shape of the final Mel spectrum graph is (1, 80, 512). In the case of conditional repair, the MIDI file needs to be converted into a pianoroll format score and segmented. After segmentation, the Mel spectrum graph of the segmented audio is aligned. Pianoroll is a format that uses a matrix to represent music score information, and its shape is exactly the same as the Mel spectrum graph. After segmentation in the same way as the Mel spectrum graph, the shape is also (1, 128, 512), where 128 represents the pitch;
[0073] Obtain the Mel spectrogram with Mask. Replace the random range of the Mel spectrogram with Gaussian noise as the Mask to obtain the Mel spectrogram with Mask, where Mask represents occlusion, that is, randomly crop gaps in the Mel spectrogram obtained by converting the audio and fill the gaps with Gaussian noise. The range of the gaps is [0, 128], and the corresponding audio length is from 0 ms to 1600 ms.
[0074] IV. Training DCT-DDPM
[0075] Feed the Mel spectrogram with Mask and the Mel spectrogram obtained by converting the audio into DCT-DDPM for training after concatenating them in the channel dimension using the concatenation function.
[0076] The structure of DCT-DDPM is a Unet structure, and its Encoder and Decoder structures are symmetric. The encoder consists of an input convolution and 3 residual modules in sequence, and the decoder consists of 3 residual modules and an output convolution in sequence. The last residual module of the encoder is connected to the first residual module of the decoder through another residual convolution module. The output of the i-th residual module of the encoder is skip-connected to the input of the (4 - i)-th residual module of the decoder, where 3 ≥ i ≥ 1. Skip connection means that the output of the i-th residual module of the encoder will be concatenated with the input of the (4 - i)-th residual module of the decoder in the channel dimension and then fed into the residual module of the decoder.
[0077] Each residual module in the encoder consists of a residual convolution module and a resampling layer connected in sequence, where the resampling layer in the encoder is downsampling. Each residual module in the decoder consists of a resampling layer and a residual convolution module connected in sequence, where the resampling layer in the decoder is upsampling.
[0078] There is a separate convolution in the first layer of the encoder and the last layer of the decoder, namely the input convolution and the output convolution, which are respectively used for: in the case of unconditional restoration, the convolution in the first layer of the encoder changes the number of channels of the features obtained by concatenating the Mel spectrogram and the Mel spectrogram with Mask in the channel dimension from 2 to 64, and the convolution in the last layer of the decoder changes the number of channels of the output data from 64 to 1; in the case of conditional restoration, the pianoroll-format score is first input into the score encoder to obtain score features, then the convolution in the first layer of the encoder changes the number of channels of the features obtained by concatenating the score features output by the score encoder and the Mel spectrogram with Mask in the channel dimension from 2 to 64, and the convolution in the last layer of the decoder changes the number of channels of the output data from 64 to 1.
[0079] Each residual convolution module includes two parallel convolution modules. The two convolution modules respectively receive the input of the corresponding feature map and time embedding. The convolution module that receives the feature map input consists of a regularization function, an activation function, and a 1×3 convolution that are connected in sequence and receive the feature map input. The convolution module that receives the time embedding input consists of an activation function, a Linear, and a 1×3 convolution that are connected in sequence and receive the time embedding input. In the encoder, the convolutions of the two convolution modules will double the channel dimension of the input features, and in the decoder, they will halve the channel dimension of the input features. Time embedding represents the time step embedding. When DDPM is training, it will add noise to the data according to a range, and the set value range is between [1, 1000].
[0080] The input feature map and time embedding are added after passing through their corresponding convolution modules respectively, and then are sequentially input into a regularization function, an activation function, and a DCTS module to obtain the output feature map. Finally, the feature map output by the DCTS module is added to the feature map input to the residual convolution module as the final output feature map.
[0081] The DCTS module includes a DCT module, a 1×1 convolution, an activation function, and an iDCT that are connected in sequence. In the residual convolution, the stride of all convolution kernels is 1, that is, the stride of the 1×1 convolution is 1. The DCT module represents the discrete cosine transform module, and the iDCT represents the inverse discrete cosine transform module.
[0082] The structure of the score encoder includes a 1×1 convolution, a Transformer layer, and a 1×1 convolution that are connected in sequence. The first convolution changes the channel dimension of the score in the pianoroll format from 1 to 128, which is the input convolution. The second convolution changes the channel dimension of the features output by the Transformer from 128 to 1, which is the output convolution.
[0083] V. Repair the audio based on the trained DCT-DDPM
[0084] When DCT-DDPM is trained, for the audio with gaps, after converting the audio to be repaired into the mel spectrogram to be repaired, in the case of unconditional repair, first randomly initialize a Gaussian noise with the same shape as the mel spectrogram to be repaired, and then concatenate the mel spectrogram to be repaired and the Gaussian noise in the channel dimension and send them into DCT-DDPM to obtain the complete mel spectrogram. Or in the case of conditional repair, concatenate the score features extracted from the score in the pianoroll format corresponding to the audio to be repaired and the mel spectrogram with a Mask in the channel dimension and send them into DCT-DDPM to obtain the complete mel spectrogram.
[0085] VI. Transforming the Mel spectrogram into the time domain
[0086] After obtaining the repaired Mel spectrogram in step 5, use a neural network vocoder or a PGHI phase reconstruction algorithm vocoder to invert the Mel spectrogram into an audio waveform. The neural network vocoder can be selected as HifiGAN, etc. Since the neural network vocoder requires a certain amount of data for training, it is recommended to use the neural network vocoder when the data volume is greater than 10 hours, otherwise use the PGHI phase reconstruction algorithm to invert the Mel spectrogram into a waveform.
[0087] Currently, the mainstream method in the field of music audio repair is the deep learning method. As a generative model that has gradually received attention in the image field recently, DDPM has the characteristics of stable training and high generation quality compared to GAN. Applying DCT-DDPM to the field of music audio repair can obtain a better repair effect than GAN. At the same time, this method proposes a way to fuse audio score information. Given the score corresponding to the audio segment, it can generate a new segment that highly restores the original music segment. Compared with traditional machine learning algorithms or generative adversarial network methods, this method can repair longer gaps. At the same time, this method can generate a new audio segment that highly restores the original segment according to the score information of the damaged segment, which is not available in previous methods.
[0088] The above are only representative embodiments among the numerous specific application scopes of the present invention, and do not constitute any limitation to the protection scope of the present invention. Any technical solutions formed by transformation or equivalent replacement fall within the scope of the rights protection of the present invention.
Claims
1. A music audio repair method based on DCT-DDPM, characterized in that, It includes the following steps: Step 1: Obtain the audio of the music and the musical score corresponding to the audio, where the musical score is a MIDI file; Step 2: Use the short-time Fourier transform and Mel filter banks to convert the audio into a Mel spectrogram, and evenly segment the Mel spectrogram. Convert the musical score into the pianoroll format, and after conversion, segment it. Align the segmented Mel spectrogram of the audio with the musical score in the pianoroll format; Step 3: Replace the random range of the Mel spectrogram with Gaussian noise as a Mask to obtain a Mel spectrogram with a Mask, where Mask represents occlusion; Step 4: In the case of unconditional restoration, splice the Mel spectrogram with a Mask and the Mel spectrogram obtained in Step 2 using a splicing function in the channel dimension and then send it to DCT-DDPM for training. In the case of conditional restoration, extract the features of the musical score in the pianoroll format corresponding to the audio to be restored, and then splice it with the Mel spectrogram with a Mask using a splicing function in the channel dimension and send it to DCT-DDPM for training. Finally, obtain the corresponding trained DCT-DDPM, where DCT-DDPM is a discrete cosine transform denoising diffusion probabilistic model; Step 5: After the DCT-DDPM training is completed, for the audio with gaps, convert the audio to be restored into a Mel spectrogram to be restored. In the case of unconditional restoration, first randomly initialize a Gaussian noise with the same shape as the Mel spectrogram to be restored, and then splice the Mel spectrogram to be restored and the Gaussian noise in the channel dimension and send it to DCT-DDPM to obtain a complete Mel spectrogram. In the case of conditional restoration, after using a musical score encoder to extract the features of the musical score in the pianoroll format corresponding to the audio to be restored, splice it with the Mel spectrogram with a Mask in the channel dimension and send it to DCT-DDPM to obtain a complete Mel spectrogram; Step 6: Use a neural network vocoder or the PGHI phase reconstruction algorithm to invert the obtained complete Mel spectrogram into an audio waveform.
2. The music audio restoration method based on DCT-DDPM according to claim 1, wherein: In Step 1, the duration of the audio of the music is at least 0.5 hours, and the sampling rate of the obtained audio is resampled to 24 kHz.
3. The method for restoring music audio based on DCT-DDPM according to claim 2, characterized in that: In Step 2, the parameters of the short-time Fourier transform are: win_size = 50 ms, fft_size = 50 ms, hop_size = 12.5 ms, where STFT represents the short-time Fourier transform, win_size represents the window size of the window function in STFT, fft_size represents how long the local data is used for Fourier transform, and hop_size represents the overlapping range of the window function when taking different windows. The values of hop_size, fft_size, and win_size result in an audio length of 12.5 ms corresponding to one frame of the Mel spectrogram; The segmentation length of the segmented Mel spectrogram is 512 frames; The musical score in pianoroll format is segmented in the same way as the mel spectrogram, and the shape of the pianoroll format musical score after segmentation is the same as that of the mel spectrogram after segmentation. Among them, the file for segmenting the pianoroll format includes pretty-midi.
4. A music audio restoration method based on DCT-DDPM according to claim 3, characterized in that: In step 3, the value range of the random range is [0, 128], and the corresponding audio length is from 0 ms to 1600 ms.
5. A method for repairing music audio based on DCT-DDPM according to claim 4, characterized in that: In step 4, the structure of DCT-DDPM is a Unet structure, including an encoder and a decoder with symmetric structures. The encoder is successively composed of an input convolution and 3 residual modules, and the decoder is successively composed of 3 residual modules and an output convolution. The last residual module of the encoder and the first residual module of the decoder are connected by another residual convolution module. The output of the i-th residual module of the encoder and the input of the (4 - i)-th residual module of the decoder are skip-connected, where 3 ≥ i ≥ 1. The skip connection means that the output of the i-th residual module of the encoder will be concatenated with the input of the (4 - i)-th residual module of the decoder in the channel dimension and then sent into the residual module of the decoder. Each residual module in the encoder is composed of a residual convolution module and a resampling layer connected in sequence. Among them, the resampling layer in the encoder is downsampling, and each residual module in the decoder is composed of a resampling layer and a residual convolution module connected in sequence. Among them, the resampling layer in the decoder is upsampling. There is a separate convolution in the first layer of the encoder and the last layer of the decoder, namely the input convolution and the output convolution, which are respectively used for: in the case of unconditional repair, the convolution in the first layer of the encoder changes the number of channels of the feature obtained by concatenating the mel spectrogram and the mel spectrogram with Mask in the channel dimension from 2 to 64, and the convolution in the last layer of the decoder changes the number of channels of the output data from 64 to 1; in the case of conditional repair, the musical score in pianoroll format is first input into the musical score encoder to obtain musical score features, and then the convolution in the first layer of the encoder changes the number of channels of the feature obtained by concatenating the musical score features output by the musical score encoder and the mel spectrogram with Mask in the channel dimension from 2 to 64, and the convolution in the last layer of the decoder changes the number of channels of the output data from 64 to 1. Each residual convolution module includes two parallel convolution modules. The two convolution modules respectively receive the input of the corresponding feature map and time embedding. The convolution module that receives the feature map input is composed of a regularization function, an activation function, and a 1×3 convolution that are connected in sequence to receive the feature map input. The convolution module that receives the time embedding input is composed of an activation function, a Linear, and a 1×3 convolution that are connected in sequence to receive the time embedding input. In the encoder, the convolution of the two convolution modules doubles the channel dimension of the input features. In the decoder, it halves the channel dimension of the input features. Time embedding represents the time step embedding. When DDPM is training, it adds noise to the data according to a range, and the set value range is between [1, 1000]. The input feature map and time embedding are added after passing through the corresponding convolution modules respectively, and then input into the regularization function, activation function, and DCTS module in sequence to obtain the output feature map. Finally, the feature map output by the DCTS module is added to the feature map input to the residual convolution module as the final output feature map. The DCTS module includes a DCT module, a 1×1 convolution, an activation function, and an iDCT that are connected in sequence. In the residual convolution, the stride of all convolution kernels is 1, that is, the stride of the 1×1 convolution is 1. DCTS represents the discrete cosine transform structure, the DCT module represents the discrete cosine transform module, and iDCT represents the inverse discrete cosine transform module. The structure of the score encoder includes a 1×1 convolution, a Transformer layer, and a 1×1 convolution that are connected in sequence. The first convolution changes the channel dimension of the score in the pianoroll format from 1 to 128, which is the input convolution. The second convolution changes the channel dimension of the features output by the Transformer from 128 to 1, which is the output convolution.
6. The music audio restoration method based on DCT-DDPM according to claim 5, wherein: In step 6, if the obtained complete mel spectrogram is more than 10 hours, use the neural network vocoder, otherwise use the PGHI phase reconstruction algorithm to invert the obtained complete mel spectrogram into a waveform. The neural network vocoder is HifiGAN.
7. A music audio repair system based on DCT-DDPM, characterized in that: Acquisition module: Acquire the audio of the music and the corresponding score of the audio, where the score is a MIDI file. Processing module: Use the short-time Fourier transform and Mel filter to convert the audio into a mel spectrogram, and evenly split the mel spectrogram. Convert the score into the pianoroll format, and after conversion, perform splitting, and align the split mel spectrogram of the audio and the score in the pianoroll format. Mask module: Replace the random range of the mel spectrogram with Gaussian noise as the Mask to obtain the mel spectrogram with Mask, where Mask represents occlusion. Training Module: In the case of unconditional repair, the Mel spectrogram with Mask and the Mel spectrogram obtained in Step 2 are concatenated in the channel dimension using a concatenation function and then fed into DCT-DDPM for training. In the case of conditional repair, the musical score features in Pianoroll format are extracted and then concatenated with the Mel spectrogram with Mask in the channel dimension using a concatenation function and then fed into DCT-DDPM for training. Finally, the corresponding trained DCT-DDPM is obtained, where DCT-DDPM is a discrete cosine transform denoising diffusion probability model; Repair Module: When DCT-DDPM training is completed, for the audio with gaps, after converting the audio to be repaired into a Mel spectrogram to be repaired, in the case of unconditional repair, first randomly initialize a Gaussian noise with the same shape as the Mel spectrogram to be repaired, and then concatenate the Mel spectrogram to be repaired and the Gaussian noise in the channel dimension and feed them into DCT-DDPM to obtain a complete Mel spectrogram. In the case of conditional repair, after using a musical score encoder to extract the musical score features of the Pianoroll format musical score corresponding to the audio to be repaired, then concatenate them with the Mel spectrogram with Mask in the channel dimension and feed them into DCT-DDPM to obtain a complete Mel spectrogram; Audio Waveform Processing Module: Use a neural network vocoder or the PGHI phase reconstruction algorithm to invert the obtained complete Mel spectrogram into an audio waveform.
8. A music audio repair system based on DCT-DDPM according to claim 7, characterized in that: In the acquisition module, the duration of the music audio is at least 0.5 hours, and the sampling rates of the acquired audio are all resampled to 24 kHz; In the processing module, the parameters of the short-time Fourier transform are: win_size = 50 ms, fft_size = 50 ms, hop_size = 12.5 ms, where STFT represents the short-time Fourier transform, win_size represents the window size of the window function in STFT, fft_size represents how long the local data is used for Fourier transform, and hop_size represents the overlapping range of the window function when taking different windows. The values of hop_size, fft_size, and win_size result in an audio length of 12.5 ms corresponding to one frame of the Mel spectrogram; The segmentation length of the segmented Mel spectrogram is 512 frames; The Pianoroll format musical score is segmented in the same way as the Mel spectrogram. The segmented Pianoroll format musical score has the same shape as the segmented Mel spectrogram. Among them, the segmented Pianoroll format files include pretty-midi; In the Mask module, the value range of the random range is [0, 128], and the corresponding audio length is from 0 ms to 1600 ms.
9. The music audio repair system based on DCT-DDPM according to claim 8, characterized in that: In the training module, the structure of DCT-DDPM is a Unet structure, including an encoder and a decoder with symmetrical structures. The encoder consists of an input convolution and three residual modules in sequence, and the decoder consists of three residual modules and an output convolution in sequence. The last residual module of the encoder is connected to the first residual module of the decoder through another residual convolution module. The output of the i-th residual module of the encoder is skip-connected to the input of the (4 - i)-th residual module of the decoder, where 3 ≥ i ≥ 1. The skip connection means that the output of the i-th residual module of the encoder will be concatenated with the input of the (4 - i)-th residual module of the decoder in the channel dimension and then sent into the residual module of the decoder; Each residual module in the encoder consists of a residual convolution module and a resampling layer connected in sequence. Among them, the resampling layer in the encoder is downsampling. Each residual module in the decoder consists of a resampling layer and a residual convolution module connected in sequence. Among them, the resampling layer in the decoder is upsampling; There is a separate convolution at the first layer of the encoder and the last layer of the decoder, namely the input convolution and the output convolution, which are respectively used for: in the case of unconditional restoration, the convolution of the first layer of the encoder changes the number of channels of the feature obtained by concatenating the mel spectrogram and the mel spectrogram with Mask in the channel dimension from 2 to 64, and the convolution of the last layer of the decoder changes the number of channels of the output data from 64 to 1; in the case of conditional restoration, the score in pianoroll format is first input into the score encoder to obtain score features, and then the convolution of the first layer of the encoder changes the number of channels of the feature obtained by concatenating the score features output by the score encoder and the mel spectrogram with Mask in the channel dimension from 2 to 64, and the convolution of the last layer of the decoder changes the number of channels of the output data from 64 to 1; Each residual convolution module includes two parallel convolution modules. The two convolution modules respectively receive the input of the corresponding feature map and time embedding. The convolution module receiving the feature map input consists of a regularization function, an activation function, and a 1×3 convolution connected in sequence; while the convolution module receiving the time embedding input consists of an activation function, a Linear, and a 1×3 convolution connected in sequence. The convolutions of the two convolution modules in the encoder will double the channel dimension of the input feature, and in the decoder, they will halve the channel dimension of the input feature. Time embedding represents the time step embedding. DDPM needs to add noise to the data according to a range during training, and the set value range is between [1, 1000]; The input feature map and time embedding are added after passing through the corresponding convolution modules respectively, and then input into the regularization function, activation function, and DCTS module in sequence to obtain the output feature map. Finally, the feature map output by the DCTS module will be added to the feature map of the input residual convolution module as the final output feature map; The DCTS module includes a DCT module, a 1×1 convolution, an activation function, and an iDCT that are connected in sequence. In the residual convolution, the stride of all convolutional kernels is 1, that is, the stride of the 1×1 convolution is 1. The DCT module represents the discrete cosine transform module, the iDCT represents the inverse discrete cosine transform module, and the DCTS represents the discrete cosine transform structure; The structure of the music score encoder includes a 1×1 convolution, a Transformer layer, and a 1×1 convolution that are connected in sequence. The first convolution changes the channel dimension of the music score in pianoroll format from 1 to 128, which is the input convolution. The second convolution changes the channel dimension of the features output by the Transformer from 128 to 1, which is the output convolution.
10. A music audio restoration system based on DCT-DDPM according to claim 9, characterized in that: In the audio waveform processing module, if the obtained complete mel spectrogram is greater than 10 hours, a neural network vocoder is used; otherwise, the PGHI phase reconstruction algorithm is used to invert the obtained complete mel spectrogram into a waveform; The neural network vocoder is HifiGAN.
Citation Information
Patent Citations
Film restoration method and device based on speech synthesis, equipment and medium
CN113345414A
Missing audio automatic restoration method based on U-Net model
CN114373469A