Audio missing fragment repairing method

By combining time-frequency domain analysis and potential diffusion probability model, using vector quantized variational autoencoder and denoising network, the problems of incoherence and poor noise environment repair in the prior art are solved, and high-quality speech repair is achieved.

CN120299467APending Publication Date: 2025-07-11DALIAN MARITIME UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510284304.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art is difficult to effectively repair longer speech missing fragments and are poorly robust in noise environments, and the repaired speech fragments are incoherent with the previous and subsequent texts.

Method used

Combining the time-frequency domain analysis of the signal and the potential diffusion probability model, the variable autoencoder and denoising network are quantized by vector quantization, using the Mel spectrum for repair, and the signal phase is repaired using the HiFiGAN vocoder.

Benefits of technology

Effective repair of longer speech missing fragments is achieved, ensuring the consistency of the repair content and the previous and subsequent speech fragments, and improving the robustness in a noisy environment, improving the subjective perception quality and objective intelligibility of the repair signal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299467A_ABST
    Figure CN120299467A_ABST
Patent Text Reader

Abstract

The invention provides an audio missing fragment repairing method, which comprises the following steps of S1, performing short-time Fourier transform on a damaged audio signal to obtain an STFT amplitude spectrum, and converting the STFT amplitude spectrum into a Mel scale through a frequency filter bank to obtain a Mel spectrum of the damaged signal; s2, using the trained vector quantization variational auto-encoder to encode the Mel spectrum of the damaged signal into a potential spatial feature with a relatively low dimension; s3, applying a diffusion model method, taking the damaged potential spatial features as conditions, splicing sampled Gaussian noise, performing T-step denoising through the trained denoising network, and outputting predicted complete potential spatial features; and S4, decoding the predicted complete potential features through an auto-encoder to obtain a predicted complete Mel spectrum, repairing a signal phase by using a sensing HiFiGAN vocoder, and outputting a repaired voice signal. According to the invention, natural and coherent restoration of long voice missing segments is realized, and the robustness of the method in a noise environment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of maritime voice communication. Specifically, it particularly relates to a method for repairing missing audio segments. Background Art

[0002] In Very High Frequency (VHF) maritime voice communication, due to the complex and changeable maritime environment, local missing of voice segments often exists in audio signals, and various noises are mixed in at the same time. Therefore, by designing relevant methods to compensate for the missing voice segments and suppress the signal background noise to improve the audio quality, the communication quality between ships will be significantly improved, ensuring stable, efficient, and reliable maritime communication.

[0003] The existing methods for repairing missing audio segments are mainly divided into two categories: one is traditional repair methods based on sparse representation, autoregressive modeling, etc., and the other is repair methods based on deep neural networks with data-driven as the core. Traditional repair methods can usually accurately repair short-duration voice missing segments, but cannot effectively repair long voice missing segments. On the other hand, although the repair methods based on deep learning can achieve the repair of long voice missing segments, they cannot guarantee the coherence of the repaired voice segments with the context, and there are often obvious artificial traces. In addition, the existing methods for repairing missing audio segments are usually based on clean voice signals, while in the actual communication environment, voice signals not only encounter information loss of local segments, but also various noises are mixed in at the same time, thus significantly affecting the performance of the methods.

[0004] The existing methods for repairing missing voice segments can usually effectively repair short-duration voice missing segments (0 - 50 ms), but the repair quality for long-time voice missing segments (100 ms and above) is poor, and it is difficult to guarantee the coherence of the repaired voice segments with the surrounding voice segments. In addition, the existing methods for repairing missing voice segments are usually based on clean voice signals, while in the actual communication environment, voice signals not only have information loss of local segments, but also various noises are mixed in at the same time. Summary of the Invention

[0005] In order to effectively repair long voice missing segments in audio signals, ensure the coherence of the repaired content with the surrounding voice segments, and improve the robustness of the method in a noisy environment, the present invention proposes a method for repairing the content of missing voice segments that can effectively repair long voice missing segments and has strong robustness in a noisy environment by combining time-frequency domain analysis of signals and a latent diffusion probability model.

[0006] The technical means adopted by the present invention are as follows:

[0007] A method for repairing missing audio segments, comprising the following steps:

[0008] S1. Perform short-time Fourier transform on the damaged audio signal to obtain the STFT magnitude spectrum, and convert the STFT magnitude spectrum to the Mel scale through a frequency filter bank to obtain the Mel spectrum of the damaged signal;

[0009] S2. Use the trained vector quantization variational autoencoder to encode the Mel spectrum of the damaged signal into a latent space feature of a lower dimension;

[0010] S3. Apply the diffusion model method. Conditional on the damaged latent space feature, by concatenating the sampled Gaussian noise, after T steps of denoising by the trained denoising network, output the predicted complete latent space feature;

[0011] S4. Decode the predicted complete latent feature to obtain the predicted complete Mel spectrum, use the perceptual HiFiGAN vocoder to repair the signal phase, and output the repaired speech signal.

[0012] Furthermore, in S2, the vector quantization variational autoencoder can be used to complete the reconstruction of the input Mel spectrum. The vector quantization variational autoencoder includes an encoder and a decoder. Both the encoder and the decoder are composed of two-dimensional convolutional layers. The encoder is used to encode the input Mel spectrum to obtain a latent feature of a lower dimension; pass the latent feature through a codebook composed of K vectors, and quantize each vector in the latent feature through nearest neighbor search to obtain the quantized latent feature; the decoder can reconstruct the input Mel spectrum by receiving the quantized latent feature.

[0013] Furthermore, the loss function of the autoencoder consists of a reconstruction loss, a vector quantization loss, and an adversarial loss; the reconstruction loss L rec is the L2 distance between the reconstructed Mel spectrogram and the real Mel spectrogram, expressed as:

[0014]

[0015] where x represents the real Mel spectrum, represents the reconstructed Mel spectrum;

[0016] The vector quantization loss L vq makes the codebook vector and the encoder output approximate each other; the vector quantization loss includes a Codebook loss and a Commitment loss. The vector quantization loss L vq is expressed as:

[0017]

[0018] where E(.) represents encoding, sg[.] represents the gradient clipping operator, and e represents the codebook vector.

[0019] The discriminator loss Ldisc It is expressed as follows:

[0020]

[0021] Where D represents the PatchGAN discriminator;

[0022] The overall loss of training the vector quantization variational autoencoder is as follows:

[0023] L = min G max D (L rec + λ vq L vq + λ disc L disc ). (4)

[0024] Where λ vq and λ disc are the coefficients of the corresponding loss terms, which are set to 1.0 and 0.5 respectively.

[0025] Furthermore, in S3, the trained denoising network is a U-shaped network composed of two-dimensional convolutional layers; when training the denoising network, the quantized latent features of the clean speech are used as the true data distribution of the forward diffusion process. Given the quantized latent features z0 of the complete mel spectrogram, Gaussian noise is added according to the current time step t, and the formula is as follows:

[0026]

[0027] During the reverse process, the quantized latent features z c of the damaged audio are used as conditional information to guide the denoising at the current time step. The reverse denoising process at time step t is expressed as follows:

[0028] p θ (z t-1 |z t ,z c ) = N(z t-1 ; μ θ (z t ,t),σ θ (z t ,t)) (6)

[0029] The loss function of the denoising network is the L1 loss between the model output and the noise introduced at each diffusion time step, and the formula is as follows:

[0030]

[0031] Furthermore, in S4, the perceptual HifiGAN vocoder introduces the SpeechVGG loss on the basis of the original HifiGAN generator loss function;

[0032] The SpeechVGG is a pre-trained model for word classification. The network architecture of SpeechVGG is based on the VGG-19 convolutional neural network and includes a two-dimensional convolutional layer, a fully connected layer, and a pooling layer;

[0033] The Speech VGG loss is represented by calculating the L1 distance between the output of the vocoder and the features of each layer obtained from the ground-truth audio through the feature extractor, as shown below:

[0034]

[0035] where P represents the number of layers for which Speech VGG extracts features.

[0036] The present invention also provides a storage medium, which includes a stored program. When the program runs, it executes any one of the above audio missing segment repair methods.

[0037] The present invention also provides an electronic device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor runs the computer program to execute any one of the above audio missing segment repair methods.

[0038] Compared with the prior art, the present invention has the following advantages:

[0039] By selecting the Mel spectrum of the audio signal as the system input and output features, the present invention analyzes and repairs the Mel spectrum of the damaged audio signal, avoiding the repair of the phase of the signal missing segment and simplifying the complexity of the repair task; by combining the vector quantization autoencoder and the denoising diffusion probability model, the input features are encoded into lower-dimensional latent space features using the autoencoder, and the mapping between the damaged features and the complete features is modeled with the help of the diffusion model, realizing the effective and contextually coherent repair of longer missing segments, while improving the robustness of the method in a noisy environment, and the lower-dimensional input and output features accelerate the convergence of the denoising network; by applying a neural vocoder to repair the signal phase, the problem that the phase of the repaired segment is not coherent with the context in some traditional repair methods is overcome, and by combining the neural vocoder with the speech fine-grained perception information modeling, the subjective perception quality and objective intelligibility of the finally repaired signal are improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0041] Figure 1 It is a flowchart of the method of the present invention.

[0042] Figure 2 It is an effect diagram of repairing a single missing segment of the present invention. Among them, a is the loss spectrogram (a single 200 ms missing); b is the reconstructed spectrogram; c is the true value spectrogram.

[0043] Figure 3 It is an effect diagram of repairing an audio with a 50% missing ratio of the present invention. Among them, a is the loss spectrogram (50% missing ratio); b is the reconstructed spectrogram; c is the true value spectrogram.

[0044] Figure 4 It is an effect diagram of repairing a noisy loss audio of the present invention. Among them, a is the loss spectrogram (including real sea background noise); b is the repaired spectrogram; c is the true value spectrogram. Detailed implementation manners

[0045] To enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0046] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned accompanying drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0047] As Figure 1 shown, the present invention provides a method for repairing missing segments of audio, including the following steps:

[0048] S1. Perform a short-time Fourier transform on the damaged audio signal to obtain the STFT magnitude spectrum, and convert the STFT magnitude spectrum to the Mel scale through a frequency filter bank to obtain the Mel spectrum of the damaged signal;

[0049] S2. Use the trained vector quantization variational autoencoder to encode the Mel spectrum of the damaged signal into a lower-dimensional latent space feature;

[0050] S3. Apply the diffusion model method, condition on the damaged latent space feature, and after T steps of denoising through the trained denoising network by concatenating the sampled Gaussian noise, output the predicted complete latent space feature;

[0051] S4. Decode the predicted complete latent feature to obtain the predicted complete Mel spectrum, use the perceptual HiFiGAN vocoder to repair the signal phase, and output the repaired speech signal.

[0052] The specific solution of the present invention is as follows:

[0053] First, perform a short-time Fourier transform on the given speech segment, use a window size of 400 points, a step size of 160 points, and a Hann window function to obtain its STFT magnitude, and then use a 64-channel filter bank to convert it to the Mel scale, where the filter bank covers a frequency range of 0 to 8 kHz. Finally, perform logarithmic normalization on the Mel spectrogram as the input of the model.

[0054] The present invention adopts segmented training, which is mainly divided into two stages. In the first stage, a vector quantization variational autoencoder is trained, and the autoencoder can complete the reconstruction of the input Mel spectrum. Among them, both the encoder and the decoder are composed of two-dimensional convolutional layers. The encoder is used to encode the input Mel spectrum to obtain a lower-dimensional latent feature, and then pass it through a codebook composed of K vectors. Each vector in the latent feature is quantized through nearest neighbor search to obtain the quantized latent feature. Finally, the decoder reconstructs the input Mel spectrum by receiving the quantized latent feature.

[0055] The loss function of the autoencoder consists of a reconstruction loss, a vector quantization loss, and an adversarial loss. Among them, the reconstruction loss L rec refers to the L2 distance between the reconstructed Mel spectrogram and the true Mel spectrogram, which is expressed as:

[0056]

[0057] where x represents the true Mel spectrum, represents the reconstructed Mel spectrum

[0058] The vector quantization loss L vqAims to make the codebook vector and the encoder output approximate each other. It consists of two parts: Codebook loss and Commitment loss, and the vector quantization loss L vq Is expressed as:

[0059]

[0060] Where E(.) represents encoding, sg[.] represents the gradient clipping operator, and e represents the codebook vector.

[0061] The discriminator loss L disc Is expressed as follows:

[0062]

[0063] Where the discriminator is a PatchGAN discriminator to make the autoencoder more accurately repair the high-frequency components of the input mel spectrogram.

[0064] The overall loss for training the vector quantization variational autoencoder is as follows:

[0065] L = min G max D (L rec + λ vq L vq + λ disc L disc ). (12)

[0066] Where λ vq and λ disc Are the coefficients of the corresponding loss terms, set to 1.0 and 0.5 respectively.

[0067] In the second stage, a denoising network is trained using the diffusion model method to repair the quantized latent features of the damaged mel spectrogram. The denoising network is a U-shaped network composed of two-dimensional convolutional layers. When training the denoising network, the quantized latent features of the clean speech are used as the true data distribution in the forward diffusion process. Given the quantized latent features z0 of the complete mel spectrogram, Gaussian noise is added according to the current time step t.

[0068]

[0069] In the reverse process, the quantized latent features z c of the damaged audio are used as conditional information to guide the denoising at the current time step. The reverse denoising process at time step t is expressed as follows:

[0070] p θ (z t-1 |z t ,z c ) = N(z t-1 ; μθ (z t ,t),σ θ (z t ,t)) (14)

[0071] The loss function of the denoising network is the L1 loss between the model output and the noise introduced at each diffusion time step, and the formula is as follows:

[0072]

[0073] The perceptual HifiGAN vocoder further improves the subjective perceptual quality and objective intelligibility of the vocoder-repaired speech by additionally introducing the SpeechVGG loss on the basis of the original HifiGAN generator loss function. Speech VGG is a pre-trained model for word classification, and its network architecture is based on the VGG-19 convolutional neural network, including two-dimensional convolutional layers, fully connected layers and pooling layers. The SpeechVGG loss is represented by calculating the L1 distance between each layer of features obtained by the vocoder output and the ground-truth audio through the feature extractor, as follows:

[0074]

[0075] where P represents the number of layers of features extracted by Speech VGG.

[0076] During the training of the first-stage autoencoder, the feature space compression coefficient r is set to 4, the codebook size K is set to 1024, the dimension d of each codebook vector is set to 4, and the Adam optimizer (β1 = 0.5, β2 = 0.9) is used to update the weights, where the initial learning rate is 4.5×10-6, the batch size is 8. In addition, to make the training process more stable, the discriminator is introduced for adversarial training after 30k iterations. In the second-stage diffusion process, a linear scheduling strategy is adopted, linearly increasing from 0.0015 to 0.0195. The denoising network is trained using the Adam optimizer (β1 = 0.5, β2 = 0.9), the initial learning rate is set to 5.0×10-6, the batch size is 8, and 100k iterations are performed.

[0077] As shown by Figure 2 , the reconstructed missing speech segments in the repaired speech can maintain coherence with the surrounding speech segments. As shown by Figure 3 , by observing the reconstructed spectrogram, it can be seen that the continuity and clarity of the repaired audio signal in the time-frequency characteristics have been significantly improved. As shown by Figure 4 , by observing the reconstructed spectrogram, it can be seen that the repaired audio signal not only recovers the missing information in the time-frequency domain, but also successfully removes the obvious noise interference.

[0078] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An audio missing segment repair method, characterized in that, It includes the following steps: S1. Perform short-time Fourier transform on the damaged audio signal to obtain the STFT magnitude spectrum, and convert the STFT magnitude spectrum to the Mel scale through a frequency filter bank to obtain the Mel spectrum of the damaged signal; S2. Use the trained vector quantization variational autoencoder to encode the Mel spectrum of the damaged signal into a lower-dimensional latent space feature; S3. Apply the diffusion model method. Conditional on the damaged latent space feature, by concatenating the sampled Gaussian noise and passing through the trained denoising network for T steps of denoising, output the predicted complete latent space feature; S4. Decode the predicted complete latent feature to obtain the predicted complete Mel spectrum, use the perceptual HiFiGAN vocoder to repair the signal phase, and output the repaired speech signal.

2. The audio missing segment repair method according to claim 1, wherein In S2, the vector quantization variational autoencoder can be used to complete the reconstruction of the input Mel spectrum. The vector quantization variational autoencoder includes an encoder and a decoder. Both the encoder and the decoder are composed of two-dimensional convolutional layers. The encoder is used to encode the input Mel spectrum to obtain a lower-dimensional latent feature; Pass the latent feature through a codebook composed of K vectors, and quantize each vector in the latent feature through nearest neighbor search to obtain the quantized latent feature; The decoder can reconstruct the Mel spectrum by receiving the quantized latent feature.

3. The method for repairing an audio missing segment according to claim 2, wherein The loss function of the autoencoder consists of a reconstruction loss, a vector quantization loss, and a discriminator loss; the reconstruction loss L rec is the L2 distance between the reconstructed Mel spectrogram and the true Mel spectrogram, expressed as: where x represents the true mel-spectrum, represents the reconstructed mel-spectrum; Vector quantization loss L vq Makes the codebook vector and the encoder output approximate each other; the vector quantization loss includes the Codebook loss and the Commitment loss, and the vector quantization loss L vq Is expressed as: Among them, E(.) represents encoding, sg[.] represents the gradient clipping operator, and e represents the codebook vector. Discriminator loss L disc is expressed as follows: Among them, D represents the PatchGAN discriminator; The overall loss of training the vector quantization variational autoencoder is as follows: L = min G max D (L rec + λ vq L vq + λ disc L disc ). (4) where λ vq and λ disc are the coefficients of the corresponding loss terms, which are set to 1.0 and 0.5 respectively.

4. The method for repairing an audio missing segment according to claim 1, wherein In S3, the trained denoising network is a U-shaped network composed of two-dimensional convolutional layers; when training the denoising network, use the quantized latent feature of the clean speech as the true data distribution in the forward diffusion process. Given the quantized latent feature z0 of the complete Mel spectrogram, add Gaussian noise according to the current time step t, and the formula is as follows: During the reverse process, the quantized latent feature z of the damaged audio c is used as conditional information to guide denoising at the current time step. The reverse denoising process at time step t is expressed as follows: p θ (z t-1 |z t ,z c ) = N(z t-1 ; μ θ (z t , t), σ θ (z t , t)) (6) The loss function of the denoising network is the L1 loss between the model output and the noise introduced at each diffusion time step, and the formula is as follows:

5. The method for repairing an audio missing segment according to claim 1, wherein In S4, the perceptual HiFiGAN vocoder introduces the SpeechVGG loss on the basis of the original HifiGAN generator loss function; The SpeechVGG is a pre-trained model for word classification. The network architecture of SpeechVGG is based on the VGG-19 convolutional neural network and includes two-dimensional convolutional layers, fully connected layers and pooling layers; The Speech VGG loss is represented by calculating the L1 distance between each layer of features obtained by passing the vocoder output and the ground truth audio through the feature extractor, as follows: Among them, P represents the number of layers for which Speech VGG extracts features.

6. A storage medium, characterized in that, The storage medium includes a stored program. When the program runs, it executes the audio missing segment repair method described in any one of claims 1 to 5.

7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, The processor runs through the computer program to execute the audio missing segment repair method described in any one of claims 1 to 5.

Citation Information

Cited By

  • Method for repairing sound quality, electronic equipment, storage medium and program product

    CN120510857A