Audio restoration method and apparatus, and program, medium and device
Through the audio restoration method of the two-stage model architecture and generative adversarial network training, the problem of audio noise and distortion being difficult to cover in the existing technology is solved, and the efficient restoration and quality improvement of voice signals in real-time communication systems are achieved.
Patent Information
- Application Number
- PCT/CN2024/139165
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-21
- Filing Date
- 2024-12-13
- Publication Date
- 2025-09-25
AI Technical Summary
Existing two-stage models are unable to effectively cover various types of audio noise and distortion, especially in real-time communication systems. Voice signals suffer from spectrum loss, packet loss, and other losses during acquisition and transmission, resulting in voice quality that is difficult to meet requirements.
A two-stage model architecture is adopted. The first-stage model downsamples and upsamples the complex spectrum through the encoder and decoder, combined with short-time Fourier transform, time series modeling module and self-attention mechanism. The second-stage model further performs audio restoration through full-band and sub-band modules, and uses generative adversarial network training to train the discriminator to assist model training.
It effectively fixes noise, color distortion, discontinuous distortion, and reverberation problems in voice signals, improves voice quality, and reduces computational complexity, meeting the requirements of real-time communication systems.
Smart Images

Figure CN2024139165_25092025_PF_FP_ABST
Abstract
Description
Audio repair method, apparatus, program, medium and equipment
[0001] This application claims priority to the Chinese invention patent application with application number 202410329857.1 filed on March 21, 2024, entitled “Audio repair method, device, program, medium and equipment”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present disclosure relates to the field of audio processing technology, and in particular to an audio repair method, an audio repair device, a computer program product, a computer-readable storage medium, and an electronic device. Background Art
[0003] Technical solutions exist in the field for audio restoration using two-stage models. However, the audio restoration results of existing two-stage models still fall short of expectations. This is because audio restoration faces various types of noise and distortion, and the distortion effects may vary across different frequency bands. Existing two-stage models struggle to address such a wide range of audio quality issues. Therefore, there is an urgent need for an audio restoration model that can address complex and diverse audio distortion issues. Summary of the Invention
[0004] To this end, the present disclosure provides an audio repair method, an audio repair device, a computer program product, a computer-readable storage medium, and an electronic device, which can repair complex and diverse audio distortion problems.
[0005] In one aspect, the present disclosure provides an audio restoration method, comprising: inputting damaged audio into a first-stage model to obtain a complex spectrum; and inputting the complex spectrum into a second-stage model to obtain repaired audio. The first-stage model includes an encoder for downsampling the complex spectrum; a decoder for upsampling the complex spectrum. The second-stage model includes a full-band module for modeling the complex spectrum across all frequency bands; and a subband module for modeling the complex spectrum across multiple subbands.
[0006] In one embodiment of the present disclosure, the first-stage model further includes: a short-time Fourier transform module for converting the damaged audio into a complex spectrum through short-time Fourier transform; and a time series modeling module for further extracting features of the complex spectrum in the time dimension.
[0007] In one embodiment of the present disclosure, the encoder includes multiple stacked downsampling modules, each including: a two-dimensional gated convolution module for performing gated convolution on a complex spectrum; a time-frequency convolution module for performing convolution on the complex spectrum in both the time and frequency dimensions; and an axial self-attention module for calculating an attention mechanism on the complex spectrum. The decoder includes multiple stacked upsampling modules, each including: a two-dimensional gated transposed convolution module for performing gated transposed convolution on the complex spectrum; a time-frequency convolution module for performing convolution on the complex spectrum in both the time and frequency dimensions; and an axial self-attention module for calculating an attention mechanism on the complex spectrum.
[0008] In one embodiment of the present disclosure, the second stage model further includes: a complex feature encoder for extracting high-dimensional features in the complex spectrum; and a complex feature decoder for restoring the high-dimensional features in the complex spectrum to low dimensions.
[0009] In one embodiment of the present disclosure, before inputting the damaged audio into the first-stage model to obtain the complex spectrum, the method further includes: training the first-stage model using a generative adversarial network training method; training the second-stage model on a specific audio repair task; training the cascaded first-stage model and the second-stage model using a generative adversarial network training method; simulating the input audios whose corresponding loss function for measuring the repair effect is still higher than a threshold after the training is completed, generating multiple simulated audios, and training the cascaded first-stage model and the second-stage model using the multiple simulated audios.
[0010] In one embodiment of the present disclosure, a first-stage model is trained using a generative adversarial network training method, including: using the first-stage model as a generator of the generative adversarial network, and using the first discriminator as a discriminator of the generative adversarial network, thereby training the first-stage model using the generative adversarial network training method; wherein the first discriminator includes multiple first sub-discriminators, the first sub-discriminator includes multiple stacked two-dimensional convolution modules, and the two-dimensional convolution module is used to convolve the amplitude spectrum of the input audio.
[0011] In one embodiment of the present disclosure, a first-stage model is trained using a generative adversarial network training method, including: using the first-stage model as the generator of the generative adversarial network and using the second discriminator as the discriminator of the generative adversarial network, thereby training the first-stage model using the generative adversarial network training method. The second discriminator includes multiple second sub-discriminators, each of which includes multiple parallel sub-band discriminator modules, each of which is used to discriminate complex spectra of different frequency bands into which the input audio is divided, and each sub-band discriminator module includes multiple stacked two-dimensional convolution modules, each of which is used to convolve the complex spectra.
[0012] In another aspect, the present disclosure provides an audio restoration device, comprising: a first-stage module for inputting damaged audio into a first-stage model to obtain a complex spectrum; and a second-stage module for inputting the complex spectrum into the second-stage model to obtain restored audio. The first-stage model includes an encoder for downsampling the complex spectrum converted from the damaged audio; and a decoder for upsampling the complex spectrum. The second-stage model includes a full-band module for modeling the complex spectrum across all frequency bands; and a subband module for modeling the complex spectrum across multiple subbands.
[0013] In one embodiment of the present disclosure, the first-stage model further includes: a short-time Fourier transform module for converting the damaged audio into a complex spectrum through short-time Fourier transform; and a time series modeling module for further extracting features of the complex spectrum in the time dimension.
[0014] In one embodiment of the present disclosure, the encoder includes multiple stacked downsampling modules, each including: a two-dimensional gated convolution module for performing gated convolution on a complex spectrum; a time-frequency convolution module for performing convolution on the complex spectrum in both the time and frequency dimensions; and an axial self-attention module for calculating an attention mechanism on the complex spectrum. The decoder includes multiple stacked upsampling modules, each including: a two-dimensional gated transposed convolution module for performing gated transposed convolution on the complex spectrum; a time-frequency convolution module for performing convolution on the complex spectrum in both the time and frequency dimensions; and an axial self-attention module for calculating an attention mechanism on the complex spectrum.
[0015] In one embodiment of the present disclosure, the second stage model further includes: a complex feature encoder for extracting high-dimensional features in the complex spectrum; and a complex feature decoder for restoring the high-dimensional features in the complex spectrum to low dimensions.
[0016] In one embodiment of the present disclosure, the device is further configured to: train the first-stage model using a generative adversarial network training method; train the second-stage model on a specific audio restoration task; train the cascaded first-stage model and the second-stage model using a generative adversarial network training method; simulate the input audios whose corresponding loss function values for measuring the restoration effect are still higher than a threshold after the training is completed, generate multiple simulated audios, and train the cascaded first-stage model and the second-stage model using the multiple simulated audios.
[0017] In one embodiment of the present disclosure, the apparatus is further configured to use the first-stage model as the generator of a generative adversarial network and the first discriminator as the discriminator of the generative adversarial network, thereby training the first-stage model using the generative adversarial network training method. The first discriminator includes a plurality of first sub-discriminators, each of which includes a plurality of stacked two-dimensional convolutional modules, which are used to convolve the amplitude spectrum of the input audio.
[0018] In one embodiment of the present disclosure, the apparatus is further configured to use the first-stage model as the generator of a generative adversarial network and the second discriminator as the discriminator of the generative adversarial network, thereby training the first-stage model using the generative adversarial network training method. The second discriminator includes multiple second sub-discriminators, each of which includes multiple parallel sub-band discriminator modules. Each sub-band discriminator module is used to discriminate complex spectra of different frequency bands into which the input audio is divided. The sub-band discriminator module includes multiple stacked two-dimensional convolution modules, each of which is used to convolve the complex spectra.
[0019] In another aspect, the present disclosure provides a computer program product, comprising a computer program, wherein the computer program implements the above-mentioned audio repair method when executed by a processor.
[0020] In another aspect, the present disclosure provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed in a computer, the computer is caused to execute the audio repair method.
[0021] On the other hand, the present disclosure provides an electronic device including a memory and a processor, wherein the memory stores an executable code, and when the processor executes the executable code, the audio repair method is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The specific embodiments of the present application are described in detail below with reference to the accompanying drawings, wherein:
[0023] FIG1 is a schematic diagram showing the overall structure of an audio restoration model according to an embodiment of the present application;
[0024] FIG2 is a schematic structural diagram of a first-stage model of an audio restoration model according to the embodiment of FIG1 ;
[0025] FIG3 is a schematic structural diagram of a downsampling module of the first stage model according to the embodiment of FIG1 ;
[0026] FIG4 is a schematic structural diagram of a two-dimensional gated convolution module in the downsampling module according to the embodiment of FIG1 ;
[0027] FIG5 is a schematic structural diagram of a time-frequency convolution module in a downsampling module according to the embodiment of FIG1 ;
[0028] FIG6 is a schematic diagram showing the structure of an axial self-attention module in the downsampling module according to the embodiment of FIG1 ;
[0029] FIG7 is a schematic structural diagram of a time series modeling module of the first stage model according to the embodiment of FIG1 ;
[0030] FIG8 is a schematic structural diagram of an upsampling module of the first stage model according to the embodiment of FIG1 ;
[0031] FIG9 is a schematic structural diagram of a two-dimensional gated transposed convolution module in the upsampling module according to the embodiment of FIG1 ;
[0032] FIG10 is a schematic structural diagram of a second-stage model of the audio restoration model according to the embodiment of FIG1 ;
[0033] FIG11 is a schematic structural diagram of a complex feature encoder of the second stage model according to the embodiment of FIG1 ;
[0034] FIG12 is a schematic structural diagram of a sub-band module of the second stage model according to the embodiment of FIG1 ;
[0035] FIG13 is a schematic structural diagram of a grouped complex convolution module of a sub-band module according to the embodiment of FIG1 ;
[0036] FIG14 is a schematic structural diagram of a grouped complex transposed convolution module of a sub-band module according to the embodiment of FIG1 ;
[0037] FIG15 is a schematic structural diagram of a full-belt module of the second-stage model according to the embodiment of FIG1 ;
[0038] FIG16 is a schematic structural diagram of a complex convolution module of the full-band module according to the embodiment of FIG1 ;
[0039] FIG17 is a schematic structural diagram of a complex transposed convolution module of the full-band module according to the embodiment of FIG1 ;
[0040] FIG18 is a schematic structural diagram of a complex feature decoder of the second stage model according to the embodiment of FIG1 ;
[0041] FIG19 is a schematic diagram showing a flow chart of an audio repair method according to an embodiment of the present application;
[0042] FIG20 is a schematic diagram showing a flow chart of an audio restoration model training method according to an embodiment of the present application;
[0043] FIG21 is a schematic structural diagram of a first sub-discriminator according to the embodiment of FIG20 ;
[0044] FIG22 is a schematic structural diagram of a second sub-discriminator according to the embodiment of FIG20 ;
[0045] FIG23 is a schematic structural diagram of an audio repair device according to an embodiment of the present application;
[0046] FIG24 shows a schematic structural diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0047] The technical solutions provided in this specification are further described in detail below in conjunction with the accompanying drawings and embodiments. It will be understood that the specific embodiments described herein are merely for explaining the relevant embodiments and are not intended to limit the embodiments. It should also be noted that, for ease of description, only portions relevant to the relevant embodiments are shown in the accompanying drawings. It should be noted that, unless there is a conflict, the embodiments of this specification and the features therein may be combined with each other.
[0048] In this document, the terms "one", "an" and other similar words are not intended to indicate that there is only one of the things described, but rather that the relevant description is only for one of the things described, and the things described may have one or more. In this document, the terms "include", "comprise" and other similar words are intended to indicate logical relationships, and cannot be regarded as indicating relationships in spatial structure. For example, "A includes B" is intended to indicate that B logically belongs to A, but does not mean that B is spatially located inside A. In addition, the meanings of the terms "include", "include" and other similar words should be regarded as open, not closed. For example, "A includes B" is intended to indicate that B belongs to A, but B does not necessarily constitute the whole of A. A may also include other elements such as C, D, and E.
[0049] In this document, the terms "first", "second" and other similar words are not intended to imply any order, quantity or importance, but are merely used to distinguish different elements. In this document, the terms "embodiment", "present embodiment", "one embodiment" and "an embodiment" do not indicate that the relevant description is only applicable to a specific embodiment, but rather indicate that these descriptions may also be applicable to one or more other embodiments. Those skilled in the art should understand that in this document, any description made for a certain embodiment can be replaced, combined, or otherwise combined with the relevant description in one or more other embodiments, and the new embodiment generated by the replacement, combination, or other combination is easily conceivable by those skilled in the art and falls within the scope of protection of this application.
[0050] The widespread adoption of real-time communication (RTC) systems has garnered increasing attention on voice signal quality. In RTC links, voice signals are subject to a variety of adverse factors at every stage, from signal acquisition, analog-to-digital conversion, signal transmission, to signal playback, severely impacting voice quality. Previous research has focused on studying noise and reverberation, two key factors affecting voice signal quality, and has proposed a variety of high-performance deep neural networks. After years of development, the noise reduction and dereverberation performance of deep neural networks has gradually reached its upper limit, leading to increasing attention on improving overall signal quality.
[0051] Audio impairments can be categorized into five categories: noise, coloration, discontinuity, volume, and reverberation. Specifically, noise distortion includes background noise, circuit noise, and coding noise; coloration is primarily due to microphone bandwidth limitations and frequency response; discontinuity distortion includes packet loss; and volume distortion primarily includes clipping, nonlinear distortion, and far-field recording.
[0052] In this field, multiple systems have demonstrated strong performance in improving speech signal quality. These systems can be categorized into three types: 1) Two-stage models, where the first stage repairs all types of distortion, and the second stage further removes residual noise and artifacts generated by the first stage; 2) Noise reduction models, which perform only noise reduction and dereverberation; and 3) Diffusion-based generative models. Diffusion-based generative models were first used in image generation, performing image generation tasks by gradually adding Gaussian noise to the image and then removing the noise in reverse order. Given their superior performance, researchers have also applied them to noise reduction tasks, achieving impressive results.
[0053] In some embodiments, a two-stage model can perform phased repair for all types of impairments, resulting in superior performance. Traditional noise reduction models can also significantly improve signal quality, but due to impairments such as band loss (for example, retaining only the portion below 8 kHz for audio sampled at 48 kHz), which can significantly affect the listening experience, traditional noise reduction models perform worse than two-stage models that simultaneously handle all types of impairments. Due to the wide variety of impairments encountered by speech, diffusion-based generative models perform poorly.
[0054] In some embodiments, the audio restoration model provided by the present disclosure is designed to address the following issues:
[0055] 1) In RTC systems, speech signals are subject to loss during acquisition and transmission, such as spectral omissions and packet loss. Restoring these speech signals requires a deep neural network to generate the missing signal components. Given the widespread use of generative adversarial networks (GANs) in speech synthesis, a discriminator can be introduced to aid in training. However, there are many different types of discriminators, and different discriminators provide different support for different deep neural networks. Therefore, it is necessary to design a discriminator suitable for the task of sound quality restoration.
[0056] 2) In the two-stage model, the main tasks of the model in each stage are different, and a suitable solution needs to be determined to train the model.
[0057] 3) In order to meet the real-time rate requirements, it is necessary to minimize the computational complexity of the model while minimizing the performance loss of the model.
[0058] To this end, according to some embodiments of the present disclosure, a two-stage model architecture for improving speech quality is proposed, which can handle problems such as noise, coloration, discontinuity, volume, and reverberation faced by speech in a noisy acoustic environment. Two discriminators are used to assist in the training of the frequency domain model, a first discriminator (multi-resolution discriminator) and a second discriminator (sub-band multi-resolution discriminator). During training, the first-stage model is pre-trained and the first-stage model parameters are frozen. Then, the two-stage model is pre-trained on the noise reduction and dereverberation tasks, and the pre-trained one-stage model is cascaded with the two-stage model for joint training. Finally, the learning rate is reduced and the two-stage model is fine-tuned.
[0059] FIG1 shows a schematic diagram of the overall structure of an audio restoration model according to an embodiment of the present application.
[0060] As shown in Figure 1, the audio restoration model consists of a first-stage model and a second-stage model. The first-stage model performs tasks such as coloration restoration, discontinuity and volume restoration, preliminary noise reduction, and preliminary dereverberation. The second-stage model performs further noise reduction, further dereverberation, and artifact removal introduced by the first-stage model.
[0061] The damaged input audio is first input into the first-stage model, processed by the first-stage model, and then input into the second-stage model. After being processed by the second-stage model, the repaired audio is obtained.
[0062] FIG2 shows a schematic structural diagram of a first-stage model of the audio restoration model according to the embodiment of FIG1 .
[0063] As shown in Figure 2, the input of the first-stage model is the complex spectrum (consisting of real and imaginary values) obtained by short-time Fourier transform (STFT) of the distorted and impaired audio. The window function, window length, window shift, and FFT (Fast Fourier Transform) length used in the Fourier transform can be, for example, a Hanning window, 20ms, 10ms, and 20ms, respectively. The output of the first-stage model is the complex spectrum.
[0064] The complex spectrum contains both amplitude and phase information of the audio in the frequency dimension. Because both information must be captured simultaneously, it must be represented using complex numbers. The complex spectrum can be converted into an amplitude spectrum, which only contains information about amplitude variations at different frequencies, and a phase spectrum, which only contains information about phase variations at different frequencies. Both the amplitude spectrum and the phase spectrum are frequency-dependent, containing information about how the audio varies with frequency. Before Fourier transforming, the original, damaged audio is a spectrum in the time dimension, containing information about how the audio varies over time.
[0065] The overall model can be divided into three parts: the encoder, the time series modeling module, and the decoder. The encoder consists of stacked frequency down-sampling modules (FDs), which primarily downsample the frequency dimension of features. The number of stacked down-sampling modules can be any suitable value, such as 3, 4, or 5. The time series modeling module uses stacked squeezed temporal convolutional modules (STCMs) to model the complex spectrum in the time dimension. The decoder consists of a real decoder (Real Decoder) and an imaginary decoder (Imaginary Decoder). The real decoder predicts the real value spectrum of the complex spectrum, while the imaginary decoder predicts the imaginary value spectrum of the complex spectrum. Both decoders have the same structure, consisting of stacked frequency up-sampling modules (FUs), which primarily upsample the frequency dimension of features. The number of stacked up-sampling modules can be any suitable value, such as 3, 4, or 5. The encoder has skip connections with the real-value decoder and the imaginary-value decoder respectively.
[0066] Specifically, after the damaged audio is input into the first-stage model, it is first short-time Fourier transformed into a complex spectrum. The complex spectrum passes through the encoder to extract high-dimensional features in the frequency dimension of the complex spectrum. The complex spectrum then enters the time series modeling module, which models the complex spectrum in the time dimension and extracts features in the time dimension. The complex spectrum is then split into a real-valued spectrum and an imaginary-valued spectrum, which are decoded by the real-value decoder and the imaginary-value decoder respectively, restoring the high-dimensional features in the complex spectrum to a low dimension. The decoded real-valued spectrum and imaginary-valued spectrum are concatenated to obtain a complete complex spectrum.
[0067] FIG3 shows a schematic structural diagram of a downsampling module of the first-stage model according to the embodiment of FIG1 .
[0068] As shown in Figure 3, the downsampling module consists of three parts: a two-dimensional gated convolution module (Gate Conv2d), a time-frequency convolution module (TFCM), and an axial self-attention module (ASA). The two-dimensional gated convolution module performs two-dimensional gated convolution on the complex spectrum of the input audio. The time-frequency convolution module performs convolution on the complex spectrum of the input audio in the time and frequency dimensions. The axial self-attention module is used for calculations related to the attention mechanism.
[0069] FIG4 shows a schematic structural diagram of a two-dimensional gated convolution module in the downsampling module according to the embodiment of FIG1 .
[0070] As shown in Figure 4, the 2D gated convolution module consists of two parallel 2D convolution modules (Conv2d). These two 2D convolution modules are subsequently connected in series with different activation functions: Sigmoid activation function and LeakyReLU activation function. After the complex spectrum is input into the 2D gated convolution module, it is split into two paths for 2D convolution. After the convolution, different activation functions are applied for calculation. The calculated results are multiplied to obtain the output complex spectrum.
[0071] FIG5 shows a schematic structural diagram of a time-frequency convolution module in a downsampling module according to the embodiment of FIG1 .
[0072] As shown in Figure 5, the time-frequency convolution module uses two-dimensional convolution to model the time and frequency dimensions, and uses dilated convolution to expand the model's receptive field in the time dimension. Specifically, the time-frequency convolution module includes a two-dimensional convolution module (Conv2d), a dilated convolution module (Dilated Conv2d), and a 1×1 two-dimensional convolution module (1x1Conv2d). The two-dimensional convolution module is followed by a regularization module (Norm) and a PReLU activation function, and the dilated convolution module is also followed by a regularization module and a PReLU activation function.
[0073] FIG6 shows a schematic structural diagram of an axial self-attention module in a downsampling module according to the embodiment of FIG1 .
[0074] As shown in Figure 6, the axial self-attention module uses two-dimensional convolution along the time and frequency dimensions to solve the attention mechanism's Q, K, and V matrices, which are used for attention-related operations. Specifically, the axial self-attention module includes three parallel 1×1 two-dimensional convolution modules, which are used to solve the Q, K, and V matrices respectively, for the attention mechanism calculations.
[0075] FIG. 7 shows a schematic structural diagram of a time series modeling module of the first stage model according to the embodiment of FIG. 1 .
[0076] As shown in Figure 7, the time series modeling module uses a stacked temporal convolutional module (STCM). In this module, the input complex spectrum first passes through a 1×1 one-dimensional convolution module for convolution in the time dimension. The complex spectrum then enters two parallel dilated convolution modules for one-dimensional convolution. One of the dilated convolution modules is preceded by a PReLU activation function and a regularization module, while the other is preceded by a PReLU activation function and a regularization module, followed by a Sigmoid activation function. The two results are multiplied and then enter a 1×1 one-dimensional convolution module for another convolution in the time dimension. The convolution result is added to the initial input complex spectrum to obtain the final output.
[0077] FIG8 shows a schematic structural diagram of an upsampling module of the first-stage model according to the embodiment of FIG1 .
[0078] As shown in Figure 8, the upsampling module is similar to the downsampling module in the encoder and consists of three parts: a two-dimensional gated transposed convolution module (Gate TransposeConv2d), a time-frequency convolution module, and an axial self-attention module. The two-dimensional gated transposed convolution module performs a two-dimensional gated transposed convolution on the complex spectrum of the input audio. The time-frequency convolution module convolves the complex spectrum of the input audio in both the time and frequency dimensions. The axial self-attention module performs calculations related to the attention mechanism. The time-frequency convolution module and the axial self-attention module in the upsampling module have the same structure as those in the downsampling module.
[0079] FIG9 shows a schematic structural diagram of a two-dimensional gated transposed convolution module in the upsampling module according to the embodiment of FIG1 .
[0080] As shown in Figure 9, the 2D Gated Transposed Convolution module consists of two parallel 2D Transposed Convolution modules (TransposeConv2d). One 2D Transposed Convolution module is followed by a sigmoid activation function, and the other 2D Transposed Convolution module is followed by a LeakyReLU activation function. After the complex spectrum is input into the 2D Gated Transposed Convolution module, it enters the two 2D Transposed Convolution modules for transposed convolution calculations. The corresponding activation functions are then substituted and the resulting results are multiplied to produce the final complex spectrum output.
[0081] FIG10 shows a schematic structural diagram of a second-stage model of the audio restoration model according to the embodiment of FIG1 .
[0082] As shown in Figure 10, the second-stage model uses S-DCCRN as the backbone network, replaces the Long Short Term Memory (LSTM) network module with a Stacked Temporal Convolution Module (STCM), and reduces the number of model parameters and computational complexity by replacing traditional convolution with depthwise separable convolution and reducing the number of convolution channels.
[0083] The input of the second-stage model is the complex spectrum output by the first-stage model. The output of the second-stage model is the restored audio. The second-stage model is generally divided into four parts: the complex feature encoder, the sub-band module, the full-band module, and the complex feature decoder.
[0084] FIG11 shows a schematic structural diagram of a complex feature encoder of the second stage model according to the embodiment of FIG1 .
[0085] As shown in Figure 11, the complex feature encoder first uses complex convolution to extract high-dimensional representations in the frequency domain, then uses dense blocks to model the time dimension of the features, and finally uses complex convolution with a convolution kernel of 1x1 to extract local information. The complex feature encoder includes a complex convolution module, a dense block, and a complex convolution module stacked in sequence, which are used to perform feature encoding and feature extraction on the complex spectrum processed by the first-stage model. For the specific structure of the complex convolution module, please refer to the diagram and related description in Figure 16, which will not be repeated here.
[0086] FIG12 shows a schematic structural diagram of a sub-band module of the second stage model according to the embodiment of FIG1 .
[0087] As shown in Figure 12, the subband module includes an encoder, a time series modeling module, and a decoder. The encoder is used to extract high-dimensional features of the complex spectrum on multiple frequency subbands. It is composed of multiple grouped complex convolution modules (Group Complex Conv2d) stacked together. The number of stacks can be any appropriate value, such as 3, 4, 5, etc. The time series modeling module is used to extract high-dimensional features in the time dimension. It can use a stacked time convolution module. The specific structure of this module can be found in the diagram and related description of Figure 7 and will not be repeated here. The decoder is used to restore high-dimensional features to low dimensions on multiple frequency subbands. It is composed of multiple grouped complex transpose convolution modules (Group Complex TransposeConv2d) stacked together. The number of stacks can be any appropriate value, such as 3, 4, 5, etc. There is a jump connection between the encoder and the decoder. The jump connection passes through the complex convolution module (Complex Conv2d) to prevent gradient disappearance.
[0088] FIG13 shows a schematic structural diagram of a grouped complex convolution module of a sub-band module according to the embodiment of FIG1 .
[0089] As shown in Figure 13, the input complex spectrum is divided into real and imaginary parts, which are used for convolution in two parallel group convolution modules (Group Conv2d). Before convolution, the real and imaginary spectra are divided into high-frequency and low-frequency parts, respectively, for group (i.e., divided into two groups) convolution. After the real and imaginary parts are grouped and convolved separately, they are concatenated to form a complete complex spectrum. The complex spectrum is then processed by the batch normalization module and substituted into the LeakyReLU activation function to obtain the final output result. In the group complex convolution module (Group Complex Conv2d), the number of groups corresponds to the number of divided subbands. The number of groups in this system is 2.
[0090] FIG14 shows a schematic structural diagram of a grouped complex transposed convolution module of a sub-band module according to the embodiment of FIG1 .
[0091] As shown in Figure 14 , the structure of the grouped complex transposed convolution module is generally similar to that of the grouped complex convolution module. The difference is that the two-dimensional convolution module of the grouped complex convolution module is replaced with a two-dimensional transposed convolution module in the grouped complex transposed convolution module. For other details about the grouped complex transposed convolution module, please refer to the description of Figure 13 and will not be repeated here.
[0092] FIG15 shows a schematic structural diagram of a full-belt module of the second-stage model according to the embodiment of FIG1 .
[0093] As shown in Figure 15, the full-band module includes an encoder, a time series modeling module, and a decoder. The encoder is used to extract high-dimensional features of the complex spectrum across all frequency bands. It is composed of multiple complex convolution modules (Complex Conv2d) stacked together. The number of stacks can be any appropriate value, such as 3, 4, 5, etc. The time series modeling module is used to extract high-dimensional features in the time dimension. It can use a stacked time convolution module. The specific structure of this module can be found in the diagram and related description of Figure 7 and will not be repeated here. The decoder is used to restore high-dimensional features to low dimensions across all frequency bands. It is composed of multiple complex transpose convolution modules (Complex TransposeConv2d) stacked together. The number of stacks can be any appropriate value, such as 3, 4, 5, etc. There is a jump connection between the encoder and decoder. The jump connection passes through the complex convolution module (Complex Conv2d) to prevent gradient disappearance.
[0094] FIG16 shows a schematic structural diagram of a complex convolution module of a full-band module according to the embodiment of FIG1 .
[0095] As shown in Figure 16, the complex convolution module consists of two parallel 2D convolution modules, one for the real and one for the imaginary convolution. After the two convolutions are complete, they are concatenated to produce the complete complex spectrum. The BatchNorm2d module then performs batch normalization, and the calculated results are substituted into the LeakyReLU activation function to produce the final output.
[0096] FIG17 shows a schematic structural diagram of a complex transposed convolution module of the full-band module according to the embodiment of FIG1 .
[0097] As shown in Figure 17 , the structure of the complex transposed convolution module is generally similar to that of the complex convolution module. The difference is that the two-dimensional convolution module of the complex convolution module is replaced by a two-dimensional transposed convolution module in the complex transposed convolution module. For other details about the complex transposed convolution module, see the description of Figure 16 and will not be repeated here.
[0098] FIG18 shows a schematic structural diagram of a complex feature decoder of the second stage model according to the embodiment of FIG1 .
[0099] As shown in Figure 18, the complex feature decoder consists of a stack of dense blocks, complex sub-pixel convolution modules, and complex convolution modules. The complex feature decoder first uses dense blocks to model the time dimension of the feature. It then uses complex sub-pixel convolution instead of complex transposed convolution to effectively avoid the checkerboard phenomenon. Finally, complex convolution is used to restore the feature to a complex spectrum.
[0100] FIG19 is a schematic flow chart showing an audio repair method according to an embodiment of the present application.
[0101] As shown in FIG. 19 , according to this embodiment, the audio repair method includes steps S1910 and S1920 , each of which is described in detail below.
[0102] S1910. Input the damaged audio into a first-stage model to obtain a complex spectrum, wherein the first-stage model includes an encoder and a decoder, the encoder is used to downsample the complex spectrum, and the decoder is used to upsample the complex spectrum.
[0103] In this embodiment, the complex spectrum is down-sampled by the encoder to extract high-dimensional features in the complex spectrum, and the complex spectrum is up-sampled by the decoder to restore the complex spectrum from high dimension to low dimension. This process can repair various distortion problems in audio.
[0104] S1920. Input the complex spectrum into the second-stage model to obtain repaired audio, wherein the second-stage model includes a full-band module and a sub-band module, the sub-band module is used to model the complex spectrum in multiple sub-bands, and the full-band module is used to model the complex spectrum in all frequency bands.
[0105] In this embodiment, by modeling the complex spectrum over the entire frequency band through the full-band module, various distortion problems presented by the overall audio can be eliminated. By modeling the complex spectrum over multiple sub-bands through the sub-band module, the distortion problems presented by the audio in each sub-band can be eliminated.
[0106] FIG20 shows a flow chart of an audio restoration model training method according to an embodiment of the present application.
[0107] As shown in Figure 20, in general, the audio restoration model training method according to this embodiment includes the following steps: using a multi-resolution discriminator and a sub-band multi-resolution discriminator to train the first-stage model; pre-training the second-stage model on noise reduction and dereverberation tasks; freezing the parameters of the pre-trained first-stage model and cascading it with the pre-trained second-stage model, and using a multi-resolution discriminator to train the cascaded model, only updating the parameters of the second-stage model at this stage; based on the bad examples in the test set, targeted simulation of these data, and then reducing the learning rate used for training, and fine-tuning the cascaded model, also freezing the parameters of the first-stage model and only updating the second-stage model.
[0108] The following describes in detail the audio restoration model training method according to this embodiment. According to this embodiment, the audio restoration model includes a first-stage model and a second-stage model, and the audio restoration model training method includes steps S2010 to S2040.
[0109] S2010, use the generative adversarial network training method to train the first stage model.
[0110] In generative adversarial network training, the first-stage model is used as the generator and trained with an appropriate discriminator. Both the first-stage and second-stage models are frequency-domain models. To enhance the model's ability to cope with various distortions, we use a discriminator that uses a spectrum as input for auxiliary training.
[0111] The discriminators used to train the first-stage model include a first discriminator and a second discriminator. The two discriminators respectively make judgments on whether the input audio is clean audio or repaired damaged audio. The losses of the two discriminators can be used as optimization targets.
[0112] As an example, to train the first-stage model using a generative adversarial network training approach, the first-stage model can be used as the generator of the generative adversarial network, and the first discriminator can be used as the discriminator of the generative adversarial network, thereby training the first-stage model using the generative adversarial network training approach. In some embodiments, the first discriminator includes multiple first sub-discriminators, each of which includes multiple stacked two-dimensional convolution modules, which are used to convolve the amplitude spectrum of the input audio.
[0113] FIG. 21 shows a schematic structural diagram of the first sub-discriminator according to the embodiment of FIG. 20 .
[0114] As shown in Figure 21, the first sub-discriminator is composed of a stack of two-dimensional convolution modules. In this embodiment, the first discriminator can be called a multi-resolution discriminator. The input of the first discriminator is the amplitude spectrum, which includes a total of 6 first sub-discriminators. The overall structure of each first sub-discriminator is the same as that shown in Figure 21, except for the number of convolution channels and the parameters of the STFT used to obtain the amplitude spectrum. The number of convolution channels of the 6 first sub-discriminators and the STFT parameters used are as follows (the number of channels described below is in the form of (a, b), where a represents the number of input channels and b represents the number of output channels).
[0115] For the first sub-discriminator, the number of convolution channels is (1, 4) (4, 8) (8, 16) (16, 32) (32, 64) (64, 128) (128, 1), and the window length and window shift of STFT are 60 and 15 sampling points respectively.
[0116] For the second first sub-discriminator, the number of convolution channels is (1, 4) (4, 8) (8, 16) (16, 32) (32, 64) (64, 128) (128, 1), and the window length and window shift of STFT are 120 and 30 sampling points respectively.
[0117] For the third first sub-discriminator, the number of convolution channels is (1, 8) (8, 16) (16, 32) (32, 64) (64, 128) (128, 256) (256, 1), and the window length and window shift of STFT are 240 and 60 sampling points respectively.
[0118] For the fourth first sub-discriminator, the number of convolution channels is (1, 8) (8, 16) (16, 32) (32, 64) (64, 128) (128, 256) (256, 1), and the window length and window shift of STFT are 480 and 120 sampling points respectively.
[0119] The fifth first sub-discriminator has the following convolution channels: (1, 16) (16, 32) (32, 64) (64, 128) (128, 256) (256, 512) (512, 1), and the STFT window length and window shift are 960 and 240 sampling points respectively.
[0120] For the sixth first sub-discriminator, the number of convolution channels is (1, 16) (16, 32) (32, 64) (64, 128) (128, 256) (256, 512) (512, 1), and the window length and window shift of STFT are 2020 and 480 sampling points respectively.
[0121] As an example, to train the first-stage model using a generative adversarial network training method, the first-stage model can be used as the generator of the generative adversarial network, and the second discriminator can be used as the discriminator of the generative adversarial network, thereby training the first-stage model using the generative adversarial network training method. In some embodiments, the second discriminator includes multiple second sub-discriminators, each of which includes multiple parallel sub-band discriminator modules. The different sub-band discriminator modules are used to discriminate the complex spectra of different frequency bands into which the input audio is divided. The sub-band discriminator module includes multiple stacked two-dimensional convolution modules, each of which is used to convolve the complex spectra.
[0122] FIG. 22 is a schematic structural diagram of a second sub-discriminator according to the embodiment of FIG. 20 .
[0123] As shown in FIG22 , in this embodiment, the second discriminator may also be referred to as a sub-band multi-resolution discriminator.
[0124] The input to the second discriminator is the complex spectrum of the divided subbands. The divided subbands can be, for example, 0-2400 Hz, 2400 Hz-6000 Hz, 6000 Hz-12000 Hz, 12000 Hz-18000 Hz, and 18000 Hz-24000 Hz. The second discriminator can include three second sub-discriminators. Each second sub-discriminator has the same structure, but differs in the STFT parameters. The STFT parameters of the three second sub-discriminators are as follows.
[0125] For the first and second sub-discriminators, the STFT window length and window shift are 512 and 128 sampling points respectively.
[0126] For the second sub-discriminator, the STFT window length and window shift are 1024 and 256 sampling points respectively.
[0127] For the third second sub-discriminator, the STFT window length and window shift are 2048 and 512 sampling points respectively.
[0128] In S2010, the first and second discriminators are updated using the discriminator loss function, and the first stage model as the generator is updated using the first stage loss function.
[0129] In this embodiment, the discriminator loss function can be as follows:
[0130] in, These are the audios after the first and second stage repairs respectively. is the output of the one-stage / two-stage repaired audio input discriminator.
[0131] In this embodiment, the generator loss function can be as follows:
[0132] Among them, S, They are clean audio, audio after one stage / two stage repair, D(S), They are the outputs obtained after the clean audio is input to the discriminator and the outputs obtained after the one-stage / two-stage repaired audio is input to the discriminator.
[0133] In this embodiment, the first-stage loss function can be as follows:
[0134] in, X、 They represent the amplitude spectrum of clean audio and the amplitude spectrum after one-stage restoration, respectively. F represents the Frobenius norm;
[0135] in, X、 They represent the amplitude spectrum of the clean audio and the amplitude spectrum after one-stage repair, respectively, and ||.||1 represents the L1 norm.
[0136] in,
[0137] X、 They represent the amplitude spectrum of the clean audio and the amplitude spectrum after one-stage repair respectively; h(x) is a function, when x>=0, h(x)=x, when x<0, h(x)=0.
[0138] S2020. Train the second-stage model on a specific audio restoration task.
[0139] As examples, specific audio restoration tasks include noise reduction and dereverberation tasks.
[0140] Pre-training on noise reduction and dereverberation tasks facilitates early training for these two more difficult tasks, making subsequent joint training easier and thus reducing the computing power required for model training overall.
[0141] In S2020, the two-stage model is updated using the second-stage pre-trained loss function.
[0142] In this embodiment, the second stage pre-training loss function can be as follows:
[0143] in, Obtained by the following formula:
[0144] in, and S represent the waveform of the repaired audio and the waveform of the clean audio respectively.
[0145] in, To compensate for the spectral compression loss, the difference between the spectrum of the enhanced audio and the spectrum of the clean audio is calculated in the frequency domain, which can be obtained by the following formula:
[0146] S2030. Use a generative adversarial network training method to train the cascaded first-stage model and the second-stage model.
[0147] In this embodiment, the cascaded first-stage model and the second-stage model serve as generators, and the first and second discriminators serve as discriminators. The cascaded first-stage model and the second-stage model are trained by the generative adversarial network training method. In this stage, only the parameters of the second-stage model are updated.
[0148] In S2030 , the first and second discriminators are updated using the discriminator loss function, and the second stage model is updated using the second stage loss function.
[0149] In this embodiment, the second stage loss function can be as follows:
[0150] S2040. Simulate the input audios whose corresponding loss function values for measuring the restoration effect are still higher than the threshold after the training is completed, generate multiple simulated audios, and train the cascaded first-stage model and the second-stage model through the multiple simulated audios.
[0151] In this embodiment, training is performed using simulated audio, or using a generative adversarial network. The cascaded first- and second-stage models serve as the generator, while the first and second discriminators serve as the discriminators. During this stage, the learning rate is reduced, the model is fine-tuned, and only the parameters of the second-stage model are updated.
[0152] In S2040 , the first and second discriminators are updated using the discriminator loss function, and the second stage model is updated using the second stage loss function.
[0153] According to the present disclosure, the provision of an encoder and decoder in the first-stage model facilitates the extraction of high-dimensional features from damaged audio, thereby repairing various audio distortion issues. Furthermore, the provision of full-band and sub-band modules in the second-stage model facilitates the repair of damaged audio across all frequency bands, enabling audio distortion issues in all frequency bands to be addressed and repaired.
[0154] FIG23 shows a schematic structural diagram of an audio repair device according to an embodiment of the present application.
[0155] As shown in Figure 23, according to this embodiment, the audio restoration model includes a first-stage model and a second-stage model. The audio restoration device 2300 includes: a first-stage module 2310 for inputting the damaged audio into the first-stage model to obtain a complex spectrum; and a second-stage module 2320 for inputting the complex spectrum into the second-stage model to obtain repaired audio. In some embodiments, the first-stage model includes an encoder for downsampling the complex spectrum converted from the damaged audio; and a decoder for upsampling the complex spectrum. The second-stage model includes a full-band module for modeling the complex spectrum across all frequency bands; and a subband module for modeling the complex spectrum across multiple subbands.
[0156] It should be noted that the audio repair device 2300 provided in the embodiment shown in FIG23 only illustrates the division of the aforementioned functional modules when executing the audio repair method. In actual applications, the aforementioned functions can be assigned to different functional modules as needed, i.e., the internal structure of the device can be divided into different functional modules to perform all or part of the functions described above. Furthermore, the audio repair device 400 provided in the above embodiment and the audio repair method embodiment shown in FIG1 each share the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0157] 24 shows a schematic diagram of an electronic device 2400 suitable for implementing the audio repair method according to the embodiment of FIG19. The electronic device shown in FIG24 is merely an example and should not limit the functionality and scope of use of the embodiments of the present application.
[0158] As shown in Figure 24, the electronic device 2400 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 2401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 2402 or a program loaded from a storage device 2408 into a random access memory (RAM) 2403. Various programs and data required for the operation of the electronic device 2400 are also stored in the RAM 2403. The processing device 2401, the ROM 2402, and the RAM 2403 are connected to each other via a bus 2404. An input / output (I / O) interface 2405 is also connected to the bus 2404.
[0159] Typically, the following devices can be connected to the I / O interface 2405: an input device 2406 including, for example, a touch screen, a touchpad, a keyboard, a mouse, etc.; an output device 2407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 2408 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 2409. The communication device 2409 can allow the electronic device 2400 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 24 shows an electronic device 2400 with various devices, it should be understood that it is not required to implement or have all the devices shown. More or fewer devices may be implemented or have alternatively. Each box shown in Figure 24 may represent one device, or may represent multiple devices as needed.
[0160] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for executing the audio repair method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication device 2409, or installed from the storage device 2408, or installed from the ROM 2402. When the computer program is executed by the processing device 2401, the above-mentioned functions defined in the audio repair method of the embodiment of the present application are performed.
[0161] The present disclosure also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed in a computer, the computer is caused to execute the audio repair method provided in this specification.
[0162] It should be noted that the computer-readable medium provided by the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the embodiments of this specification, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or device. In the embodiments of this specification, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or convey a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code embodied on the computer-readable medium may be conveyed using any suitable medium, including but not limited to wires, optical cables, RF (Radio Frequency), or any suitable combination thereof.
[0163] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device. The computer-readable medium carries one or more programs. When executed by the server, the electronic device: inputs the damaged audio into a first-stage model to obtain a complex spectrum, wherein the first-stage model includes an encoder and a decoder, the encoder is used to downsample the complex spectrum, and the decoder is used to upsample the complex spectrum; inputs the complex spectrum into a second-stage model to obtain repaired audio, wherein the second-stage model includes a full-band module and a sub-band module, the sub-band module is used to model the complex spectrum in multiple sub-bands, and the full-band module is used to model the complex spectrum in all frequency bands.
[0164] Computer program code for performing the operations of embodiments of the present specification may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0165] The various embodiments in this specification are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other, and each embodiment focuses on the differences between the other embodiments. In particular, the storage medium and computing device embodiments are described briefly because they are generally similar to the method embodiments. For relevant portions, refer to the description of the method embodiments.
[0166] Those skilled in the art will appreciate that, in one or more of the above examples, the functions described in the embodiments of the present disclosure may be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions may be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.
[0167] The specific implementation methods described above further illustrate the purpose, technical solutions, and beneficial effects of the embodiments of the present disclosure. It should be understood that the above description is only a specific implementation method of the embodiments of the present disclosure and is not intended to limit the scope of protection of the present disclosure. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the present disclosure shall be included in the scope of protection of the present disclosure.
Claims
1. An audio repair method, comprising: Input the damaged audio into the first-stage model to obtain the complex spectrum; as well as Inputting the complex spectrum into the second-stage model to obtain repaired audio; The first-stage model includes: an encoder for downsampling the complex spectrum; and a decoder for upsampling the complex spectrum; The second-stage model includes: a subband module for modeling the complex spectrum over a plurality of subbands; and A full-band module is used to model the complex spectrum over all frequency bands.
2. The method according to claim 1, wherein the first stage model further comprises: A short-time Fourier transform module, configured to transform the damaged audio into a complex spectrum through short-time Fourier transform; as well as The time series modeling module is used to further extract the features of the complex spectrum in the time dimension.
3. The method of claim 2, wherein the encoder comprises a plurality of stacked downsampling modules, the downsampling modules comprising: a two-dimensional gated convolution module, configured to perform gated convolution on the complex spectrum; A time-frequency convolution module, configured to perform convolution on the complex spectrum in the time dimension and the frequency dimension; as well as An axial self-attention module, for calculating an attention mechanism on the complex spectrum; The decoder includes a plurality of stacked upsampling modules, each of which includes: a two-dimensional gated transposed convolution module, configured to perform gated transposed convolution on the complex spectrum; a time-frequency convolution module, configured to convolve the complex spectrum in time and frequency dimensions; and An axial self-attention module is used to calculate the attention mechanism on the complex spectrum.
4. The method according to any one of claims 1 to 3, wherein the second stage model further comprises: A complex feature encoder, configured to extract high-dimensional features from the complex spectrum; as well as A complex feature decoder is used to restore the high-dimensional features in the complex spectrum to low dimensions.
5. The method according to any one of claims 1 to 4, wherein before inputting the impaired audio into the first-stage model to obtain the complex spectrum, the method further comprises: The first stage model is trained using a generative adversarial network training method; Train the second-stage model on a specific audio restoration task; Using a generative adversarial network training method to train the cascaded first-stage model and the second-stage model; as well as The input audios corresponding to the loss function for measuring the restoration effect are simulated and the values of the loss functions are still higher than the threshold after the training is completed, and multiple simulated audios are generated. The cascaded first-stage model and the second-stage model are trained by the multiple simulated audios.
6. The method according to claim 5, wherein the training of the first-stage model using a generative adversarial network training method comprises: Using the first-stage model as a generator of a generative adversarial network and using the first discriminator as a discriminator of the generative adversarial network, thereby training the first-stage model using a generative adversarial network training method; The first discriminator includes a plurality of first sub-discriminators, each of which includes a plurality of stacked two-dimensional convolution modules, and the two-dimensional convolution module is used to perform convolution on the amplitude spectrum of the input audio.
7. The method according to claim 5, wherein the training of the first-stage model using a generative adversarial network comprises: Using the first-stage model as a generator of a generative adversarial network and the second discriminator as a discriminator of the generative adversarial network, thereby training the first-stage model using a generative adversarial network training method; The second discriminator includes multiple second sub-discriminators, the second sub-discriminator includes multiple parallel sub-band discriminator modules, different sub-band discriminator modules are used to discriminate the complex spectra of different frequency bands into which the input audio is divided, and the sub-band discriminator module includes multiple stacked two-dimensional convolution modules, and the two-dimensional convolution module is used to convolve the complex spectrum.
8. An audio repair device, comprising: The first stage module is used to input the damaged audio into the first stage model to obtain the complex spectrum; as well as A second-stage module, configured to input the complex spectrum into a second-stage model to obtain a repaired audio; The first-stage model includes: An encoder configured to downsample the complex spectrum converted from the damaged audio; and a decoder for upsampling the complex spectrum; The second-stage model includes: a full-band module for modeling the complex spectrum over all frequency bands; and A subband module is used to model the complex spectrum on multiple subbands.
9. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the audio repair method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the audio repair method according to any one of claims 1 to 7.
11. An electronic device comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the audio repair method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Complex field speech enhancement method and system based on generative adversarial network and medium
CN110739002A
Voice processing method and device and device for processing voice
CN114566180A
Audio enhancement method and device, electronic equipment and readable storage medium
CN114974292A
Speech enhancement method, electronic equipment and storage medium
CN116013343A
Audio signal processing method and device, equipment and storage medium
CN116229999A
Cited By
Speech synthesis method and system of two-stage neural vocoder, terminal and medium
CN121034282A