An audio signal processing method, device, computer device and storage medium
By introducing the loss terms of the prior encoder and the posterior encoder in a vocoder system based on a variational autoencoder and combining it with the backpropagation algorithm, the problems of difficult training and poor generalization of the GAN vocoder are solved, and more efficient audio signal reconstruction and generation are achieved.
Patent Information
- Application Number
- CN202411492993.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-23
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-10-23
AI Technical Summary
Existing GAN vocoders have problems with training difficulty and poor generalization during the training process, resulting in unsatisfactory spectrum restoration effects when processing spectrums of different audio spectra.
A vocoder system based on variational autoencoder (VAE) is adopted. By introducing a loss term between the prior encoder and the posterior encoder and combining it with the backpropagation algorithm to optimize the loss function, the problems of difficult training and poor generalization in the existing technology are solved.
It improves the stability of the model, reduces the difficulty of training, enhances the generalization ability of the model, and provides more efficient audio signal reconstruction effects.
Smart Images

Figure CN119314499B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of artificial intelligence technology, and specifically relates to an audio signal processing method, apparatus, computer equipment, and storage medium. Background Art
[0002] In the field of audio signal processing, restoring phase-depleted amplitude spectra to audio has always been a difficult problem. Systems that perform this task are generally called vocoders. Traditional vocoders are typically implemented based on the Griffin-Lim algorithm, a classic algorithm in this field. This algorithm uses a random initial phase spectrum and multiple STFT-ISTFT iterations to gradually obtain a restored phase spectrum. The Griffin-Lim algorithm is simple and easy to implement, but its audio generation performance is unsatisfactory. Furthermore, the Griffin-Lim algorithm only works with linear amplitude spectra and is not applicable to Mel-frequency spectra.
[0003] In recent years, with the development of deep learning technology, a series of deep learning-based vocoder systems have emerged, among which the most effective ones are those based on Generative Adversarial Networks (GANs). GAN-based vocoders introduce a discriminator network to jointly train with the generator network for adversarial learning, thereby gradually generating more realistic audio during the learning process. However, GANs are difficult to train and require careful adjustment of training hyperparameters. Otherwise, gradient vanishing or gradient exploding problems will occur during training, affecting the final results. In addition, GAN-based vocoders have poor generalization capabilities and cannot effectively restore audio spectra outside of the training samples, making it difficult to generate audio with reasonable phases. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to propose an audio signal processing method, apparatus, computer device and storage medium to solve the problems of difficult training and poor generalization of existing GAN vocoders.
[0005] In order to solve the above technical problems, the present invention provides an audio signal processing method, which adopts the following technical solutions:
[0006] An audio signal processing method is applied to a pre-built vocoder system. The vocoder system is built based on a variational autoencoder. The vocoder system includes a priori encoder, a posteriori encoder, and a decoder. The audio signal processing method includes:
[0007] Constructing a loss function of a vocoder system, wherein the loss function of the vocoder system includes a first loss term and a second loss term, the first loss term represents an audio signal reconstruction loss, and the second loss term represents a loss between a priori encoder and a posteriori encoder;
[0008] Acquire pre-collected training data, wherein the training data is pre-collected original audio signals;
[0009] Input the original audio signal into the posterior encoder to obtain the posterior Gaussian distribution corresponding to the original audio signal;
[0010] Sampling the posterior Gaussian distribution corresponding to the original audio signal to obtain the first intermediate latent variable;
[0011] Inputting the first intermediate latent variable into a decoder for audio decoding to obtain a first restored audio signal;
[0012] Calculating a vocoder loss using a loss function of a vocoder system based on the first restored audio signal and the original audio signal;
[0013] A back-propagation algorithm is used to transfer the vocoder loss in the vocoder system, and the vocoder system is iteratively trained to obtain a pre-trained vocoder system;
[0014] An audio restoration instruction is received, a spectrum to be restored is obtained, and the spectrum to be restored is input into a pre-trained vocoder system to obtain a second restored audio signal.
[0015] In order to solve the above technical problems, the embodiment of the present application further provides an audio signal processing device, which adopts the following technical solution:
[0016] An audio signal processing device is used to run a pre-built vocoder system based on a variational autoencoder. The vocoder system includes a priori encoder, a posteriori encoder, and a decoder. The audio signal processing device includes:
[0017] A loss function construction module is used to construct a loss function of a vocoder system, wherein the loss function of the vocoder system includes a first loss term and a second loss term, the first loss term represents an audio signal reconstruction loss, and the second loss term represents a loss between a priori encoder and a posteriori encoder;
[0018] A training data acquisition module is used to acquire pre-collected training data, wherein the training data is pre-collected original audio signals;
[0019] The posterior encoding module is used to input the original audio signal into the posterior encoder to obtain the posterior Gaussian distribution corresponding to the original audio signal;
[0020] A latent variable sampling module is used to sample the posterior Gaussian distribution corresponding to the original audio signal to obtain a first intermediate latent variable;
[0021] an audio decoding module, configured to input the first intermediate latent variable into a decoder for audio decoding to obtain a first restored audio signal;
[0022] a loss calculation module, configured to calculate a vocoder loss by using a loss function of a vocoder system based on the first restored audio signal and the original audio signal;
[0023] an iterative training module for propagating vocoder loss in the vocoder system using a back-propagation algorithm and iteratively training the vocoder system to obtain a pre-trained vocoder system;
[0024] The audio restoration module is used to receive an audio restoration instruction, obtain a spectrum to be restored, and input the spectrum to be restored into a pre-trained vocoder system to obtain a second restored audio signal.
[0025] In order to solve the above technical problems, the embodiment of the present application further provides a computer device, which adopts the following technical solution:
[0026] A computer device includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of any one of the above-mentioned audio signal processing methods when executing the computer-readable instructions.
[0027] In order to solve the above technical problems, the embodiment of the present application further provides a computer-readable storage medium, which adopts the following technical solution:
[0028] A computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of any one of the above-mentioned audio signal processing methods.
[0029] Compared with the prior art, the embodiments of the present application have the following beneficial effects:
[0030] The present application discloses an audio signal processing method, device, computer equipment and storage medium, which belongs to the field of artificial intelligence technology. The present application is based on a vocoder system constructed based on a variational autoencoder (VAE), which improves the existing GAN vocoder in terms of training difficulty and poor generalization. By encoding the audio signal into a Gaussian distribution and introducing a loss term between the prior encoder and the posterior encoder, the stability of the model is effectively improved and the training difficulty is reduced. At the same time, the backpropagation algorithm is used to iteratively optimize the loss function, ensuring that the vocoder system can learn the audio signal features more efficiently, avoiding the common mode collapse problem in the GAN vocoder, and enhancing the generalization ability of the model through latent variable sampling, so that it can better handle different audio inputs and provide more accurate audio reconstruction effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the solutions in this application, a brief introduction will be given below to the drawings required for use in the description of the embodiments of this application. Obviously, the drawings described below are some embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0032] Figure 1 shows an exemplary system architecture diagram in which the present application can be applied;
[0033] Figure 2 The figure shows a framework diagram of a vocoder system based on a variational autoencoder according to the present application;
[0034] Figure 3 A flowchart of an embodiment of an audio signal processing method according to the present application is shown;
[0035] Figure 4 A schematic structural diagram of an embodiment of an audio signal processing device according to the present application is shown;
[0036] Figure 5 A schematic structural diagram of an embodiment of a computer device according to the present application is shown. DETAILED DESCRIPTION
[0037] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as commonly understood by those skilled in the art to which this application belongs. The terms used in the specification of the application are for the purpose of describing specific embodiments only and are not intended to limit this application. The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of this application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.
[0038] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0039] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.
[0040] like Figure 1 As shown, system architecture 100 may include a terminal device 101, a network 102, and a server 103. Terminal device 101 may be a laptop computer 1011, a tablet computer 1012, or a mobile phone 1013. Network 102 is a medium for providing a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0041] The user can use the terminal device 101 to interact with the server 103 via the network 102 to receive or send messages, etc. Various communication client applications can be installed on the terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0042] The terminal device 101 can be various electronic devices with a display screen and supporting web browsing. In addition to the laptop computer 1011, tablet computer 1012 or mobile phone 1013, the terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer and a desktop computer, etc.
[0043] The server 103 may be a server that provides various services, such as a background server that provides support for web pages displayed on the terminal device 101 .
[0044] It should be noted that the audio signal processing method provided in the embodiments of the present application is generally executed by a server / terminal device, and accordingly, the audio signal processing device is generally provided in the server / terminal device.
[0045] It should be understood that Figure 1 The numbers of terminal devices, networks and servers in the embodiment are merely illustrative. The above system may have any number of terminal devices, networks and servers according to implementation requirements.
[0046] In this embodiment, the present application discloses an audio signal processing method, which is applied to a pre-built vocoder system. Figure 2 , Figure 2 The framework diagram of the vocoder system based on the variational autoencoder according to the present application is shown. The vocoder system is built based on the variational autoencoder (VAE). The vocoder system includes a priori encoder, a posterior encoder and a decoder. The variational autoencoder VAE is a generative model based on probability coding and has been widely used in unsupervised learning. VAE combines the ideas of autoencoders and probability coding, and realizes a more flexible and controllable sample generation process by modeling the potential representation. VAE assumes that the potential representation obeys a prior distribution (usually a Gaussian distribution) and maps the input data to the distribution parameters of the latent space (such as mean and variance) through the encoder. The decoder samples from the latent space and maps the sampling results back to the original input space to generate new samples.
[0047] VAE mainly consists of two parts: encoder and decoder. The encoder is responsible for mapping the input data to the distribution parameters of the latent space, and the decoder is responsible for sampling from the latent space and mapping the sampling results back to the original input space.
[0048] The training goal of VAE is to maximize the likelihood of a certain data distribution, while introducing the idea of variational inference and optimizing the variational lower bound (ELBO). This is achieved by minimizing the reconstruction error (i.e., the difference between the samples generated by the decoder and the original input data) and regularizing the latent space (usually using KL divergence to measure the difference between the distribution of the encoder output and the prior distribution).
[0049] The a priori encoder generates a prior distribution for latent variables. By analyzing the structure of the input signal, it establishes a distribution hypothesis for the latent variables to ensure that the generated latent variables have reasonable distribution characteristics, which helps improve the decoder's generation performance. The posterior encoder calculates the corresponding latent variable distribution (usually a Gaussian distribution) based on the input audio signal. By learning the characteristics of the original data, the posterior encoder generates a more accurate latent variable distribution, helping the model better reconstruct the audio signal. The decoder is the module in the VAE that converts the latent variables back to the original audio signal. By inputting the latent variables and decoding them, it attempts to restore an audio signal similar to the input, completing the task of reconstructing or generating audio data.
[0050] exist Figure 2 In the figure, the left half represents the model training process, and the right half represents the model inference process. During the model training process, the spectrum to be restored is input into the prior encoder to obtain the mean μ′ and standard deviation σ′ of the prior Gaussian distribution; the original audio signal is input into the encoder to obtain the mean μ and standard deviation σ of the estimated posterior Gaussian distribution. An intermediate latent variable Z is sampled from the estimated posterior Gaussian distribution and Z is input into the decoder to obtain the restored audio signal.
[0051] During inference, the spectrum to be restored is input into the prior encoder, which obtains the mean μ′ and standard deviation σ′ of the prior Gaussian distribution. An intermediate latent variable Z′ is sampled from the prior Gaussian distribution and input into the decoder to obtain the restored audio signal. It is important to note that unlike typical VAE models, the model's prior distribution is learnable and relies on the spectrum as input. This process enhances the model's overall flexibility and robustness to different types of audio.
[0052] Continue to refer Figure 3 , shows a flow chart of an embodiment of a method for processing an audio signal according to the present application. The audio signal processing method comprises the following steps:
[0053] S301, constructing a loss function of a vocoder system, wherein the loss function of the vocoder system includes a first loss term and a second loss term, the first loss term represents the audio signal reconstruction loss, and the second loss term represents the loss between the a priori encoder and the posteriori encoder.
[0054] Specifically, in the vocoder system, the loss function is the key optimization objective. To ensure that the vocoder can effectively reconstruct and generate audio signals, the loss function of the vocoder system is usually set to include two main parts, namely the first loss term and the second loss term. The first loss term is the audio signal reconstruction loss, which measures the difference between the restored audio signal generated by the decoder and the original input audio signal. The mean square error or other distance metric functions are often used to calculate the difference, with the goal of making the restored audio signal as close as possible to the input signal. The second loss term focuses on the difference between the prior encoder and the posterior encoder. The KL divergence (Kullback-Leibler Divergence) is usually used to calculate the difference between the distributions generated by the two. This loss term aims to constrain the latent space of the model so that the distribution of the latent variables conforms to the assumptions of the prior distribution, thereby enhancing the model's generation ability and generalization performance. By optimizing these two loss terms, the vocoder system can accurately reconstruct the audio signal while generating a model with a reasonable latent variable distribution.
[0055] Furthermore, the loss function of the vocoder system is expressed as follows:
[0056] F=-log p θ (x|z)+KL(q φ (z|x)|p γ (z|c))
[0057] Where F is the loss function of the vocoder system, x is the audio sampling signal, c is the spectral feature, -log p θ (x|z) is the first loss term, which represents the difference between the sample generated by the decoder and the original input data, KL(q φ (z|x)|p γ (z|c)) is the second loss term, which is used to measure the difference between the distribution of the encoder output and the prior distribution, p θ (x|z) represents the probability of observation data x under given latent variable z, q φ (z|x) represents the posterior probability distribution of the latent variable z given the observation data x, p γ (z|c) represents the prior probability distribution of the latent variable z, and θ, φ, and γ are the parameters of the decoder, posterior encoder, and prior encoder, respectively.
[0058] Furthermore, the second loss term is expressed as follows:
[0059]
[0060] Where μ and σ are the mean and standard deviation of the posterior Gaussian distribution, and μ′ and σ′ are the mean and standard deviation of the prior Gaussian distribution.
[0061] S302: Acquire pre-collected training data, where the training data is pre-collected original audio signals.
[0062] Specifically, training a vocoder system relies on a large amount of high-quality audio data. To train a robust vocoder model, it's necessary to collect and organize audio signals as training data. Raw audio signals can be audio clips of various types and sources, such as music, speech, and ambient sound, covering a wide range of frequencies, pitches, and complexities. When collecting this data, it's important to consider the audio data's format, sampling rate, and data volume, ensuring both diversity and representativeness to ensure the model can adapt to a variety of input scenarios.
[0063] Furthermore, training data requires preprocessing, such as noise reduction, normalization, or data augmentation, to improve the model's robustness to noise and audio signal variation. The quality and quantity of training data directly impact the performance of the vocoder system, as the model uses this data to learn features, grasp the time-frequency characteristics and underlying structural patterns of audio signals, and ultimately improve its ability to generalize to unseen audio.
[0064] S303: Input the original audio signal into a posterior encoder to obtain a posterior Gaussian distribution corresponding to the original audio signal.
[0065] Specifically, in this step, the vocoder system processes the original audio signal using a posterior encoder, converting the input high-dimensional audio signal into a corresponding posterior Gaussian distribution. The posterior encoder is responsible for mapping the complex audio signal into a latent space and generating a potential representation of the audio signal by learning the time-frequency characteristics and structure of the audio signal. This representation usually exists in the form of a Gaussian distribution, defining the mean and variance of each audio segment in the latent space. The output of the posterior encoder is this posterior Gaussian distribution, which reflects the model's understanding of the input audio signal and allows the vocoder system to perform further operations on the audio signal in the latent space. The key to this step is to extract the core features of the audio signal through the posterior distribution, while introducing a certain amount of randomness, so that the model can generate new or reconstructed audio signals during the subsequent decoding process, thus possessing generative and generalization capabilities.
[0066] S304: Sampling the posterior Gaussian distribution corresponding to the original audio signal to obtain a first intermediate latent variable.
[0067] Specifically, once the posterior Gaussian distribution corresponding to the original audio signal is obtained, it is necessary to sample from this distribution to generate the first intermediate latent variable. Sampling is the process of converting the features (mean and variance) in the Gaussian distribution into a specific latent variable value. This latent variable is a simplified representation of the audio signal in the latent space. The sampling operation introduces randomness, which is also an important feature of the VAE model, allowing the generator to generate slightly different audio signals through different sampling processes. This randomness enhances the generation ability of the model while retaining the main features of the original audio signal. The sampling result of the latent variable has a lower dimension in the latent space and retains the core structural information of the input audio signal, which enables the decoding process to effectively reconstruct the audio. This latent variable can not only be used to restore the input audio signal, but also to generate new audio signals.
[0068] S305: Input the first intermediate latent variable into a decoder for audio decoding to obtain a first restored audio signal.
[0069] Specifically, after obtaining the intermediate latent variables, the next step of the system is to convert these latent variables back to audio signals through the decoder. As an important component of the vocoder system, the decoder's task is to map the low-dimensional representation (latent variables) in the latent space back to the high-dimensional audio signal space. Specifically, the decoder generates corresponding audio features based on the latent variables, and gradually restores the time domain or frequency domain features similar to the original audio signal through the learned parameters. The decoding process requires the model to effectively capture the key information in the latent variables to ensure that the restored audio signal has a high auditory similarity to the original signal. The first restored audio signal generated is not completely accurate to the original audio, but by optimizing the loss function, it can be made as close to the original audio signal as possible. The performance of the decoder directly affects the reconstruction quality of the audio signal. Therefore, during the training process, the decoder needs to be fully optimized to ensure that it has strong reconstruction capabilities and generation effects.
[0070] S306 : Calculate a vocoder loss using a loss function of a vocoder system based on the first restored audio signal and the original audio signal.
[0071] Specifically, after generating the first restored audio signal, the next step is to calculate the loss of the model through the loss function of the vocoder system. The loss function is used to measure the difference between the restored audio signal output by the model and the original audio signal input. The first loss (reconstruction loss) is used to evaluate the similarity between the restored audio signal and the original audio signal. The mean square error (MSE) or other error metrics are usually used to measure the difference between the two. The purpose is to make the restored audio as consistent as possible with the original audio. The second loss measures the difference between the distributions generated by the prior encoder and the posterior encoder, and is usually calculated using the KL divergence. By combining these two parts of the loss, the system can simultaneously optimize the reconstruction quality of the audio signal and the rationality of the latent variable distribution. The calculation results of the loss function will be used to guide the parameter update of the model to gradually improve the performance of the vocoder system in audio signal processing, enabling it to more accurately reconstruct and generate audio signals.
[0072] S307 , using a back propagation algorithm to propagate the vocoder loss in the vocoder system, and iteratively training the vocoder system to obtain a pre-trained vocoder system.
[0073] Specifically, in order to optimize the vocoder system, the model needs to be trained using the backpropagation algorithm. In this step, based on the loss calculated in the previous step, the loss is propagated back to each layer of the model through the backpropagation algorithm, and the parameters in the model are gradually updated. The backpropagation algorithm calculates the gradient of the loss function with respect to each parameter through the chain rule, indicating how to adjust the weights of the model to minimize the loss. Each component of the vocoder system (a priori encoder, a posteriori encoder, and decoder) will be optimized in this process. The iterative training and update process will be repeated many times. By continuously adjusting the model parameters, the model can gradually improve the reconstruction accuracy and generalization ability of the audio signal. After sufficient iterative training, the vocoder system can learn efficient feature representations from the training data, thereby achieving high-quality audio signal processing and reconstruction, and finally obtaining a pre-trained, stable performance vocoder system.
[0074] S308: Receive an audio restoration instruction, obtain a frequency spectrum to be restored, and input the frequency spectrum to be restored into a pre-trained vocoder system to obtain a second restored audio signal.
[0075] Specifically, after the training of the vocoder system is completed, the system can accept audio restoration instructions initiated by the user. At this time, the user can input the audio spectrum to be restored as the input signal of the model. The pre-trained vocoder system has learned the ability to process audio signals through a large amount of training data, so it can efficiently convert the input spectrum information into the corresponding restored audio signal. The input spectrum to be restored will first pass through the encoder module of the vocoder system to extract the corresponding latent variables, and then pass through the decoder to generate the final audio signal. Since the system has been fully trained, the generated second restored audio signal should have a high similarity and quality with the real audio. The core of this process is to use the generation ability of the vocoder system to reasonably reconstruct the input spectrum to achieve the purpose of efficient audio restoration.
[0076] In the above embodiments, the present application uses a vocoder system constructed based on a variational autoencoder (VAE) to improve the existing GAN vocoder's problems of difficult training and poor generalization. By encoding the audio signal into a Gaussian distribution and introducing a loss term between the prior encoder and the posterior encoder, the stability of the model is effectively improved and the training difficulty is reduced. At the same time, the backpropagation algorithm is used to iteratively optimize the loss function, ensuring that the vocoder system can learn audio signal features more efficiently, avoiding the common mode collapse problem in GAN vocoders. The model's generalization ability is also enhanced through latent variable sampling, enabling it to better handle different audio inputs and provide more accurate audio reconstruction effects.
[0077] Furthermore, after the step of inputting the original audio signal into the posterior encoder to obtain the posterior Gaussian distribution corresponding to the original audio signal, the method further includes:
[0078] performing Fourier transform on the original audio signal to obtain an original audio spectrum corresponding to the original audio signal;
[0079] inputting the original audio spectrum into the prior encoder to obtain a prior Gaussian distribution corresponding to the original audio spectrum.
[0080] In this embodiment, after inputting the original audio signal into the posterior encoder to obtain the corresponding posterior Gaussian distribution, Fourier transform is also performed on the original audio signal to obtain its spectral representation. Fourier transform is a common signal processing technique that converts time-domain audio signals into spectral information in the frequency domain, thereby revealing the components and characteristics of the audio signal at different frequencies. For audio processing, spectral representation can more intuitively reflect the frequency structure of the signal, helping the model to more effectively extract and learn the features in the audio signal. The spectrum of the original audio signal is input into the prior encoder, which generates the corresponding prior Gaussian distribution based on the spectral data. Similar to the posterior encoder, the task of the prior encoder is to learn the hidden space representation based on the spectral information, and this representation appears in the form of a Gaussian distribution. As a prior assumption of the latent variable, the prior Gaussian distribution provides a reference distribution to help the model guide the learning of the posterior distribution when processing the audio signal, ensuring that the latent variables generated by the model conform to the characteristics of the audio signal and maintain the generation ability and generalization. The prior encoder and the posterior encoder work together to jointly model the time-domain and frequency-domain characteristics of the audio signal, helping the model to better understand the structure of the audio signal.
[0081] Through the above steps, the vocoder system can establish a more accurate mapping between the time-domain and frequency-domain characteristics of the audio signal, improving the model's processing ability of the audio signal. Combined with the Gaussian distributions generated by the prior encoder and the posterior encoder, the vocoder system can better balance the reconstruction accuracy and generation quality, thereby achieving more efficient audio signal reconstruction and generation. At the same time, this double encoding structure enhances the generalization ability of the model, making it perform more stably when processing complex audio inputs.
[0082] Further, the step of performing Fourier transform on the original audio signal to obtain an original audio spectrum corresponding to the original audio signal specifically comprises:
[0083] performing a preprocessing operation on the original audio signal, wherein the preprocessing operation includes a denoising operation, a framing operation, and a windowing operation;
[0084] performing short-time Fourier transform on the preprocessed original audio signal to obtain an initial audio spectrum;
[0085] performing Mel filtering on the initial audio spectrum to obtain a Mel spectrum;
[0086] Perform logarithm operation on the Mel spectrum and perform discrete cosine transform on the Mel spectrum after logarithm operation to obtain a Mel spectrum graph;
[0087] Get the original audio spectrum corresponding to the original audio signal from the Mel spectrogram.
[0088] In this embodiment, the original audio signal is first preprocessed. Preprocessing includes denoising, framing, and windowing. The denoising operation aims to eliminate background noise and irrelevant interference signals in the audio, ensuring that the extracted spectrum is clearer. The framing operation divides the continuous audio signal into several short time segments. The windowing operation applies a windowing function before each frame segment to reduce the impact of edge effects on the spectrum, ensuring the smoothness and accuracy of the spectrum.
[0089] Next, a short-time Fourier transform (STFT) is performed on the preprocessed audio signal, converting the audio signal of each time frame into the frequency domain to generate an initial audio spectrum. The STFT processes the audio in segments, enabling it to capture the changing characteristics of the audio signal in time and frequency. The initial audio spectrum is then Mel-filtered. This involves resampling the spectrum into a Mel-spectrum that better matches the human auditory characteristics through a set of Mel filters that simulate human auditory perception. The Mel-spectrum is designed to capture the key frequency characteristics of speech and audio signals, making the spectral representation of the audio signal more consistent with the human auditory system.
[0090] After obtaining the Mel-spectrogram, a nonlinear transformation is performed by taking its logarithm, converting the amplitude spectrum into a logarithmic spectrum to further simulate the human ear's perception of sound intensity. Subsequently, a discrete cosine transform (DCT) is performed on the logarithmic Mel-spectrogram to further compress the spectral information and extract the low-dimensional features of the audio signal. This operation helps remove redundant information from the spectrum, retaining the most representative audio features, and ultimately obtaining a Mel-spectrogram. By extracting the spectrum corresponding to the original audio signal from the Mel-spectrogram, a frequency representation of the audio signal is obtained. This spectrum not only contains the frequency information of the audio but also retains its key time domain characteristics, enabling better reconstruction or generation of the audio signal.
[0091] The above steps improve the accuracy and expressiveness of audio spectrum extraction. Using processes such as Mel filtering, logarithmic transforms, and discrete cosine transforms, the vocoder system can more efficiently extract key features from audio signals, improving the model's understanding and generalization capabilities. The resulting Mel spectrogram provides a more stable and accurate input for audio signal reconstruction and generation.
[0092] Furthermore, the step of calculating the vocoder loss by using the loss function of the vocoder system based on the first restored audio signal and the original audio signal specifically includes:
[0093] Calculating an error between the first restored audio signal and the original audio signal based on the first loss term to obtain a first loss;
[0094] Performing Fourier transform on the first restored audio signal to obtain a restored audio spectrum corresponding to the first restored audio signal;
[0095] Calculate the error between the restored audio spectrum and the original audio spectrum based on the second loss term to obtain a second loss;
[0096] The sum of the first loss and the second loss is calculated to obtain the vocoder loss.
[0097] In this embodiment, the error between the first restored audio signal and the original audio signal is first calculated based on the first loss term to obtain the first loss. The first loss is usually used to measure the reconstruction error in the time domain and the difference between the restored audio and the original audio. The mean square error or other difference measurement methods are usually used. Next, the first restored audio signal is Fourier transformed to obtain its frequency domain information and obtain the corresponding restored audio spectrum. In this way, the performance of the restored signal in the frequency space can be analyzed and compared with the original audio spectrum. Subsequently, the error between the restored audio spectrum and the original audio spectrum is calculated based on the second loss term to obtain the second loss. The Mel spectrum error or other frequency domain difference measurement is usually used. This step ensures that the model can not only reconstruct the time structure of the audio, but also match the original audio in frequency distribution. Finally, the system weightedly sums the first loss and the second loss to obtain the vocoder loss, ensuring that the model is optimized in both time and frequency dimensions.
[0098] Through the above steps, the system can optimize the reconstruction quality of audio in both the time domain and the frequency domain, thereby improving the overall performance of audio restoration.
[0099] The step of calculating the error between the first restored audio signal and the original audio signal based on the first loss term to obtain the first loss specifically includes:
[0100] According to a preset audio signal sampling rule, the first restored audio signal and the original audio signal are sampled respectively to obtain audio sampling signals;
[0101] According to a preset spectrum feature sampling rule, the original audio spectrum and the restored audio spectrum are sampled respectively to obtain audio spectrum features, wherein the audio signal sampling rule and the spectrum feature sampling rule match each other;
[0102] Decoder parameters are obtained, and an error between the first restored audio signal and the original audio signal is calculated based on the decoder parameters, the audio sampling signal, and the audio spectrum characteristics to obtain a first loss.
[0103] In this embodiment, by sampling the audio signal and sampling its spectral features, a comprehensive analysis of the audio signal can be performed from both the time domain and frequency domain perspectives. First, the first restored audio signal and the original audio signal are sampled separately according to a preset audio signal sampling rule. This effectively captures the time domain differences between the restored and original signals, ensuring that important audio information is not lost during the decoding process. Second, the original audio spectrum and the restored audio spectrum are sampled according to a preset spectral feature sampling rule. This further captures subtle differences between the signals from a frequency domain perspective, particularly frequency components sensitive to the human ear, and better reflects changes in sound quality. Subsequently, by introducing decoder parameters, the audio sampling signal is combined with the spectral features to calculate the error between the restored and original audio. Based on this, a loss function, namely the first loss term, is derived. This loss term comprehensively considers the errors in both the time and frequency domains, providing a basis for optimizing the audio restoration algorithm and improving the accuracy and quality of audio restoration.
[0104] Through the above steps, the distortion phenomenon in the audio restoration process can be effectively reduced, the restoration degree of the audio signal can be improved, and the sound quality can be improved, making the restored audio more realistic.
[0105] The step of calculating the error between the restored audio spectrum and the original audio spectrum based on the second loss term to obtain the second loss specifically includes:
[0106] Get the mean and standard deviation of the posterior Gaussian distribution, and get the mean and standard deviation of the prior Gaussian distribution;
[0107] Get the posterior encoder parameters and the prior encoder parameters;
[0108] Based on the posterior encoder parameters, the prior encoder parameters, the mean and standard deviation of the posterior Gaussian distribution, and the mean and standard deviation of the prior Gaussian distribution, calculate the KL divergence between the posterior Gaussian distribution and the prior Gaussian distribution;
[0109] The second loss is obtained by taking the KL divergence as the error between the restored audio spectrum and the original audio spectrum.
[0110] In this embodiment, first, the mean and standard deviation of the posterior prior Gaussian distribution and the prior Gaussian distribution are obtained, which can be used to characterize the probability distribution of the restored audio and the original audio in different states. Then, based on the posterior encoder parameters and the prior encoder parameters, these distribution information are used to calculate the KL divergence between the posterior Gaussian distribution and the prior Gaussian distribution. KL divergence is an indicator to measure the difference between two probability distributions, which can reflect the difference in spectral characteristics between the restored audio and the original audio. By using KL divergence as a metric for spectral error, the second loss function obtained can reflect the amount of information lost in the spectral restoration process. By comprehensively considering the frequency domain characteristics of the audio signal and the corresponding probability distribution differences, further optimization of the restored audio accuracy is achieved. In addition, the introduction of the posterior encoder and the prior encoder parameters also provides richer information and constraints for the optimization of the model, which helps to improve the robustness and effect of audio restoration.
[0111] Through the above steps, the error of the audio spectrum is effectively measured by KL divergence, which can enhance the accuracy of audio spectrum restoration, thereby better preserving the detailed characteristics of the audio during the restoration process.
[0112] Furthermore, the step of obtaining a spectrum to be restored and inputting the spectrum to be restored into a pre-trained vocoder system to obtain a second restored audio signal specifically includes:
[0113] Input the spectrum to be restored into the prior encoder to obtain the prior Gaussian distribution corresponding to the spectrum to be restored;
[0114] Sampling the prior Gaussian distribution corresponding to the spectrum to be restored to obtain the second intermediate latent variable;
[0115] The second intermediate latent variable is input into the decoder for audio decoding to obtain a second restored audio signal.
[0116] In this embodiment, the spectrum to be restored is first input into the prior encoder to obtain its corresponding prior Gaussian distribution. The task of the prior encoder is to map the spectral features to the latent space and generate an ideal Gaussian distribution by learning the statistical characteristics of the spectrum. This Gaussian distribution not only reflects the characteristics of the input spectrum but also provides an important potential representation for the decoding process. Next, sampling is performed from the obtained prior Gaussian distribution to generate a second intermediate latent variable. The generation of this latent variable has a certain degree of randomness, allowing the model to explore different possibilities in the latent space, so that the decoder can generate a variety of audio signals. Finally, the second intermediate latent variable is input into the decoder for audio decoding. The decoder converts the latent variable back into an audio signal to generate a second restored audio signal. In this process, the decoder effectively reconstructs audio similar to the original signal by utilizing the information in the latent variable, ensuring high-quality restoration of the audio signal while maintaining the signal characteristics.
[0117] Through the above steps, the system can effectively convert spectral information into audio signals, ensuring that the quality of the restored audio is consistent with the original signal, thereby improving the accuracy and diversity of audio reconstruction.
[0118] In the above embodiment, the present application discloses an audio signal processing method, which belongs to the field of artificial intelligence technology. The present application is based on a vocoder system constructed based on a variational autoencoder (VAE), which improves the existing GAN vocoder in terms of difficulty in training and poor generalization. By encoding the audio signal into a Gaussian distribution and introducing a loss term between the prior encoder and the posterior encoder, the stability of the model is effectively improved and the training difficulty is reduced. At the same time, the backpropagation algorithm is used to iteratively optimize the loss function, ensuring that the vocoder system can learn the audio signal features more efficiently, avoiding the common mode collapse problem in the GAN vocoder, and enhancing the generalization ability of the model through latent variable sampling, so that it can better process different audio inputs and provide more accurate audio reconstruction effects.
[0119] In this embodiment, the audio signal processing method is executed on the electronic device (eg Figure 1 The server shown in the figure) can receive instructions or obtain data through a wired connection or a wireless connection. It should be noted that the above-mentioned wireless connection method may include but is not limited to 3G / 4G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra wideband) connection, and other wireless connection methods currently known or to be developed in the future.
[0120] It should be emphasized that in order to further ensure the privacy and security of the above-mentioned spectrum information to be restored, the above-mentioned spectrum information to be restored can also be stored in a node of a blockchain.
[0121] The blockchain referred to in this application is a new application model for computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Blockchain is essentially a decentralized database, a series of data blocks generated using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of this information (to prevent counterfeiting) and generate the next block. Blockchain can include the underlying blockchain platform, the platform product service layer, and the application service layer.
[0122] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.
[0123] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0124] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware via computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes in the above-described method embodiments. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0125] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
[0126] Further references Figure 4 , as a response to the above Figure 3 The present application provides an embodiment of an audio signal processing device. Figure 3 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0127] like Figure 4As shown, the audio signal processing apparatus 400 described in this embodiment is used to run a pre-built vocoder system, the vocoder system is built based on a variational autoencoder, and the vocoder system includes a prior encoder, a posterior encoder and a decoder. The audio signal processing apparatus 400 includes:
[0128] The loss function construction module 401 is configured to construct a loss function of the vocoder system, wherein the loss function of the vocoder system includes a first loss term and a second loss term, the first loss term represents an audio signal reconstruction loss, and the second loss term represents a loss between the prior encoder and the posterior encoder.
[0129] The training data acquisition module 402 is configured to acquire pre-collected training data, wherein the training data are pre-collected original audio signals.
[0130] The posterior encoding module 403 is configured to input the original audio signal into the posterior encoder to obtain a posterior Gaussian distribution corresponding to the original audio signal.
[0131] The latent variable sampling module 404 is configured to sample the posterior Gaussian distribution corresponding to the original audio signal to obtain a first intermediate latent variable.
[0132] The audio decoding module 405 is configured to input the first intermediate latent variable into the decoder for audio decoding to obtain a first restored audio signal.
[0133] The loss calculation module 406 is configured to calculate a vocoder loss based on the first restored audio signal and the original audio signal through the loss function of the vocoder system.
[0134] The iterative training module 407 is configured to use a back propagation algorithm to pass the vocoder loss in the vocoder system and iteratively train the vocoder system to obtain a pre-trained vocoder system.
[0135] The audio restoration module 408 is configured to receive an audio restoration instruction, acquire a to-be-restored spectrum, input the to-be-restored spectrum into the pre-trained vocoder system, and obtain a second restored audio signal.
[0136] In the above embodiment, the present application discloses an audio signal processing device, which belongs to the field of artificial intelligence technology. The present application is based on a vocoder system constructed based on a variational autoencoder (VAE), which improves the existing GAN vocoder in terms of difficulty in training and poor generalization. By encoding the audio signal into a Gaussian distribution and introducing a loss term between the prior encoder and the posterior encoder, the stability of the model is effectively improved and the difficulty of training is reduced. At the same time, the backpropagation algorithm is used to iteratively optimize the loss function, ensuring that the vocoder system can learn the audio signal features more efficiently, avoiding the common mode collapse problem in the GAN vocoder, and enhancing the generalization ability of the model through latent variable sampling, so that it can better process different audio inputs and provide more accurate audio reconstruction effects.
[0137] To solve the above technical problems, the present application also provides a computer device. Figure 5 , Figure 5 This is a basic structural block diagram of the computer device in this embodiment.
[0138] The computer device 5 includes a memory 51, a processor 52, and a network interface 53 that are interconnected through a system bus. It should be noted that the figure only shows a computer device 5 with a memory 51, a processor 52, and a network interface 53, but it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Among them, those skilled in the art can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to a microprocessor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc.
[0139] The computer device may be a desktop computer, notebook computer, PDA, cloud server, etc. The computer device may interact with the user via a keyboard, mouse, remote control, touchpad, or voice control device.
[0140] The memory 51 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic storage, magnetic disk, optical disk, etc. In some embodiments, the memory 51 can be an internal storage unit of the computer device 5, such as the hard disk or memory of the computer device 5. In other embodiments, the memory 51 can also be an external storage device of the computer device 5, such as a plug-in hard disk equipped on the computer device 5, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Of course, the memory 51 can also include both the internal storage unit of the computer device 5 and its external storage device. In this embodiment, the memory 51 is generally used to store the operating system and various application software installed on the computer device 5, such as computer-readable instructions for the audio signal processing method. In addition, the memory 51 can also be used to temporarily store various types of data that have been output or are to be output.
[0141] In some embodiments, the processor 52 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 52 is generally used to control the overall operation of the computer device 5. In this embodiment, the processor 52 is used to execute computer-readable instructions stored in the memory 51 or process data, such as computer-readable instructions for executing the audio signal processing method.
[0142] The network interface 53 may include a wireless network interface or a wired network interface. The network interface 53 is generally used to establish a communication connection between the computer device 5 and other electronic devices.
[0143] In the above embodiment, the present application discloses a computer device, which belongs to the field of artificial intelligence technology. The present application is based on a vocoder system constructed based on a variational autoencoder (VAE), which improves the existing GAN vocoder in terms of difficulty in training and poor generalization. By encoding the audio signal into a Gaussian distribution and introducing a loss term between the prior encoder and the posterior encoder, the stability of the model is effectively improved and the difficulty of training is reduced. At the same time, the backpropagation algorithm is used to iteratively optimize the loss function, ensuring that the vocoder system can learn the audio signal features more efficiently, avoiding the common mode collapse problem in the GAN vocoder, and enhancing the generalization ability of the model through latent variable sampling, so that it can better handle different audio inputs and provide more accurate audio reconstruction effects.
[0144] The present application also provides another embodiment, namely, providing a computer-readable storage medium, wherein the computer-readable storage medium stores computer-readable instructions, and the computer-readable instructions can be executed by at least one processor to enable the at least one processor to perform the steps of the audio signal processing method as described above.
[0145] In the above embodiment, the present application discloses a computer-readable storage medium, which belongs to the field of artificial intelligence technology. The present application is based on a vocoder system constructed based on a variational autoencoder (VAE), which improves the existing GAN vocoder in terms of difficulty in training and poor generalization. By encoding the audio signal into a Gaussian distribution and introducing a loss term between the prior encoder and the posterior encoder, the stability of the model is effectively improved and the training difficulty is reduced. At the same time, the backpropagation algorithm is used to iteratively optimize the loss function, ensuring that the vocoder system can learn the audio signal features more efficiently, avoiding the common mode collapse problem in the GAN vocoder, and enhancing the generalization ability of the model through latent variable sampling, so that it can better handle different audio inputs and provide more accurate audio reconstruction effects.
[0146] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0147] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0148] Obviously, the embodiments described above are only some of the embodiments of the present application, rather than all of the embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions described in the aforementioned specific embodiments, or to make equivalent replacements for some of the technical features therein. Any equivalent structure made using the contents of the present application specification and the accompanying drawings, directly or indirectly used in other related technical fields, is also within the scope of patent protection of the present application.
Claims
1. A method for processing an audio signal, characterized in that: The audio signal processing method is applied to a pre-built vocoder system, which is built based on a variational autoencoder. The vocoder system includes a priori encoder, a posteriori encoder, and a decoder. The audio signal processing method includes: Constructing a loss function of the vocoder system, wherein the loss function of the vocoder system includes a first loss term and a second loss term, the first loss term representing an audio signal reconstruction loss, and the second loss term representing a loss between the a priori encoder and the a posteriori encoder; Acquiring pre-collected training data, wherein the training data is pre-collected original audio signals; Inputting the original audio signal into the posterior encoder to obtain a posterior Gaussian distribution corresponding to the original audio signal; Sampling the posterior Gaussian distribution corresponding to the original audio signal to obtain a first intermediate latent variable; Inputting the first intermediate latent variable into the decoder for audio decoding to obtain a first restored audio signal; Calculating a vocoder loss using a loss function of the vocoder system based on the first restored audio signal and the original audio signal; Propagating the vocoder loss in the vocoder system using a back-propagation algorithm, and iteratively training the vocoder system to obtain a pre-trained vocoder system; An audio restoration instruction is received, a spectrum to be restored is obtained, and the spectrum to be restored is input into the pre-trained vocoder system to obtain a second restored audio signal.
2. The method for processing a frequency signal according to claim 1, wherein: After the step of inputting the original audio signal into the posterior encoder to obtain the posterior Gaussian distribution corresponding to the original audio signal, the method further includes: Performing Fourier transform on the original audio signal to obtain an original audio spectrum corresponding to the original audio signal; The original audio spectrum is input into the priori encoder to obtain a priori Gaussian distribution corresponding to the original audio spectrum.
3. The frequency signal processing method according to claim 2, wherein: The step of performing Fourier transform on the original audio signal to obtain an original audio spectrum corresponding to the original audio signal specifically includes: Performing a preprocessing operation on the original audio signal, wherein the preprocessing operation includes a denoising operation, a framing operation, and a windowing operation; Performing short-time Fourier transform on the preprocessed original audio signal to obtain an initial audio spectrum; Performing Mel filtering on the initial audio spectrum to obtain a Mel spectrum; Performing a logarithm operation on the Mel spectrum, and performing a discrete cosine transform on the Mel spectrum after taking the logarithm to obtain a Mel spectrum graph; An original audio spectrum corresponding to the original audio signal is obtained from the mel-spectrogram.
4. The method for processing a frequency signal according to claim 2, wherein: The step of calculating the vocoder loss by using the loss function of the vocoder system based on the first restored audio signal and the original audio signal specifically includes: calculating an error between the first restored audio signal and the original audio signal based on the first loss term to obtain a first loss; Performing Fourier transform on the first restored audio signal to obtain a restored audio spectrum corresponding to the first restored audio signal; calculating an error between the restored audio spectrum and the original audio spectrum based on the second loss term to obtain a second loss; The sum of the first loss and the second loss is calculated to obtain the vocoder loss.
5. The frequency signal processing method according to claim 4, wherein: The step of calculating the error between the first restored audio signal and the original audio signal based on the first loss term to obtain the first loss specifically includes: According to a preset audio signal sampling rule, sampling the first restored audio signal and the original audio signal respectively to obtain audio sampling signals; According to a preset spectrum feature sampling rule, the original audio spectrum and the restored audio spectrum are sampled respectively to obtain audio spectrum features, wherein the audio signal sampling rule and the spectrum feature sampling rule match each other; Decoder parameters are obtained, and an error between the first restored audio signal and the original audio signal is calculated based on the decoder parameters, the audio sampling signal, and the audio spectrum feature to obtain the first loss.
6. The method for processing a frequency signal according to claim 5, wherein: The step of calculating the error between the restored audio spectrum and the original audio spectrum based on the second loss term to obtain the second loss specifically includes: Obtaining the mean and standard deviation of the posterior Gaussian distribution, and obtaining the mean and standard deviation of the prior Gaussian distribution; Get the posterior encoder parameters and the prior encoder parameters; Calculating a KL divergence between the posterior Gaussian distribution and the prior Gaussian distribution based on the posterior encoder parameters, the prior encoder parameters, the mean and standard deviation of the posterior Gaussian distribution, and the mean and standard deviation of the prior Gaussian distribution; The KL divergence is used as the error between the restored audio spectrum and the original audio spectrum to obtain the second loss.
7. The method for processing a frequency signal according to claim 1, wherein: The step of obtaining a spectrum to be restored and inputting the spectrum to be restored into the pre-trained vocoder system to obtain a second restored audio signal specifically includes: Inputting the spectrum to be restored into the a priori encoder to obtain a priori Gaussian distribution corresponding to the spectrum to be restored; Sampling the prior Gaussian distribution corresponding to the spectrum to be restored to obtain a second intermediate latent variable; The second intermediate latent variable is input into the decoder for audio decoding to obtain a second restored audio signal.
8. An audio signal processing device, characterized in that: The audio signal processing device is used to run a pre-built vocoder system, which is built based on a variational autoencoder. The vocoder system includes a priori encoder, a posteriori encoder and a decoder. The audio signal processing device includes: a loss function construction module, configured to construct a loss function of the vocoder system, wherein the loss function of the vocoder system comprises a first loss term and a second loss term, wherein the first loss term represents an audio signal reconstruction loss, and the second loss term represents a loss between the a priori encoder and the a posteriori encoder; A training data acquisition module, configured to acquire pre-collected training data, wherein the training data is pre-collected original audio signals; a posterior encoding module, configured to input the original audio signal into the posterior encoder to obtain a posterior Gaussian distribution corresponding to the original audio signal; a latent variable sampling module, configured to sample the posterior Gaussian distribution corresponding to the original audio signal to obtain a first intermediate latent variable; an audio decoding module, configured to input the first intermediate latent variable into the decoder for audio decoding to obtain a first restored audio signal; a loss calculation module, configured to calculate a vocoder loss using a loss function of the vocoder system based on the first restored audio signal and the original audio signal; an iterative training module, configured to propagate the vocoder loss in the vocoder system using a back-propagation algorithm and iteratively train the vocoder system to obtain a pre-trained vocoder system; The audio restoration module is configured to receive an audio restoration instruction, obtain a frequency spectrum to be restored, and input the frequency spectrum to be restored into the pre-trained vocoder system to obtain a second restored audio signal.
9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of the audio signal processing method according to any one of claims 1 to 7 when executing the computer-readable instructions.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the audio signal processing method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Training method of dialogue generation model and dialogue generation method and device
CN110457457A
Method and device for realizing vocoder based on variational auto-encoder
CN111724809A