Audio signal enhancement method, device and system, computer equipment and storage medium
Through the audio enhancement model combined with the autoencoder and discriminator, the problem of signal distortion in speech enhancement is solved, and a higher quality voice enhancement effect is achieved.
Patent Information
- Application Number
- CN202311633210.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-30
- Publication Date
- 2025-07-08
AI Technical Summary
The speech enhancement method based on non-negative matrix decomposition in the prior art causes the enhanced speech signal to be distorted, affecting the speech enhancement effect.
An audio enhancement model combined with an autoencoder and a discriminator is adopted. Through the jump connection between the encoding layer and the decoding layer, noise signals are introduced and multi-scale information fusion is carried out, and the audio enhancement results are optimized with the discriminator.
Improves the robustness of the audio enhancement model and the diversity of generated samples, and improves the quality and accuracy of speech enhancement.
Smart Images

Figure CN120279926A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to an audio signal enhancement method, apparatus, system, computer device, and storage medium. Background Art
[0002] With the development of technology, the recognition, analysis, and other processing of speech can be automatically completed by a preset algorithm model, etc. The processing result of speech is closely related to the quality of the input speech. Based on high-quality speech, a more accurate processing result can be obtained. In the prior art, the enhancement of speech is usually based on non-negative matrix factorization (NMF) of the base spectrum of clean speech for training and the base spectrum adapting to noise in real time. However, this semi-supervised method will cause large signal distortion in the enhanced speech, resulting in a poor speech enhancement effect.
[0003] Currently, no effective solution has been proposed for the problem of low quality of the enhanced signal. Summary of the Invention
[0004] Based on this, it is necessary to provide an audio signal enhancement method, apparatus, system, computer device, and storage medium for the above technical problems.
[0005] In a first aspect, this application provides an audio signal enhancement method. The method includes:
[0006] Obtain an audio signal to be enhanced;
[0007] Input the audio signal to be enhanced into the encoding layer of a trained target audio enhancement model for encoding processing to obtain initial feature extraction information; based on the autoencoder in the target audio enhancement model, fuse the initial feature extraction information and the preset signal distribution sampling result to obtain target feature extraction information;
[0008] Input the target feature extraction information into the decoding layer of the target audio enhancement model for decoding processing to obtain an audio enhancement result; wherein, a skip connection is made between the encoding layer and the corresponding decoding layer through an autoencoder.
[0009] Through the above method, on the basis of promoting information transmission, gradient flow, and multi-scale information fusion, the model parameters can be further increased, and a certain degree of noise signal is introduced, thereby improving the performance of the network and making the model more robust.
[0010] In one embodiment, the target audio enhancement model is connected to a preset discriminator, and the method further includes:
[0011] Input the initial feature extraction information and the target feature extraction information into the trained discriminator;
[0012] The discriminator simultaneously discriminates the initial feature extraction information and the target feature extraction information to determine the initial feature discrimination probability for the initial feature extraction information and the target feature discrimination probability for the target feature extraction information;
[0013] The final audio enhancement result is obtained through the initial feature discrimination probability and the target feature discrimination probability.
[0014] Through the above method, the finally generated audio enhancement result can be further checked, and the network can be optimized in a timely manner, thereby improving the quality of audio enhancement processing.
[0015] In one embodiment, training the target audio enhancement model includes:
[0016] Calculation step: Based on the obtained current audio enhancement model, obtain the first encoder output for the preset audio spectrum training set and the first decoder output corresponding to the first encoder output;
[0017] Generation step: Based on the preset current discriminator, obtain the current discrimination result for the first encoder output and the first decoder output. According to the current discrimination result and the discrimination label in the audio spectrum training set, maximize the first encoder discrimination result in the current discrimination result, and optimize and train the current discriminator and the current audio enhancement model based on the maximization result; wherein, the current discrimination result includes the first encoder discrimination result, and the first encoder discrimination result corresponds to the first encoder output;
[0018] Repeat the calculation step and the generation step until the first encoder discrimination result is within the preset encoding discrimination range to obtain the trained target audio enhancement model.
[0019] By introducing a discriminator through the above method to discriminate the outputs of the encoding layer and the corresponding decoding layer, and using the entire audio enhancement model as the generator, in each iteration, the training of the audio enhancement model and the training of the discriminator can be alternately performed to make them confront each other, greatly improving the optimization quality. Further, based on the discrimination result output by the discriminator, the optimization efficiency can be analyzed more intuitively.
[0020] In one embodiment, based on the preset current audio enhancement model, obtaining the first encoder output for the obtained audio spectrum training set and the first decoder output corresponding to the first encoder output includes:
[0021] Obtain a preset audio enhancement model and a preset discriminator, and perform initialization processing to obtain the current audio enhancement model and the current discriminator; wherein, the current discriminator is respectively connected to the first encoding layer and the corresponding first decoding layer in the current audio enhancement model;
[0022] Perform encoding processing on the audio spectrum training set based on the first encoding layer in the current audio enhancement model to obtain the first encoder output;
[0023] Based on the current autoencoder in the current audio enhancement model, fuse the first encoder output with the first signal distribution sampling result to obtain the first target feature extraction information, and perform decoding processing on the first encoder output based on the first decoding layer corresponding to the first encoding layer in the current audio enhancement model to obtain the first decoding result; wherein, both the first signal distribution sampling result and the signal distribution sampling result are obtained by sampling processing a preset signal distribution.
[0024] Through the above method, not only can the training of the audio enhancement model and the training of the discriminator be alternately performed to improve the optimization quality, but further, setting the discriminator only in some levels also ensures that excessive computing resources are not consumed.
[0025] In one embodiment, the above method further includes:
[0026] Obtain a preset audio spectrum training set, and the audio spectrum training set carries audio spectrum labels;
[0027] Input the audio spectrum training set into the obtained initial audio enhancement model for training to obtain the training audio spectrum result, calculate the loss function result according to the training audio spectrum result and the audio spectrum label, and reversely transmit the gradient of the loss function result to the initial audio enhancement model for iterative training to generate the target audio enhancement model.
[0028] Through the above method, the training of the audio enhancement model can be achieved, and it can be understood that the user can adjust the required loss function according to the actual situation by himself / herself to obtain a better training effect.
[0029] In one embodiment, fusing the initial feature extraction information with the signal distribution sampling result obtained based on the signal distribution based on the trained autoencoder to obtain the target feature extraction information includes:
[0030] Perform encoding processing on the initial feature extraction information based on the autoencoding layer in the autoencoder to obtain the information weight value and the autoencoding result;
[0031] Fuse the information weight value with the signal distribution sampling result to obtain the target signal distribution sampling result, and obtain the final auto-encoding result based on the target signal distribution sampling result and the auto-encoding result;
[0032] Decode the final auto-encoding result based on the decoder layer in the auto-encoder to obtain the target feature extraction information.
[0033] Through the above method, new samples similar but not identical to the real data are generated by the auto-encoder, so as to introduce a multi-channel auto-encoder into the main architecture of the target audio enhancement model, improve the diversity and randomness of the generated samples, and thus improve its robustness in the position environment.
[0034] In a second aspect, the present application also provides an audio signal enhancement device. The device includes:
[0035] An acquisition module, configured to acquire an audio signal to be enhanced;
[0036] A calculation module, configured to input the audio signal to be enhanced into the encoding layer of the trained target audio enhancement model for encoding processing to obtain initial feature extraction information; based on the auto-encoder in the target audio enhancement model, fuse the initial feature extraction information and a preset signal distribution sampling result to obtain target feature extraction information;
[0037] A generation module, configured to input the target feature extraction information into the decoding layer in the target audio enhancement model for decoding processing to obtain an audio enhancement result; wherein, there is a skip connection between the encoding layer and the corresponding decoding layer through the auto-encoder.
[0038] In a third aspect, the present application also provides an audio signal enhancement system. The system includes an audio acquisition device and an audio signal enhancement device; wherein, the audio acquisition device is configured to send the audio signal to be enhanced to the audio signal enhancement device;
[0039] The audio signal enhancement device is configured to perform the above audio signal enhancement method based on the audio signal to be enhanced.
[0040] In a fourth aspect, the present application also provides a computer device. The computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0041] Acquire an audio signal to be enhanced;
[0042] Input the audio signal to be enhanced into the encoding layer of a trained target audio enhancement model for encoding processing to obtain initial feature extraction information; based on the autoencoder in the target audio enhancement model, fuse the initial feature extraction information and a preset signal distribution sampling result to obtain target feature extraction information;
[0043] Input the target feature extraction information into the decoding layer of the target audio enhancement model for decoding processing to obtain an audio enhancement result; wherein, a skip connection is made between the encoding layer and the corresponding decoding layer through the autoencoder.
[0044] In a fifth aspect, the present application also provides a computer-readable storage medium. On this computer-readable storage medium, a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0045] Obtain an audio signal to be enhanced;
[0046] Input the audio signal to be enhanced into the encoding layer of a trained target audio enhancement model for encoding processing to obtain initial feature extraction information; based on the autoencoder in the target audio enhancement model, fuse the initial feature extraction information and a preset signal distribution sampling result to obtain target feature extraction information;
[0047] Input the target feature extraction information into the decoding layer of the target audio enhancement model for decoding processing to obtain an audio enhancement result; wherein, a skip connection is made between the encoding layer and the corresponding decoding layer through the autoencoder.
[0048] For the above audio signal enhancement method, device, system, computer device, and storage medium, the audio signal to be enhanced is converted into a spectrogram form, and then this spectrogram is input into a trained target audio enhancement model. Compared with directly processing the audio signal, encoding and analyzing the spectrogram allows for more flexible calculations. Further, the single skip connection in the target audio enhancement model is replaced with a skip connection implemented based on an autoencoder, improving the diversity and randomness of the generated samples, thereby enhancing its robustness in unknown environments. Additionally, a discriminator is introduced in the present application, enabling adversarial training based on the discriminator to continuously improve the quality of the data generated by the target audio enhancement model. Description of the Drawings
[0049] Figure 1 It is an application environment diagram of the audio signal enhancement method in an embodiment;
[0050] Figure 2 It is a flow diagram of the audio signal enhancement method in an embodiment;
[0051] Figure 3Schematic structural diagram of a discriminator in an embodiment;
[0052] Figure 4 Schematic structural diagram of an autoencoder in another embodiment;
[0053] Figure 5 Schematic architecture diagram of an audio signal enhancement method in a preferred embodiment;
[0054] Figure 6 Block diagram of the structure of an audio signal enhancement device in an embodiment;
[0055] Figure 7 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0056] In order to make the objectives, technical solutions, and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0057] The audio signal enhancement method provided by the embodiments of the present application can be applied to an application environment as shown in Figure 1 . Among them, the terminal 102 communicates with the server 104 through a network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or can be placed in the cloud or other network servers. Obtain the audio signal to be enhanced; input the audio signal to be enhanced into the encoding layer of the target audio enhancement model for encoding processing to obtain initial feature extraction information; based on the autoencoder in the target audio enhancement model, fuse the initial feature extraction information and the preset signal distribution sampling result to obtain target feature extraction information; input the target feature extraction information into the decoding layer of the target audio enhancement model for decoding processing to obtain the audio enhancement result. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart vehicle-mounted devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers.
[0058] In one embodiment, as shown in Figure 2 , a method for enhancing an audio signal is provided. Taking the method applied to the server in Figure 1 as an example, the method includes the following steps:
[0059] Step S202, obtain the audio signal to be enhanced.
[0060] Among them, the audio signal to be enhanced includes a target audio signal with interference signals. In this application, enhancing the signal includes, but is not limited to, removing interference, improving resolution, etc.
[0061] Step S204: Input the audio signal to be enhanced into the encoding layer of the trained target audio enhancement model for encoding processing to obtain initial feature extraction information; based on the autoencoder in the target audio enhancement model, fuse the initial feature extraction information and the preset signal distribution sampling result to obtain target feature extraction information.
[0062] Among them, the above target audio enhancement model can be a model improved based on network structures such as U-net, ResUnet, RNN, etc.; the above autoencoder is used in this application to generate new samples that are similar but not exactly the same as the original input data; the above signal distribution sampling result is sampled for a preset signal distribution, and the preset signal distribution can be a standard normal distribution, a mixture of Gaussian distributions, etc. It can be understood that the audio signal to be enhanced is converted into a spectral graph and input into the target audio enhancement model for encoding, and then the audio signal to be enhanced and the signal distribution sampling result are fused based on the autoencoder to obtain the above target feature extraction information. Compared with the network structure without an autoencoder, the target audio enhancement model in this application improves the diversity and randomness of the generated samples.
[0063] Step S206: Input the target feature extraction information into the decoding layer of the target audio enhancement model for decoding processing to obtain the audio enhancement result; among them, there is a skip connection between the encoding layer and the corresponding decoding layer through the autoencoder.
[0064] Among them, there is a skip connection between the encoding layer and its corresponding decoding layer through the autoencoder. Preferably, considering the actual application effect, each encoding layer and its corresponding decoding layer can be connected through the autoencoder. It can be understood that in order to ensure the unified size of the feature extraction information between the encoding layer and the corresponding decoding layer, the structures of the autoencoders set in different layers are not the same.
[0065] Through steps S202 to S206, based on the autoencoder between the preset encoding layer and decoding layer, the initial feature extraction information and the signal distribution sampling result are fused, which can, on the basis of promoting information transmission, gradient flow, and multi-scale information fusion, further increase the model parameters and introduce a certain degree of noise signal, thereby improving the performance of the network and making the model more robust.
[0066] In one embodiment, the above method further includes:
[0067] Input the initial feature extraction information and the target feature extraction information into the trained discriminator;
[0068] Based on the discriminator, simultaneously discriminate the initial feature extraction information and the target feature extraction information to determine the initial feature discrimination probability for the initial feature extraction information and the target feature discrimination probability for the target feature extraction information;
[0069] Obtain the final audio enhancement result through the initial feature discrimination probability and the target feature discrimination probability.
[0070] Specifically, Figure 3 FIG. is a schematic structural diagram of the discriminator in an embodiment. The input of the encoder in the figure is the initial feature extraction information, and the input of the decoder is the target feature extraction information. Further, the main architecture of the discriminator can be set as DNN or CNN, etc. In this application, a discriminator is introduced and connected to the audio enhancement model. Preferably, for each encoding layer in the main architecture except the autoencoder in the audio enhancement model, and its corresponding decoding layer, a corresponding discriminator is set. The output of each encoding layer and the output of its corresponding decoding layer are both input into the discriminator for judgment. It can be understood that the discriminator is used to discriminate which of the two input samples is the real sample and which is the generated sample. Theoretically, the probability of the generated sample is close to 0, that is, "false", and the probability of the real sample is close to 1, that is, "true". In this application, it is necessary to generate a new generated sample that is similar but not exactly the same as the real sample, that is, the target feature extraction information should be very similar but not exactly the same as the initial feature extraction information. However, if it is detected that among the outputs of multiple discriminators, the discrimination probability corresponding to the real sample, that is, the initial feature extraction information, is very close to 1, and the discrimination probability for the target feature extraction information is very close to 0, it indicates that the generated target feature extraction information cannot be very similar to the real initial feature extraction information. Further, the subsequent generated audio enhancement result will also be distorted to a certain extent. It can be understood that the output of the encoding layer in this application corresponds to the above real sample, and the output of the decoding layer corresponds to the above generated sample. Through the above method, the discriminant result of the discriminator can be used to further check the finally generated audio enhancement result and optimize the network in a timely manner. Then, the above final audio enhancement result can be obtained according to the optimized target audio enhancement model, that is, the final audio enhancement result is determined by screening the audio enhancement result in the above text through the output result of the discriminator.
[0071] In one embodiment, the above method further includes:
[0072] Calculation step: Based on the obtained current audio enhancement model, obtain the first encoder output for the preset audio spectrum training set and the first decoder output corresponding to the first encoder output;
[0073] Generation step: Based on a preset current discriminator, obtain a current discrimination result for the output of the first encoder and the output of the first decoder. According to the current discrimination result and the discrimination labels in the audio spectrum training set, maximize the first encoder discrimination result in the current discrimination result, and optimize and train the current discriminator and the current audio enhancement model based on the maximization result; wherein, the current discrimination result includes the first encoder discrimination result, and the first encoder discrimination result corresponds to the output of the first encoder;
[0074] Repeat the calculation step and the generation step until the first encoder discrimination result is within a preset coding discrimination range, and obtain a trained target audio enhancement model.
[0075] Specifically, the audio enhancement model is mainly used to learn the real image distribution, so that the images generated by itself are more real to deceive the discriminator. The discriminator needs to discriminate the authenticity of the received images, which is equivalent to a process of game between the two. As time goes by, the audio enhancement model and the discriminator continuously confront each other. Eventually, the two networks reach a dynamic balance, that is, the images generated by the audio enhancement model are close to the real image distribution, and the discriminator cannot distinguish between real and fake images, and the prediction probability of the given image being true is basically close to 0.5 (equivalent to randomly guessing the category), that is, the above-mentioned preset coding discrimination range fluctuates around 0.5. The above-mentioned first encoder output is the coding result of any layer output in the current audio enhancement model and the result output by its corresponding decoding layer.
[0076] Preferably, when training the current discriminator, provide each coding layer of the current audio enhancement model and its corresponding decoding layer as inputs. In an ideal state, it is hoped that the discriminator can distinguish the difference between the data between the above-mentioned first encoder output and the first decoder output each time. The first encoder discrimination result corresponding to the first encoder output is close to 1, and the first decoder discrimination result corresponding to the first decoder output is close to 0. Then use the discrimination results of the two times to calculate the loss function of the discriminator, that is, maximize the following function:
[0077]
[0078] Similarly, there is:
[0079]
[0080] Wherein, the first item on the right side of the function equal sign represents the discriminator's judgment result on the encoder, and the second item on the right side represents the discriminator's judgment result on the data generated by the decoder. Specifically, G(z) is the fake sample generated by the decoder, p data is the real data distribution, D(x) is the real sample, denotes the expected value of the data distribution generated by the decoder over the true data distribution p data If the value is high, it indicates that the data distribution generated by the decoder is closer to the distribution of the true data. denotes the expected value of the latent variable z sampled from the prior distribution p z (z) of the latent space. This expected value represents the average of the distribution of the latent variable z under the prior distribution p z (z).
[0081] After the current discriminator and the current audio enhancement model training phase ends, update their parameters respectively to reduce the value of the loss function. Stochastic Gradient Descent (SGD), Adaptive Moment Estimation (adam), or other optimization algorithms can be used to update the parameters. Repeat the steps of audio enhancement model training and discriminator training for multiple iterations to gradually improve the performance of the audio enhancement model and the discriminator. Through the above method, a discriminator is introduced to discriminate the outputs of the encoding layer and the corresponding decoding layer, and the audio enhancement model as a whole serves as the generator. In each iteration, the training of the audio enhancement model and the discriminator can be alternately carried out to make them confront each other, greatly improving the optimization quality. Further, based on the discrimination results output by the discriminator, the optimization efficiency can be analyzed more intuitively.
[0082] In one embodiment, the above method further includes:
[0083] Obtain a preset audio enhancement model and a preset discriminator, and perform initialization processing to obtain the current audio enhancement model and the current discriminator; wherein, the current discriminator is respectively connected to the first encoding layer and the corresponding first decoding layer in the current audio enhancement model;
[0084] Encode the audio spectrum training set based on the first encoding layer in the current audio enhancement model to obtain a first encoder output;
[0085] Fuse the first encoder output with the first signal distribution sampling result based on the current autoencoder in the current audio enhancement model to obtain first target feature extraction information, and decode the first encoder output based on the first decoding layer corresponding to the first encoding layer in the current audio enhancement model to obtain a first decoding result; wherein, both the first signal distribution sampling result and the signal distribution sampling result are obtained by sampling the preset signal distribution.
[0086] Preferably, a corresponding discriminator can be set between each encoding layer and its corresponding decoding layer, and each encoding layer and its corresponding decoding layer are connected through a skip connection by an autoencoder. The above initialization process is for the parameters of the audio enhancement model and the discriminator, which can be implemented by random initialization or pre-training. It can be understood that the training of the current audio enhancement model and the training of the current discriminator cannot be carried out simultaneously. It is necessary to first complete the training of the current audio enhancement model, and then complete the training of the current discriminator according to the output of the trained current audio enhancement model. In some application scenarios, considering that the importance of feature information in different channels is different, corresponding discriminators can be set at some levels. Through the above method, not only the training of the audio enhancement model and the training of the discriminator are carried out alternately to improve the optimization quality, but further, setting discriminators only at some levels also ensures that excessive computing resources will not be consumed.
[0087] In one embodiment, the above method further includes:
[0088] Obtain a preset audio spectrum training set, and the audio spectrum training set carries audio spectrum labels;
[0089] Input the audio spectrum training set into the obtained initial audio enhancement model for training to obtain a training audio spectrum result, calculate the loss function result according to the training audio spectrum result and the audio spectrum label, and backpropagate the gradient of the loss function result to the initial audio enhancement model for iterative training to generate a target audio enhancement model.
[0090] Specifically, the above initial audio enhancement model includes an autoencoder and the main architecture of the audio enhancement model. During training, the autoencoder and the main architecture are trained simultaneously. Preferably, the above loss function includes but is not limited to reconstruction loss and KL divergence loss to guide the model to learn specific features or structures. Taking the reconstruction loss and KL divergence loss as an example, the loss function is:
[0091]
[0092] where, q φ (z|x i ) is the variational posterior distribution, p Θ (x i |z) is the generative distribution, and P(z) is the prior distribution. Represents the expected value of the distribution of the latent space variable z learned in the encoder for a given input sample xi. The first part on the right is the reconstruction loss, which is calculated by comparing the original data and the reconstructed data. The reconstruction loss can be the mean squared error loss or the cross-entropy loss, used to measure the difference between the generated data and the real data. The second part on the right is the KL divergence loss. The KL divergence is used to measure the difference between the latent distribution Z output by the encoder and the standard normal distribution. The KL divergence is used to constrain the latent vector to have a reasonable distribution property, such as the distribution being close to the standard normal distribution, so that the representation of the latent vector is more regular and interpretable, and can be specifically expressed as:
[0093]
[0094] where logσ i is the mean, μ i refers to the variance. It can be understood that the training objective of the audio enhancement model is to minimize the sum of the above reconstruction loss and KL divergence loss, so that the generated data is close to the real data, and the distribution of the latent vector is close to the standard normal distribution. Further, the autoencoder in this application can be a VAE model (Variational AutoEncoder, VAE). Generally, the loss function of VAE usually uses the variational lower bound as the training objective, but the variational lower bound is not necessarily the optimal one. Other approximate training objectives, such as IWAE (Importance Weighted Autoencoders), can be considered to obtain better generation effects.
[0095] Through the above method, the training of the audio enhancement model can be achieved, and it can be understood that the user can adjust the required loss function according to the actual situation. Further, the reconstruction loss and the KL divergence loss can be balanced in terms of weights. By adjusting the weights, an appropriate balance point can be found between the generation quality and the continuity of the latent space, so as to obtain better training results.
[0096] In one embodiment, the above method further includes:
[0097] Encoding the initial feature extraction information based on the autoencoding layer in the autoencoder to obtain the information weight value and the autoencoding result;
[0098] Fusing the information weight value with the signal distribution sampling result to obtain the target signal distribution sampling result, and obtaining the final autoencoding result based on the target signal distribution sampling result and the autoencoding result;
[0099] Decoding the final autoencoding result based on the auto-decoding layer in the autoencoder to obtain the target feature extraction information.
[0100] Specifically, Figure 4 is a schematic structural diagram of an autoencoder in an embodiment. Among them, the input is the initial feature extraction information, the output is the target feature extraction information, μ is the autoencoding result, σ is the information weight value, which can also be the standard deviation, and ε is the sampling result of the above signal distribution. The autoencoder includes an autoencoding layer that maps the input data (such as an image or feature information) to a distribution in the latent space, usually a Gaussian distribution, and then outputs two parts, one is the above information weight value, and the other is the autoencoding result. Among them, the autoencoding layer can be a CNN or a DNN. Further, in order to achieve differentiability, reparameterization is required, that is, by sampling a random vector from a preset probability distribution to obtain the above signal distribution sampling result, then multiplying it by the information weight value output by the encoder, and adding the autoencoding result to obtain the latent vector Z. Then, the latent vector Z is received by the auto-decoding layer in the autoencoder and attempts to generate reconstructed data that matches the original input data. Further, the task of the auto-decoder is to generate data from the latent vector so that the generated data is as close as possible to the real data. Similarly, the auto-decoding layer can also be a CNN or a DNN. Through the above method, the fusion processing of the target signal sampling result and the initial feature extraction information can be realized, and the reconstruction of the data can be realized. The autoencoder generates new samples that are similar but not exactly the same as the real data, so as to introduce multiple autoencoders into the main architecture of the target audio enhancement model, improve the diversity and randomness of the generated samples, and thus improve its robustness in the position environment.
[0101] This embodiment also provides a specific embodiment of an audio signal enhancement method, as Figure 5 shown, Figure 5 is a schematic architecture diagram of an audio signal enhancement method in a preferred embodiment.
[0102] First, obtain the audio signal to be enhanced, and convert the audio signal to be enhanced into the form of a spectrogram and input it into a preset target audio enhancement model.
[0103] The main architecture of the target audio enhancement model is a U-net structure. The skip connections between its encoding layer and decoding layer are replaced by new encoder and decoder paths through an autoencoder. At the same time, the inputs of both its encoding layer and decoding layer are connected to a discriminator. After the audio signal to be enhanced is input into the target audio enhancement model, it is first encoded based on the encoding layer in the main architecture of the target audio enhancement model to obtain initial feature extraction information. Then, while the initial feature extraction information is passed backward, it is fused with the signal distribution sampling result through the skip connection path with an autoencoder, and a new sample similar but not exactly the same as the original initial feature extraction information is generated based on the autoencoder, that is, the above-mentioned target feature extraction information. This new sample is output to the decoding layer in the main architecture of the target audio enhancement model and concatenated with the corresponding feature extraction information for decoding processing to obtain the above-mentioned audio enhancement result.
[0104] Furthermore, during the training process, the discriminator is used to optimize the training process. The training of the audio enhancement model and the discriminator is an alternating optimization process. First, the audio enhancement model is used to reconstruct data and learn latent representations. Then, the discriminator is used to distinguish between generated samples and real samples to minimize the classification error of the discriminator. The ultimate goal is to balance the training of the audio enhancement model and the discriminator to achieve higher-quality generated samples.
[0105] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, there is no strict order limit for the execution of these steps, and these steps can be executed in other orders. Moreover, at least some of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least some of the steps or stages in other steps or other steps.
[0106] Based on the same inventive concept, the embodiments of the present application also provide an audio signal enhancement device for implementing the above-mentioned audio signal enhancement method. The implementation solutions provided by this device to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more of the following embodiments of the audio signal enhancement device can refer to the limitations on the audio signal enhancement method in the above text and will not be repeated here.
[0107] In one embodiment, as Figure 6As shown in the figure, an audio signal enhancement device is provided, including: an acquisition module 61, a calculation module 62, and a generation module 63, where:
[0108] The acquisition module 61 is configured to acquire an audio signal to be enhanced.
[0109] The calculation module 62 is configured to input the audio signal to be enhanced into the encoding layer of a trained and complete target audio enhancement model for encoding processing to obtain initial feature extraction information; based on the autoencoder in the target audio enhancement model, fuse the initial feature extraction information and a preset signal distribution sampling result to obtain target feature extraction information.
[0110] The generation module 63 is configured to input the target feature extraction information into the decoding layer of the target audio enhancement model for decoding processing to obtain an audio enhancement result; wherein, a skip connection is made between the encoding layer and the corresponding decoding layer through the autoencoder.
[0111] Specifically, the acquisition module 61 acquires an audio signal to be enhanced, and this audio signal to be enhanced is a target signal including interference, and this interference can be environmental noise, machine noise, or distortion of the audio itself, etc., and this target signal can be human voice, target signal sound, etc. After acquiring a segment of audio signal to be enhanced, the audio signal to be enhanced is sent to the calculation module 62, and the calculation module 62 inputs the audio signal to be enhanced into the target audio enhancement model. This target audio enhancement model can be a model improved based on U-net or ResUnet, and further, the target audio enhancement model in this application introduces an autoencoder, and this autoencoder can be an improved VAE module based on the solution in this application. Initial feature extraction information is obtained based on the above target audio enhancement model, and target feature extraction information is obtained by fusing the autoencoder and the signal distribution sampling result. Then, the target feature extraction information is input into the decoding layer of the target audio enhancement model for decoding processing to obtain an audio enhancement result.
[0112] With the above device, the single skip connection in the original audio enhancement model is replaced with a path with an autoencoder, which improves the diversity and randomness of the generated samples. Further, the autoencoder can map the input data to a distribution in a latent space, thereby being able to capture the implicit structure and features of the data, and thus generating new samples similar to the original data.
[0113] Each module in the above audio signal enhancement device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above modules.
[0114] In one embodiment, an audio signal enhancement system is provided, including an audio acquisition device and the above-mentioned audio signal enhancement device; wherein, the audio acquisition device is configured to send the audio signal to be enhanced to the audio signal enhancement device; the audio signal enhancement device is configured to perform the audio signal enhancement method described above based on the audio signal to be enhanced.
[0115] Specifically, the above-mentioned audio signal enhancement system has a wide range of application environments. The audio acquisition device can be a device such as a tape recorder or a microphone for recording a target sound. Further, the above-mentioned audio signal enhancement device can be independent of the audio acquisition device, or the above-mentioned audio signal enhancement device can be integrated on the audio acquisition device, which can achieve real-time recording and enhancement of the target signal, or can end the recording after the target signal is recorded, and send the obtained audio signal to be enhanced to the audio signal enhancement device for processing, etc.
[0116] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 7 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the obtained audio signal to be enhanced and data related to the audio enhancement model and discriminator. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements an audio signal enhancement method.
[0117] Those skilled in the art can understand that Figure 7 the structure shown in
[0118] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0119] Obtain the audio signal to be enhanced;
[0120] Input the audio signal to be enhanced into the encoding layer of the trained target audio enhancement model for encoding to obtain initial feature extraction information; based on the autoencoder in the target audio enhancement model, fuse the initial feature extraction information and the preset signal distribution sampling result to obtain target feature extraction information;
[0121] Input the target feature extraction information into the decoding layer of the target audio enhancement model for decoding to obtain the audio enhancement result; among them, a skip connection is made between the encoding layer and the corresponding decoding layer through the autoencoder.
[0122] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0123] Those of ordinary skill in the art can understand that all or part of the processes in the above-described embodiment methods can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-described method embodiments. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAMs), magnetoresistive random access memories (MRAMs), ferroelectric random access memories (FRAMs), phase change memories (PCMs), graphene memories, and the like. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0124] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0125] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. An audio signal enhancement method, characterized in that, The method includes: Obtain the audio signal to be enhanced; Input the audio signal to be enhanced into the encoding layer of the trained target audio enhancement model for encoding processing to obtain initial feature extraction information; based on the autoencoder in the target audio enhancement model, fuse the initial feature extraction information and the preset signal distribution sampling result to obtain target feature extraction information; Input the target feature extraction information into the decoding layer of the target audio enhancement model for decoding processing to obtain the audio enhancement result; wherein, a skip connection is made between the encoding layer and the corresponding decoding layer through the autoencoder.
2. The method according to claim 1, wherein the target audio enhancement model is connected to a preset discriminator, characterized in that The method further includes: Input the initial feature extraction information and the target feature extraction information into the trained discriminator; Based on the discriminator, simultaneously discriminate the initial feature extraction information and the target feature extraction information to determine the initial feature discrimination probability for the initial feature extraction information and the target feature discrimination probability for the target feature extraction information; Obtain the final audio enhancement result through the initial feature discrimination probability and the target feature discrimination probability.
3. The method according to claim 1, characterized in that Training the target audio enhancement model includes: Calculation step: Based on the obtained current audio enhancement model, obtain the first encoder output for the preset audio spectrum training set, and the first decoder output corresponding to the first encoder output; Generation step: Based on the preset current discriminator, obtain the current discrimination result for the first encoder output and the first decoder output, maximize the first encoder discrimination result in the current discrimination result according to the current discrimination result and the discrimination label in the audio spectrum training set, and optimize and train the current discriminator and the current audio enhancement model based on the maximization result; wherein, the current discrimination result includes the first encoder discrimination result, and the first encoder discrimination result corresponds to the first encoder output; Repeat the calculation step and the generation step until the first encoder discrimination result is within the preset encoding discrimination range to obtain the trained target audio enhancement model.
4. The method according to claim 3, characterized in that The obtaining of the first encoder output for the obtained audio spectrum training set and the first decoder output corresponding to the first encoder output based on the preset current audio enhancement model includes: Obtain the preset audio enhancement model and the preset discriminator, and perform initialization processing to obtain the current audio enhancement model and the current discriminator; wherein, the current discriminator is respectively connected to the first encoding layer and the corresponding first decoding layer in the current audio enhancement model; Based on the first encoding layer in the current audio enhancement model, perform encoding processing on the audio spectrum training set to obtain the first encoder output; Fusing the first encoder output with the first signal distribution sampling result based on the current autoencoder in the current audio enhancement model to obtain first target feature extraction information, and performing decoding processing on the first encoder output based on the first decoding layer corresponding to the first encoding layer in the current audio enhancement model to obtain a first decoding result; wherein, both the first signal distribution sampling result and the signal distribution sampling result are obtained by sampling a preset signal distribution.
5. The method according to claim 1, wherein The method further includes: Obtaining a preset audio spectrum training set, where the audio spectrum training set carries audio spectrum labels; Inputting the audio spectrum training set into the obtained initial audio enhancement model for training to obtain a training audio spectrum result, calculating a loss function result based on the training audio spectrum result and the audio spectrum labels, and reversely transmitting the gradient of the loss function result to the initial audio enhancement model for iterative training to generate the target audio enhancement model.
6. The method according to claim 1, wherein The fusing the initial feature extraction information and a preset signal distribution sampling result based on the autoencoder in the target audio enhancement model to obtain target feature extraction information includes: Encoding the initial feature extraction information based on the autoencoding layer in the autoencoder to obtain an information weight value and an autoencoding result; Fusing the information weight value with the signal distribution sampling result to obtain a target signal distribution sampling result, and obtaining a final autoencoding result based on the target signal distribution sampling result and the autoencoding result; Performing decoding processing on the final autoencoding result based on the auto - decoding layer in the autoencoder to obtain the target feature extraction information.
7. An audio signal enhancement device, characterized in that, The apparatus includes: An obtaining module, configured to obtain an audio signal to be enhanced; A calculating module, configured to input the audio signal to be enhanced into an encoding layer of a trained - complete target audio enhancement model for encoding processing to obtain initial feature extraction information; and based on the autoencoder in the target audio enhancement model, fusing the initial feature extraction information and a preset signal distribution sampling result to obtain target feature extraction information; A generating module, configured to input the target feature extraction information into a decoding layer in the target audio enhancement model for decoding processing to obtain an audio enhancement result; wherein, there is a skip connection between the encoding layer and the corresponding decoding layer through the autoencoder.
8. An audio signal enhancement system, characterized in that, The system includes an audio acquisition device and the audio signal enhancement apparatus according to claim 7; wherein, the audio acquisition device is configured to send the audio signal to be enhanced to the audio signal enhancement apparatus; The audio signal enhancement apparatus is configured to execute the audio signal enhancement method according to claims 1 to 6 based on the audio signal to be enhanced.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 6.