Speech Denoising via Discrete Representation Learning

By combining variational autoencoder and vector quantization autoencoder, using autoregressive WaveNet decoder and new matching loss function, the instability problem of existing speech noise reduction methods in low signal-to-noise ratio scenarios is solved, and end-to-end stable noise reduction training and efficient speech clarity improvement are achieved.

CN114267366BActive Publication Date: 2025-07-04BAIDU USA LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111039819.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-10-01
Filing Date
2021-09-06
Publication Date
2025-07-04
Estimated Expiration
2041-09-06

AI Technical Summary

Technical Problem

The existing speech noise reduction method has unstable loss in low signal-to-noise ratio scenarios, and lacks an end-to-end training system, requiring pre-trained network components.

Method used

The combination of variational autoencoder (VAE) and vector quantized autoencoder (VQ-VAE) is used to generate noise reduction audio through the autoregressive WaveNet decoder, and the system is trained end-to-end using a new matching loss function to avoid explicit loss calculations, and to perform loss masking only when potential codes are inconsistent.

Benefits of technology

It realizes stable noise reduction performance in low signal-to-noise ratio scenarios, can be trained from scratch, avoids the need for pre-training networks, and improves speech clarity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114267366B_ABST
    Figure CN114267366B_ABST
Patent Text Reader

Abstract

The present application discloses a computer-implemented method for training a noise reduction system. From a comprehensive perspective, the embodiments of a new end-to-end method for audio noise reduction are developed and presented herein. As in text-to-speech systems, instead of explicitly modeling the noise components in the input signal, the embodiments directly synthesize the noise-reduced audio from a generative model (or vocoder). In one or more embodiments, for generating speech content for an autoregressive generative model, learning is performed via a variational autoencoder with a discrete latent representation. Additionally, in one or more embodiments, a new matching loss is proposed for the purpose of noise reduction, which is masked when the corresponding latent codes are different. Compared with other methods on a test dataset, the embodiments achieve competitive performance and can be trained from scratch.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to systems and methods for machine learning of computers that can provide improved computer performance, features, and uses. More specifically, the present disclosure relates to systems and methods for reducing noise in audio. Background Art

[0002] Deep neural networks have achieved great success in many fields, such as computer vision, natural language processing, text-to-speech, and many other applications. One area that has received a great deal of attention is the machine learning application of audio, especially speech noise reduction.

[0003] Speech noise reduction is an important task in audio signal processing and has been widely applied in many practical applications. The goal of speech noise reduction is to improve the clarity of noisy audio utterances. Classical methods focus on using signal processing techniques, such as filtering and spectral restoration. With the advent of deep learning, neural network-based methods have attracted increasing attention and can perform noise reduction in the time domain or frequency domain to improve performance compared with classical methods.

[0004] On the other hand, deep generative models have recently become a powerful framework for representation learning and generative tasks for various types of signals, including images, text, and audio. In deep representation learning, variational autoencoders (VAEs) have proven to be an effective tool for extracting latent representations and then facilitating downstream tasks. For audio generation, neural vocoders have achieved state-of-the-art performance in generating raw audio waveforms and have been deployed in real text-to-speech (TTS) systems.

[0005] Despite the improvements made through these different methods, they each have limitations. For example, some techniques require explicit calculation of the loss from the denoised audio to its clean counterpart at the sample level, which may become unstable in some cases. In current neural network methods, they require separate training of certain components - thus, there is no end-to-end system that can be trained as a complete system.

[0006] Therefore, what is needed is a new approach that treats the noise reduction problem as a fundamentally different type of problem and overcomes the deficiencies of current methods. Summary of the Invention

[0007] A first aspect of the present invention provides a computer-implemented method for training a noise reduction system, comprising: given a noise reduction system including a first encoder, a second encoder, a quantizer, and a decoder, and given a set of one or more clean-noisy audio pairs, where each clean-noisy audio pair includes clean audio content through a speaker and noisy audio content through a speaker: for each clean audio, using the first encoder to generate one or more consecutive latent representations of the clean audio; for each noisy audio, using the second encoder to generate one or more consecutive latent representations of the noisy audio; for each consecutive latent representation of the clean audio, using the quantizer to generate a corresponding discrete clean audio representation; for each consecutive latent representation of the noisy audio, using the quantizer to generate a corresponding discrete noisy audio representation; for each clean-noisy audio pair, inputting the discrete clean audio representation, the clean audio, and a speaker embedding representing the speaker of the clean-noisy audio pair into the decoder to generate an audio sequence prediction; calculating a loss of the noise reduction system, where the loss includes a latent representation matching loss term, and the latent representation matching loss term for time steps where the discrete clean audio representation and the discrete noisy audio representation are different is based on a distance measure between the consecutive latent representations of the clean audio and the noisy audio for the time step; and updating the noise reduction system using the loss.

[0008] A second aspect of the present invention provides a system, comprising one or more processors and a non-transitory computer-readable medium including one or more sets of instructions. The instructions, when executed by the processor, cause the processor to execute the method provided by the first aspect of the present invention.

[0009] A third aspect of the present invention provides a computer-implemented method, comprising: given an input noisy audio for noise reduction and a trained noise reduction system, the trained noise reduction system including a trained encoder, a trained quantizer, and a trained decoder: using the trained encoder to generate one or more consecutive latent representations of the input noisy audio; for the one or more consecutive latent representations, using the one or more consecutive latent representations of the input noisy audio and the trained quantizer to generate one or more discrete noisy audio representations; and generating a noise-reduced audio representation of the input noisy audio by inputting the discrete noisy audio representation into the trained decoder; wherein the noise reduction system is trained using a loss, and the loss includes a matching loss term, and the matching loss term for time steps where the discrete clean audio representation of the clean audio from the clean-noisy pair is different from the discrete noisy audio representation of the noisy audio, where the clean-noisy audio pair includes the clean audio and the corresponding noisy audio, is based on a distance measure between the consecutive latent representations of the clean audio and the noisy audio for the time step.

[0010] The present invention proposes a new matching loss for the purpose of noise reduction. When the corresponding latent codes are different, they are masked. Compared with other methods on the test dataset, the present invention achieves competitive performance and can be trained from scratch. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Reference will be made to the embodiments of the present disclosure, examples of which may be illustrated in the drawings. These figures are illustrative only and not restrictive. Although the present disclosure is generally described in the context of these embodiments, it should be understood that it is not intended to limit the scope of the present disclosure to these specific embodiments. Items in the figures may not be drawn to scale.

[0012] Figure 1 A noise reduction system according to an embodiment of the present disclosure is depicted.

[0013] Figure 2 A view of a part of the entire system according to an embodiment of the present disclosure is depicted, showing components and paths for clean audio.

[0014] Figure 3 A method for training a noise reduction system according to an embodiment of the present disclosure is depicted.

[0015] Figure 4 A trained noise reduction system according to an embodiment of the present disclosure is depicted.

[0016] Figure 5 A method for using a trained noise reduction system to generate noise-reduced audio according to an embodiment of the present disclosure is depicted.

[0017] Figure 6 A simplified block diagram of a computing device / information processing system according to an embodiment of the present disclosure is depicted. DETAILED DESCRIPTION

[0018] In the following description, for purposes of explanation, specific details are set forth in order to provide an understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these details. Additionally, those skilled in the art will recognize that the embodiments of the present disclosure described below may be implemented in a variety of ways, such as a process, apparatus, system, device, or method on a tangible computer-readable medium.

[0019] The components or modules shown in the figures are illustrative of exemplary embodiments of the present disclosure and are intended to avoid obscuring the present disclosure. It should also be understood that throughout the discussion, a component may be described as a separate functional unit that may include sub-units, but those skilled in the art will recognize that various components or portions thereof may be divided into separate components or may be integrated together, including, for example, in a single system or component. It should be noted that the functions or operations discussed herein may be implemented as components. The components may be implemented in software, hardware, or a combination thereof.

[0020] In addition, the connections between components or systems within the figures are not intended to be limited to direct connections. Instead, intermediate components may modify, reformat, or otherwise change the data between these components. Similarly, more or fewer connections may be used. It should also be noted that the terms "coupled", "connected", "communicatively coupled", "interface connected", "interface connection", or any of their derivatives should be understood to include direct connections, indirect connections through one or more intermediate devices, and wireless connections. It should also be noted that any communication, such as a signal, response, reply, confirmation, message, query, etc., may include one or more information exchanges.

[0021] References in the specification to "one or more embodiments", "preferred embodiments", "an embodiment", "multiple embodiments", etc. mean that the specific features, structures, characteristics, or functions described in connection with the embodiments are included in at least one embodiment of the present disclosure and may be in more than one embodiment. Similarly, the appearances of the above phrases throughout the specification do not necessarily all refer to the same embodiment or embodiments.

[0022] The use of certain terms throughout the specification is for illustrative purposes and should not be construed as limiting. A service, function, or resource is not limited to a single service, function, or resource; the use of these terms may refer to a grouping of related services, functions, or resources, which may be distributed or aggregated. The terms "comprising" and "including" should be understood as open terms, and any list following is an example and is not meant to be limited to the listed items. A "layer" may include one or more operations. Words such as "optimal", "optimize", etc. refer to an improvement in a result or process, and it is not required that a particular result or process has reached an "optimal" state or peak state. The use of memory, database, information repository, data storage, table, hardware, cache, etc. herein may be used to refer to one or more system components in which information may be input or otherwise recorded.

[0023] In one or more embodiments, the stop conditions may include: (1) a set number of iterations have been executed; (2) a certain processing time has been reached; (3) convergence (e.g., the difference between consecutive iterations is less than a first threshold); (4) divergence (e.g., a performance degradation); and (5) an acceptable result has been reached.

[0024] Those skilled in the art should recognize that: (1) certain steps can be optionally performed; (2) the steps may not be limited to the specific order set forth herein; (3) certain steps may be performed in a different order; (4) certain steps can be performed simultaneously.

[0025] Any headings used herein are for organizational purposes only and should not be used to limit the scope of the specification or claims. Each reference / document mentioned in the patent document is hereby incorporated by reference in its entirety.

[0026] It should be noted that any experiments and results provided herein are provided by way of example and were conducted using one or more specific embodiments under specific conditions; accordingly, these experiments and their results should not be used to limit the scope of the disclosure of the current patent document.

[0027] A. General Introduction

[0028] Embodiments herein approach the voice denoising task from a new perspective by treating it as a voice generation problem, such as in a text-to-speech system. In one or more embodiments, the denoised audio is generated autoregressively from a vocoder, such as WaveN (discussed by A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu in "WaveNet: A Generative Model for Raw Audio"), available at arxiv.org / abs / 1609.03499v2(See (2016), the entire content of which is incorporated herein by reference). This view differentiates the embodiments herein from prior methods in that the embodiments avoid the need to explicitly calculate the loss from the denoised audio to its clean counterpart at the sample level, which can become unstable in low signal-to-noise ratio (SNR) scenarios. Different from WaveNet that uses the Mel spectrogram of the original audio waveform as a conditioner, the embodiments herein directly learn the required speech information from the data. More specifically, in one or more embodiments, a vector quantization variational autoencoder (VQ-VAE) (such as described by A. vanden Oord, O. Vinyals, and K. Kavukcuoglu in "Neural Discrete Representation Learning", Advances in Neural Information Processing Systems, pages 6306–6315 (2017), which is hereby incorporated by reference in its entirety) is used to generate discrete latent representations from clean audio, which are then used as conditioners for a vocoder such as WaveNet implementation. To achieve the denoising effect, in one or more embodiments, a loss function is calculated based on the distance between the clean continuous latent representation and the noisy continuous latent representation. In one or more embodiments, to improve robustness, the loss function is further masked only when the discrete latent codes are inconsistent between the clean and noisy components. In one or more embodiments, the system embodiments do not require any pre-trained network and thus can be trained from scratch.

[0029] B. Related Work

[0030] Recent progress has shown that deep generative models may be useful tools in speech denoising. A method based on generative adversarial networks (GANs) has been proposed, where the generator outputs denoised audio and the discriminator classifies it from clean audio. Others have developed a Bayesian approach by modeling the prior function and the likelihood function via WaveNet, with each function requiring separate training. Some have used a non-causal WaveNet to generate denoised samples by minimizing the regression loss on the clean and noisy components of the predicted input signal. It has been noted that these methods can perform denoising directly in the time domain but require explicit modeling of the noise.

[0031] Some people have proposed a multi-level U-Net architecture to effectively capture long-term transient correlations in the original waveform, and their focus is on speech separation. Others have proposed a new deep feature loss function to penalize the differences in activations across multiple layers for clean audio and denoised audio; however, a pre-trained audio classification network is required, so it cannot be trained from scratch. Although some people have tried comprehensive methods for the denoising task, their methods require training two parts sequentially, where the first part needs to predict the clean mel spectrogram (or other spectral features, depending on the vocoder used), and the second part needs to use the vocoder to synthesize the denoised audio by adjusting the predictions from the first part. In contrast, the embodiments of this article are end-to-end and can be trained from scratch.

[0032] C. Denoising the Embodiments

[0033] 1. Previous Text

[0034] As a popular unsupervised learning framework, variational autoencoders (VAEs) have recently attracted increasing attention. For example, D.P. Kingma and M. Welling in "Auto-encoding Variational Bayes" (available at arxiv.org / abs / 1312.6114 preprint arXiv:1312.6114 (2013)) and D.J. Rezende, S. Mohamed, and D. Wierstra in "Stochastic Backpropagation and Approximate Inference in Deep Generative Models", in the International Conference on Machine Learning, pp. 1278-1286 (2014), both discussed variational autoencoders (each incorporated herein by reference in its entirety).

[0035] In a VAE, the encoder network corresponds to the distribution of the latent representation for a given input data and is parameterized by ; the decoder network computes the likelihood of from and is parameterized by . By defining the latent representation prior as , the objective in a VAE may be to minimize the following loss function:

[0036]

[0037] The first term in Equation (1) can be interpreted as the reconstruction loss, and the second term, the Kullback-Leibler (KL) divergence term, acts as a regularization term to minimize the posterior with the prior .

[0038] For the vector-quantized variational autoencoder (VQ-VAE), A. van den Oord, O. Vinyals, and K. Kavukcuoglum in "Neural Discrete Representation Learning" in "Advances in Neural Information Processing Systems", pages 6306–6315 (2017) (incorporated herein by reference in its entirety) show that using a discrete latent representation can learn better representations across different modalities in several unsupervised learning tasks. The encoder in a VQ-VAE embodiment can output discrete codes instead of a continuous latent representation achieved by using vector quantization (VQ), i.e., the discrete latent vector at the i-th time step can be represented as:

[0039]

[0040] where corresponds to the learnable codes in the codebook. Then, the decoder reconstructs the input from the discrete latent . In a VQ-VAE, the posterior distribution corresponds to the delta distribution, where the probability mass is only assigned to the codes returned from the vector quantizer. By assigning a uniform prior over all discrete codes, it can be shown that the KL divergence term in Equation (1) reduces to a constant. Subsequently, the loss in a VQ-VAE can be represented as:

[0041]

[0042] where represents the latent code corresponding to the input ; sg() is the stop-gradient operator, which is equal to the identity function in the forward pass and has a gradient of zero in the backpropagation phase. The in Equation (3) is a hyperparameter, and in one or more embodiments, it can be set to 0.25.

[0043] 2. Denoising via VQ-VAE Embodiment

[0044] This disclosure presents systems and methods for synthesis methods in speech denoising tasks. Figure 1 depicts a denoising system according to an embodiment of the present disclosure. As Figure 1As shown, the depicted Example 100 includes the following components: (i) two residual convolutional encoders 115 and 120 with the same or similar architecture, which are respectively applied to the noisy audio input 110 and the clean audio input 105; (ii) a vector quantizer 135; (iii) an autoregressive WaveNet decoder 150, which can be the co-pending and co-owned U.S. Patent No. 16 / 277,919, filed on February 15, 2020, titled "SYSTEMS AND METHODS FOR PARALLEL WAVE GENERATION IN END-TO-END TEXT-TO-SPEECH", and listing the inventors Wei Ping, Kainan Peng, and Jitong Chen (Case No. 28888-2269), which is hereby incorporated by reference in its entirety. A loss calculation 165 is also depicted and will be discussed in more detail below with respect to Equation (4).

[0045] In one or more embodiments, the architecture for the neural network can be as follows. Figure 2 A partial view of the overall system according to an embodiment of the present disclosure is depicted, which shows the components and paths for clean audio. Figure 2 The encoder and path for clean audio are shown; the same or similar encoder structure can be used for noisy audio, but is not depicted due to space limitations.

[0046] Encoders (105 and 110). The depicted encoders are similar to the encoders used in "Unsupervised Speech Representation Learning Using WaveNet Autoencoders" by J. Chorowski, R. J. Weiss, S. Bengio, and A. van den Oord in IEEE / ACM Transactions on Audio, Speech, and Language Processing, Vol. 27, No. 12, pp. 2041-2053, December 2019, doi: 10.1109 / TASLP.2019.2938863 (which is hereby incorporated by reference in its entirety), except that (i) instead of using the ReLU non-linearity, a leaky ReLU (α = 0.2) is used; (ii) the number of output channels is reduced from 768 to 512. Empirical observations show that these changes help to stabilize the optimization and reduce the training time without sacrificing performance.

[0047] In one or more embodiments, the original audio and their first and second derivatives are first converted to 13 standard Mel Frequency Cepstral Coefficients (MFCC) features. As Figure 2 shown, the depicted encoder network embodiment may include: (i) two residual convolutional layers with a filter size of 3; (ii) a strided convolutional layer with a stride of 2 and a filter size of 4; (iii) two residual convolutional layers with a filter size of 3; (iv) four residual fully connected layers.

[0048] Vector quantizer (135). In one or more embodiments, the number of output channels of the latent vector is first reduced to 64, and the codebook contains 512 learnable codes, each with a size of 64.

[0049] Decoder (150). In one or more embodiments, a 20-layer WaveNet with cross-entropy loss is used, where the number of channels in the softmax layer is set to 2048. The number of residual channels and skip channels can both be set to 256. The upsampling of the regulator at the sample level can be achieved by repetition. In one or more embodiments, the filter size in the convolutional layer is set to 3, and dilated blocks {1, 2, 4,..., 512} are used, corresponding to a geometric sequence with a common ratio of 2.

[0050] Return Figure 1 , the input to system 100 is a noisy audio and clean audio pair, denoted as and , respectively, including the same speech content. As described above, in one or more embodiments, Mel Frequency Cepstral Coefficients (MFCC) features are first extracted from the original audio and then passed through residual convolutional layers and fully connected layers to generate corresponding continuous latent vectors 130 and 125. Subsequently, the vector quantization introduced in equation (2) can be applied to obtain a discrete representation, i.e., 145 and 140. In one or more embodiments, during training, only the 140 is used as a conditioner for the WaveNet decoder 150. By explicitly conditioning on the speaker embedded in the decoder, the encoder can focus more on speaker-independent information, thereby better extracting phonemic content. Ultimately, the output 160 of the system corresponds to the audio sequence predicted from the autoregressive WaveNet decoder, which is trained in a teacher-forcing method with clean input as ground truth (that is, during training, the model receives time t The ground truth output is taken as the time t + 1 input).

[0051] 3. Example of noise reduction processing

[0052] To remove noise from noisy audio, the goal is to match the latent representations from the noisy and clean inputs. One motivation is that when (i) the decoder is able to use the clean latent code, i.e. as a regulator to generate high-fidelity audio, and (ii) the latent code from the noisy input is close to the latent code from the clean input, and the decoder is expected to To output high-fidelity audio. In order to design a loss function for matching, in one or more embodiments, the distance between the discrete latent representation or the continuous latent representation and the noisy branch and the clean branch can be calculated. However, in one or more embodiments, the distance between the noisy branch and the clean branch can be calculated by calculating the distance between the codes corresponding to them. and At different time steps and Between distance to use a hybrid approach. l represents the number of time steps in the potential, and M represents the number of output channels, then we get ,as well as . Subsequently, the total loss can be expressed as the sum of the VQ-VAE loss and the matching loss in equation (3) as follows:

[0053]

[0054] Note that in one or more embodiments, the matching loss (last term) in equation (4) contributes to the total loss only when the corresponding latent codes are different, resulting in more stable training. In addition, the loss function in equation (4) can be optimized from scratch, thus avoiding the need for pre-training. Another noteworthy point about equation (4) (also in Figure 1 ) is that during training, the decoder is not a function of the noisy input. Therefore, throughout the optimization process, it tends not to learn any hidden information about the noisy audio.

[0055] Example of an annealing scheme. Optimizing directly for all variables in equation (4) would quickly lead to divergence and oscillation during training. Intuitively, the reason for this phenomenon is that the latent representation of the clean input may not have enough information to capture the speech information in the initial training phase. As a result, the objective of the encoder corresponding to the noisy input (i.e., encoder 2) becomes difficult to match. Therefore, in one or more embodiments, to solve this problem, an annealing strategy can be adopted. In one or more embodiments, in equation (4), λ is introduced as a hyperparameter and annealed during training by gradually increasing it from 0 to 1. In one or more embodiments, λ can be annealed during training from 0 (or close to 0) to 1 (or close to 1) via a sigmoid function.

[0056] With this annealing strategy, the entire network can be initially trained as a VQ-VAE, where the optimization is mainly applied to the parameters involved in the path corresponding to the clean input, i.e., encoder 1 → vector quantizer → decoder, and the speaker embedding. In one or more embodiments, when the training of those components becomes stable, the matching loss can be gradually added to the optimization to minimize the distance between the noisy latent representation and the clean latent representation.

[0057] Figure 3 A method for training a noise reduction system according to an embodiment of the present disclosure is depicted. In one or more embodiments, given a noise reduction system including a first encoder, a second encoder, a quantizer, and a decoder, and given a clean-noisy audio pair including clean audio content through a speaker and noisy audio content through a speaker, the clean audio is input (305) into the first encoder to generate a continuous latent representation of the clean audio, and the noisy audio is input (305) into the second encoder to generate a continuous latent representation of the noisy audio. Then, the vector quantizer can be applied to (310) the continuous latent representations of the clean audio and the noisy audio to obtain corresponding discrete clean audio and discrete noisy audio representations respectively.

[0058] In one or more embodiments, the discrete clean audio representation, the clean audio, and the speaker embedding of the speaker representing the clean-noisy audio pair are input (315) into the decoder to generate an audio sequence prediction output.

[0059] In one or more embodiments, the loss of the noise reduction system is calculated (320), where the loss includes a measure of the distance (e.g., terms of a distance metric). The computed loss is used to update (325) the noise reduction system. In one or more embodiments, the training process can continue until a stopping condition is reached, and the trained noise reduction system can be output for noise reduction of noisy input audio.

[0060] 4. Inference Embodiments

[0061] An embodiment of the forward pass of the trained noise reduction system is shown in Figure 4 . During inference, the noisy audio 410 is used as input, and the trained noise reduction system 400 is used to generate the corresponding denoised audio 460. In one or more embodiments, the speech content is retrieved by passing the noisy audio 410 through the trained encoder 420 and the trained vector quantizer 435. The trained WaveNet decoder 450 conditions the output 445 from the vector quantizer and the speaker embedding 455 to generate the denoised audio 460. Note that in one or more embodiments, the current setup assumes that the speakers in the test set should also be present in the training set; however, it should be noted that the system can be extended to unseen speakers. On the other hand, conditioning the decoder embedding on the speaker can be helpful for tasks such as voice conversion.

[0062] Figure 5 A method for generating denoised audio using a trained noise reduction system according to an embodiment of the present disclosure is depicted. In one or more embodiments, given a trained noise reduction system including a trained encoder, a trained quantizer, and a trained decoder, and given a noisy audio for noise reduction and a speaker embedding for the speaker of the noisy audio, a sequential latent representation of the noisy audio is generated (505) using the trained encoder. In one or more embodiments, the trained quantizer is applied to (510) the sequential latent representation of the noisy audio to obtain a corresponding discrete noisy audio representation. Finally, a denoised audio representation of the noisy audio can be generated (515) by inputting the discrete noisy audio representation and the speaker embedding representing the speaker of the noisy audio into the trained decoder.

[0063] D. Computing System Embodiments

[0064] In one or more embodiments, aspects of this patent document may be directed to, may include, or may be implemented on one or more information processing systems (or computing systems). An information processing system / computing system may include any tool or aggregation of tools operable to compute, calculate, determine, classify, process, send, receive, retrieve, originate, route, switch, store, display, communicate, present, detect, record, copy, process, or utilize any form of information, intelligence, or data. For example, a computing system may be or may include a personal computer (e.g., a laptop), a tablet computer, a mobile device (e.g., a personal digital assistant (PDA), a smart phone, a phablet, a tablet, etc.), a smart watch, a server (e.g., a blade server or a rack server), a network storage device, a camera, or any other suitable device, and may vary in size, shape, performance, functionality, and price. The computing system may include random access memory (RAM), one or more processing resources such as a central processing unit (CPU) or hardware or software control logic, read only memory (ROM), and / or other types of memory. Additional components of the computing system may include one or more disk drives, one or more network ports for communicating with external devices and various input and output (I / O) devices, such as a keyboard, a mouse, a stylus, a touch screen, and / or a video display. The computing system may also include one or more buses operable to transfer communications between the various hardware components.

[0065] Figure 6 FIG. depicts a simplified block diagram of an information processing system (or computing system) according to an embodiment of the present disclosure. It will be understood that the functions shown for system 600 are operable to support various embodiments of a computing system—although it should be understood that a computing system may be configured differently and include different components, including having fewer or more components, as Figure 6 shown.

[0066] As Figure 6 shown, computing system 600 includes one or more central processing units (CPUs) 601 that provide computing resources and control the computer. The CPU 601 may be implemented with a microprocessor, etc., and may also include one or more graphics processing units (GPUs) 602 and / or a floating-point coprocessor for mathematical calculations. In one or more embodiments, one or more GPUs 602 may be incorporated within a display controller 609, such as part of one or more graphics cards. System 600 may also include a system memory 619, which may include RAM, ROM, or both.

[0067] As Figure 6As shown, multiple controllers and peripherals may also be provided. The input controller 603 represents an interface to various input devices 604 such as a keyboard, mouse, touch screen, and / or stylus. The computing system 600 may also include a storage controller 607 for interfacing with one or more storage devices 608, each of which includes a storage medium such as a magnetic tape or disk, or an optical medium that may be used to record a program of instructions for an operating system, utilities, and applications, which may include embodiments of programs implementing various aspects of the present disclosure. According to the present disclosure, the storage device 608 may also be used to store processed data or data to be processed. The system 600 may also include a display controller 609 for providing an interface to a display device 611, which may be a cathode ray tube (CRT) monitor, a thin film transistor (TFT) monitor, an organic light emitting diode, an electroluminescent panel, a plasma panel, or any other type of monitor. The computing system 600 may also include one or more peripheral device controllers or interfaces 605 for one or more peripheral devices 606. Examples of peripheral devices may include one or more printers, scanners, input devices, output devices, sensors, etc. The communication controller 614 may interface with one or more communication devices 615, which enables the system 600 to connect to remote devices through any one of a variety of networks including the Internet, cloud resources (e.g., Ethernet cloud, Fibre Channel over Ethernet (FCoE) / Data Center Bridging (DCB) cloud, etc.), local area network (LAN), wide area network (WAN), storage area network (SAN), or through any suitable electromagnetic carrier signal including infrared signals. As shown in the depicted embodiment, the computing system 600 includes one or more fans or fan trays 618 and one or more cooling subsystem controllers 617 that monitor the thermal temperature of the system 600 (or its components) and operate the fan / fan tray 618 to help regulate the temperature.

[0068] In the system shown, all of the major system components can be connected to a bus 616, which may represent more than one physical bus. However, the various system components may or may not be physically close to each other. For example, input data and / or output data can be transmitted remotely from one physical location to another. Additionally, programs implementing aspects of the present disclosure can be accessed from a remote location (e.g., a server) via a network. Such data and / or programs can be transmitted via any of a variety of machine-readable media, including, for example: magnetic media such as hard disks, floppy disks, and magnetic tape; and optical media such as compact discs (CDs) and holographic devices; magneto-optical media; and hardware devices specifically configured to store or store and execute program code, such as application specific integrated circuits (ASICs), programmable logic devices (PLDs), flash memory devices, other non-volatile memory (NVM) devices (such as 3D XPoint-based devices), and ROM and RAM devices.

[0069] Aspects of the present disclosure can utilize instructions for one or more processors or processing units to cause steps to be performed, encoded on one or more non-transitory computer-readable media. It should be noted that one or more non-transitory computer-readable media should include volatile and / or non-volatile memory. It should be noted that alternative implementations are possible, including hardware implementations or software / hardware implementations. Hardware-implemented functions can be implemented using ASICs, programmable arrays, digital signal processing circuitry, etc. Thus, the term "means" in any claim is intended to cover software and hardware implementations. Similarly, the term "computer-readable media" as used herein includes software and / or hardware on which a program including instructions is present, or a combination thereof. Given these alternative implementations, it should be understood that the figures and the accompanying description provide the functional information necessary for a person of ordinary skill in the art to write program code (i.e., software) and / or fabricate circuitry (i.e., hardware) to perform the required processing.

[0070] It should be noted that embodiments of the present disclosure may further relate to a computer product having a non-transitory, tangible computer-readable medium having computer code for performing various computer-implemented operations. The media and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well-known or available to those skilled in the relevant art. Examples of tangible computer-readable media include, for example: magnetic media such as hard disks, floppy disks, and magnetic tapes; optical media such as CDs and holographic devices; magneto-optical media; and hardware devices specially configured to store or store and execute program code, such as application specific integrated circuits (ASICs), programmable logic devices (PLDs), flash memory devices, other non-volatile memory (NVM) devices (such as 3D XPoint-based devices), and ROM and RAM devices. Examples of computer code include machine code, such as that generated by a compiler, and files containing higher-level code that is executed by a computer using an interpreter. Embodiments of the present disclosure may be implemented in whole or in part as machine-executable instructions that may be in program modules executed by a processing device. Examples of program modules include libraries, programs, routines, objects, components, and data structures. In a distributed computing environment, program modules may be physically located in local, remote, or both settings.

[0071] Those skilled in the art will recognize that no computing system or programming language is essential for the practice of the present disclosure. Those skilled in the art will also recognize that many of the above elements may be physically and / or functionally separated into modules and / or sub-modules or combined together.

[0072] Those skilled in the art will understand that the foregoing examples and embodiments are exemplary and do not limit the scope of the present disclosure. All arrangements, enhancements, equivalents, combinations, and improvements that are obvious to those skilled in the art upon reading the specification and studying the drawings are included within the true spirit and scope of the present disclosure. It should also be noted that the elements of any claim may be arranged differently, including having multiple dependencies, configurations, and combinations.

Claims

1. A computer-implemented method for training a noise reduction system, comprising: Given a noise reduction system including a first encoder, a second encoder, a quantizer, and a decoder, and given a set of one or more clean-noisy audio pairs, where each clean-noisy audio pair includes clean audio content through a speaker and noisy audio content through a speaker: For each clean audio, use the first encoder to generate one or more consecutive latent representations of the clean audio; For each noisy audio, use the second encoder to generate one or more consecutive latent representations of the noisy audio; For each consecutive latent representation of the clean audio, use the quantizer to generate a corresponding discrete clean audio representation; For each consecutive latent representation of the noisy audio, use the quantizer to generate a corresponding discrete noisy audio representation; For each clean-noisy audio pair, input the discrete clean audio representation, the clean audio, and a speaker embedding representing the speaker of the clean-noisy audio pair into the decoder to generate an audio sequence prediction; Calculate the loss of the noise reduction system, where the loss includes a latent representation matching loss term, and the latent representation matching loss term for time steps where the discrete clean audio representation and the discrete noisy audio representation are different is based on a distance measure between the consecutive latent representation of the clean audio and the consecutive latent representation of the noisy audio for the time step; and Update the noise reduction system using the loss.

2. The method according to claim 1, wherein The latent representation matching loss term further includes: An annealing term that increases from 0 or near 0 to 1 or near 1 during training.

3. The method according to claim 1, wherein, The distance measure between the consecutive latent representation of the clean audio and the consecutive latent representation of the noisy audio includes: Distance between the continuous latent representation of clean audio and the continuous latent representation of time steps. Distance.

4. The method according to claim 1, wherein The loss includes: A decoder term related to the loss of the decoder; And A quantizer term related to the loss of the quantizer.

5. The method according to claim 1, wherein, The quantizer includes one or more vector quantization variational autoencoders that convert one or more consecutive latent representations of the clean audio into corresponding one or more discrete clean audio representations and convert one or more consecutive latent representations of the noisy audio into corresponding one or more discrete noisy audio representations.

6. The method according to claim 1, further comprising: Repeating the steps of claim 1 with one or more additional sets of clean-noisy audio pairs; And In response to reaching a stop condition, output the trained noise reduction system, where the trained noise reduction system includes a trained second encoder, a trained quantizer, and a trained decoder.

7. The method according to claim 6, further comprising: Given noisy audio for noise reduction, and a speaker embedding for the speaker in the noisy audio: Use the trained second encoder to generate one or more consecutive latent representations of the noisy audio; Use one or more consecutive latent representations of the noisy audio and the trained quantizer to generate one or more discrete noisy audio representations; And Generate a noise-reduced audio representation of the noisy audio by inputting at least some of the one or more discrete noisy audio representations and the speaker embedding representing the speaker of the noisy audio into the trained decoder.

8. The method according to claim 1, wherein The decoder is an autoregressive generation model.

9. A system, comprising: One or more processors; And A non-transitory computer-readable medium comprising one or more sets of instructions that, when executed by at least one of one or more processors, cause the processor to perform the computer-implemented method of any one of claims 1-8.

10. A computer-implemented method comprising: Given input noisy audio for noise reduction and a given trained noise reduction system, the trained noise reduction system comprising a trained encoder, a trained quantizer, and a trained decoder: Inputting the input noisy audio into the trained encoder to generate one or more consecutive latent representations of the input noisy audio; Inputting the one or more consecutive latent representations of the input noisy audio into the trained quantizer to generate one or more discrete noisy audio representations; and Inputting the discrete noisy audio representations into the trained decoder to generate a noise-reduced audio representation of the input noisy audio; wherein the trained noise reduction system is trained using the method of claim 1.

Citation Information

Patent Citations

  • Systems and methods for parallel wave generation in end-to-end text-to-speech

    US20190180732A1

  • Generating music with deep neural networks

    US10068557B1

  • Systems and methods for robust speech recognition using generative adversarial networks

    US20190130903A1