Deep learning-based voice enhancement
A deep learning-based neural network model effectively suppresses noise in real-time by predicting voice presence across frequency bands, addressing the challenge of mixed audio signal noise removal with high accuracy and low latency.
Patent Information
- Application Number
- JP2023526072
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-07-14
- Filing Date
- 2021-10-29
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-10-29
AI Technical Summary
Accurately removing noise from mixed audio signals is challenging due to various forms of audio and noise types, especially in real-time applications.
A deep learning-based neural network model is employed to suppress noise by training a neural network model with a feature extraction block, encoder, and decoder to generate voice values indicating the amount of voice present in each frequency band, using look-ahead and low-latency convolutional kernels for real-time noise suppression.
The system achieves accurate and low-latency noise suppression by predicting voice presence and distribution across frequency bands, enhancing audio quality while reducing computational overhead.
Smart Images

Figure 0007711190000008 
Figure 0007711190000009 
Figure 0007711190000010
Abstract
Description
Technical Field
[0001] This application claims priority to U.S. Provisional Application No. 63 / 115,213, filed on November 18, 2020, U.S. Provisional Application No. 63 / 221,629, filed on July 14, 2021, and International Patent Application No. PCT / CN2020 / 124635, filed on October 29, 2020, and all of these are incorporated herein by reference in their entirety.
[0002] This application relates to noise reduction from audio. More specifically, the embodiments described below relate to applying a deep learning model to generate frame-based inferences from large-scale audio contexts.
Background Art
[0003] The approaches described in this section are approaches that could be advanced, and are not necessarily approaches that have been devised or pursued heretofore. Accordingly, unless otherwise noted, none of the approaches described in this section should be assumed to be prior art solely for the reason that they are included in this section.
[0004] Accurately removing noise from a mixed signal of audio and noise is generally difficult considering that there can be various forms of audio and various types of noise. Suppressing noise in real time can be particularly challenging.
Summary of the Invention
[0005] A system and related method for suppressing noise and enhancing voice are disclosed. The method includes receiving, by a processor, input audio data covering a plurality of frequency bands along a frequency dimension in a plurality of frames along a time dimension, and training, by the processor, a neural network model, the neural network model including a feature extraction block that performs a look-ahead of a specific number of frames when extracting features from the input audio data, an encoder including a first series of blocks that generate a first feature map corresponding to a gradually increasing receptive field within the input audio data along the frequency dimension, a decoder including a second series of blocks that receive the output feature map generated by the encoder as an input feature map and generate a second feature map, and a classification block that receives the second feature map and generates a voice value indicating the amount of voice present for each frequency band of the plurality of frequency bands in each frame of the plurality of frames; receiving new audio data having one or more frames; executing the neural network model on the new audio data to generate a new voice value for each frequency band of the plurality of frequency bands in each frame of the one or more frames; generating new output data that suppresses noise in the new audio data based on the new voice values; and transmitting the new output data. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] In the accompanying drawings, which include the following figures and in which like elements are referred to by like reference numerals, embodiments of the invention are shown by way of example and not by way of limitation.
Figure 1
Figure 2
Figure 3
Figure 4A
Figure 4B
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Mode for Carrying Out the Invention
[0007] In the following description, for the purpose of explanation, numerous specific details are set forth in order to provide a thorough understanding of embodiments of the present invention. However, it is apparent that the embodiments may be practiced without these specific details. Also, well-known structures and devices are shown in block diagram form in order not to unnecessarily obscure the embodiments.
[0008] The embodiments will be described in the following sections according to the following outline: 1. General Overview 2. Example Computing Environments 3. Example Computer Components 4. Functional Description 4.1. Neural Network Model 4.1.1. Feature Extraction Block 4.1.2. U-NET Block 4.1.2.1. Encryption Block 4.1.2.1.1. Depthwise Separable Convolution Using Gating 4.1.2.2. Residual Block and Recurrent Layer 4.2. Model Training 4.3. Model Execution 5. Process Example 6. Hardware Implementation **
[0009] 1. Overall Overview A system and related methods for suppressing noise and enhancing voice are disclosed. In some embodiments, the system trains a neural network model that obtains banded (banding) energy corresponding to an original noisy waveform and generates voice values indicating the amount of voice present in each band in each frame. Using these voice values, noise can be suppressed by reducing the magnitude of frequencies in frequency bands where voice is unlikely to be present. This neural network model has low latency and can be used for real-time noise suppression. The neural model has a feature extraction block that performs some lookahead. Following the feature extraction block is an encoder with stationary downsampling along the frequency domain that forms a contraction path. The convolution along the contraction path is performed with dilation coefficients that gradually increase along the time dimension. Following the encoder is a corresponding decoder with stationary upsampling along the frequency domain that forms an expansion path. The decoder receives the output feature map scaled from the encoder at the corresponding level and can thus consider all features extracted from different receptive fields along the frequency dimension when determining how much voice is present in each frequency band in each frame.
[0010] In some embodiments, at runtime, the system acquires a noisy waveform and converts it to the frequency domain covering a plurality of perceptually motivating frequency bands in each frame. The system then runs a model to obtain voice values for each frequency band in each frame. Thereafter, the system applies the voice values to the original data in the frequency domain and converts it back to an enhanced noise-suppressed waveform.
[0011] The system has various technical benefits. The system is designed to be accurate while having low latency for real-time noise suppression. The low latency is achieved by a relatively small number of relatively small convolutional kernels, such as 8 two-dimensional kernels of size 1×1 or 3×3 in a lean convolutional neural network (CNN) model. The collation and integration of the original frequency domain data into perceptually stimulating bands further reduces the amount of computation. Depthwise separable convolutions, which tend to shorten the execution time, are also applied where possible.
[0012] Accuracy is achieved by feature extraction for different receptive fields in the input data along the frequency dimension, which are used in combination to achieve dense classification. A specific feature extraction block that incorporates a look-ahead of a small number of frames, such as one or two frames, further contributes to the richness of the features. Dense blocks where the output feature maps of the convolutional layers are propagated to all subsequent convolutional layers are also applied where possible. Furthermore, the neural model can be trained to predict not only the amount of voice present for each frequency band in each frame, but also the distribution of such amounts. The prediction can be fine-tuned using the additional parameter of the distribution.
[0013] 2. Computing Environment Example FIG. 1 shows an example of a networked computer system in which various embodiments may be implemented. FIG. 1 is shown in a simplified schematic form for purposes of illustration of a clear example, and other embodiments may include more, fewer, or different elements.
[0014] In some embodiments, the networked computer system has an audio management server computer 102 (“server”), one or more sensors 104 or input devices, and one or more output devices 110, which are communicatively coupled either through a direct physical connection or via one or more networks 118.
[0015] In some embodiments, the server 102 broadly represents an instance of an application that is programmed or configured with one or more computers, virtual computing instances, and / or data structures and / or database records configured to host or execute functions related to low-latency voice enhancement by noise reduction. The server 102 can have a server farm, a cloud computing platform, a parallel computer, or any other computing facility having sufficient computing capabilities for data processing, data storage, and network communication for the functions described above.
[0016] In some embodiments, each of the one or more sensors 104 can include a microphone or other digital recording device that converts sound into an electrical signal. Each sensor is configured to transmit the detected audio data to the server 102. Each sensor may include a processor or may be integrated into a typical client device such as, for example, a desktop computer, a laptop computer, a tablet computer, a smartphone, or a wearable device.
[0017] In some embodiments, each of the one or more output devices 110 can include a speaker or other digital playback device that converts an electrical signal back into sound. Each output device is programmed to play the audio data received from the server 102. Similar to the sensors, the output devices may include a processor or may be integrated into a typical client device such as, for example, a desktop computer, a laptop computer, a tablet computer, a smartphone, or a wearable device.
[0018] One or more networks 118 can be implemented by any medium or mechanism that provides for the exchange of data between the various elements of FIG. 1. Examples of networks 118 include, but are not limited to, a cellular network communicatively coupled to a computing device via a cellular antenna, a Near Field Communication (NFC) network, a Local Area Network (LAN), a Wide Area Network (WAN), the Internet, an over-the-air link, or a satellite link, among one or more of the foregoing.
[0019] In some embodiments, the server 102 is programmed to receive input audio data corresponding to sound in a given environment from one or more sensors 104. The server 102 is then programmed to process the input audio data, which typically corresponds to a mixture of voice and noise, to estimate how much voice is present in each frame of the input data. The server 102 is also programmed to update the input audio data based on the estimate to produce cleaned-up output audio data that is expected to contain less noise than the input audio data. Further, the server 102 is programmed to send the output audio data to one or more output devices.
[0020] 3. Examples of Computer Components FIG. 2 shows an example of components of an audio management server computer according to the disclosed embodiments. This figure is for illustrative purposes only, and server 102 can have fewer or more functional components or storage components. Each of the functional components can be implemented as a software component, a general-purpose or special-purpose hardware component, a firmware component, or any combination thereof. Each of the functional components can also be coupled to one or more storage components (not shown). The storage component can be implemented using any of a relational database, an object database, a flat file system, or a JSON store. The storage component can be connected to the functional component locally or via a network using a program call, a remote procedure call (RPC) facility, or a messaging bus. The components may or may not be self-contained. Depending on implementation-specific or other considerations, the components can be functionally or physically centralized or distributed.
[0021] In some embodiments, server 102 includes a spectral conversion and banding block 204, a model block 208, an inverse banding block 212, an input spectrum multiplication block 218, and an inverse spectral conversion block 222.
[0022] In some embodiments, server 102 receives a noisy waveform. At block 204, server 102 segments the waveform through a spectral transform into a sequence of frames, e.g., a 6 - second long sequence (resulting in 300 frames) with overlapping or non - overlapping 20 - ms frames. The spectral transform can be any of a variety of transforms such as, for example, a short - time Fourier transform or a Complex Quadrature Mirror Filterbank (CQMF) transform, the latter of which tends to minimize aliasing artifacts. To ensure a relatively high frequency resolution, the number of transform kernels / filters per 20 - ms frame can be selected such that the frequency bin width is approximately 25 Hz.
[0023] In some embodiments, server 102 then converts the sequence of frames into a vector of banded energy for, e.g., 56 perceptually - stimulated bands. Each of the perceptually - stimulated bands is typically located within a frequency domain, e.g., from 120 Hz to 2000 Hz, that matches the way a human ear processes speech. Thus, capturing data within these perceptually - stimulated bands means that the speech quality is not lost for the human ear. More specifically, the squared magnitude of the output frequency bins of the spectral transform is grouped into perceptually - stimulated bands, and the number of frequency bins per band increases at higher frequencies. The grouping strategy can be “soft” where some spectral energy leaks across adjacent bands or “hard” where there is no leakage across bands.
[0024] In some embodiments, when the bin energy of a noisy frame is represented by a column vector x of size p×1, where p represents the number of bins, the conversion to a vector of banded energy can be performed by calculating y = W * x, where y is a column vector of size q×1 representing the band energy of this noisy frame, W is a banding matrix of size q×p, and q represents the number of perceptually stimulated bands.
[0025] In some embodiments, at block 208, server 102 predicts a mask value indicating the amount of voice present for each band in each frame. At block 212, server 102 converts the band mask value back to a spectral bin mask.
[0026] In some embodiments, when the band mask for y is represented by a column vector m_band of size q×1, the conversion to a bin mask can be performed by calculating m_bin = W_transpose * m_band, where m_bin is a column vector of size p×1 and W_transpose of size p×q is the transpose matrix of W. At block 218, server 102 multiplies the spectral bin mask by the spectral intensity to achieve noise masking or reduction and obtain an estimated clean spectrum. Finally, at block 222, the server converts the estimated clean spectrum back to a waveform as a waveform (relative to the noise waveform) emphasized using any method known to those skilled in the art, such as an inverse transform (e.g., inverse CQMF, etc.), which can be communicated via an output device.
[0027] 4. Functional Description 4.1. Neural Network Model FIG. 3 shows an example of a neural network model 300 for noise reduction, which represents an embodiment of block 208. In some embodiments, model 300 has a feature extraction block 308 and a block 340 based on a U-Net structure such as that described in arXiv:1505.04597v1 [cs.CV] dated May 18, 2015, with several variations as described herein. The U-Net structure has been shown to enable accurate localization of feature recognition and classification.
[0028] 4.1.1. Feature Extraction Block In some embodiments, at block 308 of FIG. 3, server 102 extracts high-level features optimized for the noise suppression task from raw band energy. FIG. 4A shows an example of a feature extraction block that represents an embodiment of block 308. FIG. 4B shows another example of a feature extraction block. As shown in configuration 400A of FIG. 4A, for example, server 102 can normalize the mean and variance of the band energy (e.g., 56 of them) in a sequence of T frames by a learnable batch normalization (BATCHNORM) layer 408 known to those skilled in the art. Alternatively, global normalization can also be pre-computed from the training set using techniques known to those skilled in the art.
[0029] In some embodiments, server 102 can take into account future information when extracting the above-described high-level features. As shown in 400A of FIG. 4A, for example, such a look-ahead can be implemented using a two-dimensional (2D) one-channel convolutional (conv2d) layer 406 having one or more kernels. The height of the kernel in the conv2d layer 406 corresponding to the number of n bands for evaluating each time can be set to a small value such as 3, for example. The kernel size along the time axis depends on how much look-ahead is desired or allowed. For example, without look-ahead, the kernel can cover the current frame and L past frames such as, for example, 2 frames, and when L future frames are allowed, the kernel size can be 2L + 1 centered on the current frame, and can be made to match the 2L + 1 frames in the input data at each time, such as 422 where L is 2 in 406. As shown in 400B of FIG. 4B, the look-ahead can also be implemented using a series of conv2d layers 410, 412, or more. In that case, each kernel has a small kernel size along the time axis. For example, L can be set to 1 for 410, 412, and all other similar layers. As a result, layer 410 can be matched with the original input data with a 2L + 1 look-ahead, such as 422 where L is 1 and leading to three kernels 428, for example, and layer 412 can be matched with the output of layer 412. The server can gradually increase the receptive field in the input data using the series of conv2d layers shown in FIG. 4B.
[0030] In some embodiments, the number of kernels in each conv2d layer can be determined based on the nature of the input audio stream, the amount of desired high-level features, the range of computing resource requirements, or other factors. For example, the number can be 8, 16, or 32. Additionally, after each of the conv2d layers within block 308, a non-linear activation function, such as a parametric rectified linear unit (PReLU), can follow, and then a separate batch normalization layer for fine-tuning the output of block 308 can follow.
[0031] In some embodiments, block 308 may be implemented using other signal processing techniques not related to artificial neural networks, such as those described in "Power-Normalized Cepstral Coefficients (PNCC) for Robust Speech Recognition" by C. Kim and R. M. Stern, IEEE / ACM Transactions on Audio, Speech, and Language Processing, vol. 24, no. 7, pp. 1315-1329, July 2016, doi: 10.1109 / TASLP.2016.2545928.
[0032] 4.1.2. U-NET Block In some embodiments, at block 340 of FIG. 3, the server 102 performs encoding of the feature data (to find more and better features) and then, before finally performing classification to determine how many voices there are, decodes to reconstruct the enhanced audio data. Thus, block 340 has a left encoder side and a right decoder connected by block 350. The encoder has one or more feature calculation blocks such as 310, 312, and 314, each followed by a frequency downsampler (DS) such as 316, 318, and 320 to form a contraction path. A dense block (DB) is an implementation of such a feature calculation block, as will be further described below. Each of the triples shown in the figure, such as (8, T, 64), includes the size of the input or output data of the feature calculation block, where the first component represents the number of channels or feature maps, the second component represents a fixed number of frames along the time dimension, and the third component represents the size along the frequency dimension. These feature calculation blocks capture higher-level features in a larger frequency context, as will be further described below. Block 350 has a feature calculation block for performing modeling that covers all perceptually stimulated bands originally available. The decoder also has one or more feature calculation blocks such as 320, 322, and 324, each followed by a frequency upsampler (US) such as 326, 328, and 330 to form an expansion path. These feature calculation blocks within the expansion path, which rely on the feature maps generated in the contraction path, are combined to project discriminative features at multiple different levels, i.e., for each band in each frame, onto a high-resolution space to obtain a dense classification, which is the mask value. As a result, the number of input channels (or feature maps) of each feature calculation block within the expansion path can be twice the number in each feature calculation block within the contraction path. However, the selection of the number of kernels in each calculation block can determine the number of output channels, which becomes the number of input channels for the next feature calculation block within the expansion path.
[0033] Server 102 generates a final mask value for each band in the frame via a classification block such as block 360 having a 1×1 2D kernel followed by a sigmoid non-linear activation function (SIGMOID).
[0034] In some embodiments, at each frequency downsampler, server 102 merges every two adjacent band energies via a conv2d layer having a kernel and a stride size of 2 along the frequency axis, via regular convolution or depthwise convolution. Alternatively, the conv2d layer may be replaced with a max-pooling layer. In either case, the width of the output feature map is halved after each frequency downsampler, thereby steadily expanding the receptive field within the input data. To enable such a continuous exponential reduction in the width of the output feature map, server 102 pads the output of block 308 to a width that is a power of 2, and that becomes the input data to block 340. This padding can be done, for example, by adding zeros to both sides of the output feature map of block 308.
[0035] In some embodiments, at each frequency upsample, server 102 uses a transposed conv2d layer corresponding to the same level of conv2d layer in the encoder to restore the original number of band energies. The depth of block 340, i.e., the number of combinations of the feature calculation block and the frequency downsampler (and equivalently, the number of combinations of the feature calculation block and the frequency upsample) may depend on the desired maximum receptive field, the amount of computing resources, or other factors.
[0036] In some embodiments, server 102 uses skip connections such as 342, 344, and 346, for example, as a path for the decoder to receive the discriminative features of the input data at multiple different levels for fine classification as described above, and connects the output of the feature calculation block in the encoder to the input of the feature calculation block in the decoder at the same level. For example, the feature map generated by block 310 is used as input data together with the feature map supplied from frequency upsampler 330 to block 324 via skip connection 346. As a result, the number of channels of the input data of each feature calculation block in the decoder becomes twice the number of channels of the input data of each dense block in the encoder.
[0037] In some embodiments, instead of simple concatenation, server 102 learns scaling multipliers for each skip connection, such as α1, α2, and α3, as shown in FIG. 3. Each α i includes N (e.g., 8) learnable parameters, which can be initialized to 1 at the start of training. Each of the learnable parameters is used to multiply the feature map generated by the corresponding feature calculation block in the encoder to generate a scaled feature map, which is then concatenated with the feature map supplied to the corresponding feature calculation block in the decoder.
[0038] In some embodiments, server 102 can replace concatenation with addition. For example, the eight feature maps generated by block 310 can be added to the eight feature maps supplied to dense block 324, respectively, and each of these eight additions is performed on a component basis. Such addition instead of concatenation reduces the number of feature maps used as input data to each feature calculation block in the decoder and reduces the overall calculation at the expense of a certain performance degradation.
[0039] 4.1.2.1. Dense Block FIG. 5 shows an example of a neural network model corresponding to one embodiment of block 310 in FIG. 3 and all other similar blocks within block 340. The neural network model is based on a DenseNet structure, such as that described in arXiv:1608.06993v5 [cs.CV] dated Jan. 28, 2018, for example, but has several variations as described herein. The DenseNet structure has been shown to mitigate the vanishing gradient problem, enhance feature propagation, promote feature reuse, and reduce the number of parameters.
[0040] In some embodiments, server 102 uses block 500 as a feature calculation block to further enhance feature propagation and dense classification. Block 500 outputs a feature map of N (e.g., 8) channels that is the same number as the number of feature map input data. Each channel also has the same time-frequency shape as the feature map of the input data. Block 500 has a series of convolutional layers, such as 520 and 530. The input data to each convolutional layer includes the concatenation of all output data from the preceding convolutional layer, thereby forming a dense connection. For example, the input data to layer 530 may include data 512, which may be the first input data or the output data from the previous convolutional layer, and data 522, which is the input data from layer 520.
[0041] In some embodiments, each convolutional layer has a bottleneck layer with one or more 1×1 2D kernels, such as layer 504, to consolidate the input data having K feature maps into a smaller number of feature maps due to the dense connection. For example, each 1×1 2D kernel can be applied to each group of K / 2N feature maps respectively to effectively add together K / 2N feature maps into one feature map, and ultimately 2N feature maps can be obtained. Alternatively, a total of 2N 1×1 2D kernels can be applied to all feature maps to generate a 2D feature map. After each 1×1 2D kernel, a non-linear activation function, such as PReLU, and / or a batch normalization layer may follow.
[0042] In some embodiments, each convolutional layer has a small conv2d layer with N kernels, such as block 506 with a 3×3 convd2d layer, followed by a bottleneck layer, to generate N feature maps. These small conv2d layers in successive convolutional layers of block 500 use dilations that increase exponentially along the time axis to model increasingly large context information. For example, the dilation coefficient used in block 506 is 1, which means no dilation in each kernel, and the dilation coefficient used in block 508 is 2, which means the kernel is dilated by a factor of 2 along the time axis and the receptive field also increases in size by a factor of 2 in each dimension.
[0043] In some embodiments, between the convolutional layers of block 500, the server 102 linearly projects the band energy into the learned space within the frequency mapping layer for more integrated output, such as that described in arXiv:1904.11148v1 [cs.SD] dated April 25, 2019. Since the same kernel may produce different effects on the same audio data depending on which frequency band the audio data is located in, some integration of such effects across different bands is useful. For example, the frequency mapping layer 580 is placed in the middle of the depth of block 500.
[0044] In some embodiments, at the end of block 500, an output tensor with N feature maps can be generated using a layer 590 similar to the bottleneck layer with another 1×1 2D kernel.
[0045] 4.1.2.1.1. Depthwise Separable Convolution Using Gating FIG. 6 shows an example of a neural network model corresponding to one embodiment of block 506 shown in FIG. 5 and all other similar blocks. In some embodiments, block 600 has a depthwise separable convolution using a non-linear activation function such as, for example, a gated linear unit (GLU). As shown in FIG. 6, the first path within the GLU has a depthwise small conv2d layer, such as, for example, a 3×3 conv2d layer 602, followed by a batch normalization (BATCHNORM) layer 604. Similarly, the second path within the GLU has a 3×3 conv2d layer 606 and a subsequent batch normalization layer 608, and then a learnable gating function, such as, for example, a sigmoid non-linear activation function (SIGMOID), follows. Similar to the dense blocks shown in FIG. 5, the small conv2d layers in successive convolutional layers of block 500 use dilations that increase exponentially along the time axis to model increasingly large context information. For example, blocks 602 and 606 within the convolutional layer corresponding to block 506 can be associated with a dilation factor of 1, and similar blocks within the next convolutional layer that may correspond to one embodiment of block 508 can be associated with a dilation factor of 2. The gating function identifies regions of the input data that are important for the task of interest. These two paths are combined by an element-wise product operator 618. A 1×1 conv2d layer 612 learns the interconnections between the output feature maps generated by the combination of the two paths as part of the depthwise separable convolution. After layer 612, a batch normalization layer 614 and a non-linear activation function 616, such as, for example, PReLU, can follow.
[0046] 4.1.2.2. Residual Blocks and Recurrent Layers FIG. 7 shows an example of a neural model corresponding to an embodiment of block 310 shown in FIG. 3 and all other similar blocks. In some embodiments, block 500 shown in FIG. 5, which also corresponds to an embodiment of block 310, can be replaced by a residual 700 block for reducing the number of connections. Block 700 has a plurality of convolutional layers such as, for example, layers 720 and 730.
[0047] In some embodiments, each convolutional layer has a bottleneck layer similar to block 504 shown in FIG. 5, such as, for example, layer 704. The bottleneck layer may also be followed by a non-linear activation such as, for example, PReLU, and / or a batch normalization layer.
[0048] In some embodiments, the convolutional layer also has a small conv2d layer similar to block 506 shown in FIG. 5, such as, for example, the 3×3 conv2d layer 706. This small conv2d block can be executed using dilation with a dilation coefficient that exponentially increases across consecutive convolutional layers. This small conv2d layer can be replaced by a depthwise separable convolution using gating as shown in FIG. 6.
[0049] In some embodiments, the convolutional layer has another 1×1 conv2d layer such as, for example, layer 708, which matches and returns the output of block 706 to the input of block 704 in terms of size, specifically in terms of the number of channels or feature maps. Then, when the output is added to the input data via the Hadamard product operator 710, the vanishing gradient problem is suppressed when training the network using backpropagation. This is because the gradient will have a direct path from the output to the input side without any multiplication in between. This conv1x1 layer can also be followed by a non-linear activation such as, for example, PReLU, and / or a batch normalization layer.
[0050] In some embodiments, block 500 shown in FIG. 5, which also corresponds to an embodiment of block 310, may be replaced by a recurrent layer having at least one recurrent neural network (RNN). Using an RNN to model long time sequences can be an efficient approach. By "efficient" it is meant that the RNN can model very long time sequences by maintaining an internal hidden state vector as a summary of all the history seen and generating an output for each new frame based on that vector. Compared to using dilation in a CNN layer, the buffer size for storing past information in an RNN is much smaller (only 1 vector compared to the 2d + 1 vectors in a CNN, where d is the dilation factor).
[0051] 4.2. Model Training In some embodiments, the training of the neural network model 208 can be performed as an end-to-end process. Alternatively, the feature extraction block 308 and the U-Net block 340 may be trained separately, and the output of applying the feature extraction block 308 to actual data can be used as training data for the U-Net block.
[0052] A variety of training data is used to train the neural network model 208 shown in FIG. 2. In some embodiments, the diversity incorporates speaker diversity by including in the training data spontaneous speech in a wide range of speaking styles with respect to speed, emotion, and other attributes. Each utterance for training may be the voice of a single speaker or a conversation between multiple speakers.
[0053] In some embodiments, diversity results from including concentrated noise data that includes reverberation data. A database such as AudioSet can be used as a seed noise database. Server 102 can filter each clip in the seed noise database using class labels indicating that there is a high likelihood that audio is present within the clip. For example, the class of "human voice" in a given ontology can be filtered. The seed noise database can be further filtered by applying any voice separation technique known to those skilled in the art to remove further clips where there is a high likelihood that voice is present. For example, any clip having at least one frame (e.g., a frame of length 100 ms) with root mean square energy exceeding a threshold (e.g., 1e-3) is removed.
[0054] In some embodiments, diversity is increased by including a wide range of intensity levels when mixing noise with voice. When creating a noisy signal, server 102 scales the clean voice signal and the noise signal to a respective predetermined maximum level and randomly adjusts each to be reduced by one of a range of dBs, such as from 0 to 30 dB, and can randomly add the adjusted clean voice signal and the adjusted noise signal according to a predetermined minimum signal-to-noise ratio. Such a wide range of loudness levels has been found to help reduce over-suppression of voice (or under-suppression of noise).
[0055] In some embodiments, diversity lies in the presence of data within different frequency bands. Server 102 can create a signal having at least a certain percentage within a particular frequency band of a particular bandwidth, such as at least 20% within a frequency band from 300 Hz to 500 Hz.
[0056] In some embodiments, server 102 trains neural network model 208 using any optimization process known to those skilled in the art, such as, for example, a stochastic gradient descent optimization algorithm in which weights are updated using the error backpropagation algorithm. Neural network model 208 can minimize the mean squared error (MSE) loss between the predicted mask and the ground truth mask for each band in each frame. The ground truth mask can be calculated as the ratio of the audio energy to the sum of the audio energy and the noise energy.
[0057] In some embodiments, since over-suppression of speech degrades speech quality more than under-suppression of speech, server 102 uses weighted MSE that assigns a larger penalty to over-suppression of speech. Since the mask values generated by neural network model 208 indicate the amount of speech present, when the predicted mask value is less than the ground truth mask value, less speech than the ground truth is predicted, and thus more speech than necessary is suppressed, leading to over-suppression of speech by the neural network model. For example, the weighted MSE can be calculated as follows:
Equation
Equation
Equation
[0058] In some embodiments, the neural network model 208 is trained to predict the distribution of speech across a plurality of different frequency bins within each band (rather than a single mask value). Specifically, the server 102 can train the model to predict the mean and variance values of a Gaussian distribution for each band in each frame, where the mean represents the best prediction of the mask value by the neural network model 208. The loss function for the Gaussian distribution is:
Number
Number
[0059] In some embodiments, the variance prediction can be interpreted as the confidence in the mean prediction for reducing the occurrence of speech over-suppression. If the mean prediction is relatively low, indicating a small amount of existing speech, and the variance prediction is relatively high, this may indicate a high likelihood of speech over-suppression, where the band mask can be scaled up. An example of a scaling function for generating a gain adjusted based on the standard deviation is:
Number
[0060] In some embodiments, assuming a Gaussian distribution for each mask, the probability of each observed (target) mask value is:
Number
[0061] 4.3. Model Execution In some embodiments, when lookahead is implemented in neural network model 208, specifically in feature extraction block 308, server 102 can accept individual frames or sets of frames as input data and generate at least mask values for each frame as output data. For each convolutional layer having a kernel size greater than 1 along the time dimension, server 102 maintains an internal buffer for storing the history required to generate the output data. The buffer can be maintained as a queue having a size equal to the receptive field of the convolutional layer along the time dimension.
[0062] 5. Process Example FIG. 8 shows an example of a process executed on an audio management server computer according to some embodiments described herein. FIG. 8 is shown in a simplified schematic form for purposes of illustration of a clear example, and other embodiments may include more, fewer, or different elements connected in various ways. Each of FIG. 8 is intended to disclose an algorithm, plan, or outline that can be used to implement one or more computer programs or other software elements that, when executed, perform the functional improvements and technological advancements described herein. Also, the flow diagrams herein are described in the same level of detail normally used to communicate to one another about algorithms, plans, or specifications that form the basis of software programs that those skilled in the art plan to code or implement using their accumulated skills and knowledge.
[0063] In some embodiments, at step 802, server 102 is programmed to receive input audio data that covers a plurality of frequency bands along the frequency dimension in a plurality of frames along the time dimension. In some embodiments, the plurality of frequency bands are perceptually stimulating bands, and at higher frequencies, cover more frequency bins.
[0064] In some embodiments, at step 804, server 102 is programmed to train a neural network model. The neural network model includes a feature extraction block that performs a look-ahead of a specific number of frames when extracting features from the input audio data, an encoder that includes a first series of blocks that generate feature maps corresponding to progressively larger receptive fields within the input audio data along the frequency dimension, a decoder that includes a second series of blocks that receive the output feature map generated by the encoder as an input feature map, and a classification block that generates an audio value indicating the amount of voice present for each frequency band of the plurality of frequency bands in each frame of the plurality of frames.
[0065] In some embodiments, the feature extraction block has a convolutional kernel of a specific size along the time dimension, and the encoder and decoder do not have convolutional kernels of a size along the time dimension that is greater than the specific size. In other embodiments, each of the feature extraction block, the first series of blocks, and the second series of blocks generates a common number of feature maps.
[0066] In some embodiments, the feature extraction block has a batch normalization layer followed by a convolutional layer having a two-dimensional convolutional kernel.
[0067] In some embodiments, each block of the first series of blocks in the encoder has a feature calculation block and a frequency downsampler. The feature calculation block has a series of convolutional layers.
[0068] In some embodiments, the output data of a convolutional layer among a series of convolutional layers is supplied to all subsequent convolutional layers among the series of convolutional layers. The series of convolutional layers implements dilations that gradually increase along the temporal dimension. In other embodiments, each of the series of convolutional layers has a depthwise separable convolutional block having a gating mechanism.
[0069] In some embodiments, each of the series of convolutional layers has a residual block having a series of convolutional blocks including a first convolutional block having an initial 1×1 two-dimensional convolutional kernel and a last convolutional block having a final 1×1 two-dimensional convolutional kernel.
[0070] In some embodiments, the output data of a feature calculation block within a block among a first series of blocks is scaled by learnable weights to form scaled output data, and the scaled output data is communicated via a skip connection to a block among a second series of blocks within a decoder.
[0071] In some embodiments, a frequency downsampler of a block within a first series of blocks has a convolutional kernel having a stride size greater than 1 along the frequency dimension.
[0072] In some embodiments, each block of a second series of blocks has a feature calculation block and a frequency upsampler. The feature calculation block within a block among the second series of blocks receives first output data from a feature calculation block within a block among the first series of blocks and second output data from a frequency upsampler of a preceding block within the second series of blocks. Then, the first output data and the second output data are concatenated or added together to form specific input data for the feature calculation block within the block among the second series of blocks.
[0073] In some embodiments, the classification block has a 1×1 two-dimensional convolutional kernel and a non-linear activation function.
[0074] In some embodiments, the neural network model further has a feature calculation block that is the output data of the encoder and the input data of the decoder.
[0075] In some embodiments, server 102 is programmed to perform training for each frequency band of a plurality of frequency bands in each frame using a function of the loss between the predicted audio value and the ground truth audio value, increasing the weight in the loss function when the predicted audio value corresponds to over-suppression of the audio, and decreasing the weight in the loss function when the predicted audio value corresponds to under-suppression of the audio. In some embodiments, the classification block further generates a distribution of the audio volume across a certain frequency band among the plurality of frequency bands in the frame, and the audio value is the average of the distribution.
[0076] In some embodiments, the input audio data has data corresponding to audio at different speeds or emotions, data containing different levels of noise, or data corresponding to different frequency bins.
[0077] In some embodiments, at step 806, server 102 is programmed to receive new audio data having one or more frames.
[0078] In some embodiments, at step 808, server 102 is programmed to execute the neural network model on the new audio data to generate new audio values for each frequency band of a plurality of frequency bands in each of the one or more frames.
[0079] In some embodiments, at step 810, server 102 is programmed to generate new output data for suppressing noise in the new audio data based on the new audio values.
[0080] In some embodiments, at step 812, server 102 is programmed to transmit new output data.
[0081] In some embodiments, server 102 is programmed to receive an input waveform. Server 102 is then programmed to convert the input waveform into raw audio data covering a plurality of frequency bins along the frequency dimension in one or more frames along the time dimension. Server 102 is then programmed to convert the raw audio data into new audio data by grouping the plurality of frequency bins into a plurality of frequency bands. Server 102 is programmed to perform inverse banding on the new audio values to generate updated audio values for each frequency bin of the plurality of frequency bins in each frame of the one or more frames. Further, server 102 is then programmed to apply the updated audio values to the raw audio data to generate new output data. Finally, server 102 is programmed to convert the new output data into an enhanced waveform.
[0082] 6. Hardware Implementation According to one embodiment, the technology described herein is implemented by at least one computing device. The technology can be implemented in whole or in part using a combination of at least one server computer and / or other computing devices coupled via a network, such as a packet data network. The computing device may be hardwired to execute the technology, may include at least one application specific integrated circuit (ASIC) or field programmable gate array (FPGA) permanently programmed to execute the technology, or may include at least one general purpose hardware processor programmed to execute the technology according to program instructions in firmware, memory, other storage, or a combination. Such a computing device may also combine custom hardwired logic, ASICs, or FPGAs with custom programming for achieving the described technology. The computing device can be a server computer, a workstation, a personal computer, a portable computer system, a handheld device, a mobile computing device, a wearable device, a body-mounted or implanted device, a smartphone, a smart appliance, an internetworking device, such as an autonomous or semi-autonomous device like a robot or an unmanned ground vehicle or aircraft, any other electronic device incorporating hardwired logic and / or program logic for implementing the described technology, one or more virtual computing machines or instances within a data center, and / or a network of server computers and / or personal computers.
[0083] FIG. 9 is a block diagram showing an example of a computer system in which an embodiment can be implemented. In the example of FIG. 9, a computer system 900 and instructions for implementing the disclosed technology in hardware, software, or a combination of hardware and software are schematically represented, for example, as boxes and circles, at the same level of detail commonly used by those skilled in the art to communicate about computer architecture and computer system implementation with which this disclosure is concerned.
[0084] The computer system 900 includes an input / output (I / O) subsystem 902 that may include a bus and / or other communication mechanism(s) for communicating information and / or instructions between components of the computer system 900 via an electronic signal path. The I / O subsystem 902 may include an I / O controller, a memory controller, and at least one I / O port. The electronic signal path is schematically represented in the figure, for example, as a line, a one-way arrow, or a two-way arrow.
[0085] At least one hardware processor 904 is coupled to the I / O subsystem 902 for processing information and instructions. The hardware processor 904 may include, for example, a general-purpose microprocessor or microcontroller, and / or a dedicated microprocessor such as, for example, an embedded system or a graphics processing unit (GPU) or a digital signal processor or an ARM processor. The processor 904 may have an integrated arithmetic logic unit (ALU) or may be coupled to a separate ALU.
[0086] Computer system 900 includes one or more units of memory 906, such as main memory, coupled to I / O subsystem 902 for electronically storing instructions and data digitally for execution by processor 904. Memory 906 may include volatile memory, such as various forms of random access memory (RAM) or other dynamic storage devices. Memory 906 may also be used to store temporary variables or other intermediate information during execution of instructions by processor 904. When such instructions are stored on a non-transitory computer-readable storage medium accessible to processor 904, computer system 900 can be customized into a dedicated machine that executes the operations defined by the instructions.
[0087] Computer system 900 further includes non-volatile memory, such as read-only memory (ROM) 908 or other static storage devices, coupled to I / O subsystem 902 for storing information and instructions for processor 904. ROM 908 may include various forms of programmable ROM (PROM), such as erasable PROM (EPROM) or electrically erasable PROM (EEPROM). A unit of persistent storage 910 can include various forms of non-volatile RAM (NVRAM), such as flash memory, or solid state storage, magnetic disk, or optical disk, such as a CD-ROM or DVD-ROM, and can be coupled to I / O subsystem 902 for storing information and instructions. Storage 910 is an example of a non-transitory computer-readable medium used to store instructions and data that, when executed by processor 904, cause a computer-implemented method to execute and implement the techniques herein.
[0088] Instructions in the memory 906, ROM 908, or storage 910 may have one or more sets of instructions organized as a module, method, object, function, routine, or call. The instructions may be organized as an application program including one or more computer programs, operating system services, or mobile apps. The instructions may be an operating system and / or system software; one or more libraries supporting multimedia, programming, or other functions; data protocol instructions or stacks implementing TCP / IP, HTTP, or other communication protocols; file processing instructions for interpreting and rendering files coded using HTML, XML, JPEG, MPEG, or PNG; user interface instructions for rendering or interpreting commands for a graphical user interface (GUI), command line interface, or text user interface; application software such as, for example, an office suite, an Internet access application, a design and manufacturing application, a graphics application, an audio application, a software engineering application, an educational application, a game, or other applications. The instructions may implement a web server, a web application server, or a web client. The instructions may be organized as a presentation layer, an application layer, and a data storage layer such as, for example, a relational database system using Structured Query Language (SQL) or NoSQL, an object store, a graph database, a flat file system, or other data storage.
[0089] Computer system 900 may be coupled to at least one output device 912 via I / O subsystem 902. In one embodiment, output device 912 is a digital computer display. Examples of displays that may be used in various embodiments include touch screen displays or light emitting diode (LED) displays or liquid crystal displays (LCD) or electronic paper displays. Computer system 900 may include other type(s) of output device 912 instead of, or in addition to, a display device. Examples of other output devices 912 include printers, ticket printers, plotters, projectors, sound cards or video cards, speakers, buzzers or piezoelectric devices or other audible devices, lamps or LEDs or LCD indicators, tactile devices, actuators or servos.
[0090] At least one input device 914 is coupled to I / O subsystem 902 to communicate signals, data, command selections, or gestures to processor 904. Examples of input device 914 include touch screens, microphones, still and video digital cameras, alphanumeric and other keys, keypads, keyboards, graphics tablets, image scanners, joysticks, clocks, switches, buttons, dials, slides, and / or various types of sensors such as force sensors, motion sensors, thermal sensors, accelerometers, gyroscopes, and inertial measurement unit (IMU) sensors, and / or various types of transceivers such as cellular or Wi-Fi, wireless, radio frequency (RF) or infrared (IR) transceivers, and global positioning system (GPS) transceivers.
[0091] Another type of input device is the control device 916, which can perform cursor control or other automated control functions such as navigation in a graphical interface on a display screen, instead of or in addition to the input function. The control device 916 can be a touch pad, mouse, trackball, or cursor direction keys for communicating direction information and command selections to the processor 904 and for controlling cursor movement on the display 912. The input device can have at least two degrees of freedom in two axes, a first axis (e.g., x) and a second axis (e.g., y), that enable the device to define a position within a plane. Other types of input devices are wired, wireless, or optical control devices such as, for example, joysticks, pen scanners, consoles, steering wheels, pedals, gear shift mechanisms, or other types of control devices. The input device 914 may include a combination of multiple different input devices, such as, for example, a video camera and a depth sensor.
[0092] In another embodiment, the computer system 900 may have an Internet of Things (IoT) device in which one or more of the output device 912, the input device 914, and the control device 916 are omitted. Alternatively, in such an embodiment, the input device 914 may have one or more cameras, motion detectors, thermometers, microphones, seismic detectors, other sensors or detectors, measurement devices, or encoders, and the output device 912 may have a dedicated display such as, for example, a single-line LED or LCD display, one or more indicators, a display panel, a meter, a valve, a solenoid, an actuator, or a servo.
[0093] If computer system 900 is a mobile computing device, input device 914 may include a GPS receiver coupled to a GPS module that triangulates with a plurality of Global Positioning System (GPS) satellites to determine and generate geolocation or position data, such as latitude-longitude values, for the geophysical location of computer system 900. Output device 912 may include hardware, software, firmware, and interfaces for generating, either alone or in combination with other application-specific data, position reporting packets, notifications, pulses or heartbeat signals, or other repeating data transmissions that define the location of computer system 900, and directing them towards host 924 or server 930.
[0094] Computer system 900 may implement the techniques described herein using customized hardwired logic, at least one ASIC or FPGA, firmware, and / or program instructions or logic that, when loaded and used or executed in combination with the computer system, cause the computer system to operate as or program the computer system to operate as a dedicated machine. According to one embodiment, the techniques herein are performed by computer system 900 in response to execution by a processor 904 of at least one sequence of at least one instruction included in main memory 906. Such instructions may be read into main memory 906 from another storage medium, such as storage 910. Execution of the sequence of instructions included in main memory 906 causes processor 904 to perform the process steps described herein. In an alternative embodiment, hardwired circuitry may be used in place of, or in combination with, software instructions.
[0095] As used herein, the term "memory medium" refers to any non-transitory medium that stores instructions and / or data that cause a machine to operate in a particular manner. Such memory media can include non-volatile media and / or volatile media. Non-volatile media includes, for example, optical disks or magnetic disks such as storage 910. Volatile media includes, for example, dynamic memory such as memory 906. Common forms of memory media include, for example, hard disks, solid state drives, flash drives, magnetic data storage media, any optical or physical data storage media, memory chips, or the like.
[0096] Memory media is different from transmission media, but may be used in conjunction with transmission media. Transmission media is involved in transferring information between memory media. For example, transmission media includes coaxial cables, copper wire, and optical fiber, including wires having a bus of the I / O subsystem 902. Transmission media can also take the form of acoustic or light waves, such as those generated in radio and infrared data communications.
[0097] Various forms of media may be involved in conveying at least one sequence of at least one instruction to the processor 904 for execution. For example, the instructions may first be carried on a magnetic disk or solid state drive of a remote computer. The remote computer may load the instructions into its dynamic memory and transmit the instructions over a communication link, such as an optical fiber or coaxial cable or telephone line, using a modem. A modem or router local to the computer system 900 may receive the data over the communication link and convert the data so that it can be read by the computer system 900. For example, a receiver, such as a radio frequency antenna or an infrared detector, may receive data carried in a radio signal or optical signal, and appropriate circuitry may provide the data to the I / O subsystem 902, such as by placing the data on a bus. The I / O subsystem 902 conveys the data to the memory 906, from which the processor 904 fetches and executes the instructions. The instructions received by the memory 906 may optionally be stored on the storage 910 either before or after execution by the processor 904.
[0098] Computer system 900 also includes a communication interface 918 coupled to bus 902. Communication interface 918 provides bi-directional data communication coupled to (one or more) network links 920 directly or indirectly connected to at least one communication network such as, for example, network 922 or a public or private cloud on the Internet. For example, communication interface 918 can be an Ethernet (registered trademark) networking interface, an integrated services digital network (ISDN) card, a cable modem, a satellite modem, or a modem that provides a data communication connection to a corresponding type of communication line such as, for example, an Ethernet (registered trademark) cable or any type of metal cable or fiber optic line or telephone line. Network 922 broadly represents a local area network (LAN), a wide area network (WAN), a campus network, an Internetwork, or any combination thereof. Communication interface 918 can have a LAN card that provides a data communication connection to a compatible LAN, or a cellular radio telephone interface wired to transmit or receive cellular data according to a cellular radio telephone wireless networking standard, or a satellite wireless interface wired to transmit or receive digital data according to a satellite wireless networking standard. In any such implementation, communication interface 918 transmits and receives electrical, electromagnetic, or optical signals on a signal path that carries a digital data stream representing various types of information.
[0099] Network link 920 typically provides electrical, electromagnetic, or optical data communication to other data devices directly or via at least one network using, for example, satellite, cellular, Wi-Fi, or BLUETOOTH (registered trademark) technology. For example, network link 920 can provide a connection to host computer 924 via network 922.
[0100] Furthermore, network link 920 can provide a connection via network 922 or a connection to another computing device through an internetworking device and / or computer operated by an Internet service provider (ISP) 926. ISP 926 provides a data communication service through a worldwide packet data communication network represented as the Internet 928. Server computer 930 can be coupled to the Internet 928. Server 930 broadly represents any computer, data center, virtual machine or virtual computing instance with or without a hypervisor, or a computer that runs a containerized program system such as DOCKER or KUBERNETES. Server 930 may be implemented using two or more computers or instances and represents an electronic digital service that is accessed and used by sending a web service request, a uniform resource locator (URL) string having parameters in an HTTP payload, an API call, an app service call, or other service calls. Computer system 900 and server 930 may form elements of a distributed computing system that includes other computers, processing clusters, server farms, or other orchestrated computers that cooperate to execute tasks or run applications or services. Server 930 may have one or more sets of instructions organized as modules, methods, objects, functions, routines, or calls. The instructions may be organized as application programs including one or more computer programs, operating system services, or mobile apps.The command may include operating system and / or system software; one or more libraries that support multimedia, programming, or other functions; data protocol instructions or stacks that implement TCP / IP, HTTP, or other communication protocols; file format processing instructions that interpret or render files coded using HTML, XML, JPEG, MPEG, or PNG; user interface instructions that render or interpret commands for a graphical user interface (GUI), command line interface, or text user interface; and application software such as an office suite, Internet access application, design and manufacturing application, graphics application, audio application, software engineering application, educational application, game, or other application. Server 930 may have a web application server that hosts a presentation layer, an application layer, and a data storage layer such as a relational database system using, for example, Structured Query Language (SQL) or NoSQL, an object store, a graph database, a flat file system, or other data storage.
[0101] Computer system 900 can send messages and receive instructions and data including program code via (one or more) networks, network link 920, and communication interface 918. In the example of the Internet, server 930 can send the requested code of an application program via Internet 928, ISP 926, local network 922, and communication interface 918. The received code may be executed by processor 904 when received and / or stored in storage 910 or other non-volatile storage for later execution.
[0102] The execution of the instructions described in this section may implement the process in the form of an instance of a running computer program consisting of program code and its current activity. Depending on the operating system (OS), the process may be composed of multiple execution threads that execute instructions simultaneously. In this context, while a computer program is a passive set of instructions, a process may be the actual execution of those instructions. Several processes may be associated with the same program; for example, opening several instances of the same program often means that two or more processes are being executed. Multitasking may be implemented to allow multiple processes to share the processor 904. Each processor 904 or each core of a processor executes one task at a time, but by programming the computer system 900 to implement multitasking, each processor may be able to switch between multiple tasks being executed without having to wait for each task to complete. In one embodiment, the switch may be executed when a task is performing an input / output operation, when the task indicates that it is switchable itself, or at a hardware interrupt. Timesharing may be implemented to enable fast response for interactive user applications by executing context switches quickly to provide the appearance of simultaneous parallel execution of multiple processes. In one embodiment, for security and reliability, the operating system may prevent direct communication between independent processes and provide a strictly mediated and controlled inter-process communication function.
[0103] 7. Extensions and Alternatives In the foregoing specification, embodiments of the present disclosure have been described with reference to numerous specific details that may vary from implementation to implementation. Accordingly, the specification and drawings are to be regarded in an illustrative rather than a limiting sense. The only thing that exclusively indicates the scope of the present disclosure and what is intended by the applicant to be the scope of the present disclosure is the literal scope and the equivalent scope of the set of claims issued from this application, including any subsequent corrections in the specific forms in which such claims arise.
Claims
1. A computer-implemented method for suppressing noise and enhancing voice, comprising: receiving, by a processor, input audio data covering a plurality of frequency bands along a frequency dimension in a plurality of frames along a time dimension; training, by the processor, a neural network model using the input audio data, the neural network model comprising: a feature extraction block that performs a look-ahead of a specific number of frames when extracting features from the input audio data; an encoder comprising a first series of blocks that generate a first feature map corresponding to successively larger receptive fields within the input audio data along the frequency dimension, each block of the first series of blocks having a feature calculation block and a frequency downsampler, the feature calculation block having a series of convolutional layers, output data of one of the series of convolutional layers being supplied to all subsequent convolutional layers of the series of convolutional layers, the series of convolutional layers implementing successively larger dilations along the time dimension; an encoder; a decoder that receives the output feature map generated by the encoder as an input feature map and includes a second series of blocks that generate a second feature map; a classification block that receives the second feature map and generates a voice value indicating the amount of voice present for each frequency band of the plurality of frequency bands in each of the plurality of frames; steps; receiving new audio data having one or more frames; executing the neural network model on the new audio data to generate a new voice value for each frequency band of the plurality of frequency bands in each of the one or more frames; generating new output data for suppressing noise in the new audio data based on the new voice values; transmitting the new output data; A computer-implemented method having the above steps.
2. receiving an input waveform; converting the input waveform into raw audio data covering a plurality of frequency bins along the frequency dimension in the one or more frames along the time dimension; Converting the raw audio data into the new audio data by grouping the plurality of frequency bins into the plurality of frequency bands; Performing inverse banding on the new audio values to generate updated audio values for each of the plurality of frequency bins in each of the one or more frames; Applying the updated audio values to the raw audio data to generate the new output data; Converting the new output data into an enhanced waveform; The computer-implemented method according to claim 1, further comprising.
3. The computer-implemented method according to claim 1 or 2, wherein the plurality of frequency bands are perceptually stimulating bands and cover more frequency bins at higher frequencies.
4. The feature extraction block has a convolutional kernel with a specific size along the time dimension, The computer-implemented method according to any one of claims 1 to 3, wherein the specific size is larger than the size along the time dimension of any convolutional kernel in the encoder or the decoder.
5. The computer-implemented method according to any one of claims 1 to 4, wherein the feature extraction block has a batch normalization layer and a convolutional layer having a subsequent two-dimensional convolutional kernel.
6. The computer-implemented method according to any one of claims 1 to 5, wherein each of the feature extraction block, the first series of blocks, and the second series of blocks generates a common number of feature maps.
7. The computer-implemented method according to any one of claims 1 to 6, wherein each of the series of convolutional layers has a depthwise separable convolutional block having a gating mechanism.
8. The computer-implemented method according to any one of claims 1 to 7, wherein each of the series of convolutional layers has a residual block having a series of convolutional blocks including a first convolutional block having an initial 1×1 two-dimensional convolutional kernel and a last convolutional block having a final 1×1 two-dimensional convolutional kernel.
9. Output data of a feature calculation block within a block of the first series of blocks is scaled by learnable weights to form scaled output data. The scaled output data is communicated via a skip connection to a block of the second series of blocks within the decoder. The computer-implemented method according to any one of claims 1 to 8.
10. The computer-implemented method according to any one of claims 1 to 9, wherein a frequency downsampler of a block within the first series of blocks has a convolutional kernel with a stride size greater than 1 along the frequency dimension.
11. The computer-implemented method according to any one of claims 1 to 10, wherein each block of the second series of blocks has a feature calculation block and a frequency upsampler.
12. A feature calculation block within a block of the second series of blocks receives first output data from a feature calculation block within a block of the first series of blocks and second output data from a frequency upsampler of a preceding block within the second series of blocks. The first output data and the second output data are concatenated or added together to form specific input data for the feature calculation block within the block of the second series of blocks. The computer-implemented method according to claim 11.
13. The computer-implemented method according to any one of claims 1 to 12, wherein the classification block has a 1×1 two-dimensional convolutional kernel and a non-linear activation function.
14. The training is performed using a function of the loss between the predicted audio value and the ground truth audio value for each frequency band of the plurality of frequency bands in each frame, and when the predicted audio value corresponds to over-suppression of the audio, the weight in the loss function is increased, and when the predicted audio value corresponds to under-suppression of the audio, the weight in the loss function is decreased. The computer-implemented method according to any one of claims 1 to 13.
15. The classification block further generates a distribution of the voice volume over a certain frequency band among the plurality of frequency bands in the frame, and the voice value is the average of the distribution. The computer-implemented method according to any one of claims 1 to 14.
16. The input audio data has data corresponding to voices of different speeds or emotions, data including different levels of noise, or data corresponding to different frequency bins. The computer-implemented method according to any one of claims 1 to 15.
17. The neural network model further has a feature calculation block that is the output data of the encoder and the input data of the decoder. The computer-implemented method according to any one of claims 1 to 16.
18. A memory, and One or more processors coupled to the memory, Receiving input audio data covering a plurality of frequency bands along the frequency dimension in a plurality of frames along the time dimension, Using the input audio data to train a neural network model, the neural network model includes A feature extraction block that performs a look-ahead of a specific number of frames when extracting features from the input audio data, An encoder including a first series of blocks that generate a first feature map corresponding to receptive fields that gradually increase in the input audio data along the frequency dimension, Each block of the first series of blocks has a feature calculation block and a frequency downsampler. The feature calculation block has a series of convolutional layers. The output data of a certain convolutional layer in the series of convolutional layers is supplied to all subsequent convolutional layers in the series of convolutional layers. The series of convolutional layers implements a gradually increasing dilation along the time dimension. Encoder, A decoder including a second series of blocks that receive the output feature map generated by the encoder as an input feature map and generate a second feature map, Receiving the second feature map, and a classification block that generates a voice value indicating the amount of existing voice for each frequency band among the plurality of frequency bands in each of the plurality of frames. Storing the neural network model, One or more processors configured to perform the above. A computer system having
19. A computer-implemented method for suppressing noise and enhancing speech, comprising: Receiving, by a processor, new audio data having one or more frames; Executing, by the processor, a trained neural network model on the new audio data, the trained neural network model being trained to receive the new audio data and generate new speech values for each frequency band of a plurality of frequency bands in each of the one or more frames, the trained neural network model comprising: A feature extraction block that performs a look-ahead of a specific number of frames when extracting features from the new audio data; An encoder including a first series of blocks that generate a first feature map corresponding to successively larger receptive fields within the new audio data along the frequency dimension; Each block of the first series of blocks having a feature calculation block and a frequency downsampler, the feature calculation block having a series of convolutional layers, output data of one convolutional layer of the series of convolutional layers being supplied to all subsequent convolutional layers of the series of convolutional layers, the series of convolutional layers implementing successively larger dilations along the time dimension, the encoder; A calculation block connecting the encoder and the decoder; The decoder including a second series of blocks that receive the output feature map generated by the encoder as an input feature map and generate a second feature map; and A classification block that receives the second feature map and generates speech values indicating the amount of speech present for each frequency band of the plurality of frequency bands in each of the plurality of frames; Steps having computer-executable instructions for; Generating new output data for suppressing noise in the new audio data based on the new speech values; Transmitting the new output data; A computer-implemented method having.
Citation Information
Patent Citations
Perceptually-based loss functions for audio encoding and decoding based on machine learning
WO2019199995A1