Audio artifact mitigation based on deep learning
By using mask blocks in speech detection and enhanced machine learning models to extract speech and mitigate artifacts, the problems of oversuppression and multiple artifact processing are solved, achieving better audio quality and perceptual effects.
Patent Information
- Application Number
- CN202380070590.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-20
- Filing Date
- 2023-07-28
- Publication Date
- 2025-05-13
AI Technical Summary
Existing machine learning methods for speech detection and enhancement have problems with oversuppression of speech, resulting in distortion or discontinuity of speech, and it is difficult to effectively alleviate multiple types of artifacts.
The combined time-frequency representation of the audio data received by the processor is employed and the speech is detected and artifacts are mitigated by a digital model, including a series of mask blocks, each mask block containing a first mask for extracting speech and a second mask for extracting residual speech masked by the first mask.
By mitigating various types of artifacts and sharpening speech without oversuppressing speech, audio quality is improved, resulting in better audio perception and user audio enjoyment.
Smart Images

Figure CN119998877A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to PCT Application No. PCT / CN2022 / 110612 filed on August 5, 2022, U.S. Provisional Application No. 63 / 424,620 filed on November 11, 2022, and European Patent Application No. 22214817.3 filed on December 20, 2022, each of which is incorporated by reference in its entirety. Technical Field
[0003] This application is related to audio processing and machine learning. Background Art
[0004] The approaches described in this section are approaches that could be pursued, but are not necessarily approaches that have been previously conceived or pursued. Therefore, unless otherwise indicated, it should not be assumed that any of the approaches described in this section qualify as prior art merely by virtue of their inclusion in this section.
[0005] In recent years, various machine learning models have been used for speech enhancement. Compared with traditional signal processing methods such as Wiener filters or spectral subtraction, machine learning methods have shown significant improvements, especially under conditions of non-stationary noise and low signal-to-noise ratio (SNR).
[0006] Existing machine learning methods for speech detection and enhancement often suffer from the problem of over-suppression of speech, which can cause speech distortion or even discontinuity. In addition, existing machine learning methods for speech detection and enhancement are usually developed to mitigate one type of artifact (e.g., noise, reverberation echo, codec effect, packet loss, or howling effect) respectively.
[0007] Improving traditional machine learning methods for speech enhancement would be helpful, especially being able to effectively reduce speech oversuppression and mitigate multiple types of artifacts in stored audio content or real-time communications. Summary of the invention
[0008] A computer-implemented method for mitigating audio artifacts is disclosed. The method includes receiving, by a processor, audio data as a joint time-frequency representation over a plurality of frames and a plurality of frequency bands. The method also includes executing, by the processor, a digital model for detecting speech from a feature vector of the audio data, the digital model including a series of mask blocks, each mask block including a first component and a second component, the first component generating a first mask for extracting speech, the second component generating a second mask for extracting residual speech masked by the first mask, and each of the first mask and the second mask including a mask value for estimating the amount of speech present in each of the plurality of frames and each of the plurality of frequency bands. In addition, the method includes transmitting information associated with the first mask generated by the series of mask blocks to a device.
[0009] The techniques described in this specification have advantages over conventional audio processing techniques. The method improves audio quality by mitigating various types of artifacts and sharpening speech without over-suppressing speech. The method utilizes a deep learning model that is configured to recognize clean speech with low latency and reduce speech over-suppression with low complexity. The improved audio quality results in better audio perception and better user audio enjoyment. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Example embodiments of the present invention are illustrated by way of example and not by way of limitation in the accompanying drawings in which like reference numerals refer to similar elements and in which:
[0011] Figure 1 An example networked computer system is shown upon which various embodiments may be implemented.
[0012] Figure 2 Example components of an audio management computer system according to the disclosed embodiments are shown.
[0013] Figure 3 A CGRU block including a convolutional neural network (CNN) block and a gated recurrent unit (GRU) block is shown.
[0014] Figure 4 A deep neural network (DNN) including an input CNN block, a CGRU block, and a mask combination block is shown.
[0015] Figure 5 An example process performed by an audio management computer system according to some embodiments of the present disclosure is shown.
[0016] Figure 6 is a block diagram illustrating a computer system upon which embodiments of the present invention may be implemented. DETAILED DESCRIPTION
[0017] In the following description, for the purpose of explanation, many specific details are listed to provide a deeper understanding of the exemplary embodiments of the present invention. However, it is apparent that the exemplary embodiments can be implemented without these specific details. In other examples, known structures and devices are shown in block diagram form to avoid unnecessarily obscuring the exemplary embodiments.
[0018] The following sections will describe the embodiments according to the following outline:
[0019] 1. Overview
[0020] 2. Sample computing environment
[0021] 3. Sample Computer Parts
[0022] 4. Functional Description
[0023] 4.1. Speech Enhancement Model Training
[0024] Feature extraction
[0025] Machine Learning Model
[0026] 4.1.3. Perceptual Loss Function
[0027] 4.2. Model Implementation for Speech Enhancement
[0028] 5. Example Process
[0029] 6. Hardware Implementation
[0030] 7. Extensions and Alternatives
[0031] 1. Overview
[0032] A system for mitigating audio artifacts is disclosed. In some embodiments, the system is programmed to establish a machine learning model that includes a series of mask blocks. Each mask block receives a feature vector of an audio segment. Each mask block includes a first component and a second component, the first component generating a first mask for extracting clean speech, and the second component generating a second mask for extracting residual speech masked by the first mask. Each mask block also generates a specific feature vector based on the first mask and the second mask, and the specific feature vector becomes a feature vector for the next mask block. The second component may include a GRU layer, and the computational complexity of the second component is lower than that of the first component that may include multiple CNN layers. In addition, the system is programmed to receive an input feature vector of an input audio segment and execute the machine learning model to obtain an output feature vector of an output audio segment, which output audio segment contains cleaner speech than the input audio segment.
[0033] In some embodiments, the first component includes a first series of CNN blocks. The first series of CNN blocks include CNN blocks with an increasing and then decreasing expansion rate in the autoencoder structure, and a trailing CNN block for classification. Each CNN block can include a CNN layer with one or more filters, and a subsequent batch normalization (BatchNorm) layer and an activation layer. The second component includes a GRU block, which may include a GRU layer, similarly, followed by a batch normalization layer and an activation layer. The mask component may also include a third component, which includes a second series of CNN blocks for combining the first mask and the second mask.
[0034] In some embodiments, the machine learning model includes an input CNN block including a CNN layer with one or more look-ahead filters for enhancing the input feature vector. The machine learning model can also include a mask combination block similar to the third component of the mask block. The mask combination block includes a third series of CNN blocks to combine a first mask generated by a first series of CNN blocks in a series of mask blocks.
[0035] The system produces a technical effect. The system solves the technical problem of improving audio data to enhance speech. The system improves audio quality by mitigating various types of artifacts and sharpening speech without over-suppressing speech. The system utilizes a deep learning model to recognize clean speech with lower latency and reduce speech over-suppression with lower complexity. The improved audio quality results in better audio perception and better audio enjoyment for the user.
[0036] 2. Sample computing environment
[0037] Figure 1 An example networked computer system is shown upon which various embodiments may be implemented. Figure 1 Examples are shown in simplified schematic form to illustrate clarity, and other embodiments may include more, fewer, or different elements.
[0038] In some embodiments, the networked computer system includes an audio management server computer 102 (“server”), one or more sensors 104 or input devices, and one or more output devices 110 , which are communicatively coupled via a direct physical connection or via one or more networks 118 .
[0039] In some embodiments, server 102 broadly represents one or more computers, virtual computing instances, and / or instances of applications that are programmed or configured with data structures and / or database records that are arranged to host or perform functions associated with audio enhancement. Server 102 may include a server farm, a cloud computing platform, a parallel computer, or any other computing facility with sufficient computing power in terms of data processing, data storage, and network communications to implement the above functions.
[0040] In some embodiments, each of the one or more sensors 104 may include a microphone or other digital recording device that converts sound into electrical signals. Each sensor is configured to transmit detected audio data to the server 102. Each sensor may include a processor, or may be integrated into a typical client device (e.g., a desktop computer, a laptop computer, a tablet computer, a smartphone, or a wearable device).
[0041] In some embodiments, each of the one or more output devices 110 may include a speaker or other digital playback device that converts electrical signals back into sound. Each output device is programmed to play audio data received from the server 102. Similar to the sensor, the output device may include a processor or may be integrated into a typical client device (e.g., a desktop computer, laptop, tablet computer, smartphone, or wearable device).
[0042] One or more networks 118 may be connected via Figure 1 The network 118 may be implemented by any medium or mechanism that provides data exchange between the various elements in the network. Examples of network 118 include, but are not limited to, one or more of a cellular network (coupled with data communication of the computing device via a cellular antenna), a near field communication (NFC) network, a local area network (LAN), a wide area network (WAN), the Internet, a terrestrial or satellite link, etc.
[0043] In certain embodiments, server 102 is programmed to receive input audio data corresponding to the sound in a given environment from one or more sensors 104. The input audio data may include multiple time-varying frames. Server 102 is programmed to process input audio data (which generally corresponds to a mixture of voice and noise or other artifacts) next, to estimate how many voices (or detect voice volume) exist in each frame of the input audio data. Server may be programmed to send final detection result to another device for downstream processing. Server may also be programmed to update input audio data based on the final detection result, to produce output audio data (the voice included therein is expected to be cleaner than the input audio data) after cleaning, and send the output audio data to one or more output devices 110.
[0044] 3. Sample Computer Parts
[0045] Figure 2 Example components of an audio management computer system according to the disclosed embodiments are shown. This figure is for illustrative purposes only, and the server 102 may include fewer or more functional components or storage components. Each functional component can be implemented as a software component, a general or special purpose hardware component, a firmware component, or any combination thereof. Each functional component can also be coupled to one or more storage components. The storage component can be implemented using any of a relational database, an object database, a flat file system, or a JavaScript Object Notation (JSON) storage. The storage component can be locally connected to the functional component or connected to the functional component through a network using a program call, a remote procedure call (RPC) facility, or a message bus. The components may be independent or non-independent. Depending on the specific implementation or other considerations, the components may be centrally deployed or may be functionally or physically distributed.
[0046] In some embodiments, the server 102 includes machine learning model training instructions 202, machine learning model execution instructions 206, and communication interface instructions 210. The server 102 also includes a database 220.
[0047] In some embodiments, the machine learning model training instructions 202 enable training of a machine learning model for detecting speech and mitigating artifacts. The machine learning model may include various artificial neural networks (ANNs) or other conversion or classification models. Training may include extracting features from training audio data, feeding given or extracted features (optionally together with expected model outputs) to a training framework to train the machine learning model, and storing the trained machine learning model. The output of the machine learning model may include an estimate of the amount of speech present in each given audio clip or an enhanced version of each given audio clip. The training framework may include an objective function intended to mitigate excessive suppression of speech.
[0048] In some embodiments, the machine learning model execution instructions 206 enable execution of the machine learning model to detect speech and mitigate artifacts. The execution may include extracting features from the new audio segment, feeding the extracted features to the trained machine learning model, and obtaining a new output by executing the trained machine learning model. The new output may include an estimate of the amount of speech in the new audio segment or an enhanced version of the new audio segment.
[0049] In some embodiments, the communication interface instructions 210 enable communication with other systems or devices over a computer network. Communication may include receiving audio data or a trained machine learning model from an audio source or other system. Communication may also include transmitting speech detection or enhancement results to other processing devices or output devices.
[0050] In some embodiments, database 220 is programmed or configured to manage storage and access to relevant data (eg, received audio data, digital models, features extracted from received audio data, or results of executing digital models).
[0051] 4. Functional Description
[0052] 4.1. Speech Enhancement Model Training
[0053] 4.1.1 Data Collection
[0054] Speech signals are often distorted by various contaminations or artifacts (such as noise or reverberation) caused by the environment or recording equipment. In some embodiments, the server 102 is programmed to establish a training data set of audio clips distorted to various degrees. The audio clips in the training data set may include additive artifacts that affect different durations or frequency bands. Ravanelli et al.'s paper entitled "Multi-task self-supervised learning for Robust Speech Recognition" on an improved version of the problem-agnostic speech encoder (PASE+) introduces an example method for mixing such artifacts with clean speech signals.
[0055] 4.1.2 Feature extraction
[0056] The training data set of audio segments is usually represented in the time domain. In some embodiments, the server 102 is programmed to convert each audio segment including waveforms over multiple frames into a joint time-frequency (TF) representation using a spectral transform (e.g., short-term Fourier transform (STFT), shift-modified discrete Fourier transform (MDFT), or complex quadratic image filter (CQMF). The joint TF representation covers multiple frames and multiple frequency bins.
[0057] In some embodiments, the server 102 is programmed to convert the TF representation into a frequency band energy vector, for example, for 56 perceptually excited frequency bands. Each perceptually excited frequency band is typically located in a frequency domain that matches the way the human ear processes speech, such as 120 Hz to 2,000 Hz, so capturing data in these perceptually excited frequency bands means that there is no loss of speech quality for the human ear. More specifically, the squared amplitudes of the output frequency bins of the spectral transform are grouped into perceptually excited frequency bands, where the number of frequency bins per band increases as the frequency increases. The grouping strategy can be "soft", that is, some spectral energy is leaked between adjacent frequency bands; it can also be "hard", that is, no energy is leaked between frequency bands. Specifically, when the bin energy of a noise frame is represented by x, x is a column vector of size p×1, where p represents the number of bins, it can be converted into a band energy vector by calculating y=W*x, where y is a column vector of size q×1 representing the band energy of the noise frame, W is a band matrix of size q×p, and q represents the number of perceptually excited frequency bands.
[0058] In some embodiments, the server 102 is then programmed to calculate the logarithm of each frequency band energy as an eigenvalue for each frame and each frequency band. Alternatively, the frequency band energy can be used directly as an eigenvalue. Thus, for each joint TF representation, an input feature vector including eigenvalues can be obtained for multiple frames and multiple frequency bands.
[0059] In some embodiments, for supervised learning, server 102 is programmed to retrieve or compute an expected mask for each joint TF representation that indicates the amount of speech present in each frame and each frequency band. The mask may be in the form of the logarithm of the ratio of the speech energy to the sum of all energies. Server 102 may include the expected mask in the training data set.
[0060] 4.1.2 Machine Learning Model
[0061] In some embodiments, server 102 is then programmed to train the machine learning model using the training data set. Figure 3 A CGRU block including a CNN block and a GRU block is shown. Figure 3 and Figure 4All aspects of (including the number of blocks, the type of blocks, or parameter values) are for illustrative purposes only. The CGRU block 300 receives an input feature vector 316 of an input audio segment and generates an output mask 312 and an output feature vector 320 corresponding to a cleaner speech. The input feature vector 316 can be in the form of (N, C, F, T), where N represents the batch size (e.g., the number of audio files), C represents the number of channels, F represents the value of the frequency dimension, and T represents the value of the time dimension. The CGRU block 300 includes a first series of CNN blocks 302 and GRU blocks 308, which generate corresponding masks 304 and 314 for separating clean speech from additive artifacts in the input audio segment while reducing excessive speech suppression. Each mask 304 and 314 can also be in the form of (N, C, F, T). The CGRU block 300 also includes a second series of CNN blocks 310 that can combine masks 304 and 314 into an output mask 312 that can be applied to an input feature vector 316 to further enhance speech, as will be discussed further below. The CGRU block 300 can serve as a component of a deep neural network (DNN), as will be discussed further below.
[0062] In some embodiments, the first series of CNN blocks 302 is intended to detect clean speech. The first series of CNN blocks 302 includes a first sub-series of dilated CNN blocks with increasing dilation rates (e.g., 1, 3, 9, 27 along the time dimension), followed by a second sub-series of dilated CNN blocks with corresponding decreasing dilation rates (e.g., 27, 9, 3, 1 along the time dimension), followed by a trailing CNN block. As shown, the first series of CNN blocks 302 receive an input feature vector 316. Each CNN block in the first series receives a feature vector of an audio segment and generates a specific feature vector of an enhanced audio segment. The output of each CNN block in the first series becomes the input of the next CNN block in the first series. The output of each CNN block in the first sub-series is combined with the input of a CNN block with the same dilation rate in the second sub-series. The combination can be performed by addition or concatenation. The dilated CNN blocks each use some relatively small filters, such as 16 3x3 filters, while the trailing CNN block uses a 1x1 filter. In other embodiments, the length of the first sub-series or the second sub-series, the dilation rate, the number of filters, and the size of each filter may vary.
[0063] Therefore, the first sub-series of CNN blocks with ever-expanding receptive fields encode feature data representing clean speech in the original audio data (to find more and better features), while the second sub-series of CNN blocks perform reconstruction on the enhanced audio data. The trailing CNN blocks linearly project the feature maps into summary feature maps that can indicate how much speech is present in the original audio data. Therefore, the first series of CNN blocks project different levels of discriminative features into a high-resolution space (i.e., at a per-band level per frame), thereby obtaining a dense classification (how much speech is present in each time frame and each frequency band), i.e., obtaining a first mask 304.
[0064] In some embodiments, the GRU block 308 is intended to detect any speech that may be overly suppressed by the first series of CNN blocks 302. The input feature vector 316 can be combined with the first mask 304 generated by the first series of CNN blocks (in Figure 3 The combination is not shown as a separate block in FIG. 3 ) to generate a residual feature vector 306, which corresponds to the portion of the input feature vector 316 that is not recognized as speech by the first mask 304. For example, the inverse I(m) of each value m of the first mask 304 (such as log(1-e m )) can be applied to the corresponding eigenvalue f of the input feature vector, such as I(m)+f, to generate the inverse of the result of applying the first mask 304 to the input feature vector 316 as the value of the residual feature vector 306.
[0065] In some embodiments, the GRU block 308 receives the residual feature vector 306 and generates a residual mask 314. The GRU block 308 may include one or more GRU layers, followed by a batch normalization layer, followed by an activation layer, which are well known to those of ordinary skill in the art. The batch normalization layer can accept a one-dimensional input (BatchNorm1d). Each GRU layer may include a certain number of nodes, the number of which is equal to the number of features in each feature vector. The GRU layer provides a gating mechanism for a recurrent neural network (RNN) that is often used to process time series data. The structure of the gating mechanism is relatively simple, and non-speech can be relatively loosely filtered, so it is particularly suitable for detecting a relatively small amount of residual speech. This relatively simple structure also helps to achieve low complexity of the machine learning model. In other embodiments, in order to further reduce complexity, the GRU layer can be replaced by a simple fully connected layer without using any recurrent connection. It can even be further simplified by using a one-dimensional filter along the time dimension instead of a two-dimensional filter in the GRU layer. The batch normalization layer usually helps to fine-tune the output of the previous layer and avoid internal covariate shifts. Activation layers often help to constrain output values to a certain range and add nonlinearity to the ANN. In some embodiments, a batch normalization layer can be placed before the activation layer.
[0066] In some embodiments, the second series of CNN blocks 310 is intended to generate an output mask 312 so that the speech is cleaner than the input audio clip. The second series of CNN blocks 310 can include two CNN blocks, the first CNN block having some relatively small filters, such as 16 3x3 filters, and the second CNN block is similar to the trailing CNN block in the first series of CNN blocks. The second series of CNN blocks 310 receives the first mask 304 and the residual mask 314, combines (such as splicing) the two masks and generates an output mask 312. In other embodiments, the length of the second series, the number of filters, and the size of each filter can vary.
[0067] Thus, the first CNN block of the second series of CNN blocks 310 identifies a particular pattern from the masks 304 and 314, and the second CNN block of the second series of CNN blocks 310 again performs classification to determine an effective way to combine the first mask 304 and the residual mask 314. The output mask 312 is then applied to the input feature vector 316 via the feature vector generator 322 to produce the output feature vector 320, which is similar to applying the inverse of the first mask 304 to the input feature vector 316 to produce the residual feature vector 306, as described above.
[0068] In some embodiments, each CNN block in the CGRU block 300 includes a CNN layer, a batch normalization layer, and an activation layer. The CNN layer includes a two-dimensional filter for causal convolution. The batch normalization layer accepts a two-dimensional input (BatchNorm2d). The activation layer can be similar to the activation layer in the GRU block.
[0069] Figure 4 4 shows a deep neural network including an input CNN block, a CGRU block, and a mask combination block. DNN 400 includes an input CNN block 402, a series of CGRU blocks 404 (each with Figure 3 300 shown in ) and a mask combination block 406. The DNN 400 receives the original feature vector 412 of the original audio segment and predicts a final mask 418 for extracting clean speech from the original audio segment.
[0070] In some embodiments, the input CNN block 402 is intended to enrich the original feature vector 412. The input CNN block 402 includes a CNN layer, a batch normalization layer, and an activation layer. In the CNN layer, each filter is applied to a small number of future frames (e.g., additional look-aheads of two frames) compared to the CNN layer in each CNN block of the CGRU block. Such a filter may be referred to as a look-ahead filter. Therefore, the input CNN block 402 receives the original feature vector 412 and generates a feature map as an improved feature vector 414. Relatively few additional look-aheads contribute to the low latency of the DNN 400 when generating rich features. In other embodiments, the input CNN block 402 is not present, and the original feature vector 412 is directly received by a series of CGRU blocks 404.
[0071] In some embodiments, the first CGRU block in the series of CGRU blocks 404 receives the improved feature vector 414 and predicts Figure 3 The mask 408 corresponding to 308 in Figure 3 The feature vector 410 corresponding to 320 in . The feature vector 410 then becomes the input of the next CGRU block, and the process continues through a series of CGRU blocks 404. As described above, the first series of CNN blocks of each CGRU block is intended to eliminate as much non-speech as possible (although some speech may also be eliminated). The GRU block of each CGRU block is intended to restore the eliminated speech (although some non-speech may also be restored). The iterative composition of the CGRU blocks in a series of CGRU blocks 404 is intended to ultimately retain as much clean speech as possible while eliminating as much non-speech as possible (including various artifacts such as echo, noise, or reverberation). Experiments have shown that, for example, four CGRU blocks can reach a point close to equilibrium. In other embodiments, the length of the CGRU block series may vary.
[0072] The mask combination block 406 is intended for efficiently combining the masks predicted by the series of CGRU blocks 404. The mask combination block 406 receives a first mask predicted by a first series of CNN blocks in each CGRU block of the series of CGRU blocks 404 and generates a final mask 418. The mask combination block 406 includes two CNN blocks, which are connected to the CNN blocks of the series of CGRU blocks 404. Figure 3 310 in the second series of CNN blocks. The first CNN block identifies a specific pattern from the first mask and the second CNN block performs classification again to determine an effective way to combine the first mask. When the second series of CNN blocks 310 takes two masks and produces an output mask 312, the mask combination block 406 takes four (or the number of CGRU blocks) masks and produces a final mask 418.
[0073] 4.1.3 Perceptual Loss Function
[0074] In some embodiments, the server 102 is programmed to train the machine learning model using an appropriate optimization method known to those skilled in the art. The optimization method is typically iterative in nature and can minimize a loss (or cost) function that measures the error between the current estimate and the ground truth. For an ANN, the optimization method can be a stochastic gradient descent method, where the weights are updated using an error back propagation algorithm.
[0075] Traditionally, objective functions or loss functions such as mean squared error (MSE) do not reflect human auditory perception well. A processed speech segment with a smaller MSE does not necessarily have high speech quality and intelligibility. Specifically, even though over-suppression of speech may have a greater perceptual effect than under-suppression of speech, and over-suppression of speech is usually treated differently from under-suppression of speech in speech enhancement applications, the objective function does not distinguish between negative detection errors (false negatives, over-suppression of speech) and positive detection errors (false positives, under-suppression of speech).
[0076] Oversuppression of speech is more detrimental to speech quality or intelligibility than undersuppression of speech. Oversuppression of speech occurs when the predicted (estimated) mask value is less than the true mask value because less speech is predicted than the true value, and thus more speech is suppressed than desired.
[0077] In some embodiments, a perceptual cost function that prevents over-suppression of speech is used in an optimization method to train a machine learning model. The perceptual cost function is nonlinear, with asymmetric penalties for over-suppression and under-suppression of speech. Specifically, the cost function assigns more penalties to negative differences between predicted mask values and true mask values, and less penalties to positive differences. Experiments have shown that, for example, the perceptual loss function outperforms MSE in reducing over-suppression of high-frequency fricatives and low-level filler pauses (such as "um" and "uh").
[0078] In some embodiments, the perceptual loss function Loss is defined as follows:
[0079] diff=y target p -y predicted p (1)
[0080] Loss = m diff -diff-1(2),
[0081] Among them, y target is the target (true value) mask value for frame and band, y predictedis the prediction mask value for the frame and band, m is a tuning parameter that can control the shape of the asymmetric penalty, and p is a power law term or scaling exponent. For example, m can be 2.6, 2.65, 2.7, etc., and p can be 0.5, 0.6, 0.7, etc. Since y predicted or target is less than 1, so a fractional value of p that is not too small (for example, greater than 0.5) will often make y predicted The smaller value than y predicted or target This fractional value of p tends to further magnify y target p and predicted p The difference between them is greater than y target and predicted The smaller the y predicted Value (corresponding to smaller y target The value) may be the result of starting with a noisy frame and continuing to be over-suppressed, resulting in y predicted The value of is smaller. If y target and predicted The difference between them is appropriately enlarged to y target p and predicted p The difference between (using too small a p value may lead to over-frequency amplification), such speech over-suppression will be more severely punished. Therefore, the power law term can be particularly helpful in improving speech over-suppression in difficult cases of noisy frames. This inherent focus on difficult cases also leads to the possibility of having a smaller machine learning model with fewer parameters. The total loss of an audio signal corresponding to multiple frequency bands and multiple frames can be calculated as the sum or average of the loss values over multiple frequency bands and multiple frames.
[0082] In some embodiments, the perceptual loss function Loss is based on MSE as follows:
[0083] diff=y target p -y predicted p (3)
[0084] w=m diff -diff-1(4)
[0085] Loss = w*diff 2 (5)
[0086] Using MSE, positive and negative diff values are penalized the same, so a negative diff value that indicates over-suppression of speech will not be penalized more than a positive diff value that indicates under-suppression of speech. 2 Significant speech undersuppression corresponding to predicted mask values that are far below the target mask value are now penalized multiple times (corresponding to larger errors).
[0087] 4.2 Model Implementation for Speech Enhancement
[0088] In some embodiments, server 102 is programmed to receive a new audio signal having one or more frames in the time domain. Server 102 then applies the machine learning method discussed in Section 4.1 to the new audio signal to generate a prediction mask indicating the amount of speech present in each frame and each frequency band in the corresponding TF representation. The application includes converting the new audio signal into a joint TF representation that initially covers multiple frames and multiple frequency bins.
[0089] In some embodiments, the server 102 is programmed to generate an improved audio signal for the new audio signal further based on the prediction mask. Given a band mask for y (obtained by applying the machine learning method discussed in Section 4.1) as a column vector m_band of size q×1, where y is a column vector of size q×1 representing the band energy of the original noise frame, q represents the number of perceptually excited bands, a conversion to a bin mask can be performed by calculating m_bin=W_transpose*m_band, where m_bin is a column vector of size p×1, p represents the number of bins, and W_transpose of size p×q is the transpose of the band matrix W of size q×p.
[0090] In some embodiments, the server 102 is programmed to apply a bin mask to the original frequency bin amplitudes in the joint TF representation to achieve noise masking or noise reduction and obtain an estimated clean spectrum. The server 102 can also convert the estimated clean spectrum back into a waveform as an enhanced waveform (relative to the noise waveform) using any method known to those skilled in the art (e.g., inverse CQMF), which can be communicated via an output device.
[0091] 5. Example Process
[0092] Figure 5 An example process performed by an audio management computer system according to some embodiments of the present disclosure is shown. Figure 5 The examples are shown in simplified schematic form to illustrate clarity, and other embodiments may include more, fewer, or different elements connected in various ways. Figure 5The intent is to disclose algorithms, plans or outlines that can be used to implement one or more computer programs or other software components. When these programs or components are executed, the functional improvements and technical advances described herein can be achieved. In addition, the level of detail described in the flowcharts herein is the same as the level of detail that is commonly used by those skilled in the art when communicating algorithms, plans or specifications to each other, which form the basis of the software programs they plan to encode or implement using their accumulated skills and knowledge.
[0093] In some embodiments, the server 102 is programmed to receive an input waveform in the time domain and convert the input waveform into raw audio data over a plurality of frequency bins and a plurality of frames. The server 102 is further programmed to convert the raw audio data into audio data by grouping the plurality of frequency bins into a plurality of frequency bands, wherein a joint time-frequency representation has an energy value for each time frame and each frequency band.
[0094] Thus, in step 502, the server 102 is programmed to receive audio data as a joint time-frequency representation over a plurality of frames and a plurality of frequency bands.
[0095] In some embodiments, server 102 is programmed to generate a feature vector from the joint time-frequency representation.
[0096] In step 504, the server 102 is then programmed to execute a digital model for detecting speech from a feature vector of the audio data. The digital model includes a series of mask blocks. Each mask block includes a first component and a second component, the first component generating a first mask for extracting speech, and the second component generating a second mask for extracting residual speech masked by the first mask. Each of the first mask and the second mask includes a mask value, which is used to estimate the amount of speech present in each of the plurality of frames and each of the plurality of frequency bands.
[0097] In some embodiments, the first component includes a series of connected CNN blocks with expansion. Each CNN block in the series of connected CNN blocks includes a CNN layer, a batch normalization layer, and an activation layer. The first component also includes a CNN layer with a 1x1 filter. In other embodiments, the second component includes a gated recurrent unit (GRU) block, which includes a GRU layer.
[0098] In some embodiments, a first component of a first mask block in the series of mask blocks receives as input the feature vector and a second component of the first mask block receives the inverse of the result of applying the first mask to the feature vector.
[0099] In some embodiments, each mask block includes a third component, the third component includes a CNN block configured to combine the first mask and the second mask into an output mask. Each mask block further includes a fourth component, which applies the output mask to a certain feature vector to generate a specific feature vector. The first mask block in a series of mask blocks receives the feature vector as a certain feature vector. Each subsequent mask block in the series of mask blocks receives the specific feature vector generated by the previous mask block as input.
[0100] In some embodiments, the digital model further includes an input CNN block comprising a CNN layer with a look-ahead filter, a batch normalization layer, and an activation layer.
[0101] In step 506, the server 102 is programmed to transmit information related to a first mask generated by the series of mask blocks to the device.
[0102] In some embodiments, the digital model further includes a mask combination block, which includes a CNN block that combines the first masks generated by the series of mask blocks into a final mask. In other embodiments, the server 102 is programmed to perform inverse banding on the mask value of the final mask to generate an updated mask value for each frequency bin in the plurality of frequency bins and each frame in the plurality of frames. The server 102 is also programmed to apply the updated mask value to the audio data to generate new output data, and convert the new output data into an enhanced waveform.
[0103] Various aspects of the embodiments disclosed herein may be understood from the following enumerated example embodiments (EEE):
[0104] EEE1. A computer-implemented method for mitigating audio artifacts, comprising: receiving, by a processor, audio data as a joint time-frequency representation over multiple frames and multiple frequency bands; executing, by the processor, a digital model for detecting speech from a feature vector of the audio data, the digital model comprising a series of mask blocks, each mask block comprising a first component and a second component, the first component generating a first mask for extracting speech, the second component generating a second mask for extracting residual speech masked by the first mask, and each of the first mask and the second mask comprising a mask value, the mask value being used to estimate the amount of speech present in each frame of the multiple frames and each frequency band of the multiple frequency bands; and transmitting information associated with the first mask generated by the series of mask blocks to a device.
[0105] EEE2. According to a computer-implemented method according to claim 1, the first component includes a convolutional neural network (CNN) block with a dilated series of connections.
[0106] EEE3. According to the computer-implemented method according to claim 2, each CNN block in the series of connected CNN blocks includes a CNN layer, a batch normalization layer and an activation layer.
[0107] EEE4. A computer-implemented method according to any one of claims 1 to 3, wherein the first component comprises a CNN layer having a 1x1 filter.
[0108] EEE5. The computer-implemented method of any one of claims 1 to 4, wherein the second component comprises a gated recurrent unit (GRU) block, the GRU block comprising a GRU layer.
[0109] EEE6. According to any one of claims 1 to 5, each mask block includes a third component, and the third component includes a CNN block configured to combine the first mask block and the second mask block into an output mask.
[0110] EEE7. According to the computer-implemented method according to claim 6, each mask block also includes a fourth component, which applies the output mask to a certain feature vector to generate a specific feature vector; the first mask block in the series of mask blocks receives the feature vector as the certain feature vector, and each subsequent mask block in the series of mask blocks receives the specific feature vector generated by the previous mask block as input.
[0111] EEE8. According to any one of claims 1 to 7, the first component of the first mask block in the series of mask blocks receives the feature vector as input, and the second component of the first mask block receives the inverse of the result of applying the first mask to the feature vector. EEE8.
[0112] EEE9. According to the computer-implemented method described in any one of claims 1 to 8, the digital model also includes an input CNN block, which includes a CNN layer with a look-ahead filter, a batch normalization layer and an activation layer.
[0113] EEE10. According to the computer-implemented method described in any one of claims 1 to 9, the digital model also includes a mask combination block, which includes a CNN block that combines the first masks generated from the series of mask blocks into a final mask.
[0114] EEE11. The computer-implemented method according to claim 10 further includes: performing inverse band processing on the mask values of the final mask to generate updated mask values for each frequency bin in a plurality of frequency bins and each frame in the plurality of frames; applying the updated mask values to the audio data to generate new output data; and converting the new output data into an enhanced waveform.
[0115] EEE12. The computer-implemented method according to any one of claims 1 to 11, further comprising: receiving an input waveform in the time domain; converting the input waveform into raw audio data on a plurality of frequency bins and the plurality of frames; converting the raw audio data into the audio data by grouping the plurality of frequency bins into the plurality of frequency bands, the joint time-frequency representation having an energy value for each time frame and each frequency band; and generating the feature vector from the joint time-frequency representation.
[0116] EEE13. The computer-implemented method according to any one of claims 1 to 12 also includes training the digital model using a loss function with a nonlinear penalty, wherein the nonlinear penalty punishes excessive speech suppression more than the penalty for insufficient speech suppression.
[0117] EEE14. A system for alleviating excessive speech suppression, the system comprising: a memory; one or more processors, the one or more processors being coupled to the memory and configured to perform the following operations: receiving audio data as a joint time-frequency representation over multiple frames and multiple frequency bands; executing a digital model for detecting speech from a feature vector of the audio data, the digital model comprising a series of mask blocks, each mask block comprising a first component and a second component, the first component generating a first mask for extracting speech, the second component generating a second mask for extracting residual speech masked by the first mask, and each of the first mask and the second mask comprising a mask value, the mask value being used to estimate the amount of speech present in each frame of the multiple frames and each frequency band of the multiple frequency bands; and transmitting information associated with the first mask generated by the series of mask blocks to a device.
[0118] EEE15. A computer-readable non-transitory storage medium storing computer-executable instructions that, when executed, implement a method for mitigating audio artifacts, the method comprising: receiving audio data as a joint time-frequency representation over multiple frames and multiple frequency bands by a processor; executing a digital model for detecting speech from a feature vector of the audio data, the digital model comprising a mask block, the mask block comprising a first series of CNN blocks and a GRU block, the first series of CNN blocks generating a first mask for extracting speech, the GRU block generating a second mask for extracting residual speech masked by the first mask, each CNN block in the first series of CNN blocks comprising a CNN layer, and the GRU block comprising a GRU layer, and each of the first mask and the second mask comprising a mask value, the mask value being used to estimate the amount of speech present in each frame in the multiple frames and each frequency band in the multiple frequency bands; and transmitting information associated with the first mask and the second mask.
[0119] EEE16. According to the computer-readable non-volatile storage medium according to claim 15, the mask block also includes an additional block, which uses the first mask and the second mask to derive a specific feature vector from a certain feature vector, and the digital model includes a series of mask blocks, the series of mask blocks includes the mask block, the first mask block in the series of mask blocks receives the feature vector, and each subsequent mask block in the series of mask blocks receives the specific feature vector generated by the previous mask block as input.
[0120] EEE17. According to the computer-readable non-transitory storage medium according to claim 16, the digital model also includes a mask combination block, which includes a CNN block that combines the first masks generated by the series of mask blocks into a final mask.
[0121] EEE18. According to the computer-readable non-transitory storage medium described in any one of claims 15 to 17, the first series of CNN blocks has an expansion rate that first increases and then decreases.
[0122] EEE19. According to any one of claims 15 to 18, the mask block further includes a CNN block that combines the first mask and the second mask into an output mask.
[0123] EEE20. According to the computer-readable non-transitory storage medium described in any one of claims 15 to 19, the digital model also includes an input CNN block, which includes a CNN layer with a look-ahead filter, a batch normalization layer and an activation layer.
[0124] 6. Hardware Implementation
[0125] According to one embodiment, the techniques described herein are implemented by at least one computing device. These techniques can be implemented in whole or in part using a combination of at least one server computer and / or other computing devices that are coupled using a network (e.g., a packet data network). To perform these techniques, the computing device can be hard-wired; or it can also include digital electronic devices, such as at least one application-specific integrated circuit (ASIC) or field-programmable gate array (FPGA), which are continuously programmed to perform the above-mentioned techniques; or it can also include at least one general-purpose hardware processor, which is programmed to perform these techniques according to program instructions in firmware, memory, other storage devices, or a combination thereof. Such computing devices can also combine hard-wired custom logic, ASICs, or FPGAs with custom programming to implement the described techniques. These computing devices may be server computers, workstations, personal computers, portable computer systems, handheld devices, mobile computing devices, wearable devices, body-mounted or implantable devices, smartphones, smart home appliances, internet-connected devices, autonomous or semi-autonomous devices (such as robots or unmanned ground or air vehicles), any other electronic device that incorporates hardwired logic and / or program logic to implement the described techniques, one or more virtual computers or instances in a data center, and / or a network of server computers and / or personal computers.
[0126] Figure 6 is a block diagram illustrating an example computer system in which embodiments of the present invention may be implemented. Figure 6 In the example of the present disclosure, computer system 600 and instructions for implementing the disclosed techniques in hardware, software, or a combination of hardware and software are represented schematically (e.g., as boxes and circles) with the same level of detail that is typically used by those of ordinary skill in the art to which the present disclosure relates when communicating about computer architecture and computer system implementations.
[0127] Computer system 600 includes an input / output (I / O) subsystem 602, which may include a bus and / or other communication mechanism for passing information and / or instructions between components of computer system 600 via electronic signal paths. I / O subsystem 602 may include an I / O controller, a memory controller, and at least one I / O port. Electronic signal paths are represented in schematic form in the figure, such as lines, unidirectional arrows, or bidirectional arrows.
[0128] At least one hardware processor 604 is coupled to the I / O subsystem 602 for processing information and instructions. For example, the hardware processor 604 may include a general-purpose microprocessor or microcontroller and / or a dedicated microprocessor, such as an embedded system or a graphics processing unit (GPU) or a digital signal processor or an ARM processor. The processor 604 may include an integrated arithmetic logic unit (ALU) or may be coupled to a separate ALU.
[0129] The computer system 600 includes one or more memory 606 units, such as a main memory, coupled to the I / O subsystem 602 for electronically digitally storing data and instructions to be executed by the processor 604. The memory 606 may include volatile memory, such as various forms of random access memory (RAM) or other dynamic storage devices. The memory 606 may also be used to store temporary variables or other intermediate information during the execution of instructions to be executed by the processor 604. These instructions, when stored in a non-transitory computer-readable storage medium accessible to the processor 604, can turn the computer system 600 into a special-purpose machine that is customized to perform the operations specified in the instructions.
[0130] Computer system 600 also includes non-transitory memory, such as read-only memory (ROM) 608 or other static storage device coupled to I / O subsystem 602, for storing information and instructions for processor 604. ROM 608 may include various forms of programmable ROM (PROM), such as erasable PROM (EPROM) or electrically erasable PROM (EEPROM). Persistent storage device 610 unit may include various forms of non-volatile RAM (NVRAM), such as FLASH memory, or solid-state storage devices, magnetic disks or optical disks, such as CD-ROM or DVD-ROM, and may be coupled to I / O subsystem 602 for storing information and instructions. Storage device 610 is an example of a non-transitory computer-readable medium that can be used to store instructions and data that, when executed by processor 604, enable execution of a computer-implemented method to perform the techniques described herein.
[0131] The instructions in the memory 606, ROM 608, or storage device 610 may include one or more groups of instructions organized as modules, methods, objects, functions, routines, or calls. The instructions may be organized into one or more computer programs, operating system services, or applications (including mobile applications). The instructions may include operating systems and / or system software; one or more libraries supporting multimedia, programming, or other functions; data protocol instructions or stacks to implement TCP / IP, HTTP, or other communication protocols; file processing instructions to interpret and render files encoded using HTML, XML, JPEG, MPEG, or PNG; user interface instructions to render or interpret commands for a graphical user interface (GUI), command line interface, or text user interface; application software, such as an office suite, an Internet access application, a design and manufacturing application, a graphics application, an audio application, a software engineering application, an educational application, a game, or other application. The instructions may implement a network server, a network application server, or a network client. The instructions may be organized into a presentation layer, an application layer, and a data storage layer, such as a relational database system using structured query language (SQL) or NoSQL, an object store, an image database, a flat file system, or other data storage.
[0132] The computer system 600 can be coupled to at least one output device 612 via the I / O subsystem 602. In one embodiment, the output device 612 is a digital computer display. Examples of displays that can be used in various embodiments include a touch screen display or a light emitting diode (LED) display or a liquid crystal display (LCD) or an electronic paper display. The computer system 600 may include other types of output devices 612 as an alternative or supplement to the display device. Examples of other output devices 612 include a printer, a ticket printer, a plotter, a projector, a sound card or a video card, a speaker, a buzzer or a piezoelectric device or other sound-generating device, a lamp or an LED or LCD indicator light, a tactile device, an actuator, or a servo.
[0133] At least one input device 614 is coupled to the I / O subsystem 602 for communicating signals, data, command selections, or gestures to the processor 604. Examples of input device 614 include touch screens, microphones, still and video digital cameras, alphanumeric and other keys, keypads, keyboards, digitizer tablets, image scanners, joysticks, clocks, switches, buttons, dials, sliders, and / or various types of sensors such as force sensors, motion sensors, thermal sensors, accelerometers, gyroscopes, and inertial measurement unit (IMU) sensors and / or various types of transceivers such as wireless transceivers, such as cellular or Wi-Fi, radio frequency (RF) or infrared (IR) transceivers, and global positioning system (GPS) transceivers.
[0134] Another type of input device is a control device 616, which can perform cursor control or other automatic control functions, such as navigating in a graphical interface on a display screen, to replace the input function or to supplement the input function. The control device 616 can be a touch pad, a mouse, a trackball, or a cursor direction key, which is used to transmit direction information and command selections to the processor 604 and control the movement of the cursor on the output device 612. The input device can have at least two degrees of freedom on two axes (a first axis (e.g., an x-axis) and a second axis (e.g., a y-axis)), which allows the device to specify a position in a plane. Another type of input device is a wired, wireless, or optical control device, such as a joystick, a wand, a console, a steering wheel, a pedal, a shift mechanism, or other types of control devices. The input device 614 can include a combination of multiple different input devices, such as a video camera and a depth sensor.
[0135] In another embodiment, computer system 600 may include an Internet of Things (IoT) device in which one or more of output device 612, input device 614, and control device 616 are omitted. Alternatively, in such an embodiment, input device 614 may include one or more cameras, motion detectors, thermometers, microphones, earthquake detectors, other sensors or detectors, measurement devices, or encoders, and output device 612 may include a dedicated display, such as a single-line LED or LCD display, one or more indicators, display panels, meters, valves, solenoids, actuators, or servos.
[0136] When the computer system 600 is a mobile computing device, the input device 614 may include a global positioning system (GPS) receiver coupled to a GPS module capable of triangulating a plurality of GPS satellites to determine and generate a geographic location or location data, such as latitude and longitude values of the geophysical location of the computer system 600. The output device 612 may include hardware, software, firmware, and interfaces for generating location report packets, notifications, pulse or heartbeat signals, or other repetitive data transmissions that may specify the location of the computer system 600 alone or in combination with other application-specific data and point to the host 624 or server 630.
[0137] The computer system 600 may implement the techniques described herein using custom hardwired logic, at least one ASIC or FPGA, firmware, and / or program instructions or logic that, when loaded and used or executed in conjunction with the computer system, cause or program the computer system to operate as a special purpose machine. According to one embodiment, the computer system 600 performs the techniques described herein in response to the processor 604 executing at least one sequence of at least one instruction contained in the main memory 606. These instructions may be read into the main memory 606 from another storage medium, such as the storage device 610. Executing the sequence of instructions contained in the main memory 606 causes the processor 604 to perform the process steps described herein. In alternative embodiments, hardwired circuitry may be used in place of or in conjunction with software instructions.
[0138] As used herein, the term "storage medium" refers to any non-transitory medium that stores data and / or instructions that cause a machine to operate in a particular manner. Such storage media may include non-volatile media and / or volatile media. For example, non-volatile media include, for example, optical or magnetic disks, such as storage device 610. Volatile media include dynamic memory, such as memory 606. Common forms of storage media include, for example, hard disks, solid-state drives, flash drives, magnetic data storage media, any optical or physical data storage media, memory chips, etc.
[0139] Storage media are distinct from, but may be used in conjunction with, transmission media. Transmission media participate in the transmission of information between storage media. For example, transmission media include coaxial cables, copper wires, and optical fibers, including the wires that make up the bus of I / O subsystem 602. Transmission media may also take the form of sound or light waves, such as those generated during radio wave and infrared data communications.
[0140] When at least one sequence of at least one instruction is transmitted to the processor 604 for execution, various forms of media may be involved. For example, the instructions may initially be carried on a disk or solid-state drive of a remote computer. The remote computer may load the instructions into its dynamic memory and send the instructions via a communication link (such as an optical fiber, a coaxial cable, or a telephone line using a modem). A modem or router local to the computer system 600 may receive data on the communication link and convert the data into data that can be read by the computer system 600. For example, a receiver (such as a radio frequency antenna or an infrared detector) may receive data carried in a wireless or optical signal, and appropriate circuitry may provide the data to the I / O subsystem 602 (such as placing the data on a bus). The I / O subsystem 602 carries the data to the memory 606, from which the processor 604 retrieves and executes the instructions. The instructions received by the memory 606 may optionally be stored on the storage device 610 before or after execution by the processor 604.
[0141] Computer system 600 also includes a communication interface 618 coupled to I / O subsystem 602. Communication interface 618 provides a two-way data communication coupling with a network link 620, which is directly or indirectly connected to at least one communication network (e.g., network 622 or a public or private cloud on the Internet). For example, communication interface 618 can be an Ethernet network interface, an integrated services digital network (ISDN) card, a cable modem, a satellite modem, or a modem for providing a data communication connection with a corresponding type of communication line (e.g., an Ethernet cable or any type of metal cable or fiber optic line or telephone line). Network 622 broadly represents a local area network, a wide area network, a campus network, an Internet network, or any combination thereof. Communication interface 618 may include a local area network card (for providing a data communication connection with a compatible local area network), or a cellular radiotelephone interface (for sending or receiving cellular data according to a cellular radiotelephone wireless network standard), or a satellite radio interface (for sending or receiving digital data according to a satellite wireless network standard). In any such implementation, communication interface 618 sends and receives electrical, electromagnetic or optical signals via signal paths that carry digital data streams representing various types of information.
[0142] The network link 620 typically uses technologies such as satellite, cellular mobile communication, Wi-Fi or Bluetooth to communicate electronically, electromagnetically or optically with other data devices directly or through at least one network. For example, the network link 620 can provide a connection to the host 624 through the network 622.
[0143] In addition, network link 620 can provide connection through network 622, or provide connection with other computing devices via Internet equipment and / or computers operated by Internet Service Provider (ISP) 626. ISP 626 provides data communication services through a global packet data communication network represented as Internet 628. Server computer 630 can be coupled to Internet 628. Server 630 broadly represents any computer, data center, virtual machine or virtual computing instance (whether or not there is a hypervisor), or a computer that executes a containerized program system such as DOCKER or Kubernetes. Server 630 can represent an electronic digital service implemented using more than one computer or instance, which is accessed and used by transmitting network service requests, uniform resource locator (URL) strings with parameters in HTTP payloads, application programming interface (API) calls, application service calls, or other service calls. Computer system 600 and server 630 can constitute elements of a distributed computing system, which includes other computers, processing clusters, server farms, or other organizations of computers that cooperate to perform tasks or execute applications or services. Server 630 may include one or more sets of instructions organized as modules, methods, objects, functions, routines, or calls. These instructions may be organized into one or more computer programs, operating system services, or applications (including mobile applications). The instructions may include operating systems and / or system software; one or more libraries to support multimedia, programming, or other functions; data protocol instructions or stacks to implement TCP / IP, HTTP, or other communication protocols; file format processing instructions to interpret or present files encoded using HTML, XML, JPEG, MPEG, or PNG; user interface instructions to present or interpret commands for graphical user interfaces, command line interfaces, or text user interfaces; application software such as office suites, Internet access applications, design and manufacturing applications, graphics applications, audio applications, software engineering applications, educational applications, games, or other applications. Server 630 may include a web application server that may host a presentation layer, an application layer, and a data storage layer, such as a relational database system using structured query language (SQL) or NoSQL, an object store, a graph database, a flat file system, or other data storage.
[0144] Computer system 600 may send messages and receive data and instructions (including program code) through the network, network link 620, and communication interface 618. In the Internet example, server 630 may send the requested code for an application program through Internet 628, ISP 626, local network 622, and communication interface 618. The received code may be executed by processor 604 as it is received and / or stored in storage device 610 or other non-volatile storage device for later execution.
[0145] The execution of the instructions described in this section may implement a process in the form of an instance of a computer program being executed, which consists of program code and its current activity. Depending on the operating system (OS), a process may consist of multiple threads of execution that execute instructions in parallel. In this case, a computer program is a passive collection of instructions, and a process may be the actual execution of these instructions. Multiple processes may be associated with the same program; for example, opening multiple instances of the same program often means that more than one process is being executed. Multitasking may be implemented to allow multiple processes to share the processor 604. Although each processor 604 or processor core executes only a single task at a time, the computer system 600 may be programmed to implement multitasking to allow each processor to switch between executing tasks without waiting for each task to complete. In one embodiment, switching may be performed when a task performs an input / output operation, when a task indicates that it can be switched, or when a hardware interrupt occurs. Time sharing may be achieved by quickly performing context switching to provide a fast response for interactive user applications, thereby providing the illusion that multiple processes are executing in parallel at the same time. In one embodiment, for security and reliability reasons, the operating system may prevent direct communication between independent processes and provide a strictly mediated and controlled inter-process communication function.
[0146] 7. Extensions and Alternatives
[0147] In the foregoing description, embodiments of the present disclosure have been described with reference to many specific details, which may vary from one implementation to another. Therefore, the description and drawings should be regarded as illustrative rather than restrictive. The only and exclusive indicator of the scope of the present disclosure and the scope of the present disclosure intended by the applicant is the literal scope of the claims issued in specific form in this application and their equivalents, including any subsequent amendments.
Claims
1. A computer-implemented method for mitigating audio artifacts, comprising: receiving, by a processor, audio data as a joint time-frequency representation over a plurality of frames and a plurality of frequency bands; executing, by the processor, a digital model for detecting speech from a feature vector of the audio data, The digital model includes a series of mask blocks, Each mask block includes a first component and a second component, wherein the first component generates a first mask for extracting speech, and the second component generates a second mask for extracting residual speech masked by the first mask, and Each of the first mask and the second mask includes a mask value for estimating an amount of speech present in each of the plurality of frames and in each of the plurality of frequency bands; as well as Information associated with the first mask produced by the series of mask blocks is transmitted to a device.
2. The computer-implemented method of claim 1 , wherein the first component comprises a convolutional neural network (CNN) block with a dilated series of connections.
3. The computer-implemented method of claim 2 , wherein each CNN block in the series of connected CNN blocks comprises a CNN layer, a batch normalization layer, and an activation layer.
4. The computer-implemented method of any one of claims 1 to 3, wherein the first component comprises a CNN layer having a 1x1 filter.
5. The computer-implemented method of any one of claims 1 to 4, the second component comprising a gated recurrent unit (GRU) block, the GRU block comprising a GRU layer.
6. The computer-implemented method of any one of claims 1 to 5, each mask block comprising a third component, the third component comprising a CNN block configured to combine the first mask block and the second mask block into an output mask.
7. The computer-implemented method of claim 6, Each mask block further includes a fourth component, the fourth component applying the output mask to a certain feature vector to generate a specific feature vector; A first mask block in the series of mask blocks receives the feature vector as the certain feature vector, and Each subsequent mask block in the series of mask blocks receives as input the particular feature vector produced by the previous mask block.
8. A computer-implemented method according to any one of claims 1 to 7, The first component of a first mask block in the series of mask blocks receives as input the feature vector, and The second component of the first mask block receives the inverse of a result of applying the first mask to the feature vector.
9. The computer-implemented method of any one of claims 1 to 8, wherein the digital model further comprises an input CNN block comprising a CNN layer with a look-ahead filter, a batch normalization layer, and an activation layer.
10. The computer-implemented method according to any one of claims 1 to 9, wherein the digital model further comprises a mask combination block, wherein the mask combination block comprises a CNN block that combines the first masks generated from the series of mask blocks into a final mask.
11. The computer-implemented method of claim 10, further comprising: performing inverse band processing on mask values of the final mask to generate updated mask values for each frequency bin of a plurality of frequency bins and each frame of the plurality of frames; applying the updated mask value to the audio data to generate new output data; as well as The new output data is converted into an enhanced waveform.
12. The computer-implemented method of any one of claims 1 to 11, further comprising: receiving an input waveform in the time domain; converting the input waveform into raw audio data over a plurality of frequency bins and the plurality of frames; converting the raw audio data into the audio data by grouping the plurality of frequency band bins into the plurality of frequency bands, The joint time-frequency representation has an energy value for each time frame and each frequency band; and The feature vector is generated from the joint time-frequency representation.
13. The computer-implemented method according to any one of claims 1 to 12 further includes training the digital model using a loss function with a nonlinear penalty, wherein the loss function with a nonlinear penalty penalizes excessive speech suppression more than the penalty for insufficient speech suppression.
14. A system for alleviating speech over-suppression, the system comprising: Memory; and One or more processors, the one or more processors being coupled to the memory and configured to perform the following operations: receiving audio data as a joint time-frequency representation over a plurality of frames and a plurality of frequency bands; executing a digital model for detecting speech from a feature vector of said audio data, The digital model includes a series of mask blocks, Each mask block includes a first component and a second component, wherein the first component generates a first mask for extracting speech, and the second component generates a second mask for extracting residual speech masked by the first mask, and Each of the first mask and the second mask includes a mask value for estimating an amount of speech present in each of the plurality of frames and in each of the plurality of frequency bands; as well as Information associated with the first mask produced by the series of mask blocks is transmitted to a device.
15. A computer-readable non-transitory storage medium storing computer-executable instructions that, when executed, implement a method of mitigating audio artifacts, the method comprising: receiving, by a processor, audio data as a joint time-frequency representation over a plurality of frames and a plurality of frequency bands; executing a digital model for detecting speech from a feature vector of said audio data, The digital model includes a mask block, the mask block includes a first series of CNN blocks and a GRU block, the first series of CNN blocks generates a first mask for extracting speech, and the GRU block generates a second mask for extracting residual speech masked by the first mask, Each CNN block in the first series of CNN blocks comprises a CNN layer, and the GRU block comprises a GRU layer, and Each of the first mask and the second mask includes a mask value for estimating an amount of speech present in each of the plurality of frames and in each of the plurality of frequency bands; as well as Information associated with the first mask and the second mask is transmitted.
16. The computer-readable non-transitory storage medium according to claim 15, The mask block further includes an additional block, wherein the additional block uses the first mask and the second mask to derive a specific feature vector from a certain feature vector, The digital model comprises a series of mask blocks, the series of mask blocks comprising the mask block, A first mask block in the series of mask blocks receives the feature vector, and Each subsequent mask block in the series of mask blocks receives as input the particular feature vector produced by the previous mask block.
17. The computer-readable non-transitory storage medium of claim 16, the digital model further comprising a mask combination block, the mask combination block comprising a CNN block that combines the first masks generated by the series of mask blocks into a final mask.
18. The computer-readable non-transitory storage medium of any one of claims 15 to 17, wherein the first series of CNN blocks have a dilation rate that first increases and then decreases.
19. The computer-readable non-transitory storage medium of any one of claims 15 to 18, wherein the mask block further comprises a CNN block that combines the first mask and the second mask into an output mask.
20. The computer-readable non-transitory storage medium of any one of claims 15 to 19, wherein the digital model further comprises an input CNN block comprising a CNN layer having a look-ahead filter, a batch normalization layer, and an activation layer.