A streaming, multi-channel, neural-emphasized front-end that is invariant in a microphone array configuration for automatic speech recognition
The multi-channel neural front-end speech enhancement model addresses the challenges of background noise and competing speech in ASR systems by using a speech cleaner and self-attention blocks to generate enhanced input speech features, thereby improving recognition accuracy in difficult acoustic conditions.
Patent Information
- Application Number
- JP2024555936
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-03-20
- Filing Date
- 2023-02-20
- Publication Date
- 2025-06-09
- Estimated Expiration
- 2043-02-20
AI Technical Summary
Existing automatic speech recognition (ASR) systems face challenges in accurately recognizing speech in background conditions with both voice-based and non-voice-based noise, particularly in reverberation, background noise, and competing speech scenarios.
A multi-channel neural front-end speech enhancement model is introduced, comprising a speech cleaner, a stack of self-attention blocks, and a masking layer. This model processes multi-channel noisy input signals and context noise signals to generate enhanced input speech features, effectively removing background interference.
The proposed model significantly improves the robustness of ASR systems by effectively separating speech from background noise and competing voices, leading to enhanced speech recognition performance even in challenging acoustic conditions.
Smart Images

Figure 0007690138000014 
Figure 0007690138000015 
Figure 0007690138000016
Abstract
Description
Technical Field
[0001] The present disclosure relates to a streaming, multi-channel, neural enhanced front-end that is invariant in a microphone array configuration for automatic speech recognition.
Background Art
[0002] The robustness of automatic speech recognition (ASR) systems has been significantly improved over the years by the emergence of neural network-based end-to-end models, large-scale training data, and improved strategies for augmenting training data. However, various conditions such as reverberation, significant background noise, and competing speech, significantly degrade the performance of automatic speech recognition ASR systems. Joint automatic speech recognition ASR models can be trained to handle these conditions.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, separating speech in background conditions that include both voice-based noise and non-voice-based noise is particularly difficult.
Means for Solving the Problems
[0005] One aspect of the present disclosure provides a multi-channel neural front-end speech enhancement model for speech recognition. The multi-channel neural front-end speech enhancement model includes a speech cleaner, a stack of self-attention blocks each having a multi-head self-attention mechanism, and a masking layer. The speech cleaner receives, as input, a multi-channel noisy input signal and a multi-channel context noise signal, and generates, as output, a single-channel cleaned input signal. The stack of self-attention blocks receives, as input, a stack input including, at the first block of the stack of self-attention blocks, the single-channel cleaned input signal output from the speech cleaner and the single-channel noisy input signal, and generates, as output from the final block of the stack of self-attention blocks, an unmasked output. The masking layer receives, as input, the single-channel noisy input signal and the unmasked output generated as output from the final block of the stack of self-attention blocks, and generates, as output, enhanced input speech features corresponding to the target utterance.
[0006] Embodiments of the present disclosure may include one or more of the following optional features. In some embodiments, the stack of self-attention blocks includes a stack of conformer blocks. In these embodiments, the stack of conformer blocks may include four conformer blocks. In some examples, the speech enhancement model is executed on data processing hardware present in the user device. Here, the user device is configured to capture the target utterance and the multi-channel context noise signal via an array of microphones of the user device. In these examples, the speech enhancement model may be agnostic with respect to the number of microphones in the microphone array.
[0007] In some embodiments, the voice cleaner generates a single-channel cleaned input signal by performing an adaptive noise cancellation algorithm, generating a total output by applying finite impulse response (FIR) filters to all channels of the multi-channel noisy input signal except the first channel of the multi-channel noisy input signal, and subtracting the total output from the first channel of the multi-channel noisy input signal. In some examples, the backend voice system is configured to process the enhanced input voice features corresponding to the target utterance. In these examples, the backend voice system comprises at least one of an automatic speech recognition (ASR) model, or an audio call application, or an audio-video call application.
[0008] In some embodiments, the voice enhancement model is jointly trained with a backend automatic speech recognition (ASR) model by using a spectral loss and an automatic speech recognition ASR loss. In these embodiments, the spectral loss can be based on the distance of the L1 loss function and the L2 loss function between the estimated ratio mask and the ideal ratio mask. Here, the ideal ratio mask is calculated by using the reverberant voice and the reverberant noise. Additionally or alternatively, the automatic speech recognition ASR loss is calculated by generating a predicted output of the automatic speech recognition ASR encoder of the enhanced voice features by using the automatic speech recognition ASR encoder of the automatic speech recognition ASR model configured to receive, as input, the enhanced voice features predicted by the voice enhancement model for the training utterances, and generating a target output of the automatic speech recognition ASR encoder of the target voice features by using the automatic speech recognition ASR encoder configured to receive, as input, the target voice features of the training utterances. Here, the step of calculating the automatic speech recognition ASR loss is based on the predicted output of the automatic speech recognition ASR encoder of the enhanced voice features and the target output of the automatic speech recognition ASR encoder of the target voice features.
[0009] Another aspect of the present disclosure provides a computer-implemented method that, when executed on data processing hardware, causes the data processing hardware to perform operations. The operations include receiving a multi-channel noisy input signal and a multi-channel context noise signal, and generating a single-channel cleaned input signal by using a voice cleaner of a voice enhancement model. The operations also include generating an unmasked output as an output from a stack of self-attention blocks of a voice enhancement model configured to receive a stacked input comprising the single-channel cleaned input signal output from the voice cleaner and the single-channel noisy input signal. Here, each self-attention block of the stack of self-attention blocks comprises a multi-head self-attention mechanism. The operations further include generating enhanced input voice features corresponding to a target utterance by using a masking layer of the voice enhancement model configured to receive the single-channel noisy input signal and the unmasked output generated as an output from the stack of self-attention blocks.
[0010] This aspect may include one or more of the following optional features. In some embodiments, the stack of self-attention blocks comprises a stack of conformer blocks. In these embodiments, the stack of conformer blocks may include four conformer blocks. In some examples, the voice cleaner, the stack of self-attention blocks, and the masking layer are executed on data processing hardware present on a user device. Here, the user device is configured to capture a target utterance and a multi-channel context noise signal via an array of microphones of the user device. In these examples, the voice enhancement model may be agnostic with respect to the number of microphones in the microphone array.
[0011] In some embodiments, the operation further comprises generating a total output by applying finite impulse response (FIR) filters to all channels of the multi-channel noisy input signal except the first channel of the multi-channel noisy input signal using an acoustic cleaner, and running an adaptive noise cancellation algorithm to generate a single-channel cleaned input signal by subtracting the total output from the first channel of the multi-channel noisy input signal. In some examples, the backend speech system is configured to process enhanced input speech features corresponding to a target utterance. In these examples, the backend speech system comprises at least one of an automatic speech recognition (ASR) model, or an audio or audio-video calling application.
[0012] In some embodiments, the voice enhancement model is jointly trained with a backend automatic speech recognition (ASR) model by using a spectral loss and an automatic speech recognition ASR loss. In these embodiments, the spectral loss can be based on the distance of the L1 loss function and the L2 loss function between the estimated ratio mask and the ideal ratio mask. Here, the ideal ratio mask is calculated by using the reverberant speech and the reverberant noise. Additionally or alternatively, the automatic speech recognition ASR loss is calculated by generating a predicted output of the automatic speech recognition ASR encoder of the enhanced speech features by using the automatic speech recognition ASR encoder of the automatic speech recognition ASR model configured to receive, as input, the enhanced speech features predicted by the voice enhancement model for the training utterances, and generating a target output of the automatic speech recognition ASR encoder of the target speech features by using the automatic speech recognition ASR encoder configured to receive, as input, the target speech features of the training utterances. Here, the step of calculating the automatic speech recognition ASR loss is based on the predicted output of the automatic speech recognition ASR encoder of the enhanced speech features and the target output of the automatic speech recognition ASR encoder of the target speech features.
[0013] Details of one or more embodiments of the present disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will be apparent from the description and drawings and from the claims.
Brief Description of the Drawings
[0014]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Modes for Carrying Out the Invention
[0015] Like reference symbols in the various drawings indicate like elements. The robustness of automatic speech recognition (ASR) systems has been significantly improved over the years by the emergence of neural network-based end-to-end models, large-scale training data, and improved strategies for augmenting training data. Nevertheless, background interference can significantly degrade the ability of an automatic speech recognition (ASR) system to accurately recognize speech directed at the ASR system. Background interference can be roughly classified into three groups: device echo, ambient noise, and competing voices. Separate automatic speech recognition (ASR) models can be trained to handle each of these background interference groups separately. However, maintaining ASR models specific to multiple tasks / conditions and switching models on-the-fly during use is not only difficult but also not practical.
[0016] Device echo can respond to the playback audio output from devices such as smart home speakers. Thus, the playback audio can be recorded as an echo, which can affect the performance of a backend audio system such as an automatic speech recognition (ASR) system. In particular, the degradation of the performance of the backend audio system is particularly severe when the playback audio includes audible speech, for example, when it includes a text-to-speech (TTS) response from a digital assistant.
[0017] Background noise with non-audio characteristics (background noise) is usually appropriately processed by using data augmentation strategies such as multi-style training (MTR) of an automatic speech recognition (ASR) model. Here, noise is added to the training data by using an indoor simulator. Then, during training, they are carefully weighted with clean data, so that the performance balance between the clean state and the noisy state is achieved. As a result, large-scale ASR models are robust to medium-level non-audio noise. However, in the presence of low signal-to-noise ratio (SNR) conditions, background noise can still affect the performance of the backend speech system.
[0018] Unlike non-audio background noise, competing speech is a major challenge for an ASR model trained to recognize a single speaker. Training an ASR model with the voices of multiple speakers can itself be problematic because it is difficult to eliminate the ambiguity of which speaker to focus on during inference. Since it is difficult to know in advance the number of users to support, using a model that recognizes multiple speakers is not optimal either. Furthermore, such multi-speaker models usually have degraded performance in a single-speaker setting, which is undesirable.
[0019] The above three classes of background interference are usually dealt with separately from each other, each using a separate modeling strategy. In recent literature, voice separation using deep clustering, permutation invariant training techniques, and speaker embedding has received much attention. When using speaker embedding, the target speaker of interest is assumed to be known a priori. Techniques developed for speaker separation are also applied to modify training data and remove non-speech noise. Acoustic echo cancellation (AEC) has also been studied, either alone or together, in the presence of ambient noise. It is well known that improving the quality of speech does not necessarily improve the performance of automatic speech recognition (ASR) because the distortion caused by non-linear processing can have an adverse effect on ASR performance. One way to reduce the mismatch between the enhanced front-end that processes incoming audio first and the resulting ASR performance is to co-train the enhanced front-end together with the back-end ASR model.
[0020] Furthermore, the application of large-scale multi-domain and multi-language ASR models continues to attract interest. Since the training data for these ASR models usually cover various acoustic and language use cases (e.g., voice search and video captioning), it has become difficult to simultaneously handle more challenging noise conditions. As a result, it is often convenient to train and maintain separate front-end feature processing models that can handle adverse conditions without combining them with the back-end ASR model.
[0021] Embodiments of this specification are directed to training a front - end voice enhancement model to improve the robustness of automatic speech recognition (ASR). This model is practical from the perspective that, especially in a streaming automatic speech recognition (ASR) setting, it is difficult, if not impossible, to know in advance which class of background interference to handle. Specifically, the front - end voice enhancement model comprises a context - enhanced neural network (CENN) that is enabled to utilize multi - channel noisy input signals and multi - channel context noise signals. In the case of voice enhancement and separation, the noise context, i.e., the audio of a few seconds before the target utterance to be recognized, conveys useful information regarding the acoustic context. The context - enhanced (enhanced) neural network CENN generates enhanced input voice features by using each neural network architecture configured to take in the noisy input and the context input. The enhanced input voice features can be passed to a back - end voice system, such as an automatic speech recognition (ASR) model, that can process the enhanced input voice features to generate a speech recognition result for the target utterance. In particular, the front - end voice enhancement model is designed to operate with a multi - channel array, but the front - end voice enhancement model itself is agnostic with respect to the number of channels of the array or its configuration.
[0022] Referring to FIG. 1, in some embodiments, system 100 includes a user 10 who, in an audio environment, conveys a spoken target utterance (spoken target at-rance 12) to an audio-responsive user device 110 (also referred to as device 110 or user device 110). The user 10 (i.e., the speaker of utterance 12) may speak the target utterance 12 as a query or command seeking a response from device 110. The user device 110 is configured to capture sound from one or more users 10, 11 within the audio environment. Here, the audio sound may refer to an audible query, a command for device 110, or a spoken utterance (spoken at-rance 12) by user 10 that functions as an audible communication captured by device 110. The audio-responsive system of device 110, or an audio-responsive system associated with device 110, may execute a query of the command by responding to the query and / or executing the command.
[0023] Various types of background interference may interfere with the ability of the backend voice system 180 to process the target utterance 12 that specifies a query or command to the device 110. As described above, the background interference includes one or more device echoes corresponding to the reproduced audio 154 already output from the user device (e.g., smart speaker) 110, competing voices 13 such as utterances other than the target utterance 12 spoken by one or more other users 11 not directed at the user device 110, and ambient noise (background noise) having non-audio characteristics such as an incoming call tone 15 from another user device 111. Embodiments herein use a multi-channel neural front-end voice enhancement model 200 (also referred to as model 200 or front-end voice enhancement model 200) executed on the device 110. The multi-channel neural front-end voice enhancement model 200 is configured to receive, as inputs, a multi-channel noisy input signal 202 having voice characteristics corresponding to the target utterance 12 and the background interference, and a multi-channel context noise signal 204, and to generate, as an output, enhanced input voice characteristics 250 corresponding to the target utterance 12 by processing the multi-channel noisy input signal 202 and the multi-channel context noise signal 204 to remove the background interference. The multi-channel noisy input signal 202 includes one or more channels 206, 206a - 206n of audio. Next, the backend voice system 180 is enabled to generate an output 182 by processing the enhanced input voice characteristics 250. In particular, the multi-channel neural front-end voice enhancement model 200 effectively removes (i.e., masks) the presence of background interference recorded by the device 110 when the user 10 speaks the target utterance 12 such that the enhanced input voice characteristics 250 provided to the backend voice system 180 convey the voice intended for the device 110 (i.e., the target utterance 12) and the output 182 generated by the backend voice system 180 is not degraded by the background interference.
[0024] In the illustrated example, the backend voice system 180 includes an automatic speech recognition (ASR) system 190. The automatic speech recognition (ASR) system 190 uses an automatic speech recognition (ASR) model 192 that processes the enhanced input speech features 250 to generate a speech recognition result (e.g., a transcription) for the target utterance 12. The automatic speech recognition (ASR) system 190 may further include a natural language understanding (NLU) module (not shown) that performs semantic interpretation on the transcription of the target utterance 12 to identify a query / command directed to the device 110. Thus, the output 182 from the backend voice system 180 may include a transcription and / or instructions for achieving the query / command identified by the natural language understanding (NLU) module.
[0025] The backend voice system 180 may alternatively or additionally include a hotword detection model (not shown) configured to detect whether the enhanced input speech features 250 include the presence of one or more hotwords / wakewords trained to be detected by the hotword detection model. For example, the hotword detection model may output a hotword detection score indicating the likelihood that the enhanced input speech features 250 corresponding to the target utterance 12 include a particular hotword / wakeword. The detection of a hotword can trigger a wake-up process that wakes up the device 110 from a sleep state. For example, the device 110 is enabled to wake up and process the hotword and / or one or more terms preceding / following the hotword.
[0026] In an additional example, the background audio system 180 comprises an audio or audio-video calling application (e.g., a video conferencing application). Here, the enhanced input audio feature 250 corresponding to the target utterance 12 is used by the audio or audio-video calling application to filter the voice of the target speaker (10) for communication to the recipient during an audio or audio-video communication session. The background audio system 180 may additionally or alternatively include a speaker identification model configured to identify the user 10 who spoke the target utterance 12 by performing speaker identification using the enhanced input audio feature 250.
[0027] In the illustrated example, device 110 captures a multi-channel noisy input signal 202 (also referred to as audio data) of a target utterance 12 spoken by user 10 in the presence of background interference originating from one or more sources other than user 10. The multi-channel noisy input signal 202 comprises one or more single-channel noisy input signals 206, 206a - 206n of audio. Device 110 may correspond to any computing device associated with user 10 and enabled to receive the multi-channel noisy input signal 202. Some examples of user device 110 include, but are not limited to, mobile devices (e.g., mobile phones, tablets, laptops, etc.), computers, wearable devices (e.g., smartwatches, smart headphones, etc.), smart appliances, Internet of Things (IoT) devices, smart speakers, etc. Device 110 comprises data processing hardware 112 and memory hardware 114 that communicates with data processing hardware 112. Memory hardware 114 stores instructions that cause data processing hardware 112 to perform one or more operations when executed by data processing hardware 112. The multi-channel neural front-end voice enhancement model 200 may be executed on data processing hardware 112. In some examples, a back-end voice system 180 is executed on data processing hardware 112.
[0028] In some examples, device 110 comprises one or more applications (i.e., software applications), and each application may utilize enhanced input voice features 250 generated by the multi-channel neural front-end voice enhancement model 200 to perform various functions within the application. For example, device 110 comprises an assistant application configured to assist user 10 with various tasks by communicating synthesized playback audio 154 to user 10.
[0029] The user device 110 further comprises (or communicates with) an audio subsystem having an array of audio capture devices (e.g., microphones) 116, 116a - 116n for capturing spoken utterances (12) within the acoustic environment and converting them into electrical signals, and an audio output device (e.g., speaker 118) for communicating audible audio signals (e.g., synthesized playback audio 154 from device 110). Each microphone 116 of the array of microphones 116 of the user device 110 is enabled to separately record the utterance (12) on a separate dedicated channel 206 of the multi - channel noisy input signal 202. For example, the user device 110 may include two microphones 116 each recording the utterance (12), and the recordings from the two microphones 116 may be combined to form a two - channel noisy input signal 202 (i.e., stereo audio or stereo). That is, the two microphones are present in the user device 110. In some examples, the user device 110 comprises three or more microphones 116. Additionally or alternatively, the user device 102 may communicate with two or more microphones 116 separate from / remote to the user device 110. For example, the user device 110 may be a mobile device disposed within a vehicle and performing wired or wireless communication (e.g., Bluetooth®) with two or more microphones 116 of the vehicle. In some configurations, the user device 110 communicates with at least one microphone 116 present in a separate device 111, which may include, but is not limited to, an in - vehicle audio system, a computing device, a speaker, or another user device. In these configurations, the separate device 111 may also communicate with one or more microphones 116 present in the user device 110.
[0030] In some examples, device 110 is configured to communicate with remote system 130 via a network (not shown). Remote system 130 may include remote resources 132 such as remote data processing hardware 134 (e.g., a remote server or CPU) and / or remote memory hardware 136 (e.g., a remote database or other storage hardware). User device 110 may utilize remote resources 132 to perform various functions related to audio processing and / or synthetic playback communication. Multi-channel neural front-end audio enhancement model 200 and backend audio system 180 may be present on device 110 (referred to as an on-device system), or may be present remotely while communicating with device 110 (e.g., may be present on remote system 130). In some examples, one or more backend audio systems 180 are present locally or on-device, while one or more other backend audio systems 180 are present remotely. In other words, one or more backend audio systems 180 that utilize enhanced input audio features 250 output from multi-channel neural front-end audio enhancement model 200 can be local or remote in any combination. For example, if the size of system 180 is quite large, or if it is a processing requirement, system 180 may be present on remote system 130. However, if device 110 can support the size or processing requirements of one or more systems 180, one or more systems 180 may be present on device 110 by using data processing hardware 112 and / or memory hardware 114. Optionally, one or more of systems 180 may be present both locally / on-device and remotely. For example, backend audio system 180 is enabled to execute on remote system 130 by default when the connection between device 110 and remote system 130 is available, but when the connection is lost or unavailable, system 180 executes locally on device 110 instead.
[0031] In some embodiments, the device 110, or a system associated with the device 110, identifies text that the device 110 communicates to the user 10 as a response to a query spoken by the user 10. Next, the device 110 uses a text-to-speech (TTS) system to convert the text into synthesized playback audio 154 that the device 110 can communicate to the user 10 as a response to the query (e.g., communicate audibly with the user 10). Once generated, the TTS system enables the device 110 to output the synthesized playback audio 154 by communicating the synthesized playback audio 154 to the device 110. For example, in response to the user 10 making an oral query about today's weather forecast, the device 110 outputs synthesized playback audio 154 saying "It is sunny today" through the speaker 118 of the device 110.
[0032] Continuing to refer to FIG. 1, when the device 110 outputs the synthesized playback audio 154, the synthesized playback audio 154 generates an echo 156 that has been captured by the audio capture device 116. The synthesized playback audio 154 corresponds to the reference audio signal. The synthesized playback audio 154 represents the reference audio signal in the example of FIG. 1, but the reference audio signal may include other types of playback audio 154 such as the media content output from the speaker 118 or the communication from a remote user with whom the user 10 is conversing via the device 110 (e.g., a voice over IP call or a video conference call). Unfortunately, in addition to the echo 156, the audio capture device 116 may also simultaneously capture the target utterance 12 spoken by the user 10 that includes a follow-up query further asking about the weather, starting with "How about tomorrow?". For example, FIG. 1 depicts that when the device 110 outputs the synthesized playback audio 154, the user 10 further asks about the weather with an uttered speech (12) starting with "How about tomorrow?" to the device 110. Here, both the uttered speech (12) and the echo 156 are captured simultaneously by the audio capture device 116, thus forming the multi-channel noisy input signal 202. In other words, the multi-channel noisy input signal 202 includes overlapping audio signals where a part of the target utterance 12 spoken by the user 10 overlaps with a part of the reference audio signal (e.g., the synthesized playback audio 154) that has already been output from the speaker 118 of the device 110. In addition to the synthesized playback audio 154, competing voices 13 spoken by another user 11 in the environment and non-audio characteristics such as the incoming call tone (ringtone) 15 from a separate user device 111 can also be captured by the audio capture device 116, and thus may contribute to the background interference that overlaps with the target utterance 12.
[0033] In FIG. 1, the back-end voice system 180 may have a problem of processing a target utterance 12 corresponding to a follow-up weather query "How about tomorrow?" in a multi-channel noisy input signal 202 due to the presence of background interference that interferes with the target utterance 12. Here, the background interference is attributed to at least one of the playback audio 154, the competing voice 13, or the non-speech ambient noise (non-speech background noise 15). To improve the robustness of the back-end voice system 180 by effectively removing (i.e., masking) the presence of background interference recorded by the device 110 when the user 10 speaks the target utterance 12, a multi-channel neural front-end voice enhancement model 200 is used.
[0034] The model 200 may perform voice enhancement by applying noise context modeling. The voice cleaner 300 of the model 200 processes a multi-channel context noise signal 204 related to a predetermined period of a noise segment captured by the audio capture device 116 before the target utterance 12 is spoken by the user 10. In some examples, the predetermined period comprises a 6-second noise segment. Thus, the multi-channel context noise signal 204 provides a noise context. In some examples, the multi-channel context noise signal 204 comprises LFBE (log-mel filter bank energy) features of the noise context signal for use as context information.
[0035] Figure 2 shows the multi-channel neural front-end voice enhancement model 200 of FIG. 1. The multi-channel neural front-end voice enhancement model 200 uses a modified version of the conformer neural network architecture that combines convolution and self-attention to model short-range and long-range interactions. The multi-channel neural front-end voice enhancement model 200 includes a voice cleaner 300, a feature stack 220, an encoder 230, and a masking layer 240. The voice cleaner 300 may execute an adaptive noise cancellation algorithm (FIG. 3). The encoder 230 may include a stack of self-attention blocks 400.
[0036] The voice cleaner 300 may be configured to receive, as input, a multi-channel noisy input signal 202 and a multi-channel context noise signal 204, and to generate, as output, a single-channel cleaned input signal 340. Here, the voice cleaner 300 includes a finite impulse response (FIR) filter for processing the multi-channel noisy input signal 202.
[0037] FIG. 3 presents an exemplary adaptive noise cancellation algorithm executed by the voice cleaner 300. Here, the voice cleaner 300 includes an FIR module 310 that includes an FIR filter, a minimization module 320, and a cancellation module 330.
[0038] In the illustrated example, for simplicity, the multi-channel noisy input signal 202 has three channels 206a - 206c, each with respective audio features captured by a separate dedicated microphone 116a - 116c of an array of three microphones 116. However, as described above, the front-end voice enhancement model 200 is agnostic with respect to the number of microphones 116 in the array of microphones 116. In other words, the multi-channel noisy input signal 202 can have one channel 206 captured by one microphone 116, two channels 206 captured by two microphones 116, or four or more channels 206 captured by four or more microphones 116, without departing from the scope of the present disclosure.
[0039] Here, the FIR module 310 generates a total output 312 by applying FIR filters to all channels 206 of the multi-channel noisy input signal 202 except the first channel 206a. In other words, the FIR module 310 does not process the first channel 206a of the multi-channel noisy input signal 202, while generating the total output 312 by applying FIR filters to the second channel 206b and the third channel 206c of the multi-channel noisy input signal 202. The minimization module 320 receives the total output 312 and the first channel 206a, and generates a "minimized output" (minimized output) 322 by subtracting the total output 312 from the first channel 206a of the multi-channel noisy input signal 202. Mathematically, the FIR filter comprises three tapped delay lines of length L that are applied to channels 206b, 206c while not being applied to channel 206a. The determination of the minimized output 322 can be expressed as follows.
[0040]
Equation
[0041] Wherein,
[0042] [Number]
[0043] is the input vector for which the time-delay short-time Fourier transform (STFT) processing of channels 206b and 206c has been performed. U m (k) is the vector of filter coefficients applied to channels 206b and 206c.
[0044] [Number]
[0045] and U m (k) can be expressed as follows.
[0046] [Number]
[0047] U m (k) = [U m (k, 0), U m (k, 1), … U m (k, N - 1)] T (3) wherein the filter coefficients can be made to minimize the output power as follows.
[0048] [Number]
[0049] Since the voice cleaner 300 is implemented in the device 110, the cancellation module 330 is enabled to use the multi-channel context noise signal 204 that occurs immediately before the utterance (12) in the multi-channel noisy input signal 202. In other words, when the utterance (12) is not present in the multi-channel noisy input signal 202, the minimization module 320 generates a minimized output 322 through adaptation during the multi-channel context noise signal 204. The adaptation may include a recursive least squares (RLS) algorithm. When the voice cleaner 300 detects the utterance (12), the filter coefficients are fixed, and the cancellation module 330 applies the last coefficients before the utterance (12) to the multi-channel noisy input signal 202 to cancel background interference, thereby generating a single-channel cleaned input signal 340 as follows.
[0050]
Number
[0051] Referring back to FIG. 2, the feature stack 220 receives as inputs the single-channel cleaned input signal 340 and a single channel 206a of the multi-channel noisy input signal 202. And the feature stack 220 is configured to generate a stack input 232. The stack input 232 includes the single-channel cleaned input signal 340 and the single channel 206a. The feature stack 220 can convert each of the single-channel cleaned input signal 340 and the single channel 206a of the multi-channel noisy input signal 202 into a 128-dimensional log-mel domain using a window size of 32 milliseconds (ms) with a step size of 10 ms. Here, four frames can be stacked at a step of 30 ms when input to the feature stack 220.
[0052] The encoder 230 receives a stacked input 232 that includes a single-channel cleaned input signal 340 and a single channel 206a of the multi-channel noisy input signal 202, and generates an unmasked output (output not masked) 480 as an output. The encoder 230 includes a stack of self-attention blocks 400 (also called block 400). Here, the first block (400) of the stack of self-attention blocks 400 receives the stacked input 232. The stacked input 232 includes a single-channel cleaned input signal 340 that has been output from the voice cleaner 300 and a single channel 206 of the multi-channel noisy input signal 202. The final block (400) of the stack of self-attention blocks 400 generates the unmasked output 480.
[0053] Each conformer block (400) may include a (first half) feed-forward layer, a self-attention layer, a convolutional layer (convolution layer), and a second (half) feed-forward layer. In some embodiments, the stack of self-attention blocks 400 includes a stack of conformer blocks (400). In these embodiments, the stack of conformer blocks (400) includes four layers of conformer blocks (400), each having 1024 units, 8 attention heads, a 15×1 convolutional kernel size, and 64-frame self-attention that enables a streaming model. An example of the conformer block (400) is described in more detail below with reference to FIG. 4.
[0054] The masking layer 240 receives, as inputs, the unmasked output 480 already output by the self-attention block 400 of the encoder 230 and a single channel 206a of the multi-channel noisy input signal 202, and is configured to generate, as an output, enhanced input speech features 250 corresponding to the target utterance 12. In some embodiments, the masking layer 240 of the model 200 includes a decoder (not shown) configured to decode the unmasked output 480 into the enhanced input speech features 250 corresponding to the target utterance 12. Here, the decoder may include a simple projection decoder having a single-layer frame-wise fully-connected network with sigmoid activation.
[0055] FIG. 4 presents an example of a block (400) from a stack of self-attention blocks 400 of the encoder 230. In the self-attention block 400, a multi-head self-attention block 420 and a convolutional layer (convolutional layer) 430 are arranged between a first half feed-forward layer 410 and a second half feed-forward layer 440, and include the first half feed-forward layer 410, the second half feed-forward layer 440, and concatenation operators 405, 405a to 405d. The first half feed-forward layer 410 generates an output 412 by processing a stack input 232 including a single-channel cleaned input signal 340 already output from the voice cleaner 300 and a single-channel noisy input signal 206a. Next, the first concatenation operator 405a generates a first concatenated input 414 by concatenating the output 412 with the stack input 232. Subsequently, the multi-head self-attention block 420 generates a noise summary 422 by receiving the first concatenated input 414. Intuitively, the role of the multi-head self-attention block 420 is to separately summarize the noise context for each input frame to be emphasized.
[0056] Next, the second concatenation operator 405b generates a second concatenation input 424 by concatenating the already output noise summary 422 to the first concatenation input 414. Subsequently, the convolutional layer 430 generates a convolutional output 432 by subsampling the second concatenation input 424, which includes the noise summary 422 of the multi-head self-attention block 420, and the first concatenation input 414. Thereafter, the third concatenation operator 405c generates a third concatenation input 434 by concatenating the convolutional output 432 to the second concatenation input 424. The third concatenation input 434 is provided as an input to the second half feed-forward layer 440, and the second half feed-forward layer 440 generates an output 442. The output 442 of the second half feed-forward layer 440 is concatenated to the third concatenation input 434 by the fourth concatenation operator 405d to generate a fourth concatenation input 444. Finally, the layernorm (layer normalization) module 450 processes the fourth concatenation input 444 from the second half feed-forward layer 440. Mathematically, the self-attention block 400 generates an output feature y by transforming the input feature x by using the modulation features m as follows.
[0057]
Number
[0058] The self-attention block 400 generates, as an output, an unmasked output 480. This unmasked output 480 is passed to the next layer of the self-attention block 400. In this way, the inputs (204, 206) are modulated by each of the self-attention blocks 400.
[0059] FIG. 5 shows an exemplary training process 500 for calculating an automatic speech recognition ASR loss 560 when the front-end voice enhancement model 200 is jointly trained with the automatic speech recognition ASR model 192. The training process 500 may be executed on the remote system 130 of FIG. 1. As shown, the training process 500 trains the multi-channel neural front-end voice enhancement model 200 with the training data set 520 by obtaining one or more training data sets 520 stored in the data store 510. The data store 510 may be present in the memory hardware 136 of the remote system 130. Each training data set 520 includes a plurality of training examples (training samples) 530, 530a - 530n. Each training example 530 may include a training utterance 532. Here, only the encoder (540) of the automatic speech recognition ASR model 192 is used to calculate the loss. The automatic speech recognition ASR loss 560 is calculated as the l2 (Euclidean) distance between the output of the automatic speech recognition ASR encoder 540 for the target features (536) of the training utterance 532 and the enhanced input speech features 250. The automatic speech recognition ASR encoder 540 is not updated during the training process 500. Specifically, the training process 500 calculates the automatic speech recognition ASR loss 560 in the following two steps. The first step is to generate a predicted output 522 of the automatic speech recognition ASR encoder 540 for the enhanced input speech features 250 by using the automatic speech recognition ASR encoder 540 of the automatic speech recognition ASR model 192. Here, the automatic speech recognition ASR encoder 540 is configured to receive, as input, the enhanced input speech features 250 predicted by the front-end voice enhancement model 200 for the training utterance 532. The second step is to generate a target output 524 of the automatic speech recognition ASR encoder 540 for the target speech features 536 by using the automatic speech recognition ASR encoder 540 configured to receive, as input, the target speech features 536 of the training utterance 532.The predicted output 522 of the enhanced input speech feature 250 and the target output 524 of the target speech feature 536 may each include a sequence of respective LFBE (Log-Mel Filter Bank Energy) features. Thereafter, the training process 500 calculates an Automatic Speech Recognition (ASR) loss 560 via a loss module 550, based on the predicted output 522 of the ASR encoder 540 of the enhanced input speech feature 250 and the target output 524 of the ASR encoder 540 of the target speech feature 536. The goal of using the ASR loss 560 is to tune the enhancement of the front-end speech enhancement model 200 to be closer to the ASR model 192. This is important for obtaining the best performance from the front-end speech enhancement model 200. By keeping the parameters of the ASR model 192 fixed, the ASR model 192 is decoupled from the front-end speech enhancement model 200. Thus, each can be trained and deployed independently of each other.
[0060] In some embodiments, the front-end speech enhancement model 200 is co-trained with the ASR model 192 of the back-end ASR system 180 by using a spectral loss and the ASR loss 560. The training target (536) for training the multi-channel neural front-end speech enhancement model 200 uses an Ideal Ratio Mask (IRM). The Ideal Ratio Mask IRM can be calculated by using reverberant speech and reverberant noise, based on the assumption that there is no correlation between speech and noise in the Mel spectrum space, as follows.
[0061]
Number
[0062] Here, X and N are the Mel spectrograms of the reverberant speech and reverberant noise, respectively. t and f represent the time and Mel frequency bin indices. The selection for estimating the ideal ratio mask IRM is based on a target restricted between [0,1], which simplifies the estimation process. Further, the automatic speech recognition ASR model 192 used for evaluation can be trained with actual and simulated reverberant data. As a result, a trained automatic speech recognition ASR model 192 that is relatively robust to reverberant speech is obtained. Therefore, the ideal ratio mask IRM derived by using reverberant speech as the target still brings a significant improvement in performance. The spectral loss L during training can be calculated based on the L1 loss and L2 loss between the ideal ratio mask IRM and the estimated ideal ratio mask IRM as
[0063]
Number
[0064] and.
[0065]
Number
[0066] During inference, the estimated ideal ratio mask IRM is scaled and floored to reduce speech distortion at the expense of reducing noise suppression. The automatic speech recognition ASR model 192 is susceptible to the effects of speech distortion and non-linear front-end processing, which is one of the main issues in improving the performance of a robust automatic speech recognition ASR model by using an enhanced front-end. Thus, this is particularly important. The enhanced features can be derived as follows.
[0067]
Number
[0068] Here, Y is the noisy Mel spectrogram.
[0069] [Number]
[0070] is an estimate of the clean Mel spectrogram. α and β are the exponential mask scalar and the mask floor. In some examples, α is set to 0.5. β is set to 0.01. The highlighted features are log-compressed and (i.e.,
[0071] [Number]
[0072] For evaluation, it can be passed to the automatic speech recognition ASR model 192. FIG. 6 includes a flowchart of an exemplary configuration of operations for method 600. Method 600 performs automatic speech recognition using a multi-channel neural front-end voice enhancement model (200). In operation 602, method 600 includes receiving a multi-channel noisy input signal 202 and a multi-channel context noise signal 204. Method 600 also includes, in operation 604, generating a single-channel cleaned input signal 340 by using the voice cleaner 300 of the voice enhancement model 200.
[0073] In operation 606, method 600 also includes generating an unmasked output 480 as an output from a stack of the self-attention block 400 of the voice enhancement model 200 configured to receive a stack input 232. Here, the stack input 232 includes a single-channel cleaned input signal 340 already output from the voice cleaner 300 and a single-channel noisy input signal 206. Here, each self-attention block 400 of the stack of self-attention blocks 400 includes a multi-head self-attention mechanism (self-attention mechanism). In operation 608, method 600 further includes generating enhanced input voice features 250 corresponding to the target utterance 12 by using the masking layer 240 of the voice enhancement model 200. Here, the masking layer 240 is configured to receive the single-channel noisy input signal 206 and the unmasked output 480 already generated as an output from the stack of self-attention blocks 400.
[0074] FIG. 7 is a schematic diagram of an exemplary computing device 700 that can be used to implement the systems and methods described herein. Computing device 700 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The components shown here, their connections and relationships, and their functions are for illustrative purposes only and are not intended to limit the embodiments of the present disclosure described and / or claimed in this document.
[0075] The computing device 700 includes a processor 710, a memory 720, a storage device 730, a high-speed interface / controller (740) connected to the memory 720 and the high-speed expansion port 750, and a low-speed interface / controller (760) connected to the low-speed bus 770 and the storage device 730. Each of the components (710, 720, 730, 740, 750, and 760) is interconnected by using various buses and may be installed on a common motherboard or exist in other ways as needed. The processor 710 (e.g., the data processing hardware 112, 134 in FIG. 1) processes instructions for execution within the computing device 700 that include instructions stored in the memory 720 or the storage device 730, enabling the display of graphical information of a graphical user interface (GUI) on an external input / output device such as a display 780 connected to the high-speed interface (740). In other embodiments, multiple memories and multiple types of memories, along with multiple processors and / or multiple buses as needed, may be used. Also, multiple computing devices 700 may be connected such that each device provides a portion of the required operations (e.g., as a server bank, a group of blade servers, or a multiprocessor system).
[0076] Memory 720 (e.g., memories 114, 136 in FIG. 1) stores information non-temporarily inside computing device 700. Memory 720 may be a computer-readable medium, a volatile memory unit(s), or a non-volatile memory unit(s). The non-temporary memory 720 may be a physical device used to store a program (e.g., an instruction sequence) or data (e.g., program state information) temporarily or permanently for use by computing device 700. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as a boot program). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.
[0077] Storage device 730 is enabled to provide large-capacity storage to computing device 700. In some embodiments, storage device 730 is a computer-readable medium. In various different embodiments, storage device 730 may be a floppy (registered trademark) disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid-state memory device, or an array of devices comprising a storage area network or other configuration of devices. In additional embodiments, a computer program product is tangibly embodied in an information carrier. The computer program product comprises instructions that, when executed, perform one or more of the methods as described above. The information carrier is a computer-readable or machine-readable medium such as memory 720, storage device 730, or the memory of processor 710.
[0078] High-speed controller 740 manages the bandwidth-intensive operations of computing device 700, and low-speed controller 760 manages the lower bandwidth-intensive operations. Such role assignments are merely examples. In some embodiments, high-speed controller 740 is coupled to high-speed expansion port 750 that can accept memory 720, display 780 (e.g., via a graphics processor or accelerator), and various expansion cards (not shown). In some embodiments, low-speed controller 760 is coupled to storage device 730 and low-speed expansion port 790. Low-speed expansion port 790 may include various communication ports (such as USB, Bluetooth (registered trademark), Ethernet (registered trademark), wireless Ethernet (registered trademark), etc.) and may be connected to one or more input / output devices such as a keyboard, a pointing device, a scanner, or a network device such as a switch or router via, for example, a network adapter.
[0079] The computing device 700 may be implemented in a plurality of different forms as shown in the figures. For example, it may be implemented as a standard server 700a, or multiple times within a group of such servers (700a), as a laptop computer 700b, or as part of a rack server system 700c.
[0080] The various embodiments of the systems and techniques described herein can be realized in digital and / or optical circuits, integrated circuits, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can be in one or more computer programs executable and / or interpretable in a programmable system comprising at least one programmable processor, at least one input device, and at least one output device, which are special or general purpose and are coupled to receive data and instructions from, and to transmit data and instructions to, a storage system.
[0081] A software application (i.e., a software resource) may refer to computer software that causes a computing device to perform a task. In some examples, a software application may be referred to as an "application", an "app", or a "program". Exemplary applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and game applications.
[0082] A non-transitory memory may be a physical device used to store programs (e.g., instruction sequences) or data (e.g., program state information) temporarily or persistently for use by a computing device. The non-transitory memory may be a volatile and / or non-volatile addressable semiconductor memory. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as a boot program). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disks or tapes.
[0083] These computer programs (also known as programs, software, software applications, or code) comprise machine instructions for a programmable processor and can be implemented in high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., magnetic disks, optical disks, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor comprising a machine-readable medium that receives the machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0084] The processes and logical flows described in this specification can be implemented by one or more programmable processors, also referred to as data processing hardware, executing one or more computer programs that act on input data and generate output to perform functions. The processes and logical flows can also be executed by special purpose logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). Processors suitable for the execution of a computer program include, by way of example, both general and special purpose processors, as well as any one or more processors of any kind of digital computer. Generally, a processor receives instructions and data from a read only memory, a random access memory, or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer also includes, or is operatively coupled to receive from or transfer data to, one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including by way of example semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0085] To interact with a user, one or more aspects of the present disclosure can be implemented on a computer having a display device, such as a CRT (cathode ray tube), LCD (liquid crystal display monitor), or touch screen, for displaying information to the user, and an optional keyboard and pointing device, such as a mouse or trackball, by which the user can input to the computer. Other types of devices can also be used to interact with the user. For example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and the input from the user can be received in any form, such as acoustic input, voice input, or tactile input. Further, the computer can interact with the user by sending and receiving documents to and from the devices used by the user, such as by sending a web page to a web browser on the user's client device in response to a received request from the web browser.
[0086] Some embodiments have been described. Nevertheless, it is understood that various modifications can be made without departing from the spirit and scope of the present disclosure. Accordingly, other embodiments are within the scope of the following claims.
Claims
1. An audio enhancement model (200) as a multi-channel neural front-end audio enhancement model (200) for speech recognition, wherein the audio enhancement model (200) causes a computer to, An audio cleaner (300), Receives, as inputs, a multi-channel noisy input signal (202) and a multi-channel context noise signal (204), and The audio cleaner (300) configured to generate, as an output, a single-channel cleaned input signal (340); and A stack of self-attention blocks (400), each having a multi-head self-attention mechanism, wherein the stack of self-attention blocks (400) Receives, as an input, a stack input (232) at a first block (400) of the stack of self-attention blocks (400), the stack input (232) comprising the single-channel cleaned input signal (340) output from the audio cleaner (300) and a single-channel noisy input signal (206), while The stack of self-attention blocks (400) configured to generate, as an output from a final block (400) of the stack of self-attention blocks (400), an unmasked output (480); and A masking layer (240), Receives, as inputs, the single-channel noisy input signal (206) and the unmasked output (480) generated as an output from the final block (400) of the stack of self-attention blocks (400), while The masking layer (240) configured to generate, as an output, an enhanced input audio feature (250) corresponding to a target utterance (12); and For causing to function, The audio enhancement model (200).
2. The stack of self-attention blocks (400) comprises a stack of conformer blocks (400), The audio enhancement model (200) according to claim 1.
3. The stack of conformer blocks (400) comprises four of the conformer blocks (400), The audio enhancement model (200) according to claim 2.
4. The audio enhancement model (200) is executed by data processing hardware (112) as the computer present in a user device (110), The user device (110) is configured to capture the target utterance (12) and the multi-channel context noise signal (204) via an array of microphones (116) of the user device (110). The voice enhancement model (200) according to any one of claims 1 to 3.
5. The voice enhancement model (200) is agnostic with respect to the number of microphones (116) of the array of microphones (116). The voice enhancement model (200) according to claim 4.
6. The voice cleaner (300) generating a total output (312) by applying a finite impulse response (FIR) filter to all channels (206) of the multi-channel noisy input signal (202) except the first channel (206) of the multi-channel noisy input signal (202); subtracting the total output (312) from the first channel (206) of the multi-channel noisy input signal (202); executing an adaptive noise cancellation algorithm to generate a single-channel cleaned input signal (340). The voice enhancement model (200) according to any one of claims 1 to 3.
7. The backend voice system (180) is configured to process the enhanced input voice feature (250) corresponding to the target utterance (12). The voice enhancement model (200) according to any one of claims 1 to 3.
8. The backend voice system (180) comprises at least one of an automatic speech recognition (ASR) model (192), an audio call application, or an audio-video call application. The voice enhancement model (200) according to claim 7.
9. The voice enhancement model (200) is jointly trained with an automatic speech recognition ASR model (192) as a backend automatic speech recognition (ASR) model (192) for causing the computer to execute backend automatic speech recognition (ASR) by using a spectral loss and an automatic speech recognition ASR loss (560). The voice enhancement model (200) according to any one of claims 1 to 3.
10. The spectral loss is based on the distances of the L1 loss function and the L2 loss function between the estimated ratio mask and the ideal ratio mask, The ideal ratio mask is calculated by using the reverberant speech and the reverberant noise, The voice enhancement model (200) according to claim 9.
11. The automatic speech recognition ASR loss (560) causes the computer to, generating a predicted output (522) of the automatic speech recognition ASR encoder (540) of the enhanced speech features (250) by using the automatic speech recognition ASR encoder (540) of the automatic speech recognition ASR model (192) configured to receive, as an input, the enhanced speech features (250) predicted by the voice enhancement model (200) for the training utterance (532); generating a target output (524) of the automatic speech recognition ASR encoder (540) of the target speech features (536) by using the automatic speech recognition ASR encoder (540) configured to receive, as an input, the target speech features (536) of the training utterance (532); and calculating the automatic speech recognition ASR loss (560) based on the predicted output (522) of the automatic speech recognition ASR encoder (540) of the enhanced speech features (250) and the target output (524) of the automatic speech recognition ASR encoder (540) of the target speech features (526); calculated by causing to execute, The voice enhancement model (200) according to claim 9.
12. A computer-implemented method (600) that, when executed on data processing hardware (112, 134), causes the data processing hardware (112, 134) to perform operations, the operations including: receiving a multi-channel noisy input signal (202) and a multi-channel context noise signal (204); generating a single-channel cleaned input signal (340) by using a voice cleaner (300) of a voice enhancement model (200); As an output from a stack of self-attention blocks (400) of the voice enhancement model (200) configured to receive a stack input (232), a step of generating an unmasked output (480), wherein the stack input (232) includes a single-channel cleaned input signal (340) already output from the voice cleaner (300) and a single-channel noisy input signal (206), and each self-attention block (400) of the stack of self-attention blocks (400) includes a multi-head self-attention mechanism, the step of generating the unmasked output (480), and, A step of generating enhanced input voice features (250) corresponding to a target utterance (12) by using a masking layer (240) of the voice enhancement model (200) configured to receive the single-channel noisy input signal (206) and the unmasked output (480) already generated as an output from a stack of self-attention blocks (400); A computer-implemented method (600) comprising the above.
13. The stack of self-attention blocks (400) includes a stack of conformer blocks (400). The computer-implemented method (600) according to claim 12.
14. The stack of conformer blocks (400) includes four of the conformer blocks (400). The computer-implemented method (600) according to claim 13.
15. The voice cleaner (300), the stack of self-attention blocks (400), and the masking layer (240) are executed by the data processing hardware (112), The data processing hardware (112) is present in a user device (110), The user device (110) is configured to capture the target utterance (12) and the multi-channel context noise signal (204) via an array of microphones (116) of the user device (110). The computer-implemented method (600) according to any one of claims 12 to 14.
16. The voice enhancement model (200) is agnostic with respect to the number of microphones (116) in the array of microphones (116). The computer-implemented method (600) according to claim 15.
17. The operation further includes using the voice cleaner (300) to generate a total output (312) by applying finite impulse response (FIR) filters to all channels (206) of the multi-channel noisy input signal (202) except the first channel (206) of the multi-channel noisy input signal (202), and executing an adaptive noise cancellation algorithm to generate a single-channel cleaned input signal (340) by subtracting the total output (312) from the first channel (206) of the multi-channel noisy input signal (202), and A computer-implemented method (600) according to any one of claims 12 to 14.
18. The backend voice system (180) is configured to process the enhanced input voice feature (250) corresponding to the target utterance (12). A computer-implemented method (600) according to any one of claims 12 to 14.
19. The backend voice system (180) comprises at least one of an automatic speech recognition (ASR) model (192), an audio call application, or an audio-video call application. A computer-implemented method (600) according to claim 18.
20. The voice enhancement model (200) is jointly trained with an automatic speech recognition (ASR) model (192) as the backend automatic speech recognition (ASR) model (192) by using a spectral loss and an automatic speech recognition ASR loss (560). A computer-implemented method (600) according to any one of claims 12 to 14.
21. The spectral loss is based on the distance between the estimated ratio mask and the ideal ratio mask using the L1 loss function and the L2 loss function, The ideal ratio mask is calculated by using reverberant speech and reverberant noise. A computer-implemented method (600) according to claim 20.
22. The automatic speech recognition ASR loss (560) is By using the automatic speech recognition ASR encoder (540) of the automatic speech recognition ASR model (192) configured to receive, as input, the emphasized speech features (250) predicted by the speech emphasis model (200) for the training utterance (532), a step of generating a prediction output (522) of the automatic speech recognition ASR encoder (540) for the emphasized speech features (250); By using the automatic speech recognition ASR encoder (540) configured to receive, as input, the target speech features (536) of the training utterance (532), a step of generating a target output (524) of the automatic speech recognition ASR encoder (540) for the target speech features (536); and A step of calculating the automatic speech recognition ASR loss (560) based on the prediction output (522) of the automatic speech recognition ASR encoder (540) for the emphasized speech features (250) and the target output (524) of the automatic speech recognition ASR encoder (540) for the target speech features (536); Calculated by; The computer-implemented method (600) according to claim 20.
Citation Information
Patent Citations
System and method for acoustic echo cancellation using deep multitask recurrent neural networks
US20200312346A1
Multistream acoustic models with dilations
US20210005182A1
Automated meeting minutes generator
US20210375289A1
Audio processing apparatus and method for denoising a multi-channel audio signal
WO2021013345A1