Combined Acoustic Echo Cancellation, Voice Enhancement, and Voice Separation for Automatic Speech Recognition
The context front-end processing model addresses the challenges of background interference in ASR systems by jointly implementing AEC, voice enhancement, and voice separation within a single neural network architecture, significantly enhancing the robustness and accuracy of speech recognition.
Patent Information
- Application Number
- JP2024508062
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-08-09
- Filing Date
- 2021-12-14
- Publication Date
- 2025-06-30
- Estimated Expiration
- 2041-12-14
AI Technical Summary
Existing automatic speech recognition (ASR) systems face challenges in accurately recognizing speech due to background interference such as device echo, background noise, and competing voices, which are typically addressed by separate modeling strategies that are not practical to switch between on-the-fly.
A context front-end processing model that jointly implements acoustic echo cancellation (AEC), voice enhancement, and voice separation within a single model, using a conformer neural network architecture that incorporates a primary encoder, a noise context encoder, a cross-attention encoder, and a decoder, and utilizes feature-wise linear modulation (FiLM) to combine input audio features with speaker embeddings.
The model effectively removes background interference, enhancing the quality of input speech features and improving the robustness of ASR systems by integrating echo cancellation, noise reduction, and speaker separation in a unified processing framework.
Smart Images

Figure 0007700365000010 
Figure 0007700365000011 
Figure 0007700365000012
Abstract
Description
Technical Field
[0001] The present disclosure relates to joint acoustic echo cancelation, voice enhancement, and voice separation for automatic speech recognition.
Summary of the Invention
Means for Solving the Problems
[0002] One aspect of the present disclosure provides a computer-implemented method for automatic speech recognition using joint acoustic echo cancelation, voice enhancement, and voice separation. When executed on data processing hardware, the computer-implemented method causes the data processing hardware to perform operations including receiving, in a context front-end processing model, input speech features corresponding to a target utterance. The operations further include receiving, in the context front-end processing model, at least one of a reference audio signal, a context noise signal including noise prior to the target utterance, or a speaker embedding including voice characteristics of a target speaker who uttered the target utterance. The operations further include processing, using the context front-end processing model, the input speech features and at least one of the reference audio signal, the context noise signal, or the speaker embedding vector to generate enhanced speech features.
[0003] Implementations of the present disclosure may include one or more of the following optional features. In some implementations, the context front-end processing model includes a conformer neural network architecture that combines convolution and self-attention to model short- and long-distance conversations. In some examples, the step of processing the input audio features and at least one of the reference audio signal, the context noise signal, or the embedding vector includes using a primary encoder to process the input audio features to generate a primary input encoding, and using a noise context encoder to process the context noise signal to generate a context noise encoding. These examples further include using a cross-attention encoder to process the primary input encoding and the context noise encoding to generate a cross-attention embedding, and decoding the cross-attention embedding into enhanced input audio features corresponding to the target utterance. In these examples, the step of processing the input audio features to generate a primary encoding may further include stacking the reference features corresponding to the reference audio signal to the input audio features to generate a primary input encoding. The input audio features and the reference features may each include respective sequences of log Mel-filterbank energy (LFBE) features.
[0004] In these examples, the step of processing the input features to generate the primary input encoding may include the step of using feature-wise linear modulation (FiLM) to combine the input audio features with the speaker embedding vector to generate the primary input encoding. Here, the step of processing the primary input encoding and the context noise encoding to generate the cross-attention embedding includes the step of using FiLM to combine the primary input encoding with the speaker embedding vector to generate the modulated primary input encoding, and the step of processing the modulated primary input encoding and the context noise encoding to generate the cross-attention embedding. Additionally or alternatively, the first encoder may include N modulated conformer blocks, the context noise encoder may include N conformer blocks, may be executed in parallel with the first encoder, and the cross-attention encoder may include M modulated cross-attention conformer blocks.
[0005] In some implementations, the data processing hardware executes a context front-end processing model and is present on the user device. The user device is configured to output the reference audio signal as playback audio via the user device's audio speaker and to capture the target utterance, the reference audio signal, and the context noise signal via one or more microphones of the user device. In some examples, the context front-end processing model is trained using spectral loss and ASR loss in conjunction with a back-end automatic speech recognition (ASR) model. In these examples, the spectral loss may be based on the L1 and L2 loss function distances between the estimated ratio mask and the ideal ratio mask. Here, the ideal ratio mask is calculated using the reverberant speech and the reverberant noise.
[0006] Furthermore, in these examples, the ASR loss can be calculated by using the ASR encoder of the ASR model, configured to receive as input the prosody speech features predicted by the context front-end processing model for the training utterance, to generate the predicted output of the ASR encoder for the prosody speech features, using the ASR encoder configured to receive as input the target speech features for the training utterance, to generate the target output of the ASR encoder for the target speech features, and calculating the ASR loss based on the predicted output of the ASR encoder for the prosody speech features and the target output of the ASR encoder for the target speech features. In some implementations, the operation further includes processing, using a backend speech system, the prosody input speech features corresponding to the target utterance. In these implementations, the backend speech system can include at least one of an automatic speech recognition (ASR) model, a hotword detection model, or an audio or audio-video call application.
[0007] Another aspect of the present disclosure provides a context front-end processing model for automatic speech recognition using combined acoustic echo cancellation, speech enhancement, and voice separation, including a primary encoder, a noise context encoder, a cross-attention encoder, and a decoder. The primary encoder receives, as input, input speech features corresponding to a target utterance and generates, as output, a primary input encoding. The noise context encoder receives, as input, a context noise signal including noise preceding the target utterance and generates, as output, a context noise encoding. The cross-attention encoder receives, as input, the primary input encoding generated as output from the primary encoder and the context noise encoding generated as output from the noise context encoder, and generates, as output, a cross-attention embedding. The decoder decodes the cross-attention embedding into prosody input speech features corresponding to the target utterance.
[0008] This aspect may include one or more of the following optional features. In some examples, the primary encoder is further configured to receive, as input, a reference feature corresponding to a reference audio signal, and generate, as output, a primary input encoding by processing the reference feature with the stacked input speech features. The input speech features and the reference features may each include respective sequences of log mel filter bank energy (LFBE) features. In some implementations, the primary encoder is further configured to receive, as input, a speaker embedding including voice characteristics of a target speaker who uttered the target utterance, and generate, as output, a primary input encoding by combining the input speech features with the speaker embedding using feature-wise linear modulation (FiLM).
[0009] In some examples, the cross-attention encoder is further configured to receive, as input, a primary input encoding modulated by a speaker embedding, where the speaker embedding includes voice characteristics of a target speaker who uttered the target utterance, and process the primary input encoding modulated by the speaker embedding and a context noise encoding to generate, as output, a cross-attention embedding. In some implementations, the primary encoder includes N modulated conformer blocks, the context noise encoder includes N conformer blocks and runs in parallel with the primary encoder, and the cross-attention encoder includes M modulated cross-attention conformer blocks. In some examples, the context front-end processing model runs on data processing hardware present on the user device. Here, the user device is configured to output the reference audio signal as playback audio via an audio speaker of the user device, and capture the target utterance, the reference audio signal, and the context noise signal via one or more microphones of the user device.
[0010] In some implementations, the context front-end processing model is trained using spectral loss and ASR loss in conjunction with a back-end automatic speech recognition (ASR) model. In these implementations, the spectral loss can be based on the L1 and L2 loss function distances between the estimated ratio mask and the ideal ratio mask. Here, the ideal ratio mask is calculated using reverberant speech and reverberant noise. Further, in these implementations, the ASR loss is calculated by receiving, as input, the enhanced speech features predicted by the context front-end processing model for the training utterance, generating the predicted output of the ASR encoder for the enhanced speech features, receiving, as input, the target speech features for the training utterance using an ASR encoder configured to receive the target speech features for the training utterance, generating the target output of the ASR encoder for the target speech features, and calculating the ASR loss based on the predicted output of the ASR encoder for the enhanced speech features and the target output of the ASR encoder for the target speech features. In some examples, the back-end voice system is configured to process the enhanced input speech features corresponding to the target utterance. In these implementations, the back-end voice system can include at least one of an automatic speech recognition (ASR) model, a hot-word detection model, or an audio or audio-video call application.
[0011] Details of one or more implementations of the present disclosure are set forth in the accompanying drawings and the following description. Other aspects, features, and advantages will be apparent from the description and drawings, and from the claims.
Background Art
[0012] The robustness of automatic speech recognition (ASR) systems has improved significantly over the years with the emergence of neural network-based end-to-end models, large-scale training data, and improved strategies for augmenting training data. Nevertheless, various conditions, such as echo, more intrusive background noise, and competing voices, can significantly degrade the performance of ASR systems. Separate ASR models may be trained to handle these conditions individually, but it is not practical to maintain multiple task / condition-specific ASR models and switch between them on-the-fly during use.
Brief Description of the Drawings
[0013]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Modes for Carrying Out the Invention
[0014] Like reference numerals in the various drawings indicate like elements.
[0015] The robustness of automatic speech recognition (ASR) systems has improved significantly over the years with the emergence of neural network-based end-to-end models, large-scale training data, and improved strategies for augmenting training data. Nevertheless, background interference can significantly degrade the ability of an ASR system to accurately recognize speech directed at the ASR system. Background interference can be broadly classified into three groups, namely, device echo, background noise, and competing voices. Separate ASR models may be trained to handle each of these background interference groups independently, but it is not practical to maintain multiple task / condition-specific ASR models and switch between models on the fly during use.
[0016] Device echo may well correspond to the playback audio output from a device such as a smart home speaker, and by doing so, the playback audio may be recorded as an echo and can affect the performance of a backend voice system such as an ASR system. Specifically, the degradation of the performance of the backend voice system is particularly severe when the playback audio includes audible speech, for example, a text-to-speech (TTS) response from a digital assistant. This problem is typically addressed by acoustic echo cancellation (AEC) techniques. The unique characteristic of AEC is that a reference signal corresponding to the playback audio is usually available and can be used for suppression.
[0017] Background noise with non-speech characteristics is generally well handled using data augmentation strategies such as multi-style training (MTR) of ASR models. Here, an indoor simulator is used to add noise to the training data, and this data is then carefully weighted with clean data during training so that a performance balance is achieved between clean and noisy conditions. As a result, large-scale ASR models are robust to moderate levels of non-speech noise. However, background noise can still affect the performance of the backend speech system when low signal-to-noise ratio (SNR) conditions exist.
[0018] Unlike non-speech background noise, competing speech is quite challenging for an ASR model trained to recognize a single speaker. Training an ASR model with multi-talker speech itself can pose problems because it is difficult to clarify which speaker to focus on during inference. Using a model that recognizes multiple speakers is also sub-optimal because it is difficult to know in advance how many users to support. Furthermore, such multi-speaker models usually result in a degradation of performance in a single-speaker setting, which is undesirable.
[0019] The three above-described classes of background interference are typically addressed independently of each other and each uses a separate modeling strategy. Voice separation has received attention in recent literature using techniques such as deep clustering, permutation invariant training, and the use of speaker embeddings. When using speaker embeddings, the target speaker is assumed to be known in advance. Techniques developed for speaker separation have also been applied to remove non-speech noise and corrections are made to the training data. AEC has also been studied independently or in combination when background noise is present. It is well known that improving voice quality does not always improve ASR performance because distortion introduced by non-linear processing can negatively impact ASR performance. One way to mitigate this is to co-train the front-end emphasis with the back-end ASR model.
[0020] Moreover, the application of large-scale multi-domain and multilingual ASR models continues to attract interest, but the training data for these ASR models typically cover a variety of acoustic and language use cases (e.g., voice search and video captioning), which makes it difficult to simultaneously address more intrusive noise conditions. As a result, it is often convenient to train and maintain separate front-end feature processing models capable of handling adverse conditions without combining them with the back-end ASR model.
[0021] The embodiments herein are directed to a context front-end processing model for improving the robustness of ASR by jointly implementing acoustic echo cancellation (AEC), voice enhancement, and voice separation modules within a single model. Particularly in streaming ASR settings, a single combined model is realistic given the position that it is difficult, if not impossible, to know in advance what classes of background interference to handle. Specifically, the context front-end processing model includes a context enhancement neural network (CENN) that can optionally utilize three different types of side context inputs, namely, a reference signal associated with the playback audio, a noise context, and a speaker embedding representing the voice characteristics of the target speaker. As will become apparent, the reference signal associated with the playback audio is necessary to provide echo cancellation, and the noise context is useful for voice enhancement. Further, the speaker embedding representing the voice characteristics of the target speaker (when available) is not only important for voice separation but also useful for echo cancellation and voice enhancement. For voice enhancement and separation, the noise context, i.e., the audio for a few seconds prior to the target utterance to be recognized, carries useful information about the acoustic context. The CENN can generate enhanced input speech features that can be passed to a back-end speech system such as an ASR model that can process the enhanced input speech features to generate a speech recognition result for the target utterance, using respective neural network architectures configured to incorporate each corresponding context side input. In particular, since the noise context and the reference features are optional context side inputs, the noise context and the reference features are assumed by the CENN to be silence signals that do not carry their respective information when not available.
[0022] Referring to FIG. 1, in some implementations, the audio environment 100 includes the user 10 communicating the spoken target utterance 12 to the voice-responsive user device 110 (also referred to as the device 110 or the user device 110). The user 10 (i.e., the speaker of the utterance 12) may speak the target utterance 12 as an inquiry or command for a response from the device 110. The device 110 is configured to capture sound from one or more users 10, 11 within the audio environment 100. Here, audible sound may refer to the spoken utterance 12 by the user 10 that functions as an audible inquiry, a command for the device 110, or an audible communication captured by the device 110. The voice-responsive system of or associated with the device 110 may handle the inquiry for the command by responding to the inquiry and / or causing the command to be executed.
[0023] Various types of background interference can interfere with the ability of the backend voice system 180 to process the target utterance 12 that specifies a query or command for the device 110. As described above, the background interference can include a device echo corresponding to the playback audio 154 output from the user device (e.g., smart speaker) 110, competing voices 13 such as voices 13 other than the target utterance 12 spoken by one or more other users 11 that are not directed at the user device 110, and background noise having non - voice characteristics. Implementations herein execute on the user device 110 and receive, as inputs, the input voice features corresponding to the target utterance 12 and one or more context signals 213, 214, 215, and generate, as an output, an enhanced input voice feature 250 corresponding to the target utterance 12 by processing the input voice feature 212 and one or more contexts 213, 214, 215, using a context front - end processing model 200. The backend voice system 180 can process the enhanced voice feature 250 to generate an output 182. In particular, the context front - end processing model 200 effectively removes the presence of background interference recorded by the device 110 when the user 10 speaks the target utterance 12 such that the enhanced voice feature 250 provided to the backend voice system 180 conveys the voice directed at the device 110 (i.e., the target utterance 12), and by doing so, the backend voice system 180 is not degraded by the background interference.
[0024] In the illustrated example, the backend voice system 180 includes an ASR system that utilizes an ASR model to process the enhanced input voice features 250 to generate a speech recognition result (e.g., a transcription) for the target utterance 12. The ASR system may further include a natural language understanding (NLU) module that performs semantic interpretation on the transcription of the target utterance 12 to identify a query / command directed to the user device 110. Thus, the output 180 from the backend voice system 180 may include a transcription and / or instructions to satisfy the query / command identified by the NLU module.
[0025] As an addition or an alternative, the backend voice system 180 may include a hotword detection model configured to detect whether the enhanced input voice features 250 include the presence of one or more hotwords / wakewords that the hotword detection model is trained to detect. For example, the hotword detection model may output a hotword detection score indicating the likelihood that the enhanced input voice features 250 corresponding to the target utterance 12 include a particular hotword / wakeword. The detection of the hotword may trigger an activation process to wake up the device 110. For example, the device 110 may wake up and process the hotword and / or one or more terms preceding / following the hotword.
[0026] In an additional example, the background voice system 180 includes an audio or audio-video call application (e.g., a video conferencing application). Here, the enhanced input voice features 250 corresponding to the target utterance 12 are used by the audio or audio-video call application to filter the voice of the target speaker 10 for communication to the recipient during an audio or audio-video communication session. As an addition or an alternative, the background voice system 180 may include a speaker identification model configured to perform speaker identification using the enhanced input voice features 250 to identify the user 10 who uttered the target utterance 12.
[0027] In the illustrated example, user device 110 captures a noisy audio signal 202 (also referred to as audio data) of target utterance 12 spoken by user 10 in the presence of background interference generated from one or more sources other than user 10. Device 110 may correspond to any computing device that is associated with user 10 and is capable of receiving the noisy audio signal 202. Some examples of user device 110 include, but are not limited to, mobile devices (e.g., mobile phones, tablets, laptops, etc.), computers, wearable devices (e.g., smartwatches), smart appliances, and Internet of Things (IoT) devices, smart speakers, etc. Device 110 includes data processing hardware 112 and memory hardware 114 that communicates with data processing hardware 112 and stores instructions that, when executed by data processing hardware 112, cause data processing hardware 112 to perform one or more operations. Context front-end processing model 200 may execute on data processing hardware 112. In some examples, backend voice system 180 executes on data processing hardware 112.
[0028] In some examples, device 110 includes one or more applications (i.e., software applications), and each application may use the enhanced input audio feature 250 generated by context front-end processing model 200 to perform various functions within the application. For example, device 110 includes an assistant application configured to communicate synthesized playback audio 154 to user 10 to assist with various tasks of user 10.
[0029] Device 110 further includes an audio subsystem having an audio capture device (e.g., a microphone) 116 for capturing and converting spoken utterances 12 within the acoustic environment 100 into an electrical signal, and an audio output device (e.g., an audio speaker) 118 for conveying an audible audio signal (e.g., a synthesized playback signal 154 from device 110). Device 110 implements a single audio capture device 116 in the illustrated example, but device 110 may implement an array of audio capture devices 116 without departing from the scope of the present disclosure, whereby one or more of the audio capture devices 116 of the array may communicate with the audio subsystem (e.g., a peripheral device of device 110) although not physically on device 110. For example, device 110 may correspond to a vehicle infotainment system that utilizes an array of microphones positioned throughout the vehicle.
[0030] In some examples, device 110 is configured to communicate with a remote system 130 via a network (not shown). Remote system 130 may include remote resources 132 such as remote data processing hardware 134 (e.g., a remote server or CPU) and / or remote memory hardware 136 (e.g., a remote database or other storage hardware). Device 110 may use remote resources 132 to perform various functionalities related to audio processing and / or synthesized playback communication. Context front-end processing model 200 and back-end audio system 180 may be present on device 110 (referred to as an on-device system), may be present remotely (e.g., on remote system 130), but communicate with device 110. In some examples, one or more back-end audio systems 180 are present locally or on-device, and one or more other back-end audio systems 180 are present remotely. In other words, one or more back-end audio systems 180 that utilize the emphasized input audio features 250 output from context front-end processing model 200 may be local or remote in any combination. For example, when system 180 is fairly large in size or processing requirements, system 180 may be present in remote system 130. However, when device 110 can support the size or processing requirements of one or more systems 180, one or more systems 180 may be present on device 110 using data processing hardware 112 and / or memory hardware 114. Optionally, one or more of systems 180 may be present both locally / on-device and remotely. For example, back-end audio system 180 may execute on remote system 130 by default when the connection between device 110 and remote system 130 is available, but when the connection is lost or unavailable, system 180 may instead execute locally on device 110.
[0031] In some implementations, the device 110 or a system associated with the device 110 identifies text that the device 110 communicates to the user 10 as a response to a query spoken by the user 10. The device 110 may then use a text-to-speech (TTS) system to convert the text into corresponding synthesized playback audio 154 for the device 110 to communicate to the user 10 (e.g., communicate audibly to the user 10) as a response to the query. Once generated, the TTS system communicates the synthesized playback audio 154 to the device 110 so that the device 110 can output the synthesized playback audio 154. For example, in response to the user 10 providing a spoken query about today's weather forecast, the device 110 outputs synthesized playback audio 154 that says "It is sunny today" at the speaker 118 of the device 110.
[0032] Continuing to refer to FIG. 1, when the device 110 outputs the synthesized playback audio 154, the synthesized playback audio 154 generates an echo 156 that is captured by the audio capture device 116. The synthesized playback audio 154 corresponds to the reference audio signal. The synthesized playback audio 154 represents the reference audio signal in the example of FIG. 1, but the reference audio signal may include other types of playback audio 154 such as media content output from the speaker 118 or communication from a remote user talking to the user 10 through the device 110 (e.g., a voice over IP call or a video conference call). Unfortunately, in addition to the echo 156, the audio capture device 116 may also simultaneously capture a target utterance 12 spoken by the user 10 that includes a follow-up query asking further about the weather by saying "How about tomorrow?". For example, FIG. 1 shows that when the device 110 outputs the synthesized playback audio 154, the user 10 asks further about the weather in the spoken utterance 12 to the device 110 by saying "How about tomorrow?". Here, both the spoken utterance 12 and the echo 156 are simultaneously captured by the audio capture device 116 to form a noisy audio signal 202. In other words, the audio signal 202 includes overlapping audio signals, where a portion of the target utterance 12 spoken by the user 10 overlaps with a portion of the reference audio signal (e.g., synthesized playback audio) 154 output from the speaker 118 of the device 110. In addition to the synthesized playback audio 154, competing voices 13 spoken by another user 11 in the environment may also be captured by the audio capture device 116 and contribute to background interference that overlaps with the target utterance 12.
[0033] In FIG. 1, the back-end voice system 180 may have problems processing the target utterance 12 corresponding to the tracking weather query "How about tomorrow?" in the noisy audio signal 202 due to background interference caused by at least one of the playback audio 154, competing voices 13, or non-audio background noise that interferes with the target utterance 12. The context front-end processing model 200 is utilized to improve the robustness of the back-end voice system 180 by co-implementing acoustic echo cancellation (AEC), voice enhancement, and a voice separation model / module within a single model.
[0034] To perform acoustic echo cancellation (AEC), the single model 200 uses the reference signal 154 being played back by the device as an input to the model 200. The reference signal 154 is assumed to be time-aligned with and the same length as the target utterance 12. In some examples, a feature extractor (not shown) extracts reference features 214 corresponding to the reference audio signal 154. The reference features 214 may include log mel filter bank energy (LFBE) features of the reference audio signal 154. Similarly, the feature extractor may extract voice input features 212 corresponding to the target utterance 12. The voice input features 212 may include LFBE features. As described in more detail below, the voice input features 212 may be stacked with the reference features 214 and provided as an input to the primary encoder 210 (FIG. 2) of the single model 200 to perform AEC. In the absence of the reference audio signal 154 being played back by the device, all-zero reference signals may be used such that only the voice input features 212 are received as an input to the primary encoder 210.
[0035] The single model 200 can further perform voice enhancement in parallel with AEC by applying noise context modeling, where the single model 200 processes a context noise signal 213 associated with a noise segment of a predetermined duration captured by an audio capture device 116 prior to a target utterance 12 spoken by a user 10. In some examples, the predetermined duration includes a 6-second noise segment. Thus, the context noise signal 213 provides a noise context. In some examples, the context noise signal 213 includes LFBE features of the noise context signal for use as context information.
[0036] Optionally, the single model 200 can further perform target speaker modeling for voice separation in conjunction with AEC and voice enhancement. Here, a speaker embedding 215 is received as input by the single model 200. The speaker embedding 215 can include voice characteristics of the target speaker 10 who uttered the target utterance 12. The speaker embedding 215 can include a d-vector. In some examples, the speaker embedding 215 is calculated using a text-independent speaker identification (TI-SID) model trained with a generalized end-to-end extended set softmax loss. The TI-SID can include three long short-term memory (LSTM) layers with 768 nodes and a projection size of 256. The output of the final frame of the last LSTM layer is then linearly transformed into a final 256-dimensional d-vector.
[0037] For training and evaluation, each target utterance can be paired with a separate "enrollment" utterance from the same speaker. The enrollment utterance can be randomly selected from a pool of available utterances of the target speaker. The d-vector is then calculated for the enrollment utterance. For most practical applications, the enrollment utterance is generally obtained by a separate offline process.
[0038] Figure 2 shows the context front - end processing model 200 of Figure 1. The context front - end processing model 200 uses a modified version of the conformer neural network architecture that combines convolution and self - attention to model short - and long - distance conversations. The model 200 includes a primary encoder 210, a noise context encoder 220, a cross - attention encoder 400, and a decoder 240. The primary encoder 210 may include N modulated conformer blocks. The noise context encoder 220 may include N conformer blocks. The cross - attention encoder 400 may include M modulated cross - attention conformer blocks. The primary and noise context encoders 210, 220 may execute in parallel. As used herein, each conformer block may use local, causal self - attention to enable streaming capabilities.
[0039] The primary encoder 210 may be configured to receive, as input, input audio features 212 corresponding to a target utterance and generate, as output, a primary input encoding 218. When a reference audio signal 154 is available, the primary encoder 210 is configured to receive, as input, reference features 214 corresponding to the reference audio signal stacked with the input audio features 212 and generate the primary input encoding by processing the reference features 214 stacked with the input audio features 212. The input audio features and the reference features may each include respective sequences of LFBE features.
[0040] The first - stage encoder 210 may be further configured to receive, as input, a speaker embedding 215 (i.e., when available) that includes the voice characteristics of the target speaker 10 who uttered the target utterance 12, and generate a primary input encoding as output by combining the input speech features 212 (or the input speech features stacked with the reference features 214) using a Feature-wise Linear Modulation (FiLM) layer. FIG. 3 provides an exemplary modulated conformable block 300 utilized by the first - stage encoder 210. Here, before each conformable block in the first - stage encoder 210, the speaker embedding 215 (e.g., d - vector) is combined with the input speech features 212 (or the stack of input speech and reference features 214) using a FiLM layer. FiLM allows the first - stage encoder 210 to condition its encoding based on the speaker embedding 215 of the target speaker 10. A residual connection is added after the FiLM layer to ensure that the architecture can be successfully implemented when there is no speaker embedding. Mathematically, the modulated conformable block 300 transforms the input feature x using the modulation feature m to produce an output feature y as follows.
[0041] [Number]
[0042] Here, h(·) and r(·) are affine transformations. FFN, Conv, and MHSA represent a feed - forward module, a convolutional module, and a multi - head self - attention module, respectively. Equation 1 shows a Feature-wise Linear Modulation (FiLM) layer with a residual connection.
[0043] Referring back to FIG. 2, the noise context encoder 220 is configured to receive, as input, a context noise signal 213 that includes noise prior to the target utterance, and generate, as output, a context noise encoding 222. The context noise signal 213 may include LFBE features of the context noise signal. The noise context encoder 220 includes a standard conformable block without modulation at the speaker embedding 215, different from the primary and cross-attention encoders 210, 400. The noise context encoder 220 does not modulate the context noise signal 213 at the speaker embedding 215, because the context noise signal 213 is associated with the acoustic noise context prior to the target utterance 12 being spoken, and thus is assumed to contain information that should be passed in sequence to the cross-attention encoder 400 to assist in noise suppression.
[0044] Continuing to refer to FIG. 2, the cross-attention encoder 400 may be configured to receive, as input, the primary input encoding 218 generated as output from the primary encoder 210 and the context noise encoding 222 generated as output from the noise context encoder 220, and generate, as output, a cross-attention embedding 480. Thereafter, the decoder 240 is configured to decode the cross-attention embedding 480 into the enhanced input speech features 250 corresponding to the target utterance 12. The context noise encoding 222 may correspond to an auxiliary input. The decoder may include a simple projection decoder having a single layer with sigmoid activation, a frame-wise fully-connected network.
[0045] As shown in FIG. 4, the cross-attention encoder 400 can utilize each set of M modulated conformable blocks that each receive, as inputs, the main input encoding 218 modulated by the speaker embedding 215 using FiLM as described in FIG. 3 and the contextual noise encoding 222 output from the noise context encoder 220. The cross-attention conformer encoder first independently processes the modulated input 218 and the auxiliary input 222 using a semi-feedforward net and a convolutional block. Subsequently, a cross-attention block is used to summarize the auxiliary input using the processed input as a query vector. Intuitively, the role of the cross-attention block is to separately summarize the noise context for each input frame to be emphasized. The summarized auxiliary features are then merged with the input using a FiLM layer, followed by a second cross-attention layer that further processes the merged features. Mathematically, if x, m, and n are the encoded input, d-vector, and encoded noise context from the previous layer, respectively, the cross-attention encoder does the following.
[0046]
Number
[0047] Accordingly, the input is modulated by each of the M conformable blocks by both the speaker embedding 215 associated with the target speaker and the contextual noise encoding 222.
[0048] In some implementations, the context front-end processing model is trained jointly with a back-end automatic speech recognition (ASR) model using spectral loss and ASR loss. The training target for training the context front-end processing model 200 uses an ideal ratio mask. The IRM is calculated as follows using reverberant speech and reverberant noise based on the assumption that the speech and noise are uncorrelated in the mel-spectrum space.
[0049]
Number
[0050] Here, X and N are the reverberant speech and reverberant noise mel spectrograms, respectively. t and c represent the time and mel frequency bin index. Since the target is bounded between [0,1], we choose to estimate the IRM and simplify the estimation process. Moreover, the ASR model is used for evaluation, trained with real and simulated reverberant data, and made relatively robust to reverberant speech. Therefore, the IRM derived using reverberant speech as the target still provides a significant gain in performance. The spectral loss during training is calculated as follows based on the L1 and L2 losses between the IRM and the estimated IRM, i.e.,
[0051]
Number
[0052] and
[0053]
Number
[0054] During inference, the estimated IRM is scaled and floored to reduce speech distortion at the expense of reducing noise suppression. This is particularly important because the ASR model is sensitive to speech distortion and non-linear front-end processing, which is one of the main difficulties in improving the performance of a robust ASR model using an emphasis front-end. The emphasis features are derived as follows.
[0055]
Number
[0056] Here, Y is a noisy mel spectrogram,
[0057]
Number
[0058] is an estimated value of a clean mel spectrogram, and α and β are the exponential mask scalar and the mask floor,
[0059]
Number
[0060] is a point - by - point multiplication. In our evaluation, α is set to 0.5 and β is set to 0.01. The highlighted features are log - compressed, that is,
[0061]
Number
[0062] and is passed to the ASR model for evaluation.
[0063] Figure 5 shows an exemplary training process 500 for calculating the ASR loss when the context front-end processing model 200 is trained jointly with the ASR model. Here, only the encoder of the ASR model is used to calculate the loss. The loss is calculated as the l2 distance between the output of the ASR encoder 510 for the target features and the highlighted features. The ASR encoder 510 is not updated during training. Specifically, the training process 500 uses the ASR encoder 510 of the ASR model configured to receive, as input, the highlighted speech features predicted by the context front-end processing model 200 for the training utterances, to generate the predicted output of the ASR encoder 510 for the highlighted speech features, and uses the ASR encoder 510 configured to receive, as input, the target speech features for the training utterances, to generate the target output of the ASR encoder 510 for the target speech features, thereby calculating the ASR loss. The predicted highlighted speech features and the target speech features may each include a respective sequence of LFBE features. Thereafter, the training process 500 calculates the ASR loss via the loss module 520 based on the predicted output of the ASR encoder 510 for the highlighted speech features and the target output of the ASR encoder 510 for the target speech features. The goal of using the ASR loss is to make the highlighting more harmonious with the ASR model, which is important for extracting the best performance from the highlighting front-end. By keeping the parameters of the ASR model fixed, the ASR model is decoupled from the context front-end processing model 200, and by doing so, each is trained and deployed independently of the other.
[0064] FIG. 6 includes a flowchart of an exemplary sequence of operations for a method 600 of performing automatic speech recognition using a context front-end processing model 200. In operation 602, method 600 includes receiving input speech features 212 corresponding to target utterance 12 in context front-end processing model 200. Method 600 also includes, in operation 604, receiving in context front-end processing model 200 at least one of reference audio signal 154, context noise signal 213 including noise prior to target utterance 12, or speaker embedding vector 215 including voice characteristics of target speaker 10 who uttered target utterance 12. In operation 606, method 600 further includes processing input speech features 212 and at least one of reference audio signal 154, context noise signal 213, or speaker embedding vector 215 using context front-end processing model 200 to generate enhanced speech features 250.
[0065] FIG. 7 is a schematic diagram of an exemplary computing device 700 that can be used to implement the systems and methods described in this document. Computing device 700 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The components shown here, their connections and relationships, and their functions are intended only as examples and are not intended to limit the implementations of the disclosure described and / or claimed in this document.
[0066] Computing device 700 includes a processor 710, a memory 720, a storage device 730, a high-speed interface / controller 740 connected to the memory 720 and a high-speed expansion port 750, and a low-speed interface / controller 760 connected to a low-speed bus 770 and the storage device 730. Each of the components 710, 720, 730, 740, 750, and 760 may be interconnected using various buses and may be mounted on a common motherboard or in other manners as appropriate. The processor 710 (i.e., data processing hardware 710 which may include either of data processing hardware 112, 134) can process instructions for execution within the computing device 700, including instructions stored in the memory 720 or on the storage device 730 for displaying graphical information about a graphical user interface (GUI) on an external input / output device such as a display 780 coupled to the high-speed interface 740. In other implementations, multiple processors and / or multiple buses may be used, along with multiple memories and multiple types of memories, as necessary. Also, multiple computing devices 700 may be connected, and each device may provide a portion of the required operations (e.g., as a server bank, a group of blade servers, or a multiprocessor system).
[0067] Memory 720 (i.e., memory hardware 720 that may include either of memory hardwares 114, 136) stores information non - transiently within computing device 700. Memory 720 may be a computer - readable medium, a volatile memory unit, or a non - volatile memory unit. The non - transient memory 720 may be a physical device used to store a program (e.g., a sequence of instructions) or data (e.g., program state information) either temporarily or persistently for use by computing device 700. Examples of non - volatile memory include, but are not limited to, flash memory and read - only memory (ROM) / programmable read - only memory (PROM) / erasable programmable read - only memory (EPROM) / electrically erasable programmable read - only memory (EEPROM) (e.g., typically used for firmware such as a boot program). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase - change memory (PCM), and disks or tapes.
[0068] Storage device 730 can provide mass storage to computing device 700. In some implementations, storage device 730 is a computer - readable medium. In various different implementations, storage device 730 can be a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid - state memory device, or an array of devices including devices in a storage area network or other configurations. In additional implementations, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more methods such as the methods described above. The information carrier is a computer or machine - readable medium such as memory 720, storage device 730, or memory on processor 710.
[0069] The high-speed controller 740 manages bandwidth-consuming operations for the computing device 700, and the low-speed controller 760 manages more bandwidth-efficient operations. Such role assignments are merely exemplary. In some implementations, the high-speed controller 740 is coupled to the memory 720, the display 780 (e.g., through a graphics processor or accelerator), and the high-speed expansion port 750 that may receive various expansion cards (not shown). In some implementations, the low-speed controller 760 is coupled to the storage device 730 and the low-speed expansion port 790. The low-speed expansion port 790 may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), but may be coupled to one or more input / output devices such as a keyboard, a pointing device, a scanner, etc., or a network device such as a switch or a router, e.g., through a network adapter.
[0070] As shown in the figure, the computing device 700 may be implemented in several different forms. For example, it may be implemented as a standard server 700a, or as a group of such servers 700a, as a laptop computer 700b, or as part of a rack server system 700c.
[0071] Various implementations of the systems and techniques described herein may be implemented in digital electronics and / or optical circuit configurations, integrated circuit configurations, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include an implementation in one or more computer programs executable and / or translatable on a programmable system including at least one programmable processor, the programmable processor being coupled to receive data and instructions from, and to transmit data and instructions to, a memory system, at least one input device, and at least one output device, which may be of special or general purpose.
[0072] A software application (i.e., a software resource) may refer to computer software that causes a computing device to perform a task. In some examples, a software application may be referred to as an “application,” an “app,” or a “program.” Exemplary applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, document processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.
[0073] A non-transitory memory may be a physical device used to temporarily or persistently store a program (e.g., a sequence of instructions) or data (e.g., program state information) for use by a computing device. The non-transitory memory may be a volatile and / or non-volatile addressable semiconductor memory. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as a boot program). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disks or tapes.
[0074] These computer programs (also known as programs, software, software applications or code) include machine instructions for a programmable processor and can be implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, non-transitory computer-readable medium, apparatus and / or device (e.g., magnetic disks, optical disks, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor including a machine-readable medium that receives the machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0075] The processes and logical flows described in this specification can be implemented by one or more programmable processors, also referred to as data processing hardware, that execute one or more computer programs to operate on input data and generate output. The processes and logical flows can also be implemented by special purpose logic circuit configurations, such as field programmable gate arrays (FPGAs) or application specific integrated circuits (ASICs). Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, as well as any one or more processors of any kind of digital computer. In general, a processor will receive instructions and data from a read only memory or a random access memory or both. Essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. In general, a computer will also include, or be operatively coupled to, one or more mass storage devices for storing data, such as magnetic, magneto-optical disks, or optical disks, or will receive data from a mass storage device, or transfer data to a mass storage device, or both. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0076] To enable interaction with a user, one or more aspects of the present disclosure can be implemented on a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, or a touch screen, and optionally a keyboard and a pointing device, such as a mouse or trackball, for the user to provide input to the computer. Other types of devices can also be used to provide interaction with the user, and for example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and the input from the user can be received in any form, including acoustic, voice, or tactile input. Further, the computer can interact with the user by sending documents to and receiving documents from the devices used by the user, for example, by sending a web page to a web browser on the user's client device in response to a request received from the web browser.
[0077] Some implementations have been described. Nevertheless, it will be understood that various modifications can be made without departing from the spirit and scope of the present disclosure. Accordingly, other implementations are within the scope of the following claims.
Description of the Reference Numerals
[0078] 110 Voice-responsive user device, device, user device 112 Data processing hardware 114 Memory hardware 116 Audio capture device 118 Voice output device, speaker 130 Remote system 132 Remote resource 134 Remote data processing hardware, data processing hardware 136 Remote memory hardware, memory hardware 180 Back-end audio system, system 200 Context front-end processing model, model 210 Primary encoder 220 Noise context encoder 240 Decoder 300 Modulated conformable block 400 Cross-attention encoder 510 ASR encoder 520 Loss module 700 Computing device 700a Server 700b Laptop computer 700c Rack server system 710 Processor, component 720 Memory, component, memory hardware 730 Storage device, component 740 High-speed interface / controller, component 750 High-speed expansion port, component 760 Low-speed interface / controller, component 770 Low-speed bus 780 Display 790 Low-speed expansion port
Claims
1. A computer-implemented method (600) for causing data processing hardware (710) to perform operations when executed on the data processing hardware (710), the operations comprising: receiving, in a context front-end processing model (200), an input voice feature (212) corresponding to a target utterance (12), a reference audio signal (154), a context noise signal (213) including noise preceding the target utterance (12), and a speaker embedding vector (215) including voice characteristics of a target speaker (10) who uttered the target utterance (12); processing, using the context front-end processing model (200), the input voice feature (212), the reference audio signal (154), the context noise signal (213), and the speaker embedding vector (215) to generate an enhanced voice feature (250), the processing comprising: processing, using a primary encoder (210), a reference feature (214) corresponding to the reference audio signal (154) and the stacked input voice feature (212) to generate a primary input encoding (218); processing, using a noise context encoder (220), the context noise signal (213) to generate a context noise encoding (222); processing, using a cross-attention encoder (400), the primary input encoding (218) and the context noise encoding (222) to generate a cross-attention embedding (480); decoding the cross-attention embedding (480) into the enhanced voice feature (250) corresponding to the target utterance (12); A computer-implemented method (600) comprising the above.
2. The computer-implemented method (600) according to claim 1, wherein the context front-end processing model (200) comprises a conformer neural network architecture that combines convolution and self-attention to model short-distance and long-distance conversations.
3. The computer-implemented method (600) according to claim 1, wherein the input voice feature (212) and the reference feature (214) each comprise respective sequences of log mel filter bank energy (LFBE) features.
4. The step of processing the input speech feature (212) to generate the main input encoding (218) includes the step of using feature-wise linear modulation (FiLM) to combine the input speech feature (212) with the speaker embedding vector (215) to generate the main input encoding (218). The step of processing the main input encoding (218) and the context noise encoding (222) to generate the cross-attention embedding (480) is using FiLM to combine the main input encoding (218) with the speaker embedding vector (215) to generate a modulated main input encoding (218), and processing the modulated main input encoding (218) and the context noise encoding (222) to generate the cross-attention embedding (480). The computer-implemented method (600) according to claim 1.
5. The primary encoder (210) comprises N modulated conformer blocks. The noise context encoder (220) comprises N conformer blocks and runs in parallel with the primary encoder (210). The cross-attention encoder (400) comprises M modulated cross-attention conformer blocks. The computer-implemented method (600) according to any one of claims 1 to 4.
6. The data processing hardware (710) runs the context front-end processing model (200) and is present on the user device (110), and the user device (110) outputs the reference audio signal (154) as playback audio via the audio speaker (118) of the user device (110), and is configured to capture the target utterance (12), the reference audio signal (154), and the context noise signal (213) via one or more microphones (116) of the user device (110). The computer-implemented method (600) according to any one of claims 1 to 5.
7. The context front-end processing model (200) uses spectral loss and ASR loss to jointly train with a back-end automatic speech recognition (ASR) model using training utterances, a context noise signal (213), reference features (214) corresponding to a reference audio signal (154), a speaker embedding (215), and target speech features for the training utterances, the computer-implemented method (600) according to any one of claims 1 to 6.
8. The ASR loss is using the ASR encoder (510) of the ASR model configured to receive, as an input, the enhanced speech features (250) predicted by the context front-end processing model (200) for the training utterance, to generate, as a predicted output, the output of the ASR encoder (510) for the enhanced speech features (250); using the ASR encoder (510) of the ASR model configured to receive, as an input, target speech features for the training utterance, to generate, as a target output, the output of the ASR encoder (510) for the target speech features; and calculating the ASR loss based on the predicted output of the ASR encoder (510) for the enhanced speech features (250) and the target output of the ASR encoder (510) for the target speech features, the computer-implemented method (600) according to claim 7.
9. The operation further includes processing, using a back-end speech system (180), the enhanced speech features (250) corresponding to the target utterance (12), the computer-implemented method (600) according to any one of claims 1 to 8.
10. The back-end speech system (180) includes an automatic speech recognition (ASR) model, a hotword detection model, or at least one of an audio or audio-video call application, the computer-implemented method (600) according to claim 9.
11. A context front-end processing model (200), receiving, as inputs, input speech features (212) corresponding to a target utterance (12) and reference features (214) corresponding to a reference audio signal (154), A primary encoder (210) configured to generate a primary input encoding (218) by processing the input speech feature (212) stacked with the reference feature (214) as an output; Receiving, as an input, a context noise signal (213) including noise prior to the target utterance (12); A noise context encoder (220) configured to generate a context noise encoding (222) as an output; Receiving, as inputs, the primary input encoding (218) generated as an output from the primary encoder (210) and the context noise encoding (222) generated as an output from the noise context encoder (220); A cross-attention encoder (400) configured to generate a cross-attention embedding (480) as an output; A context front-end processing model (200) comprising a decoder (240) configured to decode the cross-attention embedding (480) into an emphasized speech feature (250) corresponding to the target utterance (12). **Claim 12** The context front-end processing model (200) according to claim 11, wherein the input speech feature (212) and the reference feature (214) each include a respective sequence of log mel filter bank energy (LFBE) features. **Claim 13** The primary encoder (210) Receiving, as an input, a speaker embedding vector (215) including voice characteristics of a target speaker (10) who uttered the target utterance (12); The context front-end processing model (200) according to claim 11 or 12, further configured to generate the primary input encoding (218) as an output by combining the input speech feature (212) with the speaker embedding vector (215) using feature-wise linear modulation (FiLM). **Claim 14** The cross-attention encoder (400) Receiving, as an input, the primary input encoding (218) modulated by a speaker embedding vector (215) using feature-wise linear modulation (FiLM), wherein the speaker embedding vector (215) includes voice characteristics of a target speaker (10) who uttered the target utterance (12). Further configured to process the main input encoding (218) modulated by the speaker embedding vector (215) and the context noise encoding (222) to generate the cross-attention embedding (480) as an output, the context front-end processing model (200) according to any one of claims 11 to 13.
15. The primary encoder (210) comprises N modulated conformable blocks, The noise context encoder (220) comprises N conformable blocks and runs in parallel with the primary encoder (210), The cross-attention encoder (400) comprises M modulated cross-attention conformable blocks, the context front-end processing model (200) according to any one of claims 11 to 14.
16. The context front-end processing model (200) runs on data processing hardware (710) present on the user device (110), and the user device (110) Outputs the reference audio signal (154) as playback audio via the audio speaker (118) of the user device (110), Configured to capture the target utterance (12), the reference audio signal (154), and the context noise signal (213) via one or more microphones (116) of the user device (110), the context front-end processing model (200) according to any one of claims 11 to 15.
17. The context front-end processing model (200) is trained jointly with a back-end automatic speech recognition (ASR) model using spectral loss and ASR loss, using training utterances, context noise signals (213), reference features (214) corresponding to the reference audio signal (154), speaker embedding vectors (215), and target speech features for the training utterances, the context front-end processing model (200) according to any one of claims 11 to 16.
18. The ASR loss is Using the ASR encoder (510) of the ASR model configured to receive, as input, the emphasized speech features (250) predicted by the context front-end processing model (200) for the training utterance, generate the output of the ASR encoder (510) for the emphasized speech features (250) as a predicted output; Using the ASR encoder (510) configured to receive, as input, the target speech features for the training utterance, generate the output of the ASR encoder (510) for the target speech features as a target output; The context front-end processing model (200) according to claim 17, calculated by calculating the ASR loss based on the predicted output of the ASR encoder (510) for the emphasized speech features (250) and the target output of the ASR encoder (510) for the target speech features.
19. The back-end voice system (180) is configured to process the emphasized speech features (250) corresponding to the target utterance (12), the context front-end processing model (200) according to any one of claims 11 to 18.
20. The back-end voice system (180) An automatic speech recognition (ASR) model, A hotword detection model, or The context front-end processing model (200) according to claim 19, comprising at least one of an audio or audio-video call application.
Citation Information
Patent Citations
Speech recognition and model training method, device and computer readable storage medium
CN111261146A
Voice Recognition System
JP2020503570A