Generalized Automatic Speech Recognition for Integrated Acoustic Echo Cancellation, Voice Enhancement, and Voice Separation

A context front-end processing model integrates echo cancellation, voice enhancement, and voice separation, using a dropout strategy to train the ASR system for robust performance across various interference conditions, addressing the limitations of separate modeling approaches.

JP7713113B2Active Publication Date: 2025-07-24GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024555991
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-03-20
Filing Date
2023-02-19
Publication Date
2025-07-24
Estimated Expiration
2043-02-19

AI Technical Summary

Technical Problem

Existing automatic speech recognition (ASR) systems face degradation in performance due to conditions such as echo, background noise, and competing voices, which are typically addressed separately and are impractical to train for all conditions simultaneously.

Method used

A generalized automatic speech recognition model is trained using a context front-end processing model that integrates acoustic echo cancellation, voice enhancement, and voice separation, employing a context signal dropout strategy to improve robustness by simulating missing context inputs during training.

Benefits of technology

The integrated model enhances the ASR system's ability to perform effectively even when one or more context inputs are absent, improving echo cancellation, voice enhancement, and voice separation, thereby maintaining performance in diverse interference conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007713113000011
    Figure 0007713113000011
  • Figure 0007713113000012
    Figure 0007713113000012
  • Figure 0007713113000013
    Figure 0007713113000013
Patent Text Reader

Abstract

A method for training a generalized automatic speech recognition model for integrated acoustic echo cancellation, speech enhancement, and sound separation includes receiving a plurality of training utterances paired with corresponding training context signals. The training context signals include a training context noise signal including noise before the training utterances, a training reference audio signal, and a training speaker vector including voice characteristics of the speaker who spoke the training utterances. The operations also include training a contextual front-end processing model on the training utterances to learn how to predict the enhanced speech features using a context signal dropout strategy. The context signal dropout strategy drops out each of the training context signals using a predetermined probability during training of the contextual front-end processing model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to generalized automatic speech recognition for integrated acoustic echo cancellation, voice enhancement, and voice separation.

Background Art

[0002] The robustness of automatic speech recognition (ASR) systems has been significantly improved over the years by the emergence of neural network-based end-to-end models, large-scale training data, and improved strategies for augmenting training data. However, various conditions such as echo, stronger background noise, and competing voices can significantly degrade the performance of ASR systems. Integrated ASR models can be trained to handle these conditions. However, in use, integrated ASR models do not encounter all conditions that occur simultaneously. Therefore, it is not practical to train integrated ASR models for all existing conditions.

Summary of the Invention

[0003] One aspect of the present disclosure provides a computer-implemented method for training a generalized automatic speech recognition model for combined echo cancellation, voice enhancement, and voice separation. The method causes the data processing hardware to perform operations when executed on the data processing hardware. The operations include receiving a plurality of training utterances paired with corresponding training context signals. The training context signal includes a training context noise signal including noise before the corresponding training utterance, a training reference audio signal, and a training speaker vector including the voice characteristics of the target speaker who spoke the corresponding training utterance. The operations also include training a context front-end processing model for the training utterance to learn a method for predicting enhanced speech features using a context signal dropout strategy. Here, the context signal dropout strategy drops out each of the training context signals using a predetermined probability during the training of the context front-end processing model.

[0004] Embodiments of the present disclosure may include one or more of the following optional features. In some embodiments, the signal dropout strategy drops out each training context signal by replacing the corresponding context signal with all zeros. In these embodiments, replacing the training reference audio signal with all zeros includes replacing the training reference audio signal with all-zero features of the same length and feature dimension as the corresponding training utterance. Additionally or alternatively, replacing the training context noise signal includes replacing the training context noise signal with all-zero features having a predetermined length and the same feature dimension as the corresponding training utterance. Further, in these embodiments, replacing the training speaker vector includes replacing the training speaker vector with all-zero features having an all-zero vector. In some examples, the signal dropout strategy drops out each training context signal by replacing the corresponding context signal with a frame-level learned representation.

[0005] In some embodiments, the trained context front-end processing model includes a primary encoder, a noise context encoder, a cross-attention encoder, and a decoder. The primary encoder receives, as input, input speech features corresponding to a target utterance and generates, as output, a primary input encoding. The noise context encoder receives, as input, a context noise signal including noise prior to the target utterance and generates, as output, a context noise encoding. The cross-attention encoder receives, as input, the primary input encoding generated as output from the primary encoder and the context noise encoding generated as output from the noise context encoder, and generates, as output, a cross-attention embedding. The decoder decodes the cross-attention embedding into enhanced input speech features corresponding to the target utterance. In these embodiments, the primary encoder is further configured to generate the primary input encoding by receiving, as input, reference features corresponding to a reference audio signal and processing, as output, the input speech features stacked with the reference features. Alternatively, the primary encoder is further configured to generate the primary input encoding by receiving, as input, a speaker embedding including voice characteristics of a target speaker who spoke the target utterance and combining, as output, the input speech features with the speaker embedding using feature-wise linear modulation (FiLM). Additionally or alternatively, the cross-attention encoder is further configured to receive, as input, the primary input encoding modulated by the speaker embedding using FiLM. Here, the speaker embedding includes voice characteristics of a target speaker who spoke the target utterance, and processes the primary input encoding modulated by the speaker embedding and the context noise encoding to generate, as output, a cross-attention embedding. In some embodiments, the primary encoder includes N modulated conformer blocks, the noise context encoder includes N conformer blocks that execute in parallel with the primary encoder, and the cross-attention encoder includes M modulated cross-attention conformer blocks.

[0006] In some examples, the context front-end processing model is trained in integration with a back-end automatic speech recognition (ASR) model using spectral loss and ASR loss. In these examples, the spectral loss may be based on the distance of the L1 loss function and the L2 loss function between the estimated ratio mask and the ideal ratio mask. Here, the ideal ratio mask is calculated using the reverberant speech and the reverberant noise. Further, in these examples, the ASR loss is calculated by receiving, as an input, the enhanced speech features predicted by the context front-end processing model of the training utterance, receiving the predicted output of the ASR encoder of the enhanced speech features using an ASR encoder configured to receive, as an input, the target speech features of the training utterance, generating the target output of the ASR encoder of the target speech features, and calculating the ASR loss based on the predicted output of the ASR encoder of the enhanced speech features and the target output of the ASR encoder of the target speech features.

[0007] Another aspect of the present disclosure provides a system for training a generalized automatic speech recognition model for integrated echo cancellation, voice enhancement, and voice separation. The system includes data processing hardware and memory hardware that communicates with the data processing hardware. The memory hardware stores instructions that, when executed by the data processing hardware, cause the data processing hardware to perform operations including receiving a plurality of training utterances paired with corresponding training context signals. The training context signals include a training context noise signal that includes noise prior to the corresponding training utterance, a training reference audio signal, and a training speaker vector that includes voice characteristics of a target speaker who spoke the corresponding training utterance. The operations also include training a context front-end processing model for the training utterances to learn a method for predicting enhanced voice features using a context signal dropout strategy. Here, the context signal dropout strategy drops out each of the training context signals using a predetermined probability during training of the context front-end processing model.

[0008] This aspect may include one or more of the following optional features. In some embodiments, the signal dropout strategy drops out each training context signal by replacing the corresponding context signal with all zeros. In these embodiments, replacing the training reference audio signal with all zeros includes replacing the training reference audio signal with all-zero features of the same length and feature dimension as the corresponding training utterance. Additionally or alternatively, replacing the training context noise signal includes replacing the training context noise signal with all-zero features having a predetermined length and the same feature dimension as the corresponding training utterance. Additionally, in these embodiments, replacing the training speaker vector includes replacing the training speaker vector with all-zero features. In some examples, the signal dropout strategy drops out each training context signal by replacing the corresponding context signal with a frame-level learned representation.

[0009] In some embodiments, the trained context front-end processing model includes a primary encoder, a noise context encoder, a cross-attention encoder, and a decoder. The primary encoder receives, as input, input audio features corresponding to a target utterance and generates, as output, a primary input encoding. The noise context encoder receives, as input, a context noise signal including noise before the target utterance and generates, as output, a context noise encoding. The cross-attention encoder receives, as input, the primary input encoding generated as output from the primary encoder and the context noise encoding generated as output from the noise context encoder, and generates, as output, a cross-attention embedding. The decoder decodes the cross-attention embedding into enhanced input audio features corresponding to the target utterance. In these embodiments, the primary encoder is further configured to generate the primary input encoding by receiving, as input, reference features corresponding to a reference audio signal and processing, as output, the input audio features stacked with the reference features. Alternatively, the primary encoder is further configured to generate the primary input encoding by receiving, as input, a speaker embedding including voice characteristics of a target speaker who spoke the target utterance and combining, as output, the input audio features with the speaker embedding using feature-wise linear modulation (FiLM). Additionally or alternatively, the cross-attention encoder is further configured to receive, as input, the primary input encoding modulated by the speaker embedding using FiLM. Here, the speaker embedding includes voice characteristics of a target speaker who spoke the target utterance, and processes the primary input encoding modulated by the speaker embedding and the context noise encoding to generate, as output, a cross-attention embedding. In some embodiments, the primary encoder includes N modulation conformer blocks, the noise context encoder includes N conformer blocks, executes in parallel with the primary encoder, and the cross-attention encoder includes M modulation cross-attention conformer blocks.

[0010] In some examples, the context front-end processing model is trained by integrating with a back-end automatic speech recognition (ASR) model using spectral loss and ASR loss. In these examples, the spectral loss may be based on the distance between the estimated ratio mask and the ideal ratio mask using L1 and L2 loss functions. Here, the ideal ratio mask is calculated using reverberant speech and reverberant noise. Additionally, in these examples, the ASR loss is calculated by receiving the predicted output of the ASR encoder for the enhanced speech features with the enhanced speech features predicted by the context front-end processing model of the training utterance as input, generating the target output of the ASR encoder for the target speech features using an ASR encoder configured to receive the target speech features of the training utterance as input, and calculating the ASR loss based on the predicted output of the ASR encoder for the enhanced speech features and the target output of the ASR encoder for the target speech features.

[0011] Details of one or more embodiments of the present disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will be apparent from the description and drawings, as well as from the claims.

Brief Description of the Drawings

[0012]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

[0013] Like reference symbols in the various drawings refer to like elements.

[0014] The robustness of automatic speech recognition (ASR) systems has been significantly improved over the years due to the emergence of neural network-based end-to-end models, large-scale training data, and improved strategies for augmenting training data. Nevertheless, background interference can significantly degrade the performance of an ASR system in accurately recognizing utterances directed to the ASR system. Background interference can be broadly classified into three groups: device echo, ambient noise, and competing voices. Separate ASR models can be trained to handle each of these groups of background interference separately, but it is difficult and impractical to maintain ASR models specific to multiple tasks / conditions and switch models on-the-fly during use.

[0015] Device echo can be responsive to the reproduced audio output from devices such as smart home speakers, whereby the reproduced audio is recorded as an echo and can affect the performance of backend voice systems such as ASR systems. In particular, the degradation of the performance of the backend voice system is particularly severe when the reproduced audio includes audible speech, such as text-to-speech (TTS) responses from digital assistants. This problem is usually addressed via acoustic echo cancellation (AEC) techniques. The unique characteristic of AEC is that a reference signal corresponding to the reproduced audio is usually available and can be used for suppression.

[0016] Noisy background with non-speech characteristics is usually appropriately handled using data augmentation strategies such as multi-style training (MTR) of ASR models. Here, noise is added to the training data using an indoor simulator, and then they are carefully weighted with clean data during training, so as to balance the performance between the clean state and the noisy state. As a result, large-scale ASR models are robust to medium levels of non-speech noise. However, in the presence of low signal-to-noise ratio (SNR) conditions, the noisy background can still affect the performance of the backend voice system.

[0017] Unlike non-speech noisy background, competing speech is very difficult for an ASR model trained to recognize a single speaker. Training an ASR model with the voices of multiple speakers can itself be a problem because it is difficult to eliminate the ambiguity of which speaker to focus on during inference. Since it is difficult to know in advance the number of supported users, it is not optimal to use a model that recognizes multiple speakers. Furthermore, such multi-speaker models usually have degraded performance in a single-speaker setting, which is undesirable.

[0018] The three classes of background interference described above are usually addressed separately from each other, each using a separate modeling strategy. In recent literature, voice separation using techniques such as deep clustering, permutation invariant learning, and speaker embedding has received much attention. When using speaker embedding, the target speaker of interest is assumed to be known a priori. The techniques developed for speaker separation are also applied to the removal of non-speech noise by modifying the training data. AEC has also been studied alone or together in the presence of ambient noise. It is well known that improving the speech quality does not necessarily improve the ASR performance because the distortion generated by non-linear processing can have an adverse effect on the ASR performance. One way to reduce the mismatch between the extended front-end that processes the incoming audio first and the resulting ASR performance is to integrate and train the extended front-end together with the back-end ASR model.

[0019] Furthermore, due to the continued interest in the application of large-scale multi-domain and multi-language ASR models, the training data for these ASR models usually cover various acoustic and language use cases (e.g., voice search and video captioning) to simultaneously address more difficult noise conditions. As a result, it is often convenient to train and maintain a separate front-end feature processing model that can handle adverse conditions without combining it with the back-end ASR model. Additionally, although various types of data for the ASR model are available for training, the ASR model needs to function well even when one or more of the aforementioned background interference groups (e.g., device echo, ambient noise, and competing voices) are missing from the training examples.

[0020] Embodiments of this specification aim to train a context front - end processing model to improve the robustness of ASR by integrating modules for acoustic echo cancellation (AEC), voice enhancement, and voice separation into a single model. The single integrated model is practical from the perspective that, especially in streaming ASR settings, it is difficult, if not impossible, to know in advance which classes of background interference to handle. Specifically, the context front - end processing model includes a context - enhanced neural network (CENN) that can optionally utilize three different types of side - context inputs: a reference signal associated with the reproduced audio, a noise context, and a speaker embedding representing the voice characteristics of the target speaker. Embodiments of this specification more specifically aim to use a context signal dropout strategy for training the context front - end processing model to improve the performance of the model during inference when one or more context inputs are missing. As will become apparent, the reference signal related to the reproduced audio is necessary for providing echo cancellation, while the noise context is useful for voice enhancement. Additionally, the speaker embedding (when available) representing the voice characteristics of the target speaker is not only important for voice separation but also useful for echo cancellation and voice enhancement. In the case of voice enhancement and separation, the noise context, i.e., the audio in the few seconds before the target utterance to be recognized, conveys useful information about the acoustic context. The CENN can generate enhanced input audio features that can be passed to a backend audio system, such as an ASR model, which can process the enhanced input audio features using respective neural network architectures configured to take in the corresponding context - side inputs and generate the audio, and generate the speech recognition result of the target utterance. In particular, since the noise context and the reference features are optional context - side inputs, the noise context and the reference features are assumed to be respective non - informative silence signals by the CENN when not available.

[0021] Referring to FIG. 1, in some embodiments, system 100 includes a user 10 communicating a spoken target utterance 12 to an audio-responsive user device 110 (also referred to as device 110 or user device 110) in an audio environment. User 10 (i.e., the speaker of utterance 12) may speak target utterance 12 as a query or command seeking a response from device 110. Device 110 is configured to capture sound from one or more users 10, 11 within the audio environment. Here, the audio sound may refer to an utterance 12 spoken by user 10 that functions as an audible query, a command to device 110, or an audible communication captured by device 110. The device 110 or an audio-responsive system associated with device 110 may handle the query of the command by executing the query and / or command.

[0022] Various types of background interference can interfere with the ability of the backend voice system 180 to process the target utterance 12 that specifies a query or command to the device 110. As described above, background interference can include a device echo corresponding to the playback audio 154 (also referred to as the reference audio signal 154) output from the user device (e.g., smart speaker) 110, competing voices 13 such as utterances other than the target utterance 12 spoken by one or more other users 111 not directed at the user device 110, and one or more ambient noises having non-speech characteristics. In an embodiment of the present specification, a context front-end processing model 200 (also referred to as model 200), which is executed on the device 110 and is configured to receive as input the input voice features corresponding to the target utterance 12 and one or more context input features 213, 214, 215, is used to process the input voice features 212 and one or more context input features 213, 214, 215 to generate, as output, enhanced input voice features 250 corresponding to the target utterance 12. As will be described in more detail below (e.g., FIG. 5), the model 200 can be trained to improve the performance of the model 200 during inference when one or more of the context input features 213, 214, 215 are missing, using a context signal dropout strategy. Next, the backend voice system 180 can process the enhanced input voice features 250 to generate an output 182. In particular, the context front-end processing model 200 effectively removes the presence of background interference recorded by the device 110 when the user 10 speaks the target utterance 12, such that the enhanced input voice features 250 provided to the backend voice system 180 convey the voice (i.e., the target utterance 12) intended for the device 110, so that the output 182 generated by the backend voice system 180 is not degraded by the background interference.

[0023] In the illustrated example, the backend voice system 180 includes an ASR system 190 that uses an ASR model 192 to process the enhanced voice features 250 and generate a voice recognition result (e.g., a transcription) for the target utterance 12. The ASR system 190 may further include a natural language understanding (NLU) module (not shown) that performs semantic interpretation on the transcription of the target utterance 12 to identify a query / command directed to the device 110. Thus, the output 182 from the backend voice system 180 may include a transcription and / or instructions for fulfilling the query / command identified by the NLU module.

[0024] The backend voice system 180 may additionally or alternatively include a hotword detection model (not shown) configured to detect whether the enhanced input voice features 250 include the presence of one or more hotwords / wakewords trained to be detected by the hotword detection model. For example, the hotword detection model may output a hotword detection score indicating the likelihood that the enhanced input voice features 250 corresponding to the target utterance 12 include a particular hotword / wakeword. The detection of the hotword may trigger a wake-up process by which the device 110 may wake up from a sleep state. For example, the device 110 may wake up and process the hotword and / or one or more terms preceding / following the hotword.

[0025] In an additional example, the background audio system 180 includes an audio or audio-video calling application (e.g., a video conferencing application). Here, the enhanced input audio feature 250 corresponding to the target utterance 12 is used by the audio or audio-video calling application to filter the voice of the target speaker 10 for communication to the recipient during an audio or audio-video communication session. The background audio system 180 may additionally or alternatively include a speaker identification model configured to perform speaker identification using the enhanced input audio feature 250 to identify the user 10 who spoke the target utterance 12.

[0026] In the illustrated example, the device 110 captures a noisy audio signal 202 (also referred to as audio data) of the target utterance 12 spoken by the user 10 in the presence of background interference originating from one or more sources other than the user 10. The device 110 may correspond to any computing device associated with the user 10 and capable of receiving the noisy audio signal 202. Some examples of the user device 110 include, but are not limited to, mobile devices (e.g., mobile phones, tablets, laptops, etc.), computers, wearable devices (e.g., smartwatches), smart appliances, Internet of Things (IoT) devices, and smart speakers. The device 110 includes data processing hardware 112 and memory hardware 114 that communicates with the data processing hardware 112, and the memory hardware 114 stores instructions that cause the data processing hardware 112 to perform one or more operations when executed by the data processing hardware 112. The context front-end processing model 200 may be executed on the data processing hardware 112. In some examples, the back-end audio system 180 is executed on the data processing hardware 112.

[0027] In some examples, device 110 includes one or more applications (i.e., software applications), and each application may utilize the enhanced input voice feature 250 generated by the context front-end processing model 200 to perform various functions within the application. For example, device 110 includes an assistant application configured to communicate synthesized playback audio 154 to user 10 to assist user 10 with various tasks.

[0028] Device 110 further includes an audio subsystem having an audio capture device (e.g., a microphone) 116 for capturing spoken utterance 12 and converting it into an electrical signal within the acoustic environment, and an audio output device (e.g., a speaker) 118 for communicating an audible audio signal (e.g., synthesized playback signal 154 from device 110). In the illustrated example, device 110 implements a single audio capture device 116, but device 110 may implement an array of audio capture devices 116 without departing from the scope of the present disclosure. In this case, one or more of the audio capture devices 116 of the array may communicate with the audio subsystem (e.g., a peripheral device of device 110) without being physically resident on device 110. For example, device 110 may correspond to a vehicle infotainment system that utilizes an array of microphones disposed throughout the vehicle.

[0029] In some examples, device 110 is configured to communicate with remote system 130 via a network (not shown). Remote system 130 may include remote resources 132 such as remote data processing hardware 134 (e.g., a remote server or CPU) and / or remote memory hardware 136 (e.g., a remote database or other storage hardware). User device 110 may utilize remote resources 132 to perform various functions related to audio processing and / or synthetic playback communication. Context front-end processing model 200 and back-end audio system 180 may reside on device 110 (referred to as an on-device system) or may reside remotely while communicating with device 110 (e.g., may reside on remote system 130). In some examples, one or more back-end audio systems 180 reside locally or on the device, while one or more other back-end audio systems 180 reside remotely. In other words, one or more back-end audio systems 180 that utilize the enhanced input audio features 250 output from context front-end processing model 200 may be local or remote in any combination. For example, if the size or processing requirements of system 180 are fairly large, that system 180 may be resident on remote system 130. Further, if device 110 can support the size or processing requirements of one or more systems 180, one or more systems 180 may be resident on device 110 using data processing hardware 112 and / or memory hardware 114. Optionally, one or more systems 180 may be resident on both local / on-device and remote. For example, back-end audio system 180 can be executed on remote system 130 by default when the connection between device 110 and remote system 130 is available, but when the connection is lost or unavailable, system 180 is instead executed locally on device 110.

[0030] In some embodiments, the device 110 or a system associated with the device 110 identifies text that the device 110 communicates to the user 10 as a response to a query spoken by the user 10. Next, the device 110 may use a text-to-speech (TTS) system to convert the text into corresponding synthesized playback audio 154 such that the device 110 communicates (e.g., audibly communicates) with the user 10 as a response to the query. Once generated, the TTS system communicates the synthesized playback audio 154 to the device 110, enabling the device 110 to output the synthesized playback audio 154. For example, in response to the user 10 providing an oral query regarding today's weather forecast, the device 110 outputs synthesized playback audio 154 of "Today is sunny" through the speaker 118 of the device 110.

[0031] Continuing to refer to FIG. 1, when device 110 outputs synthesized playback audio 154, the synthesized playback audio 154 generates an echo 156 captured by audio capture device 116. The synthesized playback audio 154 corresponds to the reference audio signal. On the other hand, the synthesized playback audio 154 represents the reference audio signal in the example of FIG. 1, and the reference audio signal may include media content output from speaker 118, or other types of playback audio 154 such as communications from remote user 10 with whom user 10 is conversing via device 110 (e.g., Voice over IP call or video conference call). Unfortunately, in addition to echo 156, audio capture device 116 may also simultaneously capture target utterance 12 spoken by user 10, including a supplementary query that further asks about the weather by saying "How about tomorrow?". For example, in FIG. 1, since device 110 outputs synthesized playback audio 154, user 10 further asks about the weather with utterance 12 spoken by saying "How about tomorrow?" to device 110. Here, both the spoken utterance 12 and the echo 156 are simultaneously captured by audio capture device 116, forming a noisy audio signal 202. In other words, the audio signal 202 includes an overlapping audio signal where a portion of the target utterance 12 spoken by user 10 overlaps with a portion of the reference audio signal (e.g., synthesized playback audio) 154 output from speaker 118 of device 110. In addition to the synthesized playback audio 154, competing audio 13 spoken by another user 11 in the environment may also be captured by audio capture device 116 and may contribute to background interference overlapping with the target utterance 12.

[0032] In FIG. 1, the backend audio system 180 may have a problem processing the target utterance 12 corresponding to the supplementary weather query, "How about tomorrow?" due to the presence of background interference in the noisy audio signal 202 caused by at least one of the playback audio 154, competing voice 13, or non-speech ambient noise interfering with the target utterance 12. The context front-end processing model 200 is used to improve the robustness of the backend audio system 180 by integrating acoustic echo cancellation (AEC), voice enhancement, and voice separation models / modules into a single model.

[0033] To perform acoustic echo cancellation (AEC), the single model 200 uses the reference signal 154 being played by the device as an input to the model 200. The reference signal 154 is assumed to be temporally aligned with the target utterance 12 and of the same length. In some examples, a feature extractor (not shown) extracts reference features 214 corresponding to the reference audio signal 154. The reference features 214 may include log mel filter bank energy (LFBE) features of the reference audio signal 154. Similarly, the feature extractor may extract input voice features 212 corresponding to the target utterance 12. The input voice features 212 may include LFBE features. As will be described in more detail below, the input voice features 212 may be stacked with the reference features 214 and provided as an input to the primary encoder 210 (FIG. 2) of the single model 200 to perform AEC. When there is no reference audio signal 154 being played by the device, an all-zero reference signal may be used, whereby only the input voice features 212 are received as an input to the primary encoder 210.

[0034] When a single model 200 processes a context noise signal 213 related to a predetermined period of a noise segment captured by an audio capture device 116 prior to a target utterance 12 spoken by a user 10, the single model 200 may further perform voice enhancement in parallel with AEC by applying noise context modeling. In some examples, the predetermined period includes a six (6)-second noise segment. Thus, the context noise signal 213 provides a noise context. In some examples, the context noise signal 213 includes LFBE features of the noise context signal for use as context information.

[0035] Optionally, the single model 200 may further perform target speaker modeling for voice separation by integrating AEC and voice enhancement. Here, a speaker embedding 215 is received as an input by the single model 200. The speaker embedding 215 may include voice features of the target speaker 10 who spoke the target utterance 12. The speaker embedding 215 may include a d-vector. In some examples, the speaker embedding 215 is calculated using a text-independent speaker identification (TI-SID) model trained with a generalized end-to-end extended set softmax loss. The TI-SID may include three long short-term memory (LSTM) layers with 768 nodes and a projection size of 256. Next, the output of the final frame of the last LSTM layer is linearly transformed into a final 256-dimensional d-vector.

[0036] For training and evaluation, each target utterance may be paired with an individual "enrollment" utterance from the same speaker. The enrollment utterance may be randomly selected from a pool of available utterances of the target speaker. Next, the d-vector is calculated for the enrollment utterance. In most practical applications, the enrollment utterance is usually obtained via an individual offline process.

[0037] Figure 2 shows the context front - end processing model 200 of FIG. 1. The context front - end processing model 200 uses a modified version of the conformer neural network architecture that combines convolution and self - attention to model short - distance and long - distance interactions. The model 200 includes a primary encoder 210, a noise context encoder 220, a cross - attention encoder 400, and a decoder 240. The primary encoder 210 may include N modulated conformer blocks. The noise context encoder 220 may include N conformer blocks. The cross - attention encoder 230 may include M modulated cross - attention conformer blocks. The primary context encoder 210 and the noise context encoder 220 may be executed in parallel. As used herein, each conformer block may use local and causal self - attention to enable streaming capabilities.

[0038] The primary encoder 210 may be configured to receive, as input, input audio features 212 corresponding to the target utterance and generate, as output, a primary input encoding 218. When a reference audio signal 154 is available, the primary encoder 210 is configured to receive, as input, input audio features 212 with reference features 214 corresponding to the reference audio signal stacked thereon, and generate a primary input encoding by processing the input audio features 212 stacked with the reference features 214. The input audio features and the reference features may each include respective sequences of LFBE features.

[0039] The primary encoder 210 may further be configured to receive, as an input, a speaker embedding 215 (i.e., if available) that includes the voice characteristics of the target speaker (i.e., the user) 10 who spoke the target utterance 12, and as an output, generate a primary input encoding 218 by combining the input audio features 212 (or the input audio features stacked with the reference features 214) using a Feature-wise Linear Modulation (FiLM) layer 310 (Figure 3). Figure 3 provides an exemplary modulation conformable block 320 used by the primary encoder 210. Here, prior to each conformable block 320 in the primary encoder 210, the speaker embedding 215 (e.g., a d-vector) is combined with the input audio features 212 (or the stack of input audio and reference features 214) using the FiLM layer 310 to generate an output 312. FiLM enables the primary encoder 210 to adjust its encoding based on the speaker embedding 215 of the target speaker 10. A residual connection 314 is added after the FiLM layer 310, and the input audio features 212 (or the input audio features 212 stacked with the reference features 214) are combined with the output 312 of the FiLM layer 310 to generate, as an input, the modulated input features 316 of the conformable block 320 to ensure that the architecture can function well when the speaker embedding 215 is absent. Mathematically, the modulation conformable block 320 transforms the input feature x using the modulation feature m to generate the output feature y as follows. [Number]

[0040] Here, h(·) and r(·) are affine transformations. FFN, Conv, and MHSA represent a feed-forward module, a convolutional module, and a multi-head self-attention module, respectively. Equation 1 shows a Feature-wise Linear Modulation (FiLM) layer 310 with a residual connection.

[0041] Referring back to FIG. 2, the noise context encoder 220 is configured to receive, as an input, a context noise signal 213 that includes the noise before the target utterance, and to generate, as an output, a context noise encoding 222. The context noise signal 213 may include the LFBE features of the context noise signal. The noise context encoder 220 includes a standard conformable block without modulation by the speaker embedding 215, which is different from the primary and cross-attention encoders 210, 400. Since the context noise signal 213 is associated with the acoustic noise context before the target utterance 12 is spoken, it is assumed that the noise context encoder 220 does not modulate the context noise signal 213 using the speaker embedding 215, and thus includes information to be transferred to the cross-attention encoder 400 to assist in noise suppression.

[0042] Continuing to refer to FIG. 2, the cross-attention encoder 400 may be configured to receive, as inputs, the primary input encoding 218 generated as an output from the primary encoder 210, and the context noise encoding 222 generated as an output from the noise context encoder 220, and to generate, as an output, a cross-attention embedding 480. Thereafter, the decoder 240 is configured to decode the cross-attention embedding 480 into the enhanced input speech features 250 corresponding to the target utterance 12. The context noise encoding 222 may correspond to the auxiliary input. The decoder 240 may include a simple projection decoder having a single-layer frame-wise fully-connected network with sigmoid activation.

[0043] As shown in FIG. 4, the cross-attention encoder 400 can use each set of M modulated conformable blocks that respectively receive, as inputs, the main input encoding 218 modulated by the speaker embedding 215 using FiLM described in FIG. 3 and the context noise encoding 222 output from the noise context encoder 220. The cross-attention encoder 400 first independently processes the modulated input 218 and the auxiliary input 222 using a half feed-forward net 402, a first residual connection 404, a convolutional block 406, and a second residual connection 408. Specifically, the modulated input 218 is processed by a half feed-forward net 402a that generates an output 403a. Next, the first residual connection 404a combines the modulated input 218 with the output 403a of the half feed-forward net 402a to generate a modulated input feature 405a. The modulated input feature 405a is input to a convolutional block 406a that generates a convolutional output 407a. The second residual connection 408a combines the convolutional output 407a of the convolutional block 406a with the modulated input feature 405a to generate an output that includes a query vector 409a.

[0044] Similarly, the auxiliary input 222 is processed by a half feed-forward net 402b that generates an output 403b. Next, the first residual connection 404b combines the auxiliary input 222 with the output 403b of the half feed-forward net 402b to generate a modulated input feature 405b. The modulated input feature 405b is input to a convolutional block 406b that generates a convolutional output 407b. The second residual connection 408b combines the convolutional output 407b of the convolutional block 406b with the modulated input feature 405b to generate an output that includes a first key vector 409b and a first value vector 409c.

[0045] Subsequently, the multi-head cross-attention (MHCA) module 410 receives, as inputs, a query vector 409a, a first key vector 409b, and a first value vector 409c, summarizes these vectors 409a-c, and generates a noise summary 412. Intuitively, the role of the MHCA module 410 is to separately summarize the noise context for each input frame to be enhanced. The noise summary 412 output by the MHCA module 410 is then merged with the query vector 409a using a FiLM layer 420 that generates a FiLM output 422.

[0046] The multi-head self-attention (MHSA) layer 430 receives the FiLM output 422 as an input, merges the FiLM output 422 with the query vector 409a, and generates an attention output 432. A third residual connection 434 receives the query vector 409a and the attention output 432, combines the query vector 409a and the attention output 432, and generates a residual output 436. Next, the feed-forward module 440 receives the residual output 436 of the third residual connection 434 as an input and generates a feature output 442. Next, a fourth residual connection 444 combines the feature output 422 with the residual output 436 of the third residual output 434 to generate a merged input feature 446. Next, the merged input feature 446 is processed as an input by LayerNorm450, which is sent to the convolutional block 406b to generate a cross-attention embedding 480.

[0047] Mathematically, if x, m, and n are the encoded input, the d-vector, and the encoded noise context from the previous layer, respectively, the cross-attention encoder 400 performs the following.

Equation

[0048] The cross-attention encoder 400 generates, as output, a cross-attention embedding 480, which is passed to the next layer of M modulated conformable blocks, together with the d-vector m and the encoded noise context n. Thus, the input is modulated by each of the M conformable blocks by both the speaker embedding 215 related to the target speaker and the noise context encoding 222.

[0049] Figure 5 shows an exemplary training process 500 for training the context front-end processing model 200 to generate enhanced input audio features 250 when one or more of the context input features 213, 214, 215 are absent. The training process 500 can be executed on the remote system 130 of FIG. 1. As shown, the training process obtains one or more training data sets 520 stored in the data store 510 and trains the context front-end processing model 200 on the training data set 520. The data store 510 can reside on the memory hardware 136 of the remote system 130. Each training data set 520 includes a plurality of training examples 530, 530a - n, and each training example 530 can include a training utterance 532 paired with a corresponding training context signal 534, 534a - c. Specifically, the training context signal 534 includes a training context noise signal 534a that includes the noise before the corresponding training utterance 532, a training reference audio signal 534b, and a training speaker vector 534c that includes the voice characteristics of the target speaker who spoke the corresponding training utterance 532.

[0050] As described above with respect to FIG. 1, during speculation, the context front-end processing model 200 may not receive all of the context input features 213, 214, 215 at the same time. By training the context front-end processing model 200 with one or more missing training context signals 534, the context front-end processing model 200 becomes less reliant on the most relevant context input features 213, 214, 215 and more likely to utilize alternatives to the context input features 213, 214, 215. As a result, the context front-end processing model 200 can accurately predict the enhanced input audio feature 250 when one or more of the context input features 213, 214, 215 are missing. To keep the context front-end processing model 200 static, any missing training context signal 534 needs to be input to the context front-end processing model 200 in some way.

[0051] The training process 500 may also utilize a signal dropout model 550. The signal dropout model 550 receives the training context signal 534 as an input from the data store 510 and uses a context signal dropout strategy to dropout one or more of the training context signals 534 before training the context front-end processing model 200. The context signal dropout strategy of the signal dropout model 550 may include a predetermined probability (e.g., 50%, 20%, etc.) of dropping out each of the training context signals 534, and the same predetermined probability is used for each of the training context signals 534. In other words, in a given training example 530, the signal dropout model 550 may use the context signal dropout strategy to dropout the training context noise signal 534a with a predetermined probability of 50%, the training reference audio signal 534b with a predetermined probability of 50%, and the training speaker vector 534c with a predetermined probability of 50%. Similarly, in a given training example, the signal dropout model 550 may use the context signal dropout strategy to dropout the training context noise signal 534a with a predetermined probability of 20%, the training reference audio signal 534b with a predetermined probability of 20%, and the training speaker vector 534c with a predetermined probability of 20%.

[0052] In addition to the signal dropout strategy, the signal dropout model 550 can trim the length of the training context noise signal 534a to include the noise before the corresponding training utterance 532 with a uniform distribution length from zero to six (0 - 6) seconds. In other words, the signal dropout model 550 implements the signal dropout strategy and simultaneously trims the training context noise signal 534a. For example, in a given training example 530, even if the signal dropout model 550 does not dropout the training context noise signal 534a, the signal dropout model 550 can still trim the length of the training context noise signal 534a.

[0053] In some embodiments, the signal dropout model 550 drops out each training context signal 534 by using a signal dropout strategy to replace the corresponding training context signal 534 with all zeros based on a predetermined probability. In these embodiments, the signal dropout model 550 may replace the training context noise signal 534a with an all-zero feature having the same feature dimension as the corresponding training utterance 532 of a predetermined length. For example, the signal dropout strategy includes creating an all-zero feature that is six (6) seconds in length and has the same dimension as the LFBE feature. Similarly, the signal dropout model 550 may use a signal dropout strategy to replace the training reference audio signal 534b with an all-zero feature having the same length and feature dimension as the corresponding training utterance 532. Here, the feature dimension of the all-zero training reference audio signal 534b corresponds to the LFBE feature of the training reference audio signal 534b if the signal dropout strategy did not drop out the training reference audio signal 534b. Similarly, the signal dropout model 550 may use a signal dropout strategy to replace the training speaker vector 534c with an all-zero feature having an all-zero vector. Here, the training speaker vector 534c is replaced with a 256-dimensional all-zero vector. In other embodiments, the signal dropout model 550 drops out each training context signal 534 by using a signal dropout strategy to replace the corresponding training context signal 534 with a frame-level learned representation based on a predetermined probability.

[0054] In the example shown in FIG. 5, the signal dropout model 550 receives the training context signals 534a - c as inputs, and using a predetermined probability of the context signal dropout strategy, drops out the training reference audio signal 534b by replacing it with all - zero features of the same length and feature dimension as the corresponding training utterance 532. In other words, at this time step, the context front - end processing model 200 is trained only with the training context signals 534a and 534c, which approximates the conditions that the model 200 might encounter during inference when the context input features 213, 214, 215 contain only the context noise signal 213 and the speaker embedding 215.

[0055] After the signal dropout model 550 drops out the training reference audio signal 534b, the training utterance 532 and the training context signal 534 including the training reference audio signal 534b are replaced with all - zero features and dimensions, and the training reference audio signal 534b is simulated to provide a dropout for training the context front - end processing model 200. The context front - end processing model 200 receives, as inputs, the training utterance 532 and the training context signal 534 that simulates the dropout of the training reference audio signal 534b, and outputs a prediction y r is generated. The output prediction y r includes the enhanced input audio features 250 that are being tested for their accuracy. At each time step during the training process 500, the context front - end processing model 200 is additionally trained using the output prediction of the previous time step y r-1 .

[0056] FIG. 6 shows an exemplary training process 600 for calculating an ASR loss 640 when a context front-end processing model 200 is trained in integration with an ASR model 192. Here, only the encoder 620 of the ASR model 192 is used to calculate the loss. The ASR loss 640 is calculated as the l2 distance between the output of the ASR encoder 620 for the target features 540 of the training utterance 532 and the enhanced input speech features 250. The ASR encoder 620 is not updated during the training process 600. Specifically, the training process 600 uses the ASR encoder 620 of the ASR model 192 configured to receive, as input, the enhanced input speech features 250 predicted by the context front-end processing model 200 of the training utterance 532, to generate a predicted output 622 of the ASR encoder 620 for the enhanced input speech features 250, and uses the ASR encoder 620 configured to receive, as input, the target speech features 540 of the training utterance 532, to generate a target output 624 of the ASR encoder 620 for the target speech features 540, thereby calculating the ASR loss 640. The predicted output 622 of the enhanced input speech features 250 and the target output 624 of the target speech features 540 may each include respective sequences of LFBE features. Thereafter, the training process 600 calculates the ASR loss 640 via a loss module 630 based on the predicted output 622 of the ASR encoder 620 for the enhanced input speech features 250 and the target output 624 of the ASR encoder 620 for the target speech features 540. The goal of using the ASR loss 640 is to reinforce the context front-end processing model 200 to better conform to the ASR model 192, which is important for extracting the best performance from the context front-end processing model 200. By keeping the parameters of the ASR model 192 fixed, the ASR model 192 is decoupled from the context front-end processing model 200, thereby enabling each to be trained and deployed independently of each other.

[0057] In some embodiments, the context front-end processing model 200 is trained in integration with the ASR model 192 of the back-end automatic speech recognition system 180 using spectral loss and ASR loss 640. The training target 540 for training the context front-end processing model 200 uses an ideal ratio mask (IRM). The IRM is calculated using reverberant speech and reverberant noise based on the assumption that speech and noise are uncorrelated in the Mel spectrum space, as follows.

Number

Number

Number

Number

Number

[0058] During prediction, the estimated IRM is scaled and floored to reduce voice distortion at the expense of reducing noise suppression. This is particularly important because the ASR model 192 is vulnerable to voice distortion and non-linear front-end processing, which are one of the main challenges in improving the performance of a robust ASR model using an enhanced front-end. The enhanced features are derived as follows.

Number

Number

Number

[0059] FIG. 7 includes an exemplary flowchart of an exemplary arrangement of operations for a method 700 of training a generalized automatic speech recognition model using a context front-end processing model 200. At operation 702, method 700 includes receiving a plurality of training utterances 532 paired with corresponding training context signals 534, 534a-c. The training context signal 534 includes a training context noise signal 534a that includes noise prior to the corresponding training utterance 532, a training reference audio signal 534b, and a training speaker vector 534c that includes the voice characteristics of the target speaker who spoke the corresponding training utterance 532. Method 700 also includes, at operation 704, training the context front-end processing model 200 with the training utterances 532 to learn a method for predicting the enhanced voice features 250 using a context signal dropout strategy. Here, the context signal dropout strategy uses a predetermined probability of dropping out each of the training context signals 534 during training of the context front-end processing model 200 to simulate one or more of the training context signals 534 as being missing for training the model 200, and to learn a method for robustly generating the enhanced voice features 250 when any of the corresponding context input features are missing during prediction.

[0060] FIG. 8 is a schematic diagram of an exemplary computing device 800 that can be used to implement the systems and methods described in this document. Computing device 800 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The components shown here, their connections and relationships, and their functions are for purposes of illustration only and are not intended to limit the embodiments of the disclosure described and / or claimed in this document.

[0061] The computing device 800 includes a processor 810, a memory 820, a storage device 830, a high-speed interface / controller 840 connected to the memory 820 and the high-speed expansion port 850, and a low-speed interface / controller 860 connected to the low-speed bus 870 and the storage device 830. Each of the components 810, 820, 830, 840, 850, and 860 is interconnected using various buses and may be mounted on a common motherboard or exist in other ways as required. The processor 810 (e.g., the data processing hardware 112, 134 of FIG. 1) processes instructions for execution within the computing device 800, including instructions stored in the memory 820 or the storage device 830, and can display graphical information of a graphical user interface (GUI) on an external input / output device such as a display 880 connected to the high-speed interface 840. In other embodiments, multiple memories and multiple types of memories may be used, along with multiple processors and / or multiple buses as required. Also, multiple computing devices 800 may be connected so that each device provides multiple parts of multiple operations as needed (e.g., as a server bank, a group of blade servers, or a multiprocessor system).

[0062] The memory 820 (e.g., the memory hardware 114, 136 of FIG. 1) stores information non-temporarily within the computing device 800. The memory 820 may be a computer-readable medium, a volatile memory unit(s), or a non-volatile memory unit(s). The non-temporary memory 820 may be a physical device used to store programs (e.g., instruction sequences) or data (e.g., program state information) temporarily or permanently for use by the computing device 800. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as a boot program). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disks or tapes.

[0063] The storage device 830 can provide large-capacity storage to the computing device 800. In some embodiments, the storage device 830 is a computer-readable medium. In various different embodiments, the storage device 830 may be an array of devices including a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid-state memory device, or a storage area network or other configuration of devices. In additional embodiments, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more of the methods as described above. The information carrier is a computer-readable medium or a machine-readable medium such as the memory 820, the storage device 830, or the memory on the processor 810.

[0064] The high-speed controller 840 manages the bandwidth-intensive operations of the computing device 800, and the low-speed controller 860 manages the lower-bandwidth-intensive operations. Such role assignments are merely examples. In some embodiments, the high-speed controller 840 is coupled to a high-speed expansion port 850 that can accept the memory 820, the display 880 (e.g., via a graphics processor or accelerator), and various expansion cards (not shown). In some embodiments, the low-speed controller 860 is coupled to the storage device 830 and the low-speed expansion port 890. The low-speed expansion port 890 may include various communication ports (such as USB, Bluetooth, Ethernet, wireless Ethernet, etc.) and can be connected to one or more input / output devices such as a keyboard, a pointing device, a scanner, or a network device such as a switch or router via a network adapter.

[0065] As shown in the figure, the computing device 800 can be implemented in many different forms. For example, it may be implemented as a standard server 800a, or multiple times within a group of such servers 800a, as a laptop computer 800b, or as part of a rack server system 800c.

[0066] Various embodiments of the systems and techniques described herein can be realized in digital electronics and / or optical circuits, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can be special or general-purpose and can include embodiments in one or more computer programs executable and / or interpretable in a programmable system that includes at least one programmable processor, at least one input device, and at least one output device coupled to receive data and instructions from and to transmit data and instructions to a storage system.

[0067] A software application (i.e., a software resource) may refer to computer software that causes a computing device to perform a task. In some examples, a software application may be referred to as an "application", an "app", or a "program". Exemplary applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.

[0068] A non-transitory memory may be a physical device used to store a program (e.g., an instruction sequence) or data (e.g., program state information) temporarily or persistently for use by a computing device. The non-transitory memory may be a volatile and / or non-volatile addressable semiconductor memory. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as a boot program). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disks or tapes.

[0069] These computer programs (also known as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented in high-level procedural and / or object-oriented programming languages and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., magnetic disks, optical disks, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor that contains a machine-readable medium that receives the machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0070] The processes and logical flows described in this specification can be implemented by one or more programmable processors executing one or more computer programs to act on input data and generate output, also referred to as data processing hardware. The processes and logical flows can also be implemented by special purpose logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). Processors suitable for the execution of a computer program include, by way of example, both general and special purpose processors, as well as any one or more processors of any kind of digital computer. Generally, a processor receives instructions and data from a read only memory, a random access memory, or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer also includes, or is operatively coupled to receive from or transfer data to, or both, one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including by way of example semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0071] To interact with a user, one or more aspects of the present disclosure can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube), an LCD (liquid crystal display) monitor) or a touch screen for displaying information to the user, and optionally a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can input to the computer. Other types of devices can also be used to provide interaction with the user. For example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and the input from the user can be received in any form, including acoustic, speech language, or tactile input. Further, the computer can interact with the user by sending and receiving documents to and from the devices used by the user, for example, by sending a web page to a web browser on the user's client device in response to a request received from a web browser.

[0072] Some embodiments have been described. Nevertheless, it will be understood that various modifications can be made without departing from the spirit and scope of the present disclosure. Accordingly, other embodiments are within the scope of the following claims.

Claims

1. A computer-implemented method (700) for causing data processing hardware (134) to perform an operation when executed by the data processing hardware (134), the operation comprising: Receiving a plurality of training utterances (532) paired with corresponding training context signals (534, 534a-c), wherein the training context signals (534a-c) comprise: A training context noise signal (534a) including noise before the corresponding training utterance (532); A training reference audio signal (543b); A training speaker vector (534c) including voice characteristics of a target speaker who spoke the corresponding training utterance (532), and said receiving; Training a context front-end processing model (200) with the training utterances (532) to learn a method for predicting enhanced voice features (250) using a context signal dropout strategy, wherein the context signal dropout strategy drops out each of the training context signals (534) using a predetermined probability during training of the context front-end processing model (200), and said training; A computer-implemented method (700) comprising.

2. The computer-implemented method (700) according to claim 1, wherein the signal dropout strategy drops out each training context signal (534) by replacing the corresponding training context signal (534) with all zeros.

3. The computer-implemented method (700) according to claim 2, wherein replacing the training reference audio signal (543b) with all zeros comprises replacing the training reference audio signal (543b) with all-zero features having the same length and feature dimension as the corresponding training utterance (532).

4. The computer-implemented method (700) according to claim 2, wherein replacing the training context noise signal (534a) comprises replacing the training context noise signal (534a) with all-zero features having a predetermined length and the same feature dimension as the corresponding training utterance (532).

5. Replacing the training speaker vector (534c) includes replacing the training speaker vector (534c) with all-zero features having an all-zero vector, the computer-implemented method (700) according to claim 2.

6. The signal dropout strategy drops out each training context signal (534) by replacing the corresponding training context signal (534) with a frame-level learned representation, the computer-implemented method (700) according to claim 1.

7. The trained context front-end processing model (200) is a primary encoder (210), configured to receive, as input, input audio features (212) corresponding to a target utterance (12), and configured to generate, as output, a primary input encoding (218), the primary encoder (210); a noise context encoder (220), configured to receive, as input, a context noise signal (213) including noise prior to the target utterance (12), and configured to generate, as output, a context noise encoding (222), the noise context encoder (220); a cross-attention encoder (400), configured to receive, as input, the primary input encoding (218) generated as output from the primary encoder (210) and the context noise encoding (222) generated as output from the noise context encoder (220), and configured to generate, as output, a cross-attention embedding (480), the cross-attention encoder (400); a decoder (240) configured to decode the cross-attention embedding (480) into enhanced audio features (250) corresponding to the target utterance (12), the decoder (240); comprising the computer-implemented method (700) according to claim 1.

8. The primary encoder (210) further receives, as input, reference features (214) corresponding to a reference audio signal (154), and generates the primary input encoding (218) by processing the input audio features (212) stacked with the reference features (214), the computer-implemented method (700) according to claim 7.

9. The primary encoder (210) further receives, as input, a speaker embedding (215) including the voice characteristics of the target speaker (10) who spoke the target utterance (12), and generates, as output, the primary input encoding (218) by combining the input speech features (212) with the speaker embedding (215) using feature-wise linear modulation (FiLM). The computer-implemented method (700) according to claim 7, configured as such. **Claim 10** The cross-attention encoder (400) further receives, as input, the primary input encoding (218) modulated by the speaker embedding (215) using feature-wise linear modulation (FiLM), wherein the speaker embedding (215) includes the voice characteristics of the target speaker (10) who spoke the target utterance (12), processes the primary input encoding (218) modulated by the speaker embedding (215) and the context noise encoding (222), and generates, as output, the cross-attention embedding (480). The computer-implemented method (700) according to claim 7, configured to perform the above. **Claim 11** The primary encoder (210) includes N modulated conformer blocks (320). The noise context encoder (220) includes N conformer blocks and runs in parallel with the primary encoder (210). The cross-attention encoder (400) includes M modulated cross-attention conformer blocks. The computer-implemented method (700) of claim 7. **Claim 12** The computer-implemented method (700) according to claim 1, wherein the context front-end processing model (200) is integrated with the back-end automatic speech recognition (ASR) model (192) and trained using spectral loss and ASR loss (640). **Claim 13** The computer-implemented method (700) according to claim 12, wherein the spectral loss is based on the distance between the L1 loss function and the L2 loss function between the estimated ratio mask and the ideal ratio mask, and the ideal ratio mask is calculated using reverberant speech and reverberant noise. **Claim 14** For each training utterance (532), the ASR loss (640) Using the context signal dropout strategy, generate a predicted output (622) of the ASR encoder (620) for the enhanced speech features (250) using the ASR encoder (620) of the ASR model (192) configured to receive, as input, the enhanced speech features (250) predicted by the context front-end processing model (200) of the training utterance (532). Generate a target output (624) of the ASR encoder (620) for the target speech features (540) using the ASR encoder (620) configured to receive, as input, the target speech features (540) of the training utterance (532). Calculate the ASR loss (640) based on the predicted output (622) of the ASR encoder (620) for the enhanced speech features (250) and the target output (624) of the ASR encoder (620) for the target speech features (540). The computer-implemented method (700) according to claim 12, calculated by

15. A system (100) comprising data processing hardware (134) and memory hardware (136) that communicates with the data processing hardware (134) and stores instructions that, when executed on the data processing hardware (134), cause the data processing hardware (134) to receive a plurality of training utterances (532) paired with corresponding training context signals (534, 534a-c), wherein the training context signals (534a-c) include a training context noise signal (534a) including noise before the corresponding training utterance (532), a training reference audio signal (543b), and a training speaker vector (534c) including the voice characteristics of the target speaker who spoke the corresponding training utterance (532), the receiving To learn a method for predicting enhanced voice features (250) using a context signal dropout strategy, train a context front-end processing model (200) on the training utterance (532), wherein the context signal dropout strategy drops out each of the training context signals (534) using a predetermined probability during the training of the context front-end processing model (200), and execute an operation including the training. System (100). **Claim 16** The system (100) according to claim 15, wherein the signal dropout strategy drops out each training context signal (534) by replacing the corresponding training context signal (534) with all zeros. **Claim 17** The system (100) according to claim 16, wherein replacing the training reference audio signal (543b) with all zeros includes replacing the training reference audio signal (543b) with all-zero features having the same length and feature dimension as the corresponding training utterance (532). **Claim 18** The system (100) according to claim 16, wherein replacing the training context noise signal includes replacing the training context noise signal (534a) with all-zero features having a predetermined length and the same feature dimension as the corresponding training utterance (532). **Claim 19** The system (100) according to claim 16, wherein replacing the training speaker vector (534c) includes replacing the training speaker vector (534c) with all-zero features having an all-zero vector. **Claim 20** The system (100) according to claim 15, wherein the signal dropout strategy drops out each training context signal (534) by replacing the corresponding training context signal (534) with a frame-level learned representation. **Claim 21** The trained context front-end processing model (200) is a primary encoder (210), receives, as input, input voice features (212) corresponding to a target utterance (12), A primary encoder (210) configured to generate a primary input encoding (218) as an output; A noise context encoder (220), Receives, as input, a context noise signal (213) including noise before the target utterance (12), A noise context encoder (220) configured to generate a context noise encoding (222) as an output; A cross-attention encoder (400), Receives, as input, the primary input encoding (218) generated as an output from the primary encoder (210), and the context noise encoding (222) generated as an output from the noise context encoder (220), A cross-attention encoder (400) configured to generate a cross-attention embedding (480) as an output; A decoder (240) configured to decode the cross-attention embedding (480) into enhanced voice features (250) corresponding to the target utterance (12); The system (100) according to claim 15, comprising: **Claim 22** The primary encoder (210) further Receives, as input, a reference feature (214) corresponding to a reference audio signal (154), The system (100) according to claim 21, configured to generate the primary input encoding (218) by processing the input voice feature (212) stacked with the reference feature (214) as an output. **Claim 23** The primary encoder (210) further Receives, as input, a speaker embedding (215) including voice characteristics of a target speaker (10) who spoke the target utterance (12), The system (100) according to claim 21, configured to generate the primary input encoding (218) by combining the input voice feature (212) with the speaker embedding (215) using feature-wise linear modulation (FiLM) as an output. **Claim 24** The cross-attention encoder (400) further Receiving, as an input, the primary input encoding (218) modulated by speaker embedding (215) using feature-based linear modulation (FiLM), wherein the speaker embedding (215) includes voice characteristics of a target speaker (10) who spoke the target utterance (12), and the receiving; Processing the primary input encoding (218) and the context noise encoding (222) modulated by the speaker embedding (215) to generate, as an output, the cross-attention embedding (480); The system (100) according to claim 21, which is configured to perform the above.

25. The primary encoder (210) includes N modulation conformable blocks (320); The noise context encoder (220) includes N conformable blocks and executes in parallel with the primary encoder (210); The system (100) according to claim 21, wherein the cross-attention encoder (400) includes M modulation cross-attention conformable blocks.

26. The system (100) according to claim 15, wherein the context front-end processing model (200) is trained in integration with a back-end automatic speech recognition (ASR) model (192) using spectral loss and ASR loss.

27. The system (100) according to claim 26, wherein the spectral loss is based on the distance between the L1 loss function and the L2 loss function between the estimated ratio mask and the ideal ratio mask, and the ideal ratio mask is calculated using reverberant speech and reverberant noise.

28. For each training utterance (532), the ASR loss: Using the context signal dropout strategy, the ASR encoder (620) of the ASR model (192) configured to receive, as an input, the enhanced speech features (250) predicted by the context front-end processing model (200) of the training utterance (532), to generate a predicted output (622) of the ASR encoder (620) for the enhanced speech features (250); Using the ASR encoder (620) configured to receive, as an input, the target voice feature (540) of the training utterance (532), generate a target output (624) of the ASR encoder (620) for the target voice feature (540); Calculating the ASR loss (640) based on the predicted output (622) of the ASR encoder (620) for the enhanced voice feature (250) and the target output (624) of the ASR encoder (620) for the target voice feature (540); The system (100) according to claim 26, calculated by.