Hybrid Model Attention for Soft Streaming and Non-Streaming Automatic Speech Recognition

By integrating hybrid model attention with mixture model attention, ASR systems can flexibly switch between streaming and non-streaming modes, addressing inefficiencies and improving latency and accuracy in real-time transcription.

JP7693014B2Active Publication Date: 2025-06-16GOOGLE LLC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023558842
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-03-26
Filing Date
2021-12-15
Publication Date
2025-06-16
Estimated Expiration
2041-12-15

AI Technical Summary

Technical Problem

Existing automatic speech recognition (ASR) systems struggle to seamlessly switch between streaming and non-streaming modes, leading to inefficiencies in processing and increased latency, especially in applications requiring real-time transcription.

Method used

The integration of a hybrid model attention mechanism in ASR models, which allows for the use of mixture model (MiMo) attention to calculate attention probability distributions using softmax mixture components over a context window, enabling flexible switching between streaming and non-streaming modes.

Benefits of technology

This approach allows ASR models to efficiently switch between streaming and non-streaming modes without adding complexity or requiring additional training, resulting in improved latency and accuracy in real-time transcription applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007693014000004
    Figure 0007693014000004
  • Figure 0007693014000005
    Figure 0007693014000005
  • Figure 0007693014000006
    Figure 0007693014000006
Patent Text Reader

Abstract

A method (500) for an automatic speech recognition (ASR) model (200) for integrating streaming and non-streaming speech recognition, comprising receiving a sequence of acoustic frames (110). The method includes using an audio encoder (300) of the automatic speech recognition (ASR) model to generate a high-order feature representation for a corresponding acoustic frame in the sequence of acoustic frames. The method further includes using a joint network (230) of the ASR model to generate a probability distribution over possible speech recognition hypotheses at a corresponding time step based on the high-order feature representation generated by the audio encoder at the corresponding time step. The audio encoder includes a neural network that applies mixture model (MiMo) attention to compute an attention probability distribution function (PDF) using a set of mixture components of softmax over a context window.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to hybrid model attention for flexible streaming and non-streaming automatic speech recognition.

Background Art

[0002] Automatic speech recognition (ASR) systems have evolved into integrated models that directly map an audio waveform (i.e., the input sequence) to an output sentence (i.e., the output sequence) using a single neural network from multiple models where each model has a dedicated purpose. This integration results in a sequence-to-sequence approach where a sequence of audio features is given and a sequence of words (or graphemes) is generated. Using the integrated structure, all components of the model can be trained together as a single end-to-end (E2E) neural network. Here, the E2E model refers to a model whose architecture is entirely composed of neural networks. A complete neural network functions without external components and / or manually designed components (e.g., finite state transducers, lexicons, or text normalization modules). Furthermore, when training E2E models, these models generally do not require bootstrapping from decision trees or temporal alignment from separate systems. These E2E automatic speech recognition (ASR) systems have made significant progress and outperform conventional ASR systems in several common benchmarks including the word error rate (WER). The architecture of E2E ASR models mainly depends on the application. For example, in some applications involving user interaction such as voice search or on-device dictation, it is necessary for the model to perform recognition in a streaming manner. Other applications such as offline video captioning do not require the model to perform streaming and can utilize future context to improve performance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Non-Patent Literature

[0004]

Non-Patent Literature 1

Summary of the Invention

Means for Solving the Problems

[0005] One aspect of the present disclosure provides an automatic speech recognition (ASR) model for integrating streaming speech recognition and non-streaming speech recognition. The ASR model includes an audio encoder configured to receive a sequence of acoustic frames as input. The audio encoder is further configured to generate a high-order feature representation for a corresponding acoustic frame in the sequence of acoustic frames at each of a plurality of time steps. The ASR model further includes a joint network configured to receive, as input, the high-order feature representations generated by the audio encoder at each of the plurality of time steps. The joint network is further configured to generate, at each of the plurality of time steps, a probability distribution over possible speech recognition hypotheses at the corresponding time step. The audio encoder includes a neural network that applies mixture model (MiMo) attention to calculate an attention probability distribution function (PDF) using a set of softmax mixture components over a context window.

[0006] Implementations of the present disclosure may include one or more of any of the following features. In some implementations, the set of hybrid components of the MiMo attention is designed to cover all possible context use cases during inference. In other implementations, each hybrid component of the set of hybrid components includes a fully normalized softmax.

[0007] Furthermore, the ASR model may switch between the streaming mode and the non-streaming mode by adjusting the hybrid weights of the MiMo attention. In some implementations, the neural network of the audio encoder includes a plurality of conformer layers. In other implementations, the neural network of the audio encoder includes a plurality of transformer layers.

[0008] In some implementations, the ASR model receives, as input, a sequence of non-blank symbols output by the final softmax layer, and includes a label encoder configured to generate a high-density representation at each of a plurality of time steps. In these implementations, the joint network is further configured to receive, as input, the high-density representation generated by the label encoder at each of the plurality of time steps. In these implementations, the label encoder may include a neural network of a transformer layer, a conformer layer, or a long short-term memory (LSTM) layer.

[0009] Furthermore, the label encoder may include a lookup table embedding model configured to search for a high-density representation at each of the plurality of time steps. In some examples, the possible speech recognition hypotheses generated at each of the plurality of time steps correspond to a set of output labels where each output label represents a grapheme or word piece in natural language.

[0010] Another aspect of the present disclosure provides a computer-implemented method for an automatic speech recognition (ASR) model to integrate streaming speech recognition and non-streaming speech recognition. The computer-implemented method, when executed on data processing hardware, causes the data processing hardware to perform operations including receiving a sequence of acoustic frames. Further operations performed at each of a plurality of time steps include generating a high-order feature representation for a corresponding acoustic frame in the sequence of acoustic frames using an audio encoder of the ASR model. The operations further include generating a probability distribution over possible speech recognition hypotheses at a corresponding time step based on the high-order feature representation generated by the audio encoder at the corresponding time step using a joint network of the ASR model. The audio encoder includes a neural network that applies a mixture model (MiMo) attention to compute an attention probability distribution function (PDF) using a set of softmax mixture components over a context window.

[0011] This aspect may include one or more of any of the following features. In some implementations, the set of mixture components of the MiMo attention is designed to cover all possible context use cases during inference. In other implementations, each mixture component of the set of mixture components includes a fully normalized softmax.

[0012] Furthermore, the ASR model may switch between a streaming mode and a non-streaming mode by adjusting the mixture weights of the MiMo attention. In some implementations, the neural network of the audio encoder includes a plurality of conformer layers. In other implementations, the neural network of the audio encoder includes a plurality of transformer layers.

[0013] In some implementations, the operation further includes generating a high-density representation using a label encoder of an ASR model configured to receive, at each of a plurality of time steps, a sequence of non-blank symbols output by a final softmax layer of the ASR model. In these implementations, the operation of generating a probability distribution over possible speech recognition hypotheses at a corresponding time step is further based on a sequence of non-blank output symbols output by the final softmax layer at the corresponding time step. In these implementations, the label encoder may include a neural network of a transformer layer, a conformer layer, or a long short-term memory (LSTM) layer.

[0014] Further, the label encoder may include a lookup table embedding model configured to search for a high-density representation at each of a plurality of time steps. In some examples, the possible speech recognition hypotheses generated at each of the plurality of time steps correspond to a set of output labels where each output label represents a grapheme or word piece in natural language.

[0015] Details of one or more implementations of the present disclosure are set forth in the accompanying drawings and the following description. Other aspects, features, and advantages will become apparent from the description and drawings, and from the claims.

Brief Description of the Drawings

[0016]

Figure 1A

Figure 1B

Figure 2

Figure 3A

Figure 3B

Figure 3C

Figure 4

Figure 5

Figure 6

DETAILED DESCRIPTION OF THE INVENTION

[0017] Like reference symbols in the various drawings indicate like elements.

[0018] Automatic speech recognition (ASR) systems emphasize achieving not only quality / accuracy (e.g., low word error rate (WER)) but also low latency (e.g., short delay from when a user speaks until text is generated). In recent years, end-to-end (E2E) ASR models have gained significant attention for achieving state-of-the-art performance in terms of accuracy and latency. Different from traditional hybrid ASR systems that include separate acoustic, pronunciation, and language models, E2E models apply a sequence-to-sequence approach to jointly learn acoustic and language modeling in a single neural network that is trained end-to-end from training data, e.g., utterance-transcription pairs. Here, E2E models refer to models whose architecture is generally composed of neural networks. A complete neural network functions without external components and / or manually designed components (e.g., finite state transducers, lexicons, or text normalization modules). Furthermore, when training E2E models, these models generally do not require bootstrapping from decision trees or temporal alignment from separate systems.

[0019] When using an ASR system, there are cases where the ASR system needs to decode speech in a streaming manner that corresponds to displaying a description of the speech in real time or, in some cases, faster than real time when the user speaks. As an example, when an ASR system is displayed on a user computing device where direct user interaction occurs, such as a mobile phone, in an application that runs on the user device and uses the ASR system (e.g., a digital assistant application), it may be necessary to perform speech recognition in a streaming manner so that words, word pieces, and / or individual characters appear on the screen immediately after they are spoken. In addition, the user of the user device may have a low tolerance for latency. For example, when the user speaks a query asking the digital assistant to retrieve details about an appointment scheduled for the near future from a calendar application, the user expects the digital assistant to provide a response communicating the retrieved details as quickly as possible. Due to such low tolerance, efforts are being made to run the ASR system on the user device to minimize the impact of latency and inaccuracy that may negatively affect the user experience. However, attention-based sequence-to-sequence models such as the listen-attend-spell (LAS) model, which function by looking at the entire audio input sequence before generating the output text, do not allow for streaming the output when the input is received. Due to this drawback, problems may arise when introducing an attention-based sequence-to-sequence model for voice applications that are susceptible to latency and / or require real-time voice transcription. Therefore, the LAS model alone does not become an ideal model for applications that are susceptible to latency and / or applications that provide a streaming transcription function in real time when the user speaks.

[0020] Another form of sequence-to-sequence model known as the recurrent neural network transducer (RNN-T) does not use an attention mechanism and, unlike other sequence-to-sequence models that generally need to process an entire sequence (e.g., an audio waveform) to generate an output (e.g., a sentence), the RNN-T processes input samples continuously and streams output symbols. This is a particularly attractive feature for real-time communication. For example, speech recognition using RNN-T can output characters one by one during speech. Here, the RNN-T uses a feedback loop that feeds back the symbol predicted by the model to the RNN-T itself to predict the next symbol. Since the decoding of RNN-T involves beam search through a single neural network rather than a large decoder graph, the RNN-T can scale to a small fraction of the size of a server-based speech recognition model. Due to the size reduction, the RNN-T can sometimes be placed entirely on-device and run offline (i.e., without a network connection), thus avoiding reliability issues with the communication network.

[0021] The RNN-T model still lags behind conventional models of the large-scale state-of-the-art (e.g., server-based models with separate AM, PM, and LM) and attention-based sequence-to-sequence models (e.g., LAS models) in terms of quality (e.g., speech recognition accuracy often measured by word error rate (WER)) due to the inability to apply look-ahead audio context (e.g., right context) when predictive recognition is obtained. In recent years, the Transformer-Transducer (T-T) and Conformer-Transducer (C-T) models have gained more attention than the RNN-T model due to their ability to process the entire audio sequence and calculate the attention probability density function at each time step. Thus, the T-T model and the C-T model can be used for either streaming speech recognition or non-streaming speech recognition due to their ability to use look-ahead context. However, the two models are trained separately to perform streaming speech recognition and non-streaming speech recognition respectively. For example, a user uses a streaming speech recognition model such as an RNN-T model, a C-T model, or a T-T model to recognize conversational queries, and a separate non-streaming speech recognition model to recognize non-conversational queries. Generally, an application through which a user sends their utterance can be used to identify which of the streaming speech recognition model or the non-streaming speech recognition model to use for speech recognition. Since different separate speech recognition models are required to perform speech recognition depending on the application and / or query type, it is computationally costly and requires sufficient memory capacity to store each model on the user device. Even if one of the models is executable on a remote server, the additional cost and bandwidth constraints for connecting to the remote server may affect speech recognition performance and ultimately the user experience.

[0022] The implementations in this specification aim to integrate a streaming acoustic encoder and a non-streaming acoustic encoder by training an acoustic recognition model (e.g., a T-T model or a C-T model) using hybrid model attention, enabling the trained acoustic recognition model to switch between a streaming speech recognition mode and a non-streaming speech recognition mode during inference. The basic limitation of transformers and conformer-based acoustic encoders lies in the calculation of the attention probability density function (PDF) at each time step. Specifically, using a single softmax across the entire context window constrains the model to use the same context window size during both training and inference. The proposed hybrid model (MiMo) attention provides a flexible alternative that reduces the requirement to constrain the context window size during inference to the same context window size used during training. As will become apparent, MiMo attention calculates the attention PDF using a mixture model of softmaxes across the context window. The support for the mixture components is designed to cover all possible use cases during inference, such as full / entire sequence context, limited left + right context, and left-only context. Since each mixture component is a fully normalized softmax, any subset of the mixture components can be applied during inference without causing inconsistencies. Thus, C-T models and T-T models trained using MiMo attention can be used in various streaming and non-streaming modes during inference depending on the current application. In particular, MiMo attention does not add any additional complexity, parameters, or training loss to the model.

[0023] FIG. 1A and FIG. 1B are examples of the speech environments 100, 100a-100b. In the speech environment 100, the way the user 104 interacts with a computing device such as the user device 10 may be by voice input. The user device 10 (commonly also referred to as the device 10) is configured to capture voice (e.g., streaming audio data) from one or more users 104 within the speech environment 100. Here, the streaming audio data may refer to the speech 106 by the user 104 acting as an audible query, a command to the device 10, or an audible communication captured by the device 10. The speech-responsive system of the device 10 may appropriately process the query or command by answering the query and / or causing the command to be executed / fulfilled by one or more downstream applications.

[0024] The user device 10 may correspond to any computing device associated with the user 104 and capable of receiving audio data. Some examples of user devices include, but are not limited to, mobile devices (e.g., mobile phones, tablets, laptops, etc.), computers, wearable devices (e.g., smartwatches), smart home appliances, Internet of Things (IoT) devices, in-vehicle infotainment systems, smart displays, smart speakers, and the like. The user device 10 includes data processing hardware 12 and memory hardware 14 that communicates with the data processing hardware 12 and stores instructions. The instructions, when executed by the data processing hardware 12, cause the data processing hardware 12 to perform one or more operations. The user device 10 further includes an audio system 16 having an audio capture device (e.g., a microphone) 16, 16a for capturing an utterance 106 within the utterance environment 100 and converting it into an electrical signal, and an audio output device (e.g., a speaker) 16, 16b for communicating an audible audio signal (e.g., as output audio data from the device 10). The user device 10 implements a single audio capture device 16a in the illustrated example, but an array of audio capture devices 16a may be implemented without departing from the scope of the present disclosure, whereby one or more capture devices 16a within the array may not physically exist on the user device 10 but may communicate with the audio system 16.

[0025] In the speech environment 100, an automatic speech recognition (ASR) system 109 that implements an ASR model 200 trained using a hybrid model (MiMo) attention exists on the user device 10 of the user 104 and / or on a remote computing device 60 that communicates with the user device 10 via the network 40 (for example, one or more remote servers of a distributed system executed in a cloud computing environment). The user device 10 and / or the remote computing device 60 also receive the speech 106 uttered by the user 104 and captured by the audio capture device 16a, and are configured to convert the speech 106 into a corresponding digital format associated with an input acoustic frame 110 that can be processed by the ASR system 109, including an audio subsystem 108. In the example shown in FIG. 1A, the user 104 utters each speech 106, and the audio subsystem 108 converts the speech 106 into a corresponding audio data (for example, acoustic frame) 110 input to the ASR system 109. Thereafter, the model 200 receives the audio data 110 corresponding to the speech 106 as an input, and generates / predicts the corresponding transcription 120 (also referred to as the recognition result / hypothesis 120) of the speech 106 as an output.

[0026] The model 200 includes an acoustic encoder 300 trained using MiMo attention that calculates the attention PDF at each time step using a softmax mixture model over a context window. The support for the mixture components is designed to cover all possible use cases during inference, such as full / whole sequence context, limited left+right context, and left-only context. Since each mixture component is a fully normalized softmax, any subset of the mixture components can be applied during inference without causing inconsistencies.

[0027] Model 200 also includes decoders 220, 230 that enable Model 200 to operate in both a streaming mode and a non-streaming mode (unlike, for example, two separate models where each model is dedicated to either streaming mode or non-streaming mode). For example, as shown in FIG. 1A, in the digital assistant application 50 running on the user device 10, it may be necessary for speech recognition to be streamed such that words, word pieces, and / or individual characters appear on the screen immediately after being spoken. Additionally, user 104 of user device 10 may have a low tolerance for latency when issuing queries executed by digital assistant application 50. In these scenarios where the application requires minimizing latency, Model 200 may provide a streaming transcription function in real time when user 104 makes utterance 106.

[0028] The user device 10 and / or the remote computing device 60 also execute a user interface generator 107 configured to present an expression of the transcription 120 of the utterance 106 to the user 104 of the user device 10. As will be described in more detail below, the user interface generator 107 may stream and display the speech recognition result 120a, which is then displayed as the final speech recognition result 120. In some configurations, the transcription 120 output from the ASR system 109 is processed by, for example, a natural language understanding (NLU) module executed on the user device 10 or the remote computing device 60 to execute the user command / query specified by the utterance 106. Additionally or alternatively, a text-to-speech system (not shown) (e.g., executed on any combination of the user device 10 or the remote computing device 60) may convert the transcription 120 into synthesized speech for audible output by the user device 10 and / or another device.

[0029] In the example of FIG. 1A, the user 104 in the speech environment 100a interacts with a program or application 50 (e.g., digital assistant application 50a) of the user device 10 that uses the ASR system 109. For example, FIG. 1A shows that the user 104 is communicating with the digital assistant application 50a, and the digital assistant application 50a is displaying a digital assistant interface 18 on the screen of the user device 10, showing a conversation between the user 10 and the digital assistant of the digital assistant application 50a. In this example, the user 104 asks the digital assistant application 50a, "What time does the concert start tonight?" This question from the user 104 is a speech 106 captured by the audio capture device 16a and processed by the audio system 16 of the user device 10. In this example, the audio system 16 receives the speech 106 and converts it into an acoustic frame 110 that is input to the ASR system 109.

[0030] Continuing with the above example, the model 200 encodes the acoustic frame 110 using an audio encoder 300 (i.e., FIG. 2) while receiving the acoustic frame 110 corresponding to the speech 106 when the user 104 speaks, and then decodes the encoded representation of the acoustic frame 110 into a streaming speech recognition result 120a using decoders 220, 230 (FIG. 2). The user interface generator 107 presents the representation of the streaming speech recognition result 120a of the speech 106 to the user 104 of the user device 10 in a streaming manner via the digital assistant interface 18, such that words, word pieces, and / or individual characters appear on the screen immediately after they are spoken. The user interface generator 107 also presents the representation of the final speech recognition result 120 of the speech 106 to the user 104 of the user device 10 via the digital assistant interface 18.

[0031] In the example shown in FIG. 1A, the digital assistant application 50a may respond to a question presented by the user 104 using natural language processing. Natural language processing generally refers to the process of interpreting written language (e.g., partial speech recognition result 120a and / or final speech recognition result 120b) and determining whether the written language prompts some action. In this example, the digital assistant application 50a uses natural language processing to recognize that the question from the user 104 pertains to the user's environment and more specifically to the song being played near the user. By recognizing these details using natural language processing, the automated assistant returns a response 19 to the user's query. The response 19 states, "The opening is at 6:30 PM." In some configurations, natural language processing is performed on a remote computing device 60 that communicates with the data processing hardware 12 of the user device 10.

[0032] Figure 1B is another example of speech recognition using the ASR system 109 in the speech environment 100b. As shown in this example, the user 104 displays the voicemail application interfaces 18, 18b on the screen of the user device 10 and interacts with the voicemail applications 50, 50b that transcribe the voicemail left for the user 104 by Jane Doe into text. In this example, latency is not important, but the accuracy of the transcription when processing long-tail proper nouns or rare words is important. The model 200 of the ASR system 109 can utilize the full context of the audio by waiting until all the acoustic frames 110 corresponding to the voicemail are generated. This voicemail scenario also illustrates how the model 200 can handle long-form utterances, which is often due to voicemails being multiple sentences or, in some cases, several paragraphs. The ability to handle long-form utterances is particularly advantageous compared to other ASR models such as a two-pass model using an LAS decoder. The reason is that in this two-pass model, long-form problems (e.g., a high word deletion rate for long-form utterances) occur when applied to long-form conditions.

[0033] Continuing to refer to Figure 1B, as described with respect to Figure 1A, the model 200 encodes the acoustic frames 110 using the audio encoder 300 that receives the acoustic frames 110. Compared to the example of Figure 1A where the audio encoder 300 operates in streaming mode using only the left context, the audio encoder 300 switches to operate in non-streaming mode using both the left and right contexts. Here, the mixing weights of the audio encoder 300 trained using MiMo attention may be set to appropriate values associated with the non-streaming mode. After the model 200 receives all of the acoustic frames 110 and encodes them using the audio encoder 300, it provides the audio encoder output as an input to the decoders 220, 230 to generate the final speech recognition result 120b.

[0034] Referring to FIG. 2, the model 200 may perform end-to-end (E2E) speech recognition by integrating an acoustic model, a pronunciation model, and a language model into a single neural network, and may not require a lexicon or a separate text normalization component. Various structures and optimization mechanisms can improve accuracy and shorten the model training time. The model 200 may include a Transformer-Transducer (T-T) or Conformer-Transducer (C-T) model architecture, which complies with the latency constraints associated with interactive applications. An exemplary T-T model is described in U.S. Patent Application No. 17 / 210,465 filed on March 23, 2021, the content of which is incorporated herein by reference in its entirety. An exemplary C-T model is described in "Conformer: Convolution-augmented Transformer for Speech Recognition," arxiv.org / pdf / 2005.08100, the content of which is incorporated herein by reference in its entirety. The model 200 has a small computational footprint and lower memory requirements than conventional ASR architectures, such that the model architecture is suitable for performing speech recognition entirely on the user device 10 (e.g., communication with the remote server 60 is not required). The model 200 includes an audio encoder 300, a label encoder (e.g., a prediction network) 220, and a joint network 230. The label encoder 220 and the joint network 230 collectively form a decoder. The audio encoder 300 is generally similar to an acoustic model (AM) in a conventional ASR system and includes a neural network having a plurality of conformer layers or transformer layers. For example, the audio encoder 300 is a sequence x = (x1, x2, ..., x of d-dimensional feature vectors (e.g., acoustic frames 110 (FIG. 1)) T )(

[0035] [Number]

[0036] ) is read and high-order feature representations are generated at each time step. These high-order feature representations are ah1, ..., ah T shown as.

[0037] Similarly, the label encoder 220 may also include a neural network of a transformer layer, a conformer layer, a long short-term (LSTM) memory layer, or a lookup table embedding model. The lookup table embedding model, similar to a language model (LM), encodes the predicted label history of the sequence of non-blank symbols y0, ..., y ui-1 previously output by the final softmax layer 240 into a high-density representation Ih u for processing as.

[0038] Finally, by the T-T or C-T model architecture, the representations generated by the audio encoder 300 and the label encoder 220 are combined by the joint network 230 using a high-density layer J u,t . The joint network 230 then outputs P(z u,t |x,t,y1, …, y u-1) That is, it predicts the distribution over the next output symbols. In other words, the joint network 230 generates a probability distribution over the possible speech recognition hypotheses at each output step (e.g., time step). Here, the "possible speech recognition hypotheses" correspond to a set of output labels (also referred to as "speech units") that represent graphemes (e.g., symbols / characters) or word pieces in the natural language where each output label is specified. For example, if the natural language is English, the set of output labels may include 27 symbols, e.g., one label for each of the 26 letters in the English alphabet and one label for specifying space. Thus, the joint network 230 may output a set of values indicating the likelihood of occurrence of each output label in a predetermined set of output labels. This set of values can be a vector and can indicate a probability distribution over the set of output labels. In some cases, the output labels are graphemes (e.g., individual characters, and optionally punctuation marks and other symbols), but the set of output labels is not so limited. For example, the set of output labels can include word pieces and / or whole words in addition to and / or instead of graphemes. The output distribution of the joint network 230 can include posterior probability values for each of the different output labels. Thus, if there are 100 different output labels representing different graphemes or other symbols respectively, the output z u,t of the joint network 230 can include 100 different probability values, one for each output label. Then, in the beam search process (e.g., by the softmax layer 240) for determining the transcription 120, the probability distribution can be used to select scores and assign them to candidate orthographic elements (e.g., graphemes, word pieces, and / or words).

[0039] The softmax layer 240 may select, using any technique, the output label / symbol having the highest probability in the distribution as the next output symbol predicted by the model 200 at the corresponding output step. In this way, the model 200 does not make the assumption of conditional independence, and the prediction of each symbol is conditioned not only on the acoustics but also on the sequence of labels output so far.

[0040] Embodiments herein aim to replace the multi-head attention mechanism in the transformer or conformer layer of the audio encoder 300 with MiMo attention to increase flexibility when switching between streaming mode and non-streaming mode during inference. MiMo attention is used to calculate the output sequence (y0, ..., y T-1 ) of feature vectors from the input sequence (x0, ..., x T-1 ) of feature vectors by generating a PDF over a plurality of T time steps. For the plurality of T time steps for generating y k at the current time step, unnormalized attention scores

[0041]

Number

[0042] are shown. In this formulation, it is not assumed how these scores are generated. For example, they could have been calculated by taking the dot product between the query vector for time step k and all T key vectors. Conventional attention uses a single softmax to convert the attention scores to a PDF over T time steps, whereas MiMo attention instead applies a mixture of M softmaxes to calculate the attention PDF as follows.

[0043]

Number

[0044] In the above formula, w0, ..., w M-1 are the mixing weights. Advantageously, the mixture model allows for determining / setting the support of the softmax mixture components during training to accommodate all possible context use cases during inference. This flexibility is not available with standard / conventional attention.

[0045] In a non-limiting example, M softmaxes are set for two mixture components including a first component over k - L, k time steps and a second component over [k+1, k+R]. In other words, the first softmax acts over the left + central context of L+1 time steps, and the second softmax acts over the context to the right of R time steps relative to the central frame at time step k. The schematic diagrams 301 in FIGS. 3A - 3C show attention probability matrices 301a, 301b, 301c of size T×T using MiMo attention for the simple case of M = 2 mixture components. The T×T attention probability matrix 301a in FIG. 3A is written as a mixture model composed of a left-only (causal, shown by horizontal dotted line) attention matrix and a right-only (anti-causal, shown by vertical dotted line) attention matrix. MiMo attention decomposes this banded diagonal matrix into a weighted sum of a lower banded diagonal matrix for the left + central context and an upper banded diagonal matrix for the right context. All matrices are row probabilities (i.e., the sum of each row is 1) and thus represent appropriately normalized attention PDFs.

[0046] During inference, it is possible to use both left and right contexts (non-streaming mode), left-only context (streaming mode) by setting the mixing weight w0 = 1, or right-only context by setting w1 = 1. It can be easily seen that MiMo attention is a general framework that can accommodate all types of contexts and use cases during inference.

[0047] Even if the ASR model is trained using MiMo attention, the training loss does not change, nor is any additional complexity added. In fact, adding noise to the mixing weights during training may be useful. In the case of M = 2 described above, u~uniform(0, w1) may be sampled for each training batch and used when perturbing the mixing weights as w0 + u and w1 - u.

[0048] Figure 4 shows table 400 of WER for test clean and test other for an exemplary dataset. Table 400 shows the WER of the baseline and MiMo streaming / non-streaming conformers for Librispeech 960. In table 400, (L, R) indicates the size of the left + center and right contexts. All checkpoints are selected based on the dev-clean WER. Two baseline ASR models are trained: a left-only context model with a context size of 65 and a non-streaming model with a left context of 65 and a right context of 64. The baseline non-streaming (65, 64) model results in a lower WER than the corresponding streaming (65, 0) model when using the aligned non-streaming (65, 64) inference graph.

[0049] The model 200 may be trained on a remote server. The trained model 200 may be pushed to a user device for performing on-device speech recognition.

[0050] FIG. 5 is a flowchart of an exemplary configuration of operations for a method 500 of integrating streaming speech recognition and non-streaming speech recognition. Data processing hardware 12 (FIG. 1A) may execute instructions stored on memory hardware 14 (FIG. 1A) to implement the exemplary configuration of operations for method 500. Method 500 includes, at operation 502, receiving a sequence of acoustic frames 110. Method 500 may then proceed to operations 504 and 506, which are executed at each of a plurality of time steps.

[0051] Method 500 includes, at operation 504, using an audio encoder 300 of an automatic speech recognition (ASR) model 200 to generate a higher-order feature representation for a corresponding acoustic frame 110 in the sequence of acoustic frames 110. Method 500 includes, at operation 506, using a joint network (230) of the ASR model to generate a probability distribution over possible speech recognition hypotheses at a corresponding time step based on the higher-order feature representation generated by the audio encoder 300 at the corresponding time step. Here, the audio encoder 300 includes a neural network that applies a mixture model (MiMo) attention and calculates an attention probability distribution function (PDF) using a set of softmax mixture components over a context window.

[0052] A software application (i.e., a software resource) may refer to computer software that causes a computing device to perform tasks. In some instances, a software application may be referred to as an “application,” an “app,” or a “program.” Exemplary applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, document creation applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.

[0053] A non-transitory memory may be a physical device used to temporarily or permanently store a program (e.g., a sequence of instructions) or data (e.g., program state information) for use by a computing device. The non-transitory memory may be a volatile and / or non-volatile addressable semiconductor memory. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., generally used for firmware such as a boot program). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), as well as disks or tapes.

[0054] FIG. 6 is a schematic diagram of an exemplary computing device 600 that may be used to implement the systems and methods described in this document. Computing device 600 is intended to represent various forms of digital computers, such as a laptop, desktop, workstation, personal digital assistant, server, blade server, mainframe, and other appropriate computers. The components shown here, their connections and relationships, and their functions are exemplary only and are not intended to limit the implementation forms of the inventions described and / or claimed in this document.

[0055] The computing device 600 includes a processor 610, a memory 620, a storage device 630, a high-speed interface / controller 640 connected to the memory 620 and the high-speed expansion port 650, and a low-speed interface / controller 660 connected to the low-speed bus 670 and the storage device 630. Each of the components 610, 620, 630, 640, 650, and 660 may be interconnected using various buses and may be mounted on a common motherboard or otherwise mounted as appropriate. The processor 610 processes instructions executed within the computing device 600, including instructions stored in the memory 620 or on the storage device 630, to display graphical information about a graphical user interface (GUI) on an external input / output device such as a display 680 coupled to the high-speed interface 640. In other implementations, multiple processors and / or multiple buses may be used, along with multiple memories and multiple types of memories, as appropriate. Also, multiple computing devices 600 may be connected and each device may provide a part of the required operations (e.g., as a server bank, a group of blade servers, or a multiprocessor system).

[0056] Memory 620 stores information non - transiently within computing device 600. Memory 620 may be a computer - readable medium, a volatile memory unit, or a non - volatile memory unit. The non - transient memory 620 may be a physical device used to store a program (e.g., a sequence of instructions) or data (e.g., program state information) temporarily or persistently for use by computing device 600. Examples of non - volatile memory include, but are not limited to, flash memory and read - only memory (ROM) / programmable read - only memory (PROM) / erasable programmable read - only memory (EPROM) / electrically erasable programmable read - only memory (EEPROM) (e.g., generally used for firmware such as a boot program). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase - change memory (PCM), and disks or tapes.

[0057] Storage device 630 can provide a mass storage device for computing device 600. In some implementations, storage device 630 is a computer - readable medium. In various different implementations, storage device 630 may be a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory, or other similar solid - state memory devices, or an array of devices including devices within a storage area network or other configurations. In additional implementations, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more of the methods as described above. The information carrier is a computer or machine - readable medium such as memory 620, storage device 630, or memory on processor 610.

[0058] The high-speed controller 640 manages bandwidth-intensive operations for the computing device 600, and the low-speed controller 660 manages lower bandwidth-intensive operations. Such duty assignments are merely exemplary. In some implementations, the high-speed controller 640 is coupled to the memory 620, coupled to the display 680 (e.g., via a graphics processor or accelerator), and coupled to a high-speed expansion port 650 that can receive various expansion cards (not shown). In some implementations, the low-speed controller 660 is coupled to the storage device 630 and the low-speed expansion port 690. The low-speed expansion port 690 may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), and may be coupled to one or more input / output devices such as a keyboard, pointing device, scanner, or networking devices such as switches or routers, e.g., via a network adapter.

[0059] The computing device 600 may be implemented in several different forms as shown. For example, the computing device 600 may be implemented as a standard server 600a, or multiple times as a group of such servers 600a, or may be implemented as a laptop computer 600b, or may be implemented as part of a rack server system 600c.

[0060] The various implementations of the systems and techniques described in this specification can be realized in digital electronics and / or optical circuits, integrated circuits, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementations in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which processor may be special purpose or general purpose and may be coupled to receive and transmit data and instructions between a memory system, at least one input device, and at least one output device.

[0061] These computer programs (also called programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented in high-level procedural and / or object-oriented programming languages, and / or in assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., magnetic disks, optical disks, memory, Programmable Logic Device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives the machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0062] The processes and logic flows described in this specification can be executed by one or more programmable processors, which are also referred to as data processing hardware, and which function by executing one or more computer programs to process input data and generate outputs. The processes and logic flows can also be executed by dedicated logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). Processors suitable for executing a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. In general, a processor receives instructions and data from a read only memory or a random access memory or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. In general, a computer also includes one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or is operatively coupled to receive data from, transfer data to, or both, such devices, but a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and memory devices, including by way of example semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices, magnetic disks, such as internal hard disks or removable disks, magneto-optical disks, and CD ROM and DVD-ROM disks. The processor and the memory can be assisted by, or incorporated in, dedicated logic circuitry.

[0063] To enable interaction with a user, one or more aspects of the present disclosure can be implemented on a computer having a display device for presenting information to the user, such as a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touch screen, and optionally a keyboard and a pointing device by which the user can provide input to the computer, such as a mouse or trackball. Other types of devices can also be used to enable interaction with the user. For example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and the input from the user can be received in any form including acoustic, speech, or tactile input. Additionally, the computer can send and receive documents to and from devices used by the user, for example, by sending a web page to a web browser on the user's client device in response to a request received from the web browser to interact with the user.

[0064] Some implementations have been described. Nevertheless, it will be understood that various modifications may be made without departing from the spirit and scope of the present disclosure. Accordingly, other implementations are within the scope of the following claims.

Explanation of Reference Numerals

[0065] 10 User device 12 Data processing hardware 14 Memory hardware 16 Audio system 16, 16a Audio capture device 16, 16b Voice output device 18 Digital assistant interface 18, 18b Voice mail application interface 19 Response 40 Network 50, 50a Digital Assistant Application 50, 50b Voice Mail Application 60 Remote Computing Device, Remote Server 100, 100a, 100b Speech Environment 104 User 106 Speech 107 User Interface Generator 108 Audio Subsystem 109 ASR System 110 Acoustic Frame 120 Transcription, Recognition Result / Hypothesis 120a Speech Recognition Result, Partial Speech Recognition Result 120b Final Speech Recognition Result 200 ASR Model 220, 230 Decoder 220 Label Encoder 230 Joint Network 240 Final Softmax Layer 300 Acoustic Encoder, Audio Encoder 301a, 301b, 301c Attention Probability Matrix 400 Table 500 Method 600 Computing Device 600a Server 600b Laptop Computer 600c Rack Server System 610 Processor 620 Memory 630 Storage Device 640 High-Speed Interface / Controller 650 High-Speed Expansion Port 660 Low-Speed Interface / Controller 670 Low-Speed Bus 680 Display 690 Low-Speed Expansion Port

Claims

1. An automatic speech recognition (ASR) model (200) implemented in data processing hardware for integrating streaming speech recognition and non-streaming speech recognition, wherein the ASR model (200) is an audio encoder (300) that receives a sequence of acoustic frames (110) as input, and is configured to cause the data processing hardware to generate, for each of a plurality of time steps, a high-order feature representation for a corresponding acoustic frame (110) in the sequence of acoustic frames (110); and an audio encoder (300) is a joint network (230) that receives, as input, the high-order feature representation generated by the audio encoder (300) at each of the plurality of time steps, and is configured to cause the data processing hardware to generate, for each of the plurality of time steps, a probability distribution over possible speech recognition hypotheses at the corresponding time step; and a joint network (230), wherein the audio encoder (300) comprises a neural network that applies a mixture model (MiMo) attention to calculate an attention probability distribution function (PDF) using a set of softmax mixture components over a context window. The ASR model (200).

2. The ASR model (200) according to claim 1, wherein the set of mixture components of the MiMo attention is designed to cover all possible context use cases during inference.

3. The ASR model (200) according to claim 1 or 2, wherein each mixture component of the set of mixture components comprises a fully normalized softmax.

4. The ASR model (200) according to any one of claims 1 to 3, which switches between a streaming mode and a non-streaming mode by adjusting the mixing weights of the MiMo attention.

5. The neural network of the audio encoder (300) of the ASR model (200) according to any one of claims 1 to 4, comprising a plurality of conformer layers.

6. The neural network of the audio encoder (300) of the ASR model (200) according to any one of claims 1 to 5, comprising a plurality of transformer layers.

7. A label encoder (220), receiving, as an input, a sequence of non-blank symbols output by the final softmax layer (240), further comprising a label encoder (220) configured to cause the data processing hardware to generate a high-density representation at each of the plurality of time steps, The ASR model (200) according to any one of claims 1 to 6, wherein the joint network (230) is further configured to cause the data processing hardware to receive, as an input, the high-density representation generated by the label encoder (220) at each of the plurality of time steps.

8. The ASR model (200) according to claim 7, wherein the label encoder (220) comprises a neural network of a transformer layer, a conformer layer, or a long short-term memory (LSTM) layer.

9. The ASR model (200) according to claim 7 or 8, wherein the label encoder (220) comprises a lookup table embedding model configured to search for the high-density representation at each of the plurality of time steps.

10. The possible speech recognition hypotheses generated in each of the plurality of time steps correspond to a set of output labels where each output label represents a grapheme or word piece in natural language, the ASR model (200) according to any one of claims 1 to 9. **Claim 11** A computer-implemented method (500) which, when executed on data processing hardware (12), causes the data processing hardware (12) to perform operations, the operations being receiving a sequence of acoustic frames (110); in each of a plurality of time steps using an audio encoder (300) of an automatic speech recognition (ASR) model (200) to generate a high-order feature representation for a corresponding acoustic frame (110) in the sequence of acoustic frames (110); using a joint network (230) of the ASR model (200) to generate a probability distribution over possible speech recognition hypotheses at the corresponding time step based on the high-order feature representation generated by the audio encoder (300) at the corresponding time step. The audio encoder (300) comprises a neural network that applies a mixture model (MiMo) attention to compute an attention probability distribution function (PDF) using a set of softmax mixture components over a context window, the computer-implemented method (500). **Claim 12** The set of mixture components of the MiMo attention is designed to cover all possible context use cases during inference, the computer-implemented method (500) according to claim 11. **Claim 13** Each mixture component of the set of mixture components comprises a fully normalized softmax, the computer-implemented method (500) according to claim 11 or 12. **Claim 14** The ASR model (200) switches between a streaming mode and a non-streaming mode by adjusting the mixing weights of the MiMo attention, the computer-implemented method (500) according to any one of claims 11 to 13.

15. The neural network of the audio encoder (300) comprises a plurality of conformer layers, the computer-implemented method (500) according to any one of claims 11 to 14.

16. The neural network of the audio encoder (300) comprises a plurality of transformer layers, the computer-implemented method (500) according to any one of claims 11 to 15.

17. The operation is In each of the plurality of time steps, further including the operation of generating a high-density representation using the label encoder (220) of the ASR model (200) configured to receive a sequence of non-blank symbols output by the final softmax layer (240) of the ASR model (200), The operation of generating the probability distribution over possible speech recognition hypotheses at the corresponding time step is further based on the sequence of non-blank symbols output by the final softmax layer (240) at the corresponding time step, the computer-implemented method (500) according to any one of claims 11 to 16.

18. The label encoder (220) comprises a neural network of a transformer layer, a conformer layer, or a long short-term memory (LSTM) layer, the computer-implemented method (500) according to claim 17.

19. The label encoder (220) comprises a lookup table embedding model configured to search for the high-density representation in each of the plurality of time steps, the computer-implemented method (500) according to claim 17 or 18.

20. The computer-implemented method (500) according to any one of claims 11 to 19, wherein the possible speech recognition hypotheses generated at each of the plurality of time steps correspond to a set of output labels, each output label representing a grapheme or word piece in natural language.

Citation Information

Patent Citations

  • Pointer Sentinel Mixed Architecture

    JP2019537809A

  • Real-time voice recognition method based on cutting attention, device, apparatus and computer readable storage medium

    JP2020112787A

  • Speech recognition system and method for using a speech recognition system

    JP2021507312A

  • Transformer transducer: one model unifying streaming and non-streaming speech recognition

    US11741947B2

  • Speech recognition with sequence-to-sequence models

    US20200027444A1