Transformer Transducer: A Model for Integrating Streaming and Non-Streaming Speech Recognition

A single Transformer-Transducer model integrates streaming and non-streaming speech recognition by using multiple Transformer layers with varying look-ahead audio contexts, addressing the balance between accuracy and latency in ASR systems.

JP7679468B2Active Publication Date: 2025-05-19GOOGLE LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023520502
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-10-05
Filing Date
2021-03-19
Publication Date
2025-05-19
Estimated Expiration
2041-03-19

AI Technical Summary

Technical Problem

Current automatic speech recognition (ASR) systems face challenges in balancing accuracy and latency, with streaming models often being inaccurate despite low latency, and non-streaming models having high latency but higher accuracy.

Method used

A single Transformer-Transducer model is introduced to integrate streaming and non-streaming speech recognition, utilizing an audio encoder, a label encoder, and a joint network. This model includes multiple Transformer layers with different look-ahead audio contexts for low-latency and high-latency decoding branches.

Benefits of technology

The integrated model achieves both low-latency and high-accuracy speech recognition by using a single model for streaming and non-streaming tasks, reducing the need for separate models and minimizing computational and memory costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007679468000002
    Figure 0007679468000002
  • Figure 0007679468000003
    Figure 0007679468000003
  • Figure 0007679468000004
    Figure 0007679468000004
Patent Text Reader

Abstract

The transformer-transducer model (200) includes an audio encoder (300), a label encoder (220), and a joint network (230). The audio encoder receives a sequence of acoustic frames (110) and generates a high-order feature representation for each acoustic frame at each of a plurality of time steps. The label encoder receives a sequence of non-blank symbols output by a softmax layer (240) and generates a dense representation at each of a plurality of time steps. The joint network receives the high-order feature representation and the dense representation at each of the plurality of time steps and generates a probability distribution over possible speech recognition hypotheses. The audio encoder of the model further includes a neural network having an initial stack (310) of transformer layers (400) trained with zero look-ahead audio context and a final stack (320) of transformer layers (400) trained with a variable look-ahead audio context.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the use of an integrated model for streaming speech recognition and non-streaming speech recognition.

Background Art

[0002] Automatic speech recognition (ASR), which is the process of obtaining an audio input and transcribing it into text, is a very important technology used by mobile devices and other devices. Generally, ASR attempts to provide an accurate transcription of what a person has said by obtaining an audio input (e.g., speech) and transcribing the audio input into text. Modern ASR models continue to improve in both accuracy (e.g., low word error rate (WER)) and latency (e.g., the delay between a user's speech and the transcription) based on the continuous development of deep neural networks. When using current ASR systems, it is required that the ASR system decode speech in a streaming manner that is equivalent to or in some cases faster than real time and is also accurate. However, one issue in developing deep learning-based ASR models is that streaming models may be inaccurate although they have low latency. Conversely, non-streaming models have high latency but generally result in higher accuracy.

Summary of the Invention

Means for Solving the Problems

[0003] One aspect of the present disclosure provides a single Transformer-Transducer model for integrating streaming speech recognition and non-streaming speech recognition. The single Transformer-Transducer model includes an audio encoder, a label encoder, and a joint network. The audio encoder receives a sequence of acoustic frames as input and generates a high-order feature representation for the corresponding acoustic frames within the sequence of acoustic frames at each of a plurality of time steps. The label encoder receives a sequence of non-blank symbols output by a final softmax layer as input and generates a dense representation at each of a plurality of time steps. The joint network receives as input the high-order feature representation generated by the audio encoder and the dense representation generated by the label encoder at each of a plurality of time steps, and is configured to generate a probability distribution over possible speech recognition hypotheses at the corresponding time step. The audio encoder of the model further includes a neural network having a plurality of Transformer layers. The plurality of Transformer layers includes an initial stack of Transformer layers each trained using a zero look ahead audio context and a final stack of Transformer layers each trained using a variable look ahead audio context.

[0004] Implementations of the present disclosure may include one or more of any of the following features. In some implementations, each transformer layer of the audio encoder includes a normalization layer, a masked multi-head attention layer with relative position encoding, a residual connection, a stacking / unstacking layer, and a feed-forward layer. In these implementations, the stacking / unstacking layer may be configured to adjust the frame rate of the corresponding transformer layer to adjust the processing time by a single transformer-transducer model during training and inference. In some examples, the initial stack of transformer layers includes more transformer layers than the final stack of transformer layers. In some examples, during training, a variable look-ahead audio context is uniformly sampled for each transformer layer in the final stack of transformer layers.

[0005] In some implementations, the model further includes a low-latency decoding branch configured to decode the corresponding speech recognition result for an utterance input from audio data encoded using a first look-ahead audio context, and a high-latency decoding branch configured to decode the corresponding speech recognition result for an utterance input from audio data encoded using a second look-ahead audio context. Here, the second look-ahead audio context includes look-ahead audio with a longer duration than the first look-ahead audio context. In these implementations, an initial stack of the transformer layer may apply a zero look-ahead audio context to calculate shared activations used by both the low-latency decoding branch and the high-latency decoding branch, and a final stack of the transformer layer may apply the first look-ahead audio context to calculate low-latency activations used by the low-latency decoding branch but not used by the high-latency decoding branch, and the final stack of the transformer layer may apply the second look-ahead audio context to calculate high-latency activations used by the high-latency decoding branch but not used by the low-latency decoding branch. In some additional implementations, the first look-ahead audio context includes a zero look-ahead audio context.

[0006] In some examples, the low-latency decoding branch and the high-latency decoding branch are executed in parallel to decode the corresponding speech recognition results for the input utterance. In these examples, the corresponding speech recognition results decoded by the high-latency decoding branch for the input utterance are delayed by a duration based on the difference between the second look-ahead audio context and the first look-ahead audio context with respect to the corresponding speech recognition results decoded by the low-latency decoding branch for the input utterance. Additionally or alternatively, the low-latency decoding branch may be configured to stream the corresponding speech recognition results as partial speech recognition results when the input utterance is received by a single transformer-transducer model, and the high-latency decoding branch may be configured to output the corresponding speech recognition results as the final transcription after the single transformer-transducer model has received the complete input utterance.

[0007] In some implementations, the input utterance is directed to an application, and the duration of the second look-ahead audio context used by the high-latency decoding branch to decode the corresponding speech recognition results for the input utterance is based on the type of application to which the input utterance is directed. In some examples, the label encoder includes a neural network having a plurality of transformer layers. Alternatively, the label encoder may include a bigram embedding lookup decoder model. The single transformer-transducer model may be executed on a client device or on a server-based system.

[0008] Another aspect of the present disclosure provides a computer-implemented method that, when executed on data processing hardware, performs operations including receiving, as input to a transformer-transducer model, audio data corresponding to an utterance, and performing streaming speech recognition and non-streaming speech recognition in parallel on the audio data using the transformer-transducer model. In a low-latency branch of the transformer-transducer model, the operations also include encoding the audio data using a first look-ahead audio context while receiving the audio data corresponding to the utterance, decoding the audio data encoded using the first look-ahead audio context into a partial speech recognition result for the input utterance, and streaming the partial speech recognition result for the input utterance. In a high-latency branch of the transformer-transducer model, the operations include encoding the audio data using a second look-ahead audio context after the audio data corresponding to the utterance has been received, decoding the audio data encoded using the second look-ahead audio context into a final speech recognition result for the input utterance, and replacing the streamed partial speech recognition result with the final speech recognition result.

[0009] This aspect may include one or more of any of the following features. In some implementations, the operation further includes an audio encoder including a neural network having a plurality of transformer layers. The plurality of transformer layers include an initial stack of transformer layers each trained using zero look-ahead audio context and a final stack of transformer layers trained using variable look-ahead audio context. Each transformer layer may include a normalization layer, a masked multi-head attention layer with relative position encoding, a residual connection, a stacking / unstacking layer, and a feed-forward layer. Here, the stacking / unstacking layer may be configured to adjust the processing time by a single transformer-transducer model during training and inference by changing the frame rate of the corresponding transformer layer.

[0010] In some examples, the initial stack of transformer layers includes more transformer layers than the final stack of transformer layers. In some implementations, during training, a variable look-ahead audio context is uniformly sampled for each transformer layer in the final stack of transformer layers. In some examples, the initial stack of transformer layers may apply a zero look-ahead audio context to compute shared activations used by both the low-latency branch and the high-latency branch, and the final stack of transformer layers may apply a first look-ahead audio context to compute low-latency activations used by the low-latency branch but not by the high-latency decoding branch, and the final stack of transformer layers may apply a second look-ahead audio context to compute high-latency activations used by the high-latency branch but not by the low-latency branch. The first look-ahead audio context may include a zero look-ahead audio context. In some implementations, the final speech recognition result decoded by the high-latency branch for the input utterance is delayed by a duration based on the difference between the second look-ahead audio context and the first look-ahead audio context relative to the partial speech recognition result decoded by the low-latency branch for the input utterance.

[0011] In some examples, the operation further includes receiving an application identifier indicating the type of application to which the input utterance is directed, and setting the duration of the second look-ahead audio context based on the application identifier. In some implementations, the transformer-transducer model includes a label encoder including a neural network having a plurality of transformer layers. In some examples, the transformer-transducer model includes a label encoder including a bigram embedding lookup decoder model. The data processing hardware executes the transformer-transducer model and is present on a client device or a server-based system.

[0012] Details of one or more implementations of the present disclosure are set forth in the accompanying drawings and the following description. Other aspects, features, and advantages will become apparent from the description, the drawings, and the claims.

Brief Description of the Drawings

[0013]

Figure 1

Figure 2

Figure 3A

Figure 3B

Figure 4

Figure 5

Figure 6

Modes for Carrying Out the Invention

[0014] Like reference numerals in the various drawings indicate like elements.

[0015] An automatic speech recognition (ASR) system aims to achieve not only high quality / accuracy (e.g., low word error rate (WER)) but also low latency (e.g., short delay between a user's utterance and the resulting transcription). Recently, end-to-end (E2E) ASR models have been widely used to achieve state-of-the-art performance in terms of accuracy and latency. In contrast to conventional hybrid ASR systems that include separate acoustic, pronunciation, and language models, E2E models apply a sequence-to-sequence approach to jointly learn acoustic and language modeling in a single neural network that is trained end-to-end from training data, e.g., utterance-transcription pairs. Here, an E2E model refers to a model whose architecture is entirely composed of neural networks. A complete neural network functions without external components and / or manually designed components (e.g., finite state transducers, dictionaries, or text normalization modules). Also, when training E2E models, these models generally do not require bootstrapping from decision trees nor time alignment from separate systems.

[0016] Examples of sequence-to-sequence models include "attention-based" models and "listen-attend-spell" (LAS) models. The LAS model uses a listener component, an attender component, and a speller component to transcribe speech into text. Here, the listener is a recurrent neural network (RNN) encoder that receives an audio input (e.g., a time-frequency representation of speech input) and maps the audio input to a higher-order feature representation. The attender pays attention to the higher-order features and learns to align the input features with the predicted sub-word units (e.g., graphemes or word pieces). The speller is an attention-based RNN decoder that generates a character sequence from the input by generating a probability distribution over a hypothesized set of words. Attention-based models such as the LAS model generally process the entire sequence (e.g., an audio waveform) and use the audio context for the entire sentence before generating an output (e.g., a sentence) based on it, so these models provide the output as a non-streaming transcription.

[0017] Furthermore, when currently using an ASR system, there may be a requirement for the ASR system to decode speech in a streaming manner that corresponds to displaying a description of the speech in real time or, in some cases, faster than real time when the user is speaking. As an example, when the ASR system is displayed on a user computing device that experiences direct interaction with the user, such as a mobile phone, an application (e.g., a digital assistant application) that runs on the user device and uses the ASR system needs to stream speech recognition so that words, word pieces, and / or individual characters are displayed on the screen as they are spoken. Also, the user of the user device may have a low tolerance for latency. For example, when the user speaks a query asking the digital assistant to retrieve detailed information from the calendar application to check the schedule, the user expects the digital assistant to provide a response conveying the retrieved detailed information as quickly as possible. Due to such low tolerance, the ASR system tries hard to operate on the user device to minimize the impact of latency and inaccuracy, which may negatively affect the user experience. However, attention-based sequence-to-sequence models, such as LAS models, which function by reexamining the entire sequence of input audio before generating the output text, do not allow streaming of the output when the input is received. Due to this drawback, problems may arise when developing attention-based sequence-to-sequence models for voice applications that are susceptible to latency and / or require real-time voice transcription. As a result, the LAS model alone is not an ideal model for applications that are susceptible to latency and / or applications that need to implement the streaming transcription function in real time when the user is speaking.

[0018] Another form of the sequence-to-sequence model called the Recurrent Neural Network Transducer (RNN-T) does not use an attention mechanism and, unlike other sequence-to-sequence models that generally need to process the entire sequence (audio waveform) to generate an output (e.g., a sentence), the RNN-T processes input samples continuously and streams output symbols. This is particularly attractive for real-time communication. For example, speech recognition by RNN-T may output characters one by one in response to an utterance. Here, the RNN-T uses a feedback loop that sends the symbol predicted by the model back to the RNN-T itself to predict the next symbol. Decoding the RNN-T involves beam search by a single neural network rather than a large decoder graph, so the RNN-T may scale to a fraction of the size of a server-based speech recognition model. With the size reduction, it may be possible to deploy the entire RNN-T on a device and it may operate offline (i.e., without a network connection), thus avoiding reliability issues with the communication network.

[0019] The RNN-T model still lags behind large-scale state-of-the-art conventional models (e.g., server-based models with separate AM, PM, and LM) and attention-based sequence-to-sequence models (e.g., LAS models) in terms of quality (e.g., speech recognition accuracy often measured by word error rate (Wer)) because it cannot apply look-ahead audio context (e.g., trailing context) when predicting recognition results. To compensate for such a lag in speech recognition accuracy, recently, there has been a focus on the development of a two-pass recognition system that includes a first-pass component of the RNN-T network and a second-pass component of the LAS network that re-evaluates the recognition results generated during the first pass. In this design, the two-pass model benefits from the streaming characteristics of the low-latency RNN-T model, while the accuracy of the RNN-T model is improved by the second pass incorporating the LAS network. The LAS network increases latency compared to the RNN-T model alone, but the increase in latency is considered to be negligible and complies with the latency constraints for on-device operation.

[0020] The RNN-T model that uses long short-term memory (LSTM) to provide a sequence encoder is suitable for providing a streaming transcription function and applications that are generally sensitive to latency for recognizing conversational queries (e.g., "Please set a timer", "Please remind me to buy milk", etc.), but has limited ability to look-ahead audio context, and thus tends to delete words when recognizing long utterances. In this specification, long utterances include speech-based queries non-conversational queries where the user dictates part of an email, message, document, social media post, or other content.

[0021] Often, a user uses a streaming speech recognition model, such as an RNN-T model, to recognize conversational queries and a separate non-streaming speech recognition model to recognize non-conversational queries. Generally, an application in which the user is sending an utterance can be used to identify whether to use a streaming speech recognition model or a non-streaming speech recognition model for speech recognition. Since different separate speech recognition models are required to perform speech recognition depending on the application and / or query type, it is computationally costly and requires sufficient memory capacity to store each model on the user device. Even if one of the models is executable on a remote server, additional costs and bandwidth constraints for connecting to the remote server can affect speech recognition performance and ultimately the user experience.

[0022] Embodiments herein are directed to a single Transformer-Transducer (T-T) model for integrating streaming speech recognition tasks and non-streaming speech recognition tasks. As will become apparent, the T-T model may realize positive attributes of the RNN-T model, such as a streaming transcription function, low-latency speech recognition, a small computational footprint, and low memory requirements, without being affected by the aforementioned drawbacks of the RNN-T model. That is, the T-T model may be trained with variable look-ahead audio context, thereby applying a look-ahead audio context of sufficient duration when performing speech recognition on an input utterance. Also, the T-T model may implement a y-architecture that provides in parallel a low-latency branch for decoding a streaming partial speech recognition result for an input utterance and a high-latency branch for decoding a final speech recognition result for the same input utterance.

[0023] FIG. 1 is an example of an acoustic environment 100. In the acoustic environment 100, the way for a user 104 to interact with a computing device such as a user device 10 may be via voice input. The user device 10 (commonly also referred to as the device 10) is configured to capture voice (e.g., streaming audio data) from one or more users 104 within the acoustic environment 100. Here, the streaming audio data may refer to an audible query, a command for the device 10, or an utterance 106 by the user 104 that serves as an audible communication captured by the device 10. The voice-responsive system of the device 10 may process the query or command by answering the query and / or executing / completing the command by one or more downstream applications.

[0024] The user device 10 may correspond to any computing device associated with the user 104 and can receive audio data. Some examples of the user device 10 include, but are not limited to, mobile devices (e.g., mobile phones, tablets, laptops, etc.), computers, wearable devices (e.g., smartwatches), smart home appliances, Internet of Things (IoT) devices, vehicle infotainment systems, smart displays, smart speakers, and the like. The user device 10 includes data processing hardware 12 and memory hardware 14 that communicates with the data processing hardware 12, and stores instructions that cause the data processing hardware 12 to perform one or more operations when executed by the data processing hardware 12. The user device 10 further includes an audio system 16 having audio capture devices (e.g., microphones) 16, 16a for capturing utterances 106 within the acoustic environment 100 and converting them into electrical signals, and audio output devices (e.g., speakers) 16, 16b for communicating audible audio signals (e.g., as output audio data from the device 10). Although the user device 10 implements a single audio capture device 16a in the illustrated example, the user device 10 may implement an array of audio capture devices 16a without departing from the scope of the present disclosure, and one or more of the capture devices 16a within the array may communicate with the audio system 16 without physically existing on the user device 10.

[0025] In the acoustic environment 100, an automatic speech recognition (ASR) system 118 implementing a transformer-transducer (T-T) model 200 exists on the user device 10 of the user 104 and / or on a remote computing device 60 (e.g., one or more remote servers of a distributed system executed in a cloud-computing environment) that communicates with the user device 10 via the network 40. The user device 10 and / or the remote computing device 60 also includes an audio subsystem 108 configured to receive an utterance 106 by the user 104 captured by the audio capture device 16a and convert the utterance 106 into a corresponding digital format related to an input acoustic frame 110 that can be processed by the ASR system 118. In the illustrated example, the user emits each utterance 106, and the audio subsystem 108 converts the utterance 106 into a corresponding audio data (e.g., acoustic frame) 110 input to the ASR system 118. Thereafter, the T-T model 200 receives the audio data 110 corresponding to the utterance 106 as an input and generates / predicts a corresponding transcription 120 (e.g., recognition result / hypothesis) of the utterance 106 as an output. As will be described in more detail below, the T-T model 200 may be trained using a variable look-ahead audio context such that the T-T model 200 sets different durations for each of the look-ahead audio contexts when performing speech recognition depending on how much latency the query specified by the utterance 106 perceives during inference and / or how much latency the user 106 tolerates. For example, a digital assistant application 50 running on the user device 10 may need to stream speech recognition such that words, word pieces, and / or individual characters are displayed on the screen when uttered. Also, the user 104 of the user device 10 may have a low tolerance for latency when issuing a query executed by the digital assistant application 50.In such a scenario, when it is preferred to minimize the speech recognition latency, the T-T model 200 may apply zero or minimal look-ahead audio context (also referred to as "following context") to provide the streaming transcription function in real time while the user 104 is uttering the utterance 106. On the other hand, when the user has a higher tolerance for speech recognition latency and / or the recognized utterance 106 involves long-form audio, the same T-T model 200 applies a look-ahead audio context duration sufficient to provide an accurate transcription 120, and the latency may increase based on the look-ahead audio context duration. Thus, the ASR system 118 may implement a single T-T model 200 for a plurality of different speech recognition tasks, providing both a streaming transcription function and a non-streaming transcription function without the need to utilize separate ASR models for each task.

[0026] In some implementations, the T-T model 200 performs both streaming speech recognition and non-streaming speech recognition in parallel on the audio data 110. For example, in the illustrated example, the T-T model 200 uses a first decoding branch (i.e., the low-latency branch 321 (FIG. 3B)) to perform streaming speech recognition on the audio data 110 to generate partial speech recognition results 120, 120a, and in parallel, uses a second decoding branch (i.e., the high-latency branch 322 (FIG. 3B)) to perform non-streaming speech recognition on the same audio data 110 to generate final speech recognition results 120, 120b. In particular, the first decoding branch may use a first look-ahead audio context that can be set to zero (or about 240 milliseconds) to generate the partial speech recognition result 120a, while the second decoding branch may use a second look-ahead audio context having a longer duration than the first look-ahead audio context to generate the final speech recognition result 120b. Thus, the final speech recognition result 120b for the input utterance 106 may be delayed relative to the partial speech recognition result 120a for the input utterance by a duration based on the difference between the second look-ahead audio context and the first look-ahead audio context.

[0027] The user device 10 and / or the remote computing device 60 also execute a user interface generator 107 configured to present the representation of the transcription 120 of the utterance 106 to the user 104 of the user device 10. As will be described in detail below, the user interface generator 107 may display the partial speech recognition result 120a in a streaming manner during time 1 and then display the final speech recognition result 120b during time 2. In some configurations, the transcription 120 output from the ASR system 118 is processed, for example, by a natural language understanding (NLU) module executed on the user device 10 or the remote computing device 60 to execute the user command / query specified by the utterance 106. Additionally or alternatively, a text-to-speech system (not shown) (for example, executed on any combination of the user device 10 or the remote computing device 60) may convert the transcription into synthetic speech for audible output by the user device 10 and / or another device.

[0028] In the illustrated example, the user 104 interacts with a program or application 50 of the user device 10 that uses the ASR system 118 (for example, a digital assistant application 50). For example, FIG. 1 shows the user 104 communicating with the digital assistant application 50, and the digital assistant application 50 displaying a digital assistant interface 18 on the screen of the user device 10, showing a conversation between the user 104 and the digital assistant application 50. In this example, the user 104 asks the digital assistant application 50, "What time is the concert tonight?" This question from the user 104 is an utterance 106 that is captured by the audio capture device 16a and processed by the audio system 16 of the user device 10. In this example, the audio system 16 receives the utterance 106 and converts it into an acoustic frame 110 that is input to the ASR system 118.

[0029] Continuing with the example, the T-T model 200 receives an acoustic frame 110 corresponding to utterance 106 when user 104 speaks, encodes the acoustic frame 110 using a first look-ahead audio context, and then decodes the encoded acoustic frame into a partial speech recognition result 120a using the first look-ahead audio context. During time 1, the user interface generator 107 presents, via the digital assistant interface 18, an expression of the partial speech recognition result 120a of utterance 106 to user 104 of user device 10 in a streaming manner, whereby words, word pieces, and / or individual characters are displayed on the screen as they are spoken. In some examples, the first look-ahead audio context is equal to zero.

[0030] In parallel, after all of the acoustic frames 110 corresponding to the utterance 106 have been received, the T-T model 200 encodes all of the acoustic frames 110 corresponding to the utterance 106 using the second look-ahead audio context, and decodes the acoustic frames 110 using the second look-ahead audio context into the final speech recognition result 120b. The duration of the second look-ahead audio context may be 1.2 seconds, 2.4 seconds, or any other duration. In some examples, an indication such as an endpoint indicating that the user 104 has finished the utterance 106 triggers the T-T model 200 to encode all of the acoustic frames 110 using the second look-ahead audio context. During time 2, the user interface generator 107 presents, via the digital assistant interface 18, the representation of the final speech recognition result 120b of the utterance 106 to the user 104 of the user device 10. In some implementations, the user interface generator 107 replaces the representation of the partial speech recognition result 120a with the representation of the final speech recognition result 120b. For example, assuming that the final speech recognition result 120b is more accurate than the partial speech recognition result 120a generated without using the look-ahead audio context, the final speech recognition result 120b, which is ultimately displayed as the transcription 120, may correct words that may have been misrecognized in the partial speech recognition result 120a. In this example, the streaming partial speech recognition result 120a output by the T-T model 200 and displayed on the screen of the user device 10 at time 1 provides the user 104 with a response that the associated latency is low and that the query of the user 104 is being processed, while the final speech recognition result 120b output by the T-T model 200 and displayed on the screen at time 2 uses the look-ahead audio context to improve the speech recognition quality in terms of accuracy but increases the latency. However, since the partial speech recognition result 120a is displayed as soon as the user utters the utterance 106, the higher latency associated with the generation and the display of the final final speech recognition result is not noticed by the user 104.

[0031] In the example shown in FIG. 1, the digital assistant application 50 may respond to questions presented by the user 104 using natural language processing. Natural language processing generally refers to the process of interpreting written language (e.g., partial speech recognition result 120a and / or final speech recognition result 120b) and determining whether the written language prompts some action. In this example, the digital assistant application 50 uses natural language processing to recognize that the question from the user 104 is about the user's schedule, and more specifically, a question about a concert on the user's schedule. By recognizing such detailed information through natural language processing, the automated assistant returns a response 19 to the user's query, in this case, the response 19 presents "The venue will open at 6:30 PM and the concert will start at 8:00." In some configurations, natural language processing is performed on a remote server 60 that communicates with the data processing hardware 12 of the user device 10.

[0032] Referring to FIG. 2, the T-T model 200 may enable end-to-end (E2E) speech recognition by incorporating an acoustic model, a pronunciation model, and a language model into a single neural network, in which case neither a dictionary nor a separate text normalization component is required. Various structures and optimization mechanisms can improve accuracy and shorten the model training time. The T-T model 200 includes a Transformer-Transducer (T-T) model architecture that meets the latency constraints associated with interactive applications. The T-T model 200 has a small computational footprint and fewer memory requirements than conventional ASR architectures, such that the T-T model architecture is suitable for performing speech recognition on the entire user device 10 (e.g., communication with the remote server 60 is not required). The T-T model 200 includes an audio encoder 300, a label encoder 220, and a joint network 230. The audio encoder 300 is generally similar to the acoustic model (AM) in a conventional ASR system and includes a neural network having a plurality of Transformer layers 400 (FIGS. 3A, 3B, and 4). For example, the audio encoder 300 reads a sequence x = (x 1 , x 2 ,..., x T ) of d-dimensional feature vectors (e.g., acoustic frames 110 (FIG. 1)), where

[0033]

Number

[0034] and the audio encoder 300 generates a high-order feature representation at each time step. This high-order feature representation is denoted as ah 1 ,..., ah T .

[0035] Similarly, the label encoder 220 may also include a neural network of the transformer layer or a lookup table embedding model, and the lookup table embedding model, similar to the language model (LM), is a sequence of non-blank symbols y 0 , ..., y ui-1 that has been output by the final softmax layer 240 so far, and processes it as a dense representation Ih u that encodes the predicted label history. In an implementation where the label encoder 220 includes a neural network of the transformer layer, each transformer layer may include a normalization layer, a masked multi-head attention layer with relative position encoding, a residual connection, a feed-forward layer, and a dropout layer. In these implementations, the label encoder 220 may include two transformer layers. In an implementation where the label encoder 220 includes a lookup table embedding model with a bi-gram label context, the embedding model is configured to learn a d-dimensional weight vector for each possible bi-gram label context, where d is the dimension of the outputs of the audio encoder 300 and the label encoder 200. In some examples, the total number of parameters in the embedding model is N 2 ×d, where N is the vocabulary size for the labels. Here, the learned weight vectors are then used as the embedding of the bi-gram label context in the T-T model 200 to generate fast label encoder 220 execution time.

[0036] Finally, in the T-T model architecture, the representations generated by the audio encoder 300 and the label encoder 220 are combined by the joint network 230 using a dense layer J u,t . The joint network 230 then outputs P(z u,t |x,t,y 1 , ..., y u-1) That is, it predicts the distribution over the following output symbols. In other words, the joint network 230 generates a probability distribution over possible speech recognition hypotheses at each output step (e.g., time step). Here, "possible speech recognition hypotheses" corresponds to a set of output labels (also called "speech units") each represented by a grapheme (e.g., symbol / character) or word piece in a specified natural language. For example, when the natural language is English, the set of output labels may include 27 symbols, i.e., one label for each of the 26 characters in the English alphabet and one label for the space. Thus, the joint network 230 may output a set of values indicating the likelihood of occurrence of each of a predetermined set of output labels. This set of values can be a vector (e.g., a one-hot vector) and can represent a probability distribution over the set of output labels. In some cases, the output labels are graphemes (e.g., individual characters, and optionally punctuation marks and other symbols), but the set of output labels is not so limited. For example, the set of output labels can include word pieces and / or whole words in addition to or instead of graphemes. The output distribution of the joint network 230 can include posterior probability values for each of the different output labels. Thus, if there are 100 different output labels each representing a different grapheme or other symbol, the output z u,t of the joint network 230 can include 100 different probability values, one for each output label. Then, the probability distribution can be used to select scores (e.g., by the softmax layer 240) in a beam search process and assign them to candidate orthographic elements (e.g., graphemes, word pieces, and / or words) to determine the transcription 120.

[0037] The softmax layer 240 may use any technique to select, as the next output symbol predicted by the T-T model 200 in the corresponding output step, the output label / symbol having the highest probability in the distribution. In this way, the T-T model 200 does not make a conditional independence assumption, and the prediction of each symbol is conditioned on not only the acoustics but also the sequence of labels that have been output so far.

[0038] Referring to FIG. 3A, in some implementations, the plurality of transformer layers 400 of the audio encoder 300 of the T-T model 200 includes an initial stack 310 of transformer layers 400 and a final stack 320 of transformer layers 400. Each transformer layer 400 in the initial stack 310 may be trained using zero look-ahead audio context, while each transformer layer 400 in the final stack 320 may be trained using variable look-ahead audio context. The initial stack 310 of the T-T model 200 may have more transformer layers 400 than the final stack 320 of transformer layers 400. For example, the initial stack 310 of transformer layers 400 may include 15 transformer layers 400, while the final stack 320 of transformer layers 400 may include 5 transformer layers 400. However, the respective number of transformer layers 400 in each of the initial stack 310 and the final stack 320 is not limited. Thus, in the examples herein, although it may be described that the initial stack 310 includes 15 transformer layers 400 and the final stack 320 of transformer layers includes 5 transformer layers 400, the initial stack 310 of transformer layers 400 may include fewer or more than 15 transformer layers 400, and the final stack 320 of transformer layers 400 may include fewer or more than 5 transformer layers 400. Further, the total number of transformer layers 400 utilized by the audio encoder 300 may be fewer or more than 20 transformer layers 400.

[0039] The T-T model 200 is trained on a training dataset of audio data corresponding to utterances paired with corresponding transcriptions. The training of the T-T model 200 includes the remote server 60, and the trained T-T model 200 may be pushed to the user device 10. Training the final stack 320 of the transformer layer 400 using a variable look-ahead audio context includes setting the preceding context of the transformer layer 400 in the final stack 320 to be constant and sampling the succeeding context length of each layer in the final stack 320 from a given distribution within the training dataset. Here, the sampled succeeding context length corresponds to the duration of the look-ahead audio context sampled from the given distribution. As will become apparent, the sampled succeeding context length specifies a mask for self-attention applied by the corresponding masked multi-head attention layer 406 (FIG. 4) of each transformer layer 400 in the final stack 320.

[0040] In some implementations, during the training of the audio encoder 300, a variable look-ahead audio context is uniformly sampled for each transformer layer 400 in the final stack 320 of the transformer layers 400. For example, the variable look-ahead audio context that is uniformly sampled for each transformer layer 400 in the final stack 320 of the transformer layers 400 may include durations of zero, 1.2 seconds, and 2.4 seconds. Such durations of the look-ahead audio context are non-limiting and may include different durations for each, and / or may include sampling of look-ahead audio contexts with more or fewer than three durations for the final stack 320 of the transformer layers 400. In these implementations during training, each look-ahead audio context configuration for the transformer layer 400 may be specified for look-ahead audio contexts of each different duration. Further, the audio encoder 300 of the T-T model 200 may be trained using an output delay of four acoustic frames 110. For example, continuing with the above example where the initial stack 310 of 15 transformer layers 400 is trained using a zero look-ahead audio context and the final stack 320 of 5 transformer layers 400 is trained using a variable look-ahead context, the first look-ahead context configuration specified for the zero look-ahead audio context may include [0]×19+[4], the second look-ahead context configuration specified for the 1.2-second look-ahead audio context may include [0]×15+[8]×5, and the third look-ahead audio context configuration specified for the 2.4-second look-ahead audio context may include [0]×15+

[16] ×5. The numbers within the parentheses in each of the look-ahead audio context configurations indicate the number of look-ahead audio frames corresponding to the specified duration of the look-ahead audio context. In the above example, the first look-ahead context configuration of [0]×19+[4] used to evaluate the zero look-ahead audio context has the final transformer layer 400 apply an output delay of four audio frames 110.Here, the output delay of the four audio frames corresponds to a 240 - millisecond look - ahead audio context.

[0041] In some additional implementation forms where instead of training the initial stack 310 of the transformer layer 400 using a zero look - ahead audio context, only the final stack 320 of the transformer layer 400 is trained using a variable look - ahead audio context, all of the transformer layers 400 of the audio encoder 300 of the T - T model 200 are trained using a variable look - ahead audio context. For example, continuing with the above example where look - ahead contexts of duration zero, 1.2 seconds, and 2.4 seconds are uniformly sampled during training, the first look - ahead context configuration specified for the zero look - ahead audio context may include [0]×19+[4] (for example, [4] specifies that the final transformer layer 400 applies an output delay of four audio frames corresponding to a 240 - millisecond look - ahead audio context), the second look - ahead context configuration specified for the 1.2 - second look - ahead audio context may include [2]×20, and the third look - ahead context configuration specified for the 2.4 - second look - ahead audio context may include [4]×20. That is, the output delay of two audio frames applied by each of the 20 transformer layers 400 evaluates a given distribution of the training data set for the 1.2 - second look - ahead audio context, and the output delay of four audio frames applied by each of the 20 transformer layers 400 evaluates another distribution of the training data set for the 2.4 - second look - ahead audio context.

[0042] FIG. 3B shows an example of a T-T model 200 having an audio encoder 300 arranged in a y-architecture to enable streaming speech recognition and non-streaming speech recognition to be performed in parallel on an utterance 106 with a single T-T model 200 input. The y-architecture of the audio encoder 300 is formed by parallel low-latency branches 321 and high-latency branches 322 each extending from an initial stack 310 of transformer layers 400. As described above with reference to FIG. 3A, the initial stack 310 of transformer layers 400 may be trained using zero look-ahead audio context, and the final stack 320 of transformer layers 400 may be trained using variable look-ahead audio context. Thus, during inference, the final stack 320 of transformer layers 400 provides a low-latency branch 321 (also referred to as a low-latency decoding branch 321) by applying a first look-ahead audio context and may provide a high-latency branch 322 (also referred to as a "high-latency decoding branch 322") by applying a second look-ahead audio context related to a look-ahead audio context having a duration longer than the first look-ahead audio context.

[0043] In some examples, the first look-ahead audio context is zero or a minimum output delay (e.g., 240 milliseconds) and reduces the word error rate due to the constrained alignment applied during training. The second look-ahead audio context may include any duration of look-ahead audio context. For example, the second look-ahead audio context may include a duration of 1.2 seconds or 2.4 seconds.

[0044] In the illustrated example, the initial stack 310 of the transformer layer 400 of the audio encoder 300 receives audio data 110 (e.g., acoustic frames) corresponding to the utterance 106 input by the user 104 and captured by the user device 10. While receiving the audio data 110, the initial stack 310 of the transformer layer 400 (e.g., 15 transformer layers 400) may calculate a shared activation 312 used by both the low-latency branch 321 and the high-latency branch 322 by applying a zero look-ahead audio context. Thereafter, the final stack 320 of the transformer layer 400 (e.g., 5 transformer layers 400) uses the shared activation 312 calculated by the initial stack 310 of the transformer layer 400 while receiving the audio data 110 when the user 104 is uttering the input utterance 106, applies a first look-ahead audio context, and calculates a low-latency activation 323 that is used by the low-latency branch 321 but not by the high-latency branch 322. This low-latency activation 323 may be decoded via the joint network 230 and the softmax 240 to provide a partial speech recognition result 120a for the input utterance 106. Thus, the low-latency branch 321 uses the first look-ahead audio context to encode the audio data 110 while receiving the audio data 110 corresponding to the input utterance 106, and decodes the encoded audio data 110 (i.e., represented by the low-latency activation 323) into the partial speech recognition result 120a. The partial speech recognition result 120a may be streamed in real time to be displayed on the user device 10 (FIG. 1) when the input utterance 106 is uttered by the user 104.

[0045] After the utterance 106 input by the user 104 is completed and all the audio data 110 is received, the final stack 320 of the transformer layer 400 uses the shared activation 312 calculated by the initial stack 310 of the transformer layer 400, applies the second look-ahead audio context, and calculates a high-latency activation 324 that is used by the high-latency branch 322 but not by the low-latency branch 321. This high-latency activation 324 may be decoded via the joint network 230 and the softmax 240 to provide the final speech recognition result 120b for the input utterance 106. Here, the final speech recognition result 120b decoded by the high-latency branch 322 for the input utterance 106 is delayed by a duration based on the difference between the second look-ahead audio context and the first look-ahead audio context with respect to the partial speech recognition result 120a decoded by the low-latency branch 321. Therefore, after the audio data 110 corresponding to the input utterance 106 is received (for example, after the user 104 completes the utterance 106), the high-latency branch 322 encodes the audio data 110 using the second look-ahead audio context and decodes the encoded audio data 110 (that is, the audio data represented by the high-latency activation 324) into the final speech recognition result 120b. The final speech recognition result 120b replaces the streamed partial speech recognition result 120a. In particular, since the partial speech recognition result 120a is streamed in real time, the user 104 may not perceive the latency caused by the final speech recognition result 120b. However, the final speech recognition result 120b that benefits from the look-ahead audio context may correct the recognition errors present in the partial speech recognition result 120a, thereby making the final speech recognition result 120b more suitable for query interpretation by downstream NLU modules and / or applications. Although not shown, a separate re-determination model (for example, a LAS model) may re-evaluate the candidate hypotheses decoded by the high-latency decoding branch 322.

[0046] In some examples, the duration of the second look-ahead audio context is based on the type of application that the input utterance 106 is directed to (e.g., digital assistant application 50). For example, the audio encoder 300 may receive an application identifier 52 indicating the type of application that the input utterance is directed to, and the audio encoder may set the duration of the second look-ahead audio context based on the application identifier 52. The type of application that the input utterance 106 is directed to may serve as context information indicating the user's tolerance for speech recognition latency, speech recognition accuracy, whether the input utterance corresponds to a conversational query (e.g., a short utterance) or a non-conversational query (e.g., a long utterance such as dictation), or any other information that can be derived to optimize / adjust the duration of the second look-ahead audio context applied.

[0047] FIG. 4 shows an exemplary transformer layer 400 between multiple transformer layers of the audio encoder 300. Here, between each time step, the initial transformer layer 400 receives the corresponding acoustic frame 110 as input and generates a corresponding output representation / embedding 450 that is received as input by the next transformer layer 400. That is, each transformer layer 400 following the initial transformer layer 400 may receive an input embedding 450 corresponding to the output representation / embedding output as output by the immediately preceding transformer layer 400. The final transformer layer 400 (e.g., the last transformer layer within the final stack 320) generates a higher-order feature representation ah t (FIG. 2) for the corresponding acoustic frame 110 at each of the plurality of time steps.

[0048] The input to the label encoder 220 (FIG. 2) is the sequence of non-blank symbols y 0 ,..., y ui-1It may include a vector indicating (e.g., a one-hot vector). Thus, when the label encoder 220 includes a transformer layer, the initial transformer layer may receive the input embedding 111 by passing the one-hot vector through a lookup table.

[0049] Each transformer layer 400 of the audio encoder 300 includes a normalization layer 404, a masked multi-head attention layer 406 with relative position encoding, a residual connection 408, a stacking / unstacking layer 410, and a feed-forward layer 412. The masked multi-head attention layer 406 with relative position encoding provides a flexible way to control the amount (i.e., duration) of the look-ahead audio context used by the T-T model 200. Specifically, after the normalization layer 404 normalizes the acoustic frames 110 and / or the input embedding 111, the masked multi-head attention layer 406 projects the input to a certain value for all heads. Then, the masked multi-head layer 406 may mask the attention scores to the preceding frames of the current acoustic frame 110 to generate an output conditioned only on the previous acoustic frames 110. Next, the weighted average values for all heads are concatenated and passed to the dense layer 416, where the residual connection 414 is added to the normalized input and the output of the dense layer 416 to form the final output of the multi-head attention layer 406 with relative position encoding. The residual connection 408 is added to the output of the normalization layer 404 and provided as an input to each of the masked multi-head attention layer 406 or the feed-forward layer 412. The stacking / unstacking layer 410 can be used to change the frame rate for each transformer layer 400 to accelerate training and inference.

[0050] The feed-forward layer 412 applies the normalization layer 404 and then is applied in sequence to the dense layer 1 420, the normalization linear layer (ReLu) 418, and the dense layer 2 416. Relu 418 is used as the activation for the output of the dense layer 1 420. Similar to the multi-head attention layer 406 with relative position encoding, a residual connection 414 of the output from the normalization layer 404 is added to the output of the dense layer 2 416.

[0051] FIG. 5 includes a flowchart of an exemplary configuration of operations for a method 500 of integrating streaming speech recognition and non-streaming speech recognition using a single transformer-transducer (T-T) model 200. In operation 502, method 500 includes receiving audio data 110 corresponding to utterance 106 as an input to the T-T model 200. In operation 504, the method further includes using the T-T model 200 to perform streaming speech recognition and non-streaming speech recognition in parallel on the audio data 110.

[0052] In the low-latency branch 321 of the T-T model 200, method 500 includes, in operation 506, encoding the audio data 110 using a first look-ahead audio context while receiving the audio data 110 corresponding to the utterance 106. Method 500 also includes, in operation 508, decoding the audio data 110 encoded using the first look-ahead audio context into a partial speech recognition result 120a for the input utterance 106. In operation 510, method 500 further includes streaming the partial speech recognition result 120a for the input utterance 106.

[0053] In the high-latency branch 322 of the T-T model 200, method 500 includes, in operation 512, encoding audio data 110 corresponding to utterance 106 using a second look-ahead audio context after the audio data 110 has been received. Method 500 also includes, in operation 514, decoding the audio data 110 encoded using the second look-ahead audio context into a final speech recognition result 120b for the input utterance 106. In operation 516, method 500 further includes replacing the streamed partial speech recognition result 120a with the final speech recognition result 120b.

[0054] FIG. 6 is a schematic diagram of an exemplary computing device 600 that may be used to implement the systems and methods described herein. Computing device 600 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The components shown here, their connections and relationships, and their functions are intended to be exemplary only, and are not intended to limit the implementations of the invention described and / or claimed herein.

[0055] Computing device 600 includes a processor 610, a memory 620, a storage device 630, a high-speed interface / controller 640 connected to the memory 620 and a high-speed expansion port 650, and a low-speed interface / controller 660 connected to a low-speed bus 670 and the storage device 630. Each of the components 610, 620, 630, 640, 650, and 660 is interconnected using various buses and may be mounted on a common motherboard or attached in other ways as required. The processor 610 can process instructions executed within the computing device 600, including instructions stored in the memory 620 or on the storage device 630 for displaying graphical information about a graphical user interface (GUI) on an external input / output device such as a display 680 coupled to the high-speed interface 640. In other implementations, multiple processors and / or multiple buses may be used, along with multiple memories and multiple types of memory, as required. Also, multiple computing devices 600 may be connected so that each device performs a part of the required operations (e.g., a server bank, a group of blade servers, or a multi-processor system).

[0056] Memory 620 stores information non-temporarily within computing device 600. Memory 620 may be a computer-readable medium, a volatile memory unit, or a non-volatile memory unit. The non-temporary memory 620 may be a physical device used to temporarily or persistently store a program (e.g., a sequence of instructions) or data (e.g., program state information) so that it can be used by computing device 600. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (commonly used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disks or tapes.

[0057] Storage device 630 can provide a mass storage device for computing device 600. In some implementations, storage device 630 is a computer-readable medium. In various different implementations, storage device 630 may be a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid state memory device, or an array of devices including devices within a storage area network or other configuration. In additional implementations, a computer program product is actually embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more of the methods as described above. The information carrier is a computer or machine-readable medium such as memory 620, storage device 630, or memory on processor 610.

[0058] The high-speed controller 640 manages bandwidth-intensive operations for the computing device 600, while the low-speed controller 660 manages lower bandwidth-intensive operations. Such an allocation of duties is merely exemplary. In some implementations, the high-speed controller 640 is coupled to the memory 620, the display 680 (e.g., via a graphics processor or accelerator), and the high-speed expansion port 650. The high-speed expansion port 650 may accept various expansion cards (not shown). In some implementations, the low-speed controller 660 is coupled to the storage device 630 and the low-speed expansion port 690. The low-speed expansion port 690 may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), and may be coupled to one or more input / output devices such as a keyboard, a pointing device, a scanner, or a networking device such as a switch or router, e.g., via a network adapter.

[0059] The computing device 600 may be implemented in several different ways as shown. For example, the computing device 600 may be implemented as a standard server 600a, or multiple times within a group of such servers 600a, as a laptop computer 600b, or as part of a rack server system 600c.

[0060] Various implementations of the systems and techniques described herein can be realized in digital electrical and / or optical circuits, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementations within one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor. The programmable processor may be dedicated or general purpose and may be coupled to receive and transmit data and instructions between a memory system, at least one input device, and at least one output device.

[0061] A software application (i.e., a software resource) may refer to computer software that causes a computing device to perform a task. In some examples, a software application may be referred to as an "application", an "app", or a "program". Exemplary applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, document processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.

[0062] A non-transitory memory may be a physical device used to temporarily or persistently store a program (e.g., a sequence of instructions) or data (e.g., program state information) so as to be usable by a computing device. The non-transitory memory may be a volatile and / or non-volatile addressable semiconductor memory. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (commonly used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), as well as disks or tapes.

[0063] These computer programs (also known as programs, software, software applications or code) include machine instructions for a programmable processor and can be implemented in high-level procedural languages and / or object-oriented programming languages, and / or assembly / machine language. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, non-transitory computer-readable medium, apparatus and / or device (e.g., magnetic disks, optical disks, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives the machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0064] The processes and logical flows described in this specification, also referred to as data processing hardware, can be executed by one or more programmable processors that execute one or more computer programs to act on input data and generate output. The processes and logical flows can also be executed by special purpose logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). Processors suitable for executing computer programs include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. In general, a processor receives instructions and data from a read only memory or a random access memory or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. In general, a computer also includes one or more mass storage devices, such as magnetic disks, magneto - optical disks, or optical disks, for storing data, or is operatively coupled to receive data from one or more mass storage devices or transfer data to one or more mass storage devices or both. However, a computer need not have such devices. Computer - readable media suitable for storing computer program instructions and data include all forms of non - volatile memory, media and memory devices, including by way of example semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices, magnetic disks, such as internal hard disks or removable disks, magneto - optical disks, and CD ROM and DVD ROM disks. The processor and the memory can be assisted by, or incorporated in, special purpose logic circuitry.

[0065] To enable interaction with a user, one or more aspects of the present disclosure can be implemented on a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touch screen, and optionally a keyboard and a pointing device, such as a mouse or trackball, by which the user can provide input to the computer. Other types of devices can also be used to enable interaction with the user. For example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and the input from the user can be received in any form, including acoustic input, voice input, or tactile input. Also, the computer can interact with the user by sending documents to the device used by the user and receiving documents from that device, for example, by sending a web page to a web browser on the user's client device in response to a request received from the web browser.

[0066] Some implementations have been described. Nevertheless, it will be understood that various modifications may be made without departing from the spirit and scope of the present disclosure. Accordingly, other implementations are within the scope of the following claims.

Description of Reference Numerals

[0067] 1 or 2 hours 10 User device 12 Data processing hardware 14 Memory hardware 16 Audio system 16a Voice input device 16b Voice output device 18 Digital assistant interface 19 Response 20 Network 50 Digital assistant application 52 Application identifier 60 Remote computing device 100 Acoustic environment 104 User 106 Utterance 107 User interface generator 108 Audio subsystem 110 Acoustic frame, audio data 111 Input embedding 118 ASR system 120 Transcription 120a Partial speech recognition result 120b Final speech recognition result 200 Transformer-Transducer (T-T) model 220 Label encoder 230 Joint network 240 Softmax layer 300 Audio encoder 310 Initial stack 312 Shared activation 320 Final stack 321 Low-latency branch 322 High-latency branch 323 Low-latency activation 324 High-latency activation 400 Transformer layer 404 Normalization layer 406 Masked multi-head attention layer 408, 414 Residual connection 410 Stacking / unstacking layer 412 Feed-forward layer 418 Normalized linear layer 416, 420 Dense layer 450 Output representation / embedding 600 Computing device 600a Server 600b Laptop computer 600c Rack Server System 610 Processor 620 Memory 630 Storage Device 640 High-Speed Interface / Controller 650 High-Speed Expansion Port 660 Low-Speed Interface / Controller 670 High-Speed Bus 680 Display 690 Low-Speed Expansion Port

Claims

1. A computer program for causing a computer to operate as a single transformer-transducer model (200) that integrates streaming and non-streaming speech recognition, the single transformer-transducer model (200) comprising: An audio encoder (300), receiving as input a sequence of acoustic frames (110); At each of a plurality of time steps, generate a high-level feature representation for a corresponding acoustic frame (110) in the sequence of acoustic frames (110). An audio encoder (300) configured to: A label encoder (220), receiving as input the sequence of non-blank symbols output by the final softmax layer (240); generating a dense representation at each of said plurality of time steps A label encoder (220) configured as follows: A joint network (230), receiving as input the high-level feature representation generated by the audio encoder (300) at each of the plurality of time steps and the dense representation generated by the label encoder (220) at each of the plurality of time steps; generating, at each of the plurality of time steps, a probability distribution over possible speech recognition hypotheses at the corresponding time step; and a joint network (230) configured to: The audio encoder (300) comprises a neural network having a number of transformer layers (400), the number of transformer layers (400) comprising: an initial stack (310) of transformer layers (400), each trained with a zero look-ahead audio context; a final stack (320) of transformer layers (400), each trained with a variable look-ahead audio context.

2. Each transformer layer (400) of the audio encoder (300) A normalization layer (404); a masked multi-head attention layer with relative position coding (406); Residual connection (408); a stacking / unstacking layer (410); A computer program product as claimed in claim 1 , further comprising:

3. 3. The computer program product of claim 2, wherein the stacking / unstacking layer is configured to modify a frame rate of the corresponding transformer layer to accommodate processing time by the single transformer-transducer model during training and inference.

4. 4. The computer program of claim 1, wherein the initial stack of transformer layers comprises more transformer layers than the final stack of transformer layers.

5. 5. The computer program product of claim 1, wherein during training, the variable look-ahead audio context is uniformly sampled for each transformer layer (400) in the final stack (320) of transformer layers (400).

6. a low-latency decoding branch (321) configured to decode a corresponding speech recognition result (120) for the input utterance (106) from the encoded audio data (110) using a first look-ahead audio context; and a high-latency decoding branch (322) configured to decode corresponding speech recognition results (120) for the input utterance (106) from audio data (110) encoded using a second look-ahead audio context, the second look-ahead audio context including look-ahead audio of a longer duration than the first look-ahead audio context.

7. the initial stack (310) of the transformer layer (400) applies a zero look-ahead audio context to calculate shared activations (312) used by both the low-latency decoding branch (321) and the high-latency decoding branch (322); the final stack (320) of the transformer layer (400) applies the first look-ahead audio context to calculate low-latency activations (323) used by the low-latency decoding branch (321) but not by the high-latency decoding branch (322); 7. The computer program product of claim 6, wherein the final stack (320) of the transformer layer (400) applies the second look-ahead audio context to calculate high-latency activations (324) that are used by the high-latency decoding branch (322) but not by the low-latency decoding branch (321).

8. 8. The computer program product of claim 6 or 7, wherein the first lookahead audio context comprises a lookahead audio context of zero.

9. 9. The computer program product of claim 6, wherein the low latency decoding branch (321) and the high latency decoding branch (322) are executed in parallel to decode the corresponding speech recognition result (120) for the input utterance (106).

10. 10. The computer program product of claim 9, wherein a corresponding speech recognition result (120) decoded by the high latency decoding branch (322) for the input utterance (106) is delayed by a duration based on a difference between the second look-ahead audio context and the first look-ahead audio context relative to a corresponding speech recognition result (120) decoded by the low latency decoding branch (321) for the input utterance (106).

11. the low latency decoding branch (321) is configured to stream the corresponding speech recognition result (120) as a partial speech recognition result (120) when the input utterance (106) is received by the single transformer-transducer model (200); 11. The computer program product of claim 6, wherein the high-latency decoding branch (322) is configured to output the corresponding speech recognition result as a final transcription (120) after the single transformer-transducer model (200) receives the complete input utterance (106).

12. The input speech (106) is directed to an application (50); 12. The computer program product of claim 6, wherein a duration of the second look-ahead audio context used by the high latency decoding branch to decode the corresponding speech recognition result for the input utterance is based on a type of the application to which the input utterance is intended.

13. The computer program product of claim 1 , wherein the label encoder (220) comprises a neural network having a plurality of transformer layers (400).

14. The computer program product of claim 1 , wherein the label encoder (220) comprises a bigram embedding lookup decoder model.

15. the single transformer-transducer model (200) is executed on a client device (10); or 15. The computer program of claim 1, wherein the single transformer-transducer model (200) is executed on a server-based system (60).

16. A computer-implemented method (500) that, when executed on data processing hardware (12), causes the data processing hardware (12) to perform operations, the operations including: receiving audio data (110) corresponding to speech (106) as input to a transformer-transducer model (200), the transformer-transducer model (200) comprising an audio encoder (300) comprising a neural network having a plurality of transformer layers (400), the plurality of transformer layers (400) comprising: an initial stack (310) of transformer layers (400), each trained with a zero look-ahead audio context; a final stack (320) of transformer layers (400), each trained with a variable look-ahead audio context; and performing streaming and non-streaming speech recognition in parallel on the audio data (110) using the transformer-transducer model (200), said performing including: In the low latency branch (321) of the transformer-transducer model (200), During reception of the audio data (110) corresponding to the utterance (106), encoding the audio data (110) using a first look-ahead audio context; decoding the audio data (110) encoded using the first look-ahead audio context into a partial speech recognition result (120) for the utterance (106); streaming the partial speech recognition results (120) for the utterance (106); In the high latency branch (322) of the transformer-transducer model, after the audio data (110) corresponding to the utterance (106) is received, encoding the audio data (110) using a second look-ahead audio context; decoding the audio data (110) encoded using the second look-ahead audio context into a final speech recognition result (120) for the utterance (106); The computer-implemented method (500) is performed by replacing the streamed partial speech recognition results (120) with the final speech recognition results (120).

17. Each transformer layer (400) A normalization layer (404); a masked multi-head attention layer with relative position coding (406); Residual connection (408); a stacking / unstacking layer (410); and a feedforward layer.

18. 20. The computer-implemented method of claim 17, wherein the stacking / unstacking layer is configured to modify a frame rate of the corresponding transformer layer to accommodate processing time by a single transformer-transducer model during training and inference.

19. 19. The computer-implemented method (500) of any one of claims 16 to 18, wherein the initial stack (310) of transformer layers (400) comprises more transformer layers (400) than the final stack (320) of transformer layers (400).

20. 20. The computer-implemented method (500) of any one of claims 16 to 19, wherein during training, the variable look-ahead audio context is uniformly sampled for each transformer layer (400) in the final stack (320) of transformer layers (400).

21. the initial stack (310) of the transformer layer (400) applies a zero look-ahead audio context to calculate a shared activation (312) used by both the low latency branch (321) and the high latency branch (322); the final stack (320) of the transformer layer (400) applies the first look-ahead audio context to compute low-latency activations (323) used by the low-latency branch (321) but not the high-latency branch (322); 21. The computer-implemented method (500) of any one of claims 16 to 20, wherein the final stack (320) of a transformer layer (400) applies the second look-ahead audio context to calculate high-latency activations (324) that are used by the high-latency branch (322) but not the low-latency branch (321).

22. 22. The computer-implemented method (500) of any one of claims 16 to 21, wherein the first lookahead audio context comprises a zero lookahead audio context.

23. 23. The computer-implemented method (500) of any one of claims 16 to 22, wherein the final speech recognition result (120) decoded by the high latency branch (322) for the utterance (106) is delayed by a duration based on a difference between the second look-ahead audio context and the first look-ahead audio context relative to the partial speech recognition result (120) decoded by the low latency branch (321) for the utterance (106).

24. The operation includes: receiving an application identifier (52) indicating a type of application (50) to which the utterance (106) is directed; 24. The computer-implemented method (500) of any one of claims 16 to 23, further comprising: setting a duration of the second look-ahead audio context based on the application identifier (52).

25. 25. The computer-implemented method (500) of any one of claims 16 to 24, wherein the transformer-transducer model (200) comprises a label encoder (220) including a neural network having a plurality of transformer layers (400).

26. 25. The computer-implemented method (500) of any one of claims 16 to 24, wherein the transformer-transducer model (200) comprises a label encoder (220) including a bigram embedding lookup decoder model.

27. The data processing hardware (12) executes the transformer-transducer model (200); Client device (10), or 27. The computer-implemented method (500) of any one of claims 16 to 26, residing on a server-based system (60).

Citation Information

Patent Citations

  • System and method for end-to-end speech recognition with triggered attention

    WO2020195068A1