Tied and reduced RNN-T

The RNN-T model addresses the challenge of real-time ASR in mobile devices by using a prediction network and joint network to provide immediate and accurate speech recognition results, enhancing user experience through reduced latency and improved accuracy.

JP7716491B2Active Publication Date: 2025-07-31GOOGLE LLC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023558608
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-03-23
Filing Date
2021-05-26
Publication Date
2025-07-31
Estimated Expiration
2041-05-26

AI Technical Summary

Technical Problem

Modern automatic speech recognition (ASR) systems face challenges in achieving low latency and high accuracy while decoding speech in real-time, especially in mobile devices where immediate display of recognized words is crucial.

Method used

A recurrent neural network transducer (RNN-T) model that includes a prediction network and a joint network, utilizing a shared embedding matrix and multi-head attention mechanism to generate embedded representations and probability distributions over speech recognition hypotheses, allowing for streaming speech recognition with reduced latency and improved accuracy.

Benefits of technology

The RNN-T model enables low-latency, streaming speech recognition with improved accuracy by generating initial results quickly and refining them with additional processing, ensuring high-quality transcription without noticeable delays.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007716491000015
    Figure 0007716491000015
  • Figure 0007716491000016
    Figure 0007716491000016
  • Figure 0007716491000017
    Figure 0007716491000017
Patent Text Reader

Abstract

The RNN-T model (200) includes a prediction network (300) configured to receive a sequence of non-blank symbols (301) at each of a plurality of time steps following the initial time step. For each non-blank symbol, the prediction network is also configured to generate an embedding representation (306) of the corresponding non-blank symbol using a shared embedding matrix (304), assign a respective position vector (308) to the non-blank symbol, and weight the embedding representation in proportion to the similarity between the embedding representation and the respective position vector. The prediction network is also configured to generate a single embedding vector (350) at the corresponding time step. The RNN-T model also includes a joint network (230) configured to receive the single embedding vector generated as an output from the prediction network at the corresponding time step at each of a plurality of time steps following the initial step and generate a probability distribution over possible speech recognition hypotheses.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a timed and reduced recurrent neural network transducer (RNN-T) model.

Background Art

[0002] Modern automatic speech recognition (ASR) systems emphasize providing not only high quality (e.g., low word error rate (WER)) but also low latency (e.g., reduction of the delay from when a user speaks until characters are produced). Furthermore, when using an ASR system currently, it is required that the ASR system decodes speech in a streaming manner corresponding to real time, or in some cases, decodes speech faster than real time. As an example, when an ASR system is placed on a mobile phone where direct user interaction processing is performed, it may be necessary for an application on the mobile phone using the ASR system to stream voice recognition, such that words appear on the screen immediately after they are spoken. Here, there may also be a low tolerance for latency for the mobile phone user. Due to this low tolerance, efforts have been made to execute on a mobile device such that speech recognition minimizes the effects of latency and inaccuracy that can negatively impact the user experience.

Summary of the Invention

Means for Solving the Problems

[0003] One aspect of the present disclosure provides a recurrent neural network transducer (RNN-T) model that includes a prediction network configured to receive, at each of a plurality of time steps following an initial time step, a sequence of non-blank symbols output by a final softmax layer as input. The prediction network is also configured to generate, for each non-blank symbol in the sequence of non-blank symbols received as input at a corresponding time step, an embedded representation of the corresponding non-blank symbol using a shared embedding matrix, assign a respective position vector to the corresponding non-blank symbol, and weight the embedded representation in proportion to the similarity between the embedded representation and the respective position vector. The prediction network is also configured to generate, as output, a single embedded vector at a corresponding time step, the single embedded vector being based on a weighted average of the weighted embedded representations. The RNN-T model also includes a joint network configured to receive, at each of a plurality of time steps following an initial step, a single embedded vector generated as output from the prediction network at a corresponding time step and generate a probability distribution over possible speech recognition hypotheses at the corresponding time step.

[0004] Implementations of the present disclosure may include one or more of any of the following features. In some implementations, the RNN-T model further includes an audio encoder configured to receive a sequence of acoustic frames as input and generate, for each of a plurality of time steps, a higher-order feature representation for the corresponding acoustic frame in the sequence of acoustic frames. Here, the joint network is further configured to receive, for each of the plurality of time steps, the higher-order feature representation generated at the corresponding time step by the audio encoder as input. In some examples, weighting the embedding representations in proportion to the similarity between the embedding representations and their respective positional vectors includes weighting the embedding representations in proportion to the cosine similarity between the embedding representations and their respective positional vectors. The sequence of non-blank symbols output by the final softmax layer includes word pieces. Optionally, the sequence of non-blank symbols output by the final softmax layer may include graphemes. Each of the embedding representations may include the same dimensionality as each of the positional vectors. In some implementations, the sequence of non-blank symbols received as input is limited to the N previous non-blank symbols output by the final softmax layer. In these implementations, N may be equal to 2. Alternatively, N may be equal to 5.

[0005] In some examples, the prediction network includes a multi-head attention mechanism that shares a shared embedding matrix across its heads. In these examples, at each of a plurality of time steps following an initial time step, the prediction network, at each head of the multi-head attention mechanism and for each non-blank symbol in a sequence of non-blank symbols received as input at a corresponding time step, uses the shared embedding matrix to generate an embedding representation of the corresponding non-blank symbol that is the same as the embedding representations generated at each other head of the multi-head attention mechanism, assigns to the corresponding non-blank symbol a respective position vector that is different from the respective position vectors assigned to the corresponding non-blank symbols at each other head of the multi-head attention mechanism, and is configured to weight the embedding representation in proportion to the similarity between the embedding representation and the respective position vectors. Here, the prediction network also generates, as an output from the corresponding head of the multi-head attention mechanism, a respective weighted average of the weighted embedding representations of the sequence of non-blank symbols, and is configured to generate, as an output, a single embedding vector at the corresponding time step by averaging the respective weighted averages output from the corresponding head of the multi-head attention mechanism. In these examples, the multi-head attention mechanism may include four heads. Optionally, the prediction network may tie the dimension of the shared embedding matrix to the dimension of the output layer of the joint network.

[0006] Another aspect of the present disclosure provides a computer-implemented method that, when executed on data processing hardware, causes the data processing hardware to perform operations. At each of a plurality of time steps following an initial time step, the operations include: receiving as input to a recurrent neural network transducer (RNN-T) model a sequence of non-blank symbols output by a final softmax layer; for each non-blank symbol in the sequence of non-blank symbols received as input at the corresponding time step, generating, by a prediction network, an embedding representation of the corresponding non-blank symbol using a shared embedding matrix; assigning, by the prediction network, a respective position vector to the corresponding non-blank symbol; and weighting, by the prediction network, the embedding representation in proportion to the similarity between the embedding representation and the respective position vector. At each of a plurality of time steps following the initial time step, the operations also include generating, by a joint network of the RNN-T model, a single embedding vector at the corresponding time step as output from the prediction network; and generating, by the joint network of the RNN-T model, a probability distribution over possible speech recognition hypotheses at the corresponding time step using the single embedding vector generated as output from the prediction network at the corresponding time step. The single embedding vector is based on a weighted average of the weighted embedding representations.

[0007] Implementations of the present disclosure may include one or more of any of the following features. In some implementations, the operations further include receiving, as input to an audio encoder, a sequence of acoustic frames; generating, by the audio encoder, for each of a plurality of time steps, a higher-order feature representation for a corresponding acoustic frame in the sequence of acoustic frames; and receiving, for each of the plurality of time steps, as input to a joint network, the higher-order feature representation generated at the corresponding time step by the audio encoder. In some examples, weighting an embedding representation in proportion to a similarity between the embedding representation and its respective position vector includes weighting the embedding representation in proportion to a cosine similarity between the embedding representation and its respective position vector.

[0008] The sequence of non-blank symbols output by the final softmax layer may include word pieces. In some cases, the sequence of non-blank symbols output by the final softmax layer may include graphemes. Each of the embedding representations may include the same dimensionality as each of the position vectors. In some implementations, the sequence of non-blank symbols received as input is limited to the N previous non-blank symbols output by the final softmax layer. In these implementations, N may be equal to 2. Alternatively, N may be equal to 5. In some cases, the prediction network may tie the dimension of a shared embedding matrix to the dimension of the output layer of the joint network.

[0009] In some examples, the prediction network includes a multi-head attention mechanism that shares a shared embedding matrix across each of its heads. The multi-head attention mechanism may include four heads. In these examples, at each of a plurality of time steps following the initial time step, the operations may further include: for each non-blank symbol in a sequence of non-blank symbols received as input at each head of the multi-head attention mechanism and at the corresponding time step, generating, by the prediction network, an embedding representation of the corresponding non-blank symbol using the shared embedding matrix that is the same as an embedding representation generated at each other head of the multi-head attention mechanism; assigning, by the prediction network, a respective position vector to the corresponding non-blank symbol that differs from the respective position vector assigned to the corresponding non-blank symbol at each other head of the multi-head attention mechanism; and weighting, by the prediction network, the embedding representation in proportion to the similarity between the embedding representation and the respective position vector. Here, at each of a plurality of time steps following the initial time step, and at each head of the multi-head attention mechanism, the operations may also include generating, by the predictive network, a weighted average of each of the weighted embedding representations of the sequence of non-blank symbols as an output from the corresponding head of the multi-head attention mechanism. Thereafter, at each of a plurality of time steps following the initial time step, the operations may further include generating, as an output from the predictive network, a single embedding vector for the corresponding time step by averaging the respective weighted averages output from the corresponding head of the multi-head attention mechanism.

[0010] The details of one or more implementations of the disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will become apparent from the description and drawings, and from the claims. [Brief explanation of the drawings]

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

DETAILED DESCRIPTION OF THE INVENTION

[0012] Like reference symbols in the various drawings indicate like elements.

[0013] 1 is an example of a speech environment 100. In the speech environment 100, a user 104 may interact with a computing device, such as a user device 10, through voice input. The user device 10 (also commonly referred to as a device 10) is configured to capture speech (e.g., streaming audio data) from one or more users 104 in the speech environment 100. Here, streaming audio data may refer to utterances 106 by the users 104 that serve as audible queries, commands to the device 10, or audible communications captured by the device 10. A voice-enabled system of the device 10 may appropriately process the query or command by answering the query and / or having the command executed / completed by one or more downstream applications.

[0014] The user device 10 may correspond to any computing device associated with a user and capable of receiving audio data. Some examples of the user device 10 include, but are not limited to, a mobile device (e.g., a mobile phone, a tablet, a laptop, etc.), a computer, a wearable device (e.g., a smart watch), a smart appliance, an Internet of Things (IoT) device, an in-vehicle infotainment system, a smart display, a smart speaker, etc. The user device 10 includes data processing hardware 12 and memory hardware 14 in communication with the data processing hardware 12, storing instructions that, when executed by the data processing hardware 12, cause the data processing hardware 12 to perform one or more operations. The user device 10 further includes an audio system 16 having audio capture devices (e.g., microphones) 16, 16a for capturing and converting speech 106 in the speech environment 100 into electrical signals, and speech output devices (e.g., speakers) 16, 16b for communicating audible audio signals (e.g., as output audio data from the device 10). Although the user device 10 implements a single audio capture device 16a in the illustrated example, it may implement an array of audio capture devices 16a without departing from the scope of this disclosure, and one or more capture devices 16a in the array may not be physically present on the user device 10 but may communicate with the audio system 16.

[0015] In the speech environment 100, an automatic speech recognition (ASR) system 118 that implements a recurrent neural network transducer (RNN-T) model 200 and an optional rescorer 180 exists on the user device 10 of the user 104 and / or on a remote computing device 60 (e.g., one or more servers of a distributed system executed in a cloud computing environment) that communicates with the user device 10 via the network 40. The user device 10 and / or the remote computing device 60 also receive the speech 106 spoken by the user 104 and captured by the audio capture device 16a, and convert the speech 106 into a corresponding digital format associated with the input acoustic frame 110 that can be processed by the ASR system 118, including an audio subsystem 108 configured to do so. In the illustrated example, the user emits each speech 106, and the audio subsystem 108 converts the speech 106 into a corresponding audio data (e.g., acoustic frame) 110 for input to the ASR system 118. Thereafter, the RNN-T model 200 receives the audio data 110 corresponding to the speech 106 as input and generates / predicts the corresponding transcription 120 (e.g., recognition result / hypothesis) of the speech 106 as output. In the illustrated example, the RNN-T model 200 may perform streaming speech recognition to generate initial speech recognition results 120, 120a, and the rescorer 180 may update (rescore) the initial speech recognition result 120a to generate final speech recognition results 120, 120b.

[0016] The user device 10 and / or the remote computing device 60 also execute a user interface generator 107 configured to present a representation of a transcription 120 of the utterance 106 to the user 104 of the user device 10. As described in more detail below, the user interface generator 107 may stream initial speech recognition results 120a for time 1, followed by final speech recognition results 120b for time 2. In some configurations, the transcription 120 output from the ASR system 118 is processed by a natural language understanding (NLU) module executing, for example, on the user device 10 or the remote computing device 60, to execute the user command / query specified by the utterance 106. Additionally or alternatively, a text-to-speech system (not shown) (e.g., executing on any combination of the user device 10 or the remote computing device 60) may convert the transcription into synthesized speech for audible output by the user device 10 and / or another device.

[0017] In the illustrated example, a user 104 interacts with a program or application 50 (e.g., a digital assistant application 50) on a user device 10 that uses an ASR system 118. For example, FIG. 1 shows a user 104 communicating with the digital assistant application 50 and the digital assistant application 50 displaying a digital assistant interface 18 on the screen of the user device 10, illustrating a conversation between the user 104 and the digital assistant application 50. In this example, the user 104 asks the digital assistant application 50, "What time does the concert start tonight?" This question from the user 104 is utterance 106 that is captured by an audio capture device 16a and processed by an audio system 16 of the user device 10. In this example, the audio system 16 receives the utterance 106 and converts it into acoustic frames 110 for input to the ASR system 118.

[0018] Continuing with this example, the RNN-T model 200 receives the acoustic frame 110 corresponding to the utterance 106 spoken by the user 104, encodes the acoustic frame 110, and then decodes the encoded acoustic frame 110 into the initial speech recognition result 120a. During time 1, the user interface generator 107 streams, via the digital assistant interface 18, the representation of the initial speech recognition result 120a of the utterance 106 to the user 104 of the user device 10 such that words, word pieces, and / or individual characters appear on the screen immediately after being spoken. In some examples, the first look-ahead audio context is equal to zero.

[0019] During time 2, the user interface generator 107 presents the representation of the final speech recognition result 120b of the utterance 106 to the user 104 of the user device 10, which is rescored by the rescorer 180, via the digital assistant interface 18. In some implementations, the user interface generator 107 replaces the representation of the initial speech recognition result 120a presented at time 1 with the representation of the final speech recognition result 120b presented at time 2. Here, time 1 and time 2 may include corresponding timestamps when the user interface generator 107 presents the respective speech recognition results. In this example, the timestamp of time 1 indicates that the user interface generator 107 presents the initial speech recognition result 120a at a time earlier than the final speech recognition result 120b. For example, since the final speech recognition result 120b is presumed to be more accurate than the initial speech recognition result 120a, the final speech recognition result 120b, which is ultimately displayed as the transcription 120, may correct words that may have been misrecognized in the initial speech recognition result 120a. In this example, the streaming initial speech recognition result 120a output by the RNN-T model 200 and displayed on the screen of the user device 10 at time 1 has low latency and provides the user 104 with a response indicating that the user's query is being processed, while the final speech recognition result 120b output by the rescorer 180 and displayed on the screen at time 2 utilizes additional speech recognition models and / or language models to improve the speech recognition quality in terms of accuracy but increases the latency. However, since the initial speech recognition result 120a is displayed when the user utters the utterance 106, the higher latency associated with the generation and final display of the final recognition result is not noticed by the user 104.

[0020] In the example shown in FIG. 1 , digital assistant application 50 may respond to a question posted by user 104 by using natural language processing. Natural language processing generally refers to the process of interpreting written language (e.g., initial speech recognition result 120a and / or final speech recognition result 120b) and determining whether the written language prompts an action. In this example, digital assistant application 50 uses natural language processing to recognize that the question from user 104 is about the user's schedule, more specifically, a concert on the user's schedule. By recognizing these details through natural language processing, the automated assistant returns a response 19 to the user's query, which in this case states, "Doors open at 6:30 PM, and the concert starts at 8 PM." In some configurations, natural language processing occurs on a remote server 60 in communication with data processing hardware 12 of user device 10.

[0021] Referring to FIG. 2, an exemplary frame alignment-based transducer model 200 includes a recurrent neural network transducer (RNN-T) model architecture that adheres to latency constraints associated with interactive applications. The RNN-T model 200 offers a small computational footprint and utilizes lower memory requirements than traditional ASR architectures, making the RNN-T model architecture suitable for performing speech recognition entirely on the user device 102 (e.g., no communication with a remote server is required). The RNN-T model 200 includes an encoder network 210, a prediction network 300, and a joint network 230. The prediction network 300 and the joint network 230 collectively constitute an RNN-T decoder. The encoder network 210 is generally similar to an acoustic model (AM) in traditional ASR systems and includes a recurrent network of stacked long short-term memory (LSTM) layers. For example, the encoder receives a sequence of d-dimensional feature vectors (e.g., acoustic frames 110 (FIG. 1)) x = (x1, x2, ..., x T), where

[0022]

number

[0023] At each time step, the encoder generates a high-dimensional feature representation, which is

[0024]

number

[0025] is shown as:

[0026] Similarly, the prediction network 300 is also an LSTM network, and like the language model (LM), it computes the sequence of non-blank symbols 301, y 0 , ..., y ui-1 Expressing

[0027]

number

[0028] As explained in more detail below, the representation

[0029]

number

[0030] 350 contains a single embedding vector. In particular, the sequence of non-blank symbols 301 received at the prediction network 300 captures linguistic dependencies between the non-blank symbols 301 predicted during previous time steps to help the joint network 230 predict the probability of the next output symbol or blank symbol during the current time step. As will be explained in more detail below, to contribute techniques for reducing the size of the prediction network 300 without sacrificing the accuracy / performance of the RNN-T model 200, the prediction network 300 uses the non-blank symbols y that are limited to the N previous non-blank symbols 301 output by the final softmax layer 240. ui-n , ..., y ui-1 A limited historical sequence of

[0031] Finally, using an RNN-T model architecture, the representations generated by the encoder and predictor networks 210, 300 are combined by a joint network 230. The joint network then:

[0032]

number

[0033] Predict this, which is a distribution over the next output symbol. In other words, the joint network 230 generates, at each output step (e.g., time step), a probability distribution over possible speech recognition hypotheses. Here, "possible speech recognition hypotheses" each correspond to a set of output labels that represent symbols / characters in the specified natural language. For example, when the natural language processing is in English, the set of output labels may include 27 symbols. For example, one label corresponds to each of the 26 letters in the English alphabet, and one label designates a space. Thus, the joint network 230 may output a set of values indicating the likelihood of occurrence of each output label in a given set of output labels. This set of values may be a vector and can represent a probability distribution over the set of output labels. In some cases, the output labels are graphemes (e.g., individual characters, and in some cases punctuation marks and other symbols), but the set of output labels is not so limited. For example, the set of output labels can include word pieces and / or whole words in addition to or instead of graphemes. The output distribution of the joint network 230 can include posterior probability values for each of the different output labels. Thus, if there are 100 different output labels representing different graphemes or other symbols respectively, the output y of the joint network 230 i can include 100 different probability values, one corresponding to each output label. In that case, the probability distribution can be used to select scores to assign scores to candidate notation elements (e.g., graphemes, word pieces, and / or words) in a beam search process (e.g., by the softmax layer 240) for determining the transcription 120.

[0034] The softmax layer 240 may use any technique to select the output label / symbol with the highest probability in the distribution as the next output symbol predicted by the RNN-T model 200 at the corresponding output step. In this way, the RNN-T model 200 does not make a conditional independence assumption; rather, the prediction of each symbol is conditioned not only on the acoustics but also on the sequence of labels output so far. The RNN-T model 200 does not assume that the output symbols are independent of future acoustic frames 110, thereby allowing the RNN-T model 200 to be used in a streaming manner.

[0035] In some examples, the encoder network 210 of the RNN-T model 200 consists of eight 2048-dimensional LSTM layers, each followed by a 640-dimensional projection layer. In other implementations, the encoder network 210 includes a conformer or transformer layer network. The prediction network 220 may have two 2048-dimensional LSTM layers, each followed by a 640-dimensional projection layer and a 128-unit embedding layer. Finally, the joint network 230 may also have 640 hidden units. The softmax layer 240 may consist of a unified word piece or grapheme set generated using all unique word pieces or graphemes in the training data. When the output symbols / labels include word pieces, the set of output symbols / labels may include 4096 different word pieces. When the output symbols / labels include graphemes, the set of output symbols / labels may include fewer than 100 different graphemes.

[0036] FIG. 3 shows the non-blank symbol y y , limited to the N previous non-blank symbols 301 a - 301 n output by the final softmax layer 240. ui-n , ..., y ui-11 illustrates a prediction network 300 of an RNN-T model 200 receiving as input a sequence of non-blank symbols 301a-301n. In some examples, N is equal to 2. In other examples, N is equal to 5, although this disclosure is non-limiting and N may be equal to any integer. The sequence of non-blank symbols 301a-301n represents the initial speech recognition result 120a (FIG. 1). In some implementations, the prediction network 300 includes a multi-head attention mechanism 302, which shares a shared embedding matrix 304 across each of its heads 302A-302H. In one example, the multi-head attention mechanism 302 includes four heads. However, any number of heads may be used by the multi-head attention mechanism 302. Notably, the multi-head attention mechanism significantly improves performance with minimal increase in model size. As explained in more detail below, each head 302A-302H contains its own row of position vectors, and instead of increasing the model size by concatenating the outputs 318A-318H from all heads, the outputs 318A-318H are averaged by a head averaging module 322.

[0037] Referring to the first head 302A of the multi-head attention mechanism 302, head 302A uses a shared embedding matrix 304 to generate a non-blank symbol yi received as input at a corresponding time step from the plurality of time steps. ui-n , ..., y ui-1 For each non-blank symbol 301 in the sequence, a corresponding embedded representation 306, 306a-306n (e.g.,

[0038]

number

[0039] generates it. In particular, since the shared embedding matrix 304 is shared across all heads of the multi-head attention mechanism 302, all other heads 302B - 302H all generate the same corresponding embedding representation 306 for each non-blank symbol. Head 302A also assigns a respective position vector PV ui-n , ..., y ui-1 to each corresponding non-blank symbol in the sequence of, for example, Aa~An 308, 308Aa - 308An (for example,

[0040]

Number

[0041] ). Each respective position vector PV308 assigned to each non-blank symbol indicates the position in the history of the sequence of non-blank symbols (for example, the N previous non-blank symbols output by the final softmax layer 240). For example, the first position vector PV Aa is assigned to the latest position in the history, and the last position vector PV An is assigned to the last position in the history of the N previous non-blank symbols output by the final softmax layer 240. In particular, each of the embedding representations 306 may include the same dimension (i.e., dimension size) as each of the position vectors PV308.

[0042] The corresponding embedding representations generated by the shared embedding matrix 304 for each non-blank symbol 301 in the sequence of non-blank symbols 301a - 301n, y ui-n , ..., y ui-1 are the same in all of the heads 302A - 302H of the multi-head attention mechanism 302, but each head 302A - 302H defines a different set / row of the position vectors 308. For example, the first head 302A defines the rows of the position vectors PV Aa~An 308Aa - 308An, and the second head 302B defines the position vectors PV Ba~Bn308 Ba~Bn The Hth head 302H defines a position vector PV Ha~Hn 308 Ha~Hn define another different row of

[0043] For each non-blank symbol in the received sequence of non-blank symbols 301a-301n, the first head 302A also weights the corresponding embedded representation 306 via the weighting layer 310 in proportion to the similarity between the corresponding embedded representation and its assigned respective position vector PV 308. In some examples, the similarity may include cosine similarity (e.g., cosine distance). In the illustrated example, the weighting layer 310 outputs a sequence of weighted embedded representations 312, 312Aa-312An, each associated with a corresponding embedded representation 306 weighted in proportion to its assigned respective position vector PV 308. In other words, the weighted embedded representation 312 output by the weighting layer 310 for each embedded representation 306 may correspond to a dot product between the embedded representation 306 and its assigned respective position vector PV 308. The weighted embedded representation 312 may be interpreted as an attenuation for the embedded representation in proportion to how similar it is to the position associated with its assigned respective position vector PV 308. To increase computational speed, the prediction network 300 includes a non-recurrent layer, so that the sequence of weighted embeddings 312Aa-312An are not concatenated but instead averaged by a weighted average module 316 to produce a weighted average 318A of the weighted embeddings 312Aa-312An as output from the first head 302A, expressed as follows:

[0044]

number

[0045] In Equation 1, h represents the index of the head 302, n represents the position in the context, and e represents the embedding dimension. In addition, in Equation 1, H, N, and d eincludes the size of the corresponding dimension. The position vector PV 308 does not need to be trainable and may contain random values. In particular, even though the weighted embeddings 312 are averaged, the position vector PV 308 can potentially preserve position history state, alleviating the need for recurrent connections at each layer of the prediction network 300.

[0046] The operations described above with respect to the first head 302A are similarly performed by each of the other heads 302B-302H of the multi-head attention mechanism 302. Due to the different sets of position vectors PV 308 defined by each head 302, the weighting layer 310 outputs a sequence of weighted embedding representations 312Ba-312Bn, 312Ha-312Hn for each of the other heads 302B-302H that differs from the sequence of weighted embedding representations 312Aa-312An for the first head 302A. The weighted average module 316 then generates a weighted average 318A-318H of the corresponding weighted embedding representations 312 of the sequences of non-blank symbols as output from each of the other corresponding heads 302B-302H.

[0047] In the illustrated example, the prediction network 300 includes a head averaging module 322 that averages weighted averages 318A-318H output from corresponding heads 302A-302H. A projection layer 326 with SWISH may receive as input an output 324 from the head averaging module 322 corresponding to the average of the weighted averages 318A-318H and generate as output a projected output 328. A final layer normalization 330 normalizes the projected output 328 to produce a single embedding vector Pu at corresponding time steps from multiple time steps. i 350. The prediction network 300 generates a single embedding vector Pu at each of multiple time steps following the initial time step. i Generates only 350.

[0048] In some configurations, the prediction network 300 does not implement the multi-head attention mechanism 302 and only performs the operations described above for the first head 302A. In these configurations, the weighted average 318A of the weighted embedding representations 312Aa to 312An is simply passed through the projection layer 326 and the layer normalization 330 to provide a single embedding vector Pu i 350.

[0049] Referring again to FIG. 2, the joint network 230 receives a single embedding vector Pu i 350 from the prediction network 300 and receives the higher-level feature representation

[0050] [Number]

[0051] from the encoder 210. The joint network 230 generates a probability distribution

[0052] [Number]

[0053] over the possible speech recognition hypotheses at the corresponding time step. Here, the possible speech recognition hypotheses each correspond to a set of output labels that represent symbolic characters in the specified natural language. The probability distribution over the speech recognition hypotheses

[0054] [Number]

[0055] shows the probabilities for the final speech recognition result 120b (Figure 1). That is, the joint network 230 determines a probability distribution for the final speech recognition result 120b using a single embedding vector 350 based on a sequence of non-blank symbols (e.g., the initial speech recognition result 120a). The final softmax layer 240 receives the probability distribution for the final speech recognition result 120b and selects the output symbol / label with the highest probability to generate a transcription.

[0056] The RNN-T model 200 determines the initial speech recognition result 120a in a streaming manner and determines the final speech recognition result 120b using the previous non-blank symbols from the initial speech recognition result 120a. Therefore, the final speech recognition result 120b is estimated to be more accurate than the initial speech recognition result 120a. That is, the final speech recognition result 120b takes into account the previous non-blank symbols. Thus, since the initial speech recognition result 120a does not take into account the previous non-blank symbols, the final speech recognition result 120b is estimated to be more accurate. Further, the rescorer 180 (Figure 1) may update the initial speech recognition result 120a based on the final speech recognition result 120b and provide the transcription 120 to the user 104 via the user interface generator 107.

[0057] In some implementations, parameter tying between the prediction network 300 and the joint network 230 is applied to further reduce the size of the RNN-T decoder, i.e., the prediction network 300 and the joint network 230. Specifically, for the vocabulary size |V| and the embedding dimension d e the shared embedding matrix 304 in the prediction network is

[0058]

Number

[0059] is. On the other hand, the last hidden layer has a dimension size d in the joint network 230 hincluding, the feed-forward projection weights from the hidden layer to the output logits are

[0060]

Number

[0061] and the vocabulary contains an extra blank token. Thus, the feed-forward layer corresponding to the last layer of the joint network 230 has a weight matrix [d h , |V|]. The embedding dimension d e of the prediction network 300 is sized to the dimension d h of the last hidden layer of the joint network 230. By tying them, the feed-forward projection weights of the joint network 230 and the shared embedding matrix 304 of the prediction network 300 can share all their weights for all non-blank symbols via a simple transpose transformation. Since the two matrices share all their values, the RNN-T decoder only needs to store the values in memory once instead of storing two individual matrices. By setting the size of the embedding dimension d e equal to the size of the hidden layer dimension d h , the RNN-T decoder reduces the number of parameters to a value equal to the product of the embedding dimension d e and the vocabulary size |V|. This weight tying corresponds to a regularization technique.

[0062] Figure 4 is a plot showing the word error rate (WER) versus the number of parameters of the RNN-T decoder. Here, plot 400 in Figure 4 shows the number of parameters (shown by the solid line) for the timed RNN-T decoder 410, the number of parameters for the untimed RNN-T decoder 420 (shown by the dotted line), and the long short-term memory (LSTM) network 430 (shown by the dashed line). Specifically, plot 400 shows the sizes of the prediction network 300 and the joint network 230 having timed outputs and embedded representations as well as the sizes of the prediction network 300 and the joint network 230 not having timed outputs and embedded representations. Plot 400 shows the embedding dimension de that varies to perform a sweep over the model size. As shown in Figure 4, the untimed RNN-T decoder 420 includes five measurements including embedding dimensions de of 64, 320, 640, 960, and 1280. Here, the timed RNN-T decoder 410 includes three measurements including embedding dimensions of 640, 960, and 1280. In the case of the untimed RNN-T decoder 420, the last hidden layer of the joint network 230 always includes a dimension d h of size 640 (dimension d h is not shown in plot 400). In the case of the timed RNN-T decoder 410, plot 4 also shows that the dimension d h of the last hidden layer of the joint network 230 (dimension d h is not shown in plot 400) is equal to the size of the embedding dimension de of the prediction network 300 so that the size and performance of the RNN-T decoder are more affected by the change in the dimension of the last hidden layer. Therefore, the results shown by plot 400 indicate that weight tying is more parameter efficient, thereby achieving better performance with fewer parameters. In addition, in a sufficiently large model using weight tying, the same word error rate as the conventional RNN-T decoder using the LSTM network 430 is reached.

[0063] FIG. 5 is a flowchart of an exemplary configuration of operations for a computer-implemented method 500 for executing a timed and reduced RNN-T model 200. At each of a plurality of time steps following an initial time step, method 500 performs operations 502-512. At operation 502, method 500 receives, as input to prediction network 300 of recurrent neural network transducer RNN-T model 200, a sequence of non-blank symbols 301, 301a-301n y ui-n , ..., y ui-1 output by final softmax layer 240. For each non-blank symbol in the sequence of non-blank symbols received as input during a corresponding time step, method 500 performs operations 504-508. At operation 504, method 500 includes generating, by prediction network 300, an embedded representation 306 of a corresponding non-blank symbol using shared embedding matrix 304. At operation 506, method 500 includes assigning, by prediction network 300, respective position vectors PV Aa~An 308, 308Aa-308An to the corresponding non-blank symbols. At operation 508, method 500 includes weighting, by prediction network 300, embedded representation 306 in proportion to the similarity between embedded representation 306 and respective position vectors 308.

[0064] At operation 510, method 500 includes generating, as output from prediction network 300, a single embedded vector 350 at a corresponding time step. Here, the single embedded vector 350 is based on a weighted average 318A-318H of weighted embedded representations 312Aa-312An. At operation 512, method 500 uses, by joint network 230 of RNN-T model 200, the single embedded vector 350 generated as output from prediction network 300 at a corresponding time step to generate a probability distribution over possible speech recognition hypotheses at the corresponding time step

[0065]

number

[0066] This includes generating:

[0067] 6 is a schematic diagram of an exemplary computing device 600 that may be used to implement the systems and methods described herein. Computing device 600 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown, their connections and relationships, and their functionality are merely exemplary and do not limit the implementation of the inventions described and / or claimed herein.

[0068] The computing device 600 includes a processor 610, a memory 620, a storage device 630, a high-speed interface / controller 640 connected to the memory 620 and the high-speed expansion port 650, and a low-speed interface / controller 660 connected to the low-speed bus 670 and the storage device 630. Each of the components 610, 620, 630, 640, 650, and 660 may be interconnected using various buses and may be mounted on a common motherboard or otherwise appropriately mounted. The processor 610 processes instructions executed within the computing device 600, including instructions stored in the memory 620 or on the storage device 630, to display graphical information about a graphical user interface (GUI) on an external input / output device such as a display 680 coupled to the high-speed interface 640. In other implementations, multiple processors and / or multiple buses may be used, along with multiple memories and multiple types of memory, as appropriate. Also, multiple computing devices 600 may be connected to each device that provides a portion of the required operations (e.g., as a server bank, a group of blade servers, or a multiprocessor system).

[0069] The memory 620 stores information non-transiently within the computing device 600. The memory 620 may be a computer-readable medium, a volatile memory unit, or a non-volatile memory unit. The non-transient memory 620 may be a physical device used to temporarily or permanently store programs (e.g., sequences of instructions) or data (e.g., program state information) for use by the computing device 600. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., commonly used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), as well as disk or tape.

[0070] The storage device 630 can provide mass storage for the computing device 600. In some implementations, the storage device 630 is a computer-readable medium. In various different implementations, the storage device 630 may be a floppy disk device, a hard disk device, an optical disk device, or an array of devices including a tape device, a flash memory, or other similar solid-state memory device, or a device in a storage area network or other configuration. In additional implementations, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer- or machine-readable medium, such as the memory 620, the storage device 630, or memory on the processor 610.

[0071] The high-speed controller 640 manages bandwidth-intensive operations for the computing device 600, and the low-speed controller 660 manages lower-bandwidth intensive operations. Such duty allocation is merely exemplary. In some implementations, the high-speed controller 640 is coupled to the memory 620, coupled to the display 680 (e.g., via a graphics processor or accelerator), and coupled to a high-speed expansion port that may accept various expansion cards (not shown). In some implementations, the low-speed controller 660 is coupled to the storage device 630 and to a low-speed expansion port 690. The low-speed expansion port 690 may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet) and may be coupled to one or more input / output devices, such as a keyboard, a pointing device, a scanner, or a networking device such as a switch or router, e.g., via a network adapter.

[0072] Computing device 600 may be implemented in several different forms, as shown, for example, as a standard server 600a, or multiple such servers 600a, or as a laptop computer 600b, or as part of a rack server system 600c.

[0073] Various implementations of the systems and techniques described herein can be realized in digital electronics and / or optical circuits, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementations in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which processor may be special purpose or general purpose and may be coupled to receive and transmit data and instructions between a memory system, at least one input device, and at least one output device.

[0074] A software application (i.e., a software resource) may refer to computer software that causes a computing device to perform tasks. In some examples, a software application may be referred to as an “application,” an “app,” or a “program.” Exemplary applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, document creation applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.

[0075] These computer programs (also known as programs, software, software applications, or code) contain machine instructions for a programmable processor and may be implemented in high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0076] The processes and logic flows described in this specification can be executed by one or more programmable processors, which are also referred to as data processing hardware, and which function by executing one or more computer programs to process input data and generate output. The processes and logic flows can also be executed by dedicated logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). Examples of processors suitable for executing computer programs include both general purpose and special purpose microprocessors, and any one or more processors of any kind of digital computer. In general, a processor receives instructions and data from a read only memory or a random access memory or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. In general, a computer also includes one or more storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, operatively coupled to receive data from, transfer data to, or both of these devices. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and memory devices, including by way of example semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices, magnetic disks, such as internal hard disks or removable disks, magneto-optical disks, and CD ROM and DVD-ROM disks. The processor and the memory can be assisted by, or incorporated in, dedicated logic circuitry.

[0077] To enable interaction with a user, one or more aspects of the present disclosure can be implemented on a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube), LCD (liquid crystal display), monitor, or touch screen, and optionally a keyboard and a pointing device by which the user can provide input to the computer, such as a mouse or trackball. Other types of devices can also be used to enable interaction with the user. For example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and the input from the user can be received in any form including acoustic, speech, or tactile input. Additionally, the computer can communicate documents to and from the devices used by the user, and can interact with the user, for example, by sending a web page to a web browser on the user's client device in response to a request received from the web browser.

[0078] Some implementations have been described. Nevertheless, it will be understood that various modifications may be made without departing from the spirit and scope of the present disclosure. Accordingly, other implementations are within the scope of the following claims.

Explanation of Reference Numerals

[0079] 10 User device 12 Data processing hardware 14 Memory hardware 16, 16a Audio capture device 16, 16b Voice output device 19 Response 40 Network 50 Digital assistant application 60 Remote computing device 100 Speech environment 104 User 106 Speech 107 User interface generator 108 Audio subsystem 110 Acoustic frame 118 ASR system 120 Transcription 120a, 120b Speech recognition results 170 User interface generator 180 Rescorer 200 Recurrent neural network transducer RNN-T model 210 Encoder network 230 Joint network 240 Final softmax layer 300 Prediction network 301 Non-blank symbol 302 Multi-head attention mechanism 302A~302H Heads 302B~302H Other heads 304 Shared embedding matrix 306, 306a~306n Embedding representations 308, 308Aa~308An Position vectors 310 Weight layer 312, 312Aa~312An Weighted embedding representations 316 Weighted average module 318A~318H Outputs 322 Head average module 324 Output 326 Projection layer 328 Projected output 350 Embedding vector 400 Plot 410 Timed RNN-T decoder 420 Untimed RNN-T decoder 500 Method 600 Computing device 600a Server 600b Laptop computer 600c Rack Server System 610 Processor 620 Memory 630 Storage Device 640 High-Speed Interface / Controller 650 High-Speed Expansion Port 660 Low-Speed Interface / Controller 670 Low-Speed Bus 680 Display 690 Low-Speed Expansion Port

Claims

1. A system comprising: data processing hardware; and memory hardware that communicates with the data processing hardware and stores instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations, the operations including executing a recurrent neural network transducer (RNN-T) model (200), the RNN-T model comprising: receiving, as input, a sequence of acoustic frames; an audio encoder configured to generate, for each of a plurality of time steps, a higher-order feature representation for a corresponding acoustic frame in the sequence of acoustic frames; for each of the plurality of time steps subsequent to an initial time step: receiving, as input, a sequence of non-blank symbols (301) output by a final softmax layer (240); for each non-blank symbol (301) in the sequence of non-blank symbols (301) received as input at the corresponding time step: generating an embedding representation (306) of the corresponding non-blank symbol (301) using a shared embedding matrix (304); assigning a respective position vector (308) to the corresponding non-blank symbol (301); weighting the embedding representation (306) in proportion to a similarity between the embedding representation (306) and the respective position vector (308); generating, as output, a single embedding vector (350) at the corresponding time step, the single embedding vector (350) being generated based on a weighted average (318) of the weighted embedding representations (312); for each of the plurality of time steps subsequent to the initial time step: receiving, as input, the single embedding vector (350) generated as output from the prediction network (300) at the corresponding time step; a joint network (230) configured to generate a probability distribution over possible speech recognition hypotheses at the corresponding time step; A system. **Claim 2**: The system according to claim 1, wherein the joint network (230) is further configured to receive, as input, the high-level feature representation generated by the audio encoder at the corresponding time step in each of the plurality of time steps. **Claim 3** Weighting the embedded representation (306) in proportion to the similarity between the embedded representation (306) and the respective position vectors (308) includes weighting the embedded representation (306) in proportion to the cosine similarity between the embedded representation (306) and the respective position vectors (308). The system according to claim 1 or 2. **Claim 4** The sequence of non-blank symbols output by the final softmax layer (240) includes word pieces. The system according to any one of claims 1 to 3. **Claim 5** The sequence of non-blank symbols (301) output by the final softmax layer (240) includes graphemes. The system according to any one of claims 1 to 4. **Claim 6** Each of the embedded representations includes the same dimensional size as each of the position vectors. The system according to any one of claims 1 to 5. **Claim 7** The sequence of non-blank symbols received as input is limited to N previous non-blank symbols output by the final softmax layer. The system according to any one of claims 1 to 6. **Claim 8** N is equal to 2. The system according to claim 7. **Claim 9** N is equal to 5. The system according to claim 7. **Claim 10** The prediction network includes a multi-head attention mechanism that shares the shared embedding matrix across each head of the multi-head attention mechanism (302). The system according to any one of claims 1 to 9. **Claim 11** The prediction network (300) at each of the plurality of time steps following the initial time step, at each head of the multi-head attention mechanism (302), for each non-blank symbol (301) in the sequence of non-blank symbols (301) received as input at the corresponding time step. Using the shared embedding matrix (304), generate the embedding representation (306) of the corresponding non-blank symbol (301) that is the same as the embedding representations (306) generated in each other head (302) of the multi-head attention mechanism (302). Assign to the corresponding non-blank symbol (301) a respective position vector (308) that is different from the respective position vectors (308) assigned to the corresponding non-blank symbol (301) in each other head (302) of the multi-head attention mechanism (302). Weight the embedding representation (306) in proportion to the similarity between the embedding representation (306) and the respective position vectors (308). Generate, as an output from the corresponding head (302) of the multi-head attention mechanism (302), a respective weighted average (318) of the weighted embedding representations (312) of the sequence of non-blank symbols (301). The system according to claim 10, configured to generate, as an output, a single embedding vector (350) at the corresponding time step by averaging the respective weighted averages (318) output from the corresponding head (302) of the multi-head attention mechanism (302). **Claim 12** The system according to claim 10 or 11, wherein the multi-head attention mechanism (302) comprises four heads (302). **Claim 13** The system according to any one of claims 1 to 12, wherein the prediction network ties the dimension of the shared embedding matrix (304) to the dimension of the output layer of the joint network (230). **Claim 14** When executed on data processing hardware (12), cause the data processing hardware (12) to in each of a plurality of time steps following an initial time step receive, as an input to a prediction network (300) of a recurrent neural network transducer (RNN-T) model (200), a sequence of non-blank symbols (301) output by a final softmax layer (240). For each non-blank symbol (301) in the sequence of non-blank symbols (301) received as input at the corresponding time step, generating, by the prediction network (300), an embedded representation (306) of the corresponding non-blank symbol (301) using a shared embedding matrix (304); assigning, by the prediction network (300), a respective position vector (308) to the corresponding non-blank symbol (301); weighting, by the prediction network (300), the embedded representation (306) in proportion to the similarity between the embedded representation (306) and the respective position vector (308); generating, as an output from the prediction network (300), a single embedded vector (350) at the corresponding time step, wherein the single embedded vector (350) is based on a weighted average (318) of the weighted embedded representations (312); causing a computer-implemented method (500) to perform an operation comprising: generating, by a joint network (230) of the RNN-T model (200), a probability distribution over possible speech recognition hypotheses at the corresponding time step using the single embedded vector (350) generated as an output from the prediction network (300) at the corresponding time step. **Claim 15** The operation further comprises: receiving, as input to an audio encoder (210) of the RNN-T model (200), a sequence of acoustic frames (110); generating, by the audio encoder (210), at each of the plurality of time steps, a high-order feature representation for a corresponding acoustic frame (110) in the sequence of acoustic frames; receiving, as input to the joint network (230), the high-order feature representation generated by the audio encoder (210) at the corresponding time step. The computer-implemented method (500) according to claim 14. **Claim 16** The step of weighting the embedded representation (306) in proportion to the similarity between the embedded representation (306) and the respective position vectors (308) includes the step of weighting the embedded representation (306) in proportion to the cosine similarity between the embedded representation (306) and the respective position vectors (308), the computer-implemented method (500) according to claim 14 or 15.

17. The sequence of non-blank symbols (301) output by the final softmax layer (240) includes word pieces, the computer-implemented method (500) according to any one of claims 14 to 16.

18. The sequence of non-blank symbols (301) output by the final softmax layer (240) includes writing elements, the computer-implemented method (500) according to any one of claims 14 to 17.

19. Each of the embedded representations (306) includes the same dimensional size as each of the position vectors (308), the computer-implemented method (500) according to any one of claims 14 to 18.

20. The sequence of non-blank symbols (301) received as input is limited to N previous non-blank symbols (301) output by the final softmax layer (240), the computer-implemented method (500) according to any one of claims 14 to 19.

21. N is equal to 2, the computer-implemented method (500) according to claim 20.

22. N is equal to 5, the computer-implemented method (500) according to claim 20.

23. The prediction network (300) comprises a multi-head attention mechanism (302), and the multi-head attention mechanism (302) shares the shared embedding matrix (304) across each head (302) of the multi-head attention mechanism (302), the computer-implemented method (500) according to any one of claims 14 to 22.

24. The operation is, in each of the plurality of time steps following the initial time step, in each head (302) of the multi-head attention mechanism (302), For each non - blank symbol (301) in the sequence of the non - blank symbols (301) received as input in the corresponding time step, generating, by the prediction network (300), using the shared embedding matrix (304), an embedding representation (306) of the corresponding non - blank symbol (301) that is the same as the embedding representations (306) generated in each other head (302) of the multi - head attention mechanism (302); assigning, by the prediction network (300), to the corresponding non - blank symbol (301), respective position vectors (308) that are different from the respective position vectors (308) assigned to the corresponding non - blank symbol (301) in each other head (302) of the multi - head attention mechanism (302); weighting, by the prediction network (300), the embedding representation (306) in proportion to the similarity between the embedding representation (306) and the respective position vectors (308); generating, by the prediction network (300), as an output from the corresponding head (302) of the multi - head attention mechanism (302), a weighted average (318) of each of the weighted embedding representations (312) of the sequence of non - blank symbols (301); generating, as an output from the prediction network (300), a single embedding vector (350) in the corresponding time step by averaging the respective weighted averages (318) output from the corresponding head (302) of the multi - head attention mechanism (302), the computer - implemented method (500) according to claim 23, further comprising. **Claim 25** The computer - implemented method (500) according to claim 23 or 24, wherein the multi - head attention mechanism (302) comprises four heads (302). **Claim 26** The computer - implemented method (500) according to any one of claims 14 to 25, wherein the prediction network (300) ties the dimension of the shared embedding matrix (304) to the dimension of the output layer of the joint network (230).

Citation Information

Patent Citations

  • Neural Machine Translation System

    JP2019537096A

  • Speech recognition device, speech recognition method and speech recognition program

    JP2021039216A

  • Minimum word error rate training for attention-based sequence-to-sequence models

    US20200043483A1

  • Method and device for speech recognition

    US20200234713A1

  • Language processing using a neural network

    US20210049327A1