4-bit Conformer with Accurate Quantization Training for Speech Recognition

Quantization-aware training with native integer arithmetic addresses resource constraints in mobile devices by enabling efficient, low-latency ASR models with integer bitwidths, ensuring accurate real-time speech recognition.

JP7698154B2Active Publication Date: 2025-06-24GOOGLE LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024556057
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-03-21
Filing Date
2023-03-20
Publication Date
2025-06-24
Estimated Expiration
2043-03-20

AI Technical Summary

Technical Problem

Mobile devices face limitations in resources, which restrict the size of automatic speech recognition (ASR) models, leading to challenges in achieving low latency and high accuracy in real-time speech recognition due to the need for smaller model sizes.

Method used

Implementing quantization-aware training with native integer arithmetic to train ASR models, allowing for efficient conversion to integer fixed bitwidths, such as 4-bit or 8-bit integers, while maintaining accuracy and reducing computational requirements.

Benefits of technology

The method enables ASR models to operate efficiently on mobile devices with minimal latency and accuracy loss, supporting real-time speech recognition without the need for additional conversion steps, thus enhancing user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007698154000010
    Figure 0007698154000010
  • Figure 0007698154000011
    Figure 0007698154000011
  • Figure 0007698154000012
    Figure 0007698154000012
Patent Text Reader

Abstract

A method (500) for training a model includes obtaining a plurality of training samples (152), each of which comprises a respective speech utterance (152) and a respective text utterance (154) representing a transcription of the respective speech utterance. An automatic speech recognition (ASR) model (200) is trained on the plurality of training samples using quantization-aware training with native integer arithmetic. The trained automatic speech recognition ASR model is quantized to an integer target fixed bit-width (162). The quantized trained automatic speech recognition ASR model comprises a plurality of weights (202). Each weight of the plurality of weights comprises an integer having a target fixed bit-width. The quantized trained automatic speech recognition ASR model is provided to a user device (10).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to accurate quantization training for speech recognition.

Background Art

[0002] Modern automatic speech recognition (ASR) systems focus on providing not only high quality (e.g., low word error rate (WER)), but also low latency (e.g., short delay from when a user speaks until the transcription appears). Furthermore, today, when using an automatic speech recognition ASR system, there is a need to decode (decode) utterances (utterances) in a streaming format in which the automatic speech recognition ASR system responds in real time or even faster than real time. To explain, when an automatic speech recognition ASR system is deployed in a mobile phone that experiences direct user interactivity, the application of the mobile phone using the automatic speech recognition ASR system may require speech recognition to be streamed so that words are displayed on the screen as soon as they are spoken. Here, the user of the mobile phone may also have a low tolerance for latency. This low tolerance causes speech recognition to attempt to run on a mobile device in a way that minimizes the impact of latency and inaccuracy, which can negatively affect the user experience.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, the resources of mobile phones are often limited, which limits the size of the automatic speech recognition ASR model.

Means for Solving the Problem

[0005] One aspect of the present disclosure provides a method for training an automatic speech recognition (ASR) model. The computer-implemented method causes the data processing hardware to perform operations when executed on the data processing hardware. The operations include obtaining a plurality of training samples. Each of the plurality of training samples includes each speech utterance and each text utterance representing a transcription of each speech utterance. The method includes training an automatic speech recognition ASR model for the plurality of training samples using quantization-aware training with native integer arithmetic. The method also includes quantizing the trained automatic speech recognition ASR model to an integer target fixed bitwidth. The quantized trained automatic speech recognition ASR model includes a plurality of weights. Each of the plurality of weights includes an integer having the target fixed bitwidth. The method includes providing the quantized trained automatic speech recognition ASR model to a user device.

[0006] Embodiments of the present disclosure may include one or more of the following optional features. In some embodiments, the target fixed bitwidth is "4". In some examples, the automatic speech recognition ASR model further includes a plurality of activation functions. Each of the plurality of activation functions may include an integer having the target fixed bitwidth. In other examples, the automatic speech recognition ASR model further includes a plurality of activation functions, and each of the plurality of activation functions includes an integer having a fixed bitwidth greater than the target fixed bitwidth. In still other examples, the automatic speech recognition ASR model further includes a plurality of activation functions, and each of the plurality of activation functions includes a floating point value.

[0007] Optionally, the step of quantizing a trained automatic speech recognition ASR model comprises determining a scale factor based on an estimated maximum value of an axis (axis) to be quantized and a target fixed bit width. In some embodiments, the automatic speech recognition ASR model comprises one or more multi-head attention layers. In some of these embodiments, one or more multi-head attention layers comprise one or more conformer layers or one or more transformer layers. The automatic speech recognition ASR model may include a plurality of encoders and a plurality of decoders, and the step of quantizing the automatic speech recognition ASR model may include a step of quantizing the plurality of encoders and a step of not quantizing the plurality of decoders. In some examples, the automatic speech recognition ASR model comprises an audio encoder, and the audio encoder comprises a cascaded (connected) encoder having a first causal (coastal) encoder and a second non-causal (non-coastal) encoder.

[0008] Other aspects of the present disclosure provide a system for training an automatic speech recognition (ASR) model. The system includes data processing hardware and memory hardware communicatively coupled to the data processing hardware. The memory hardware stores instructions that, when executed by the data processing hardware, cause the data processing hardware to perform operations. The operations include obtaining a plurality of training samples. Each of the plurality of training samples includes each respective speech utterance and each respective text utterance representing a transcription of each respective speech utterance. The method includes training an automatic speech recognition (ASR) model with respect to the plurality of training samples using quantization-aware training with native integer arithmetic. The method also includes quantizing the trained automatic speech recognition (ASR) model to a target fixed bitwidth of integers. The quantized trained automatic speech recognition (ASR) model includes a plurality of weights. Each of the plurality of weights includes an integer having the target fixed bitwidth. The method includes providing the quantized trained automatic speech recognition (ASR) model to a user device.

[0009] This aspect may include one or more of the following optional features. In some embodiments, the target fixed bitwidth is "4". In some examples, the automatic speech recognition (ASR) model further includes a plurality of activation functions, and each of the plurality of activation functions may include an integer having the target fixed bitwidth. In other examples, the automatic speech recognition (ASR) model further includes a plurality of activation functions, and each of the plurality of activation functions includes an integer having a fixed bitwidth greater than the target fixed bitwidth. In yet other examples, the automatic speech recognition (ASR) model further includes a plurality of activation functions, and each of the plurality of activation functions includes a floating-point value.

[0010] Optionally, the step of quantizing the trained automatic speech recognition (ASR) model comprises determining a scale factor based on an estimated maximum value of the axis to be quantized and a target fixed bit width. In some embodiments, the automatic speech recognition (ASR) model comprises one or more multi-head attention layers. In some of these embodiments, the one or more multi-head attention layers comprise one or more conformer layers or one or more transformer layers. The automatic speech recognition (ASR) model may include a plurality of encoders and a plurality of decoders, and the step of quantizing the automatic speech recognition (ASR) model may include a step of quantizing the plurality of encoders and a step of not quantizing the plurality of decoders. In some examples, the automatic speech recognition (ASR) model comprises an audio encoder, and the audio encoder comprises a cascaded encoder having a first causal encoder and a second non-causal encoder.

[0011] Details of one or more embodiments of the present disclosure are set forth in the accompanying drawings and the following description. Other aspects, features, and advantages will be apparent from the description and drawings and from the claims.

Brief Description of the Drawings

[0012]

Fig. 1A

Fig. 1B

Fig. 2

Fig. 3

Fig. 4

Fig. 5

Fig. 6

DETAILED DESCRIPTION OF THE INVENTION

[0013] Like reference symbols in the various drawings refer to like elements. Due to the rapid growth of voice search and speech interactive features, automatic speech recognition (ASR) has become an essential component for user interactive services and devices (e.g., search engines and voice-enabled search on smartphones). Many of the latest automatic speech recognition ASR applications are developed based on end-to-end models, which have shown a significantly smaller model size and a substantial improvement in recognition performance compared to conventional hybrid systems. Improving latency and model size without sacrificing recognition quality has been actively pursued to benefit live automatic speech recognition ASR applications on both server-side and on-device models.

[0014] Quantization is a technique that reduces the computational and memory costs of an automatic speech recognition (ASR) model by representing weights and / or activation functions with low-precision data types (e.g., 8-bit integers) instead of conventional 32-bit floating-point values. Among the latest model quantization methods, "post-training quantization" (PTQ) using 8-bit integers (int8) is a popular and easy-to-use technique that has been successfully applied in many applications. However, one of the drawbacks of such a technique is the risk of performance degradation due to a decrease in accuracy. Another limitation of post-training quantization (PTQ) is that it cannot control the quantization of the model. For example, post-training quantization (PTQ) may not support 4-bit integer (int4) quantization or customized quantization of a selected set of layers.

[0015] Embodiments of this specification include a model trainer that trains an automatic speech recognition (ASR) model by using native "quantization-aware training" (QAT) with native integer arithmetic. In contrast to some methods that use "fake" quantization-aware training (QAT) (i.e., methods that convert floating-point numbers to integers by using conversion after using floating-point arithmetic), native quantization-aware training (QAT) uses native integer arithmetic to perform quantization operations (e.g., matrix multiplication) and generates a model that has no difference in accuracy during training and inference. That is, "fake quantization" may have a numerical difference between the training mode (i.e., using floating-point arithmetic) and the inference mode (i.e., using integer arithmetic) when the floating-point arithmetic during training does not fit within the bits of the mantissa.

[0016] Embodiments of this specification include a model trainer that trains an automatic speech recognition (ASR) model using native quantization-aware training (QAT). This approach ensures that "what is trained is useful." That is, in native integer arithmetic, there is no numerical difference between the forward propagation (forward pass) of training and inference. Therefore, the trained model may be executed in multiple applications, such as both in the cloud (e.g., on a tensor processing unit (TPU)) or in a model application where the performance is the same. The model trainer minimizes the number of arithmetic operations used for quantization, thus shortening the training time compared to conventional techniques. The automatic speech recognition (ASR) model may include one or more multi-head attention layers, such as one or more conformer layers and / or one or more transformer layers.

[0017] Figure 1A is an example of a system 100 operating in a voice environment 101. In the voice environment 101, the way a user 104 interacts with a computing device, such as user device 10, may be by voice input. User device 10 (commonly also referred to as device 10) is configured to capture sound (e.g., streaming audio data) from one or more users 104 within the voice environment 100. Here, the streaming audio data may refer to an utterance 106 spoken by user 104 that functions as an audible query, a command to device 10, or an audible communication already captured by device 10. The speech-enabled system of device 10 may handle the query or command by answering the query and / or by having one or more downstream applications perform the command.

[0018] The user device 10 can correspond to any computing device that is associated with the user 104 and is capable of receiving audio data. Some examples of the user device 10 include, but are not limited to, mobile devices (e.g., mobile phones, tablets, laptops, etc.), computers, wearable devices (e.g., smartwatches), smart appliances, Internet of Things (IoT) devices, in-vehicle infotainment systems, smart displays, smart speakers, and the like. The user device 10 includes data processing hardware 12 and memory hardware 14 that communicates with the data processing hardware 12, and stores instructions that cause the data processing hardware 12 to perform one or more operations when executed by the data processing hardware 12. Further, the user device 10 includes an audio system 16 that includes audio capture devices (e.g., microphones) 16, 16a for capturing an utterance 106 spoken within the acoustic environment 100 and converting it into an electrical signal, and an utterance output device (e.g., a speaker) 16, 16b for communicating an audible audio signal (e.g., as output audio data from the device 10). In the illustrated example, the user device 10 implements a single audio capture device 16a, but the user device 10 can implement an array of audio capture devices 16a without departing from the scope of the present disclosure, such that one or more capture devices 16a within the array can communicate with the audio system 16 without being physically present in the user device 10.

[0019] In the acoustic environment 100, the automatic speech recognition (ASR) system 118 has a model 200 (such as a recurrent neural network - transducer (RNN - T) model or other conformer transducer model / multi - path model, etc.) existing on the user device 10 of the user 104 and / or on a remote computing device 60 that communicates with the user device 10 via the network 40 (for example, one or more remote servers of a distributed system running in a cloud computing environment). The remote computing device includes data processing hardware 62 and memory hardware 64. The user device 10 and / or the remote computing device 60 also includes an audio subsystem 108, which not only receives the utterance 106 spoken by the user 104 and captured by the audio capture device 16a, but is also configured to convert the utterance 106 into a corresponding digital format associated with the input acoustic frame 110 that can be processed by the automatic speech recognition ASR system 118. In the example shown, the user is speaking each utterance 106, and the audio subsystem 108 converts (converts) the utterance 106 into the corresponding audio data (for example, acoustic frame) 110 and inputs it to the automatic speech recognition ASR system 118. Thereafter, the model 200 receives the audio frame 110 (i.e., audio data) corresponding to the utterance 106 as input and generates / predicts the corresponding transcription 120 (for example, speech recognition result / hypothesis) of the utterance 106 as output.

[0020] The user device 10 and / or the remote computing device 60 also execute a user interface generator 107 configured to present the representation of the transcription 120 of the utterance 106 to the user 104 of the user device 10. As will be described in more detail below, the user interface generator 107 may display the speech recognition result 120 in a streaming format. In some configurations, the transcription 120 output from the automatic speech recognition ASR system 118 is processed by a natural language understanding (NLU) module that executes on, for example, the user device 10 or the remote computing device 60 to execute the user command / query specified by the utterance 106. Additionally or alternatively, a text-to-speech system (not shown) (e.g., executed in any combination of the user device 10 or the remote computing device 60) is enabled to convert the transcription (transcript, transcription) into a synthetic voice for audible output by the user device 10 and / or other devices.

[0021] In the illustrated example, user 104 interacts with a program or application 50 (e.g., digital assistant application 50) of user device 10 that uses automatic speech recognition ASR system 118. For example, FIG. 1A depicts user 104 communicating with digital assistant application 50 and digital assistant application 50 displaying digital assistant interface 18 on the screen of user device 10 to show a conversation between user 104 and digital assistant application 50. In this example, user 104 asks digital assistant application 50, "What time is the concert tonight?" This question from user 104 is a spoken utterance 106 that is captured by audio capture device 16a and processed by audio system 16 of user device 10. In this example, audio system 16 receives the spoken utterance 106 and converts it to acoustic frame 110 for input to automatic speech recognition ASR system 118.

[0022] Referring now to FIG. 1B, remote computing device 60 trains model 200 of FIG. 1A by executing model trainer 150. Model trainer 150 obtains a plurality of training samples 152, 152a - 152n (e.g., from memory hardware 64). Each training sample 152 comprises a spoken training utterance 154 (i.e., a sequence of input audio features) and a corresponding text utterance 156 representing a transcription 156 of utterance 154. Model trainer 150 trains model 200 with respect to the plurality of training samples 152 using quantization-aware training (QAT) by native integer arithmetic. As discussed in further detail below, model trainer 150 uses quantization-aware training by determining a scale factor 160 during training. Model trainer 150 quantizes model 200 to an integer fixed bitwidth 162 during or after training. Integer fixed bitwidth 162 represents the number of bits allocated to each native integer arithmetic operation. For example, if the integer fixed bitwidth 162 is "8", model 200 is quantized to 8-bit integers (i.e., int8). In other examples, if the integer fixed bitwidth 162 is "4", model 200 is quantized to 4-bit integers (i.e., int4). Other examples such as 6-bit integers are possible. Integer fixed bitwidth 162 may be settable by a user and may depend on the use case of the model.

[0023] Conventional quantization-aware techniques rely on "fake" quantization-aware training QAT. For example, many common systems use tf.quantization.fake_quant_ *The model is quantized by using operations. These operations are used during server-side inference (server-side inference), but in the case of on-device models (such as user devices like smartphones), it is necessary to convert (convert) the fake quantization operations to integer operations by using conversion operations (such as TFlite). Therefore, these conventional techniques require this additional conversion step to convert the fake quantization operations to actual integer operations (integer operations), so existing application programming interfaces (APIs) only support the estimation of the minimum and maximum values for each channel in the last dimension. In some use cases where the channel dimension is not the last dimension, this is not ideal. In these cases, the conventional techniques permute the dimensions of the input tensor to make the channel dimension the last dimension, and then it is necessary to "fake" quantize the tensor by using the application programming interface API. Finally, it is necessary to permute the dimensions back to the original order of the input tensor. These additional permutation operations increase the training time. In contrast, the model trainer 150 uses native integer operations (such as tf operations). As a result, the model 200 can be used for training and inference in both mobile and tensor processing unit TPU applications. Furthermore, by using hardware-supported integer operations (such as matrix multiplication), the training can be further accelerated.

[0024] During quantization, the model trainer 150 reduces the size of the model 200 by adjusting the size of one or more weights 202 and / or activation functions 204 of the model 200. Conventionally, both the weights 202 and the activation functions 204 of an automatic speech recognition (ASR) model occupy 32-bit space and are often represented by float32 values that require complex calculations for processing. In "post-training quantization" (PTQ), these float32 values can be "clipped" or rounded to reduce accuracy and memory requirements. In contrast, the model trainer 150 uses native integer operations to represent the weights 202 of the model 200 by using integers of a size determined by a fixed bitwidth 162. In some examples, the model trainer 150 quantizes the weights 202 and the activation functions 204 for each same fixed bitwidth 162 (e.g., 4 bits or 8 bits). In other examples, the model trainer 150 quantizes the weights 202 and the activation functions by using different fixed bitwidths 162. For example, the model trainer 150 quantizes the activation function 204 using a fixed bitwidth 162 that is larger than the fixed bitwidth 162 of the weights 202 (e.g., 4 bits for the weights 202 and 8 bits for the activation function 204). In still other examples, the model trainer 150 quantizes the weights 202 while not quantizing the activation function 204 (e.g., the activation function 204 is represented by a floating-point value such as float32).

[0025] The model trainer 150 may quantize only a part of the model 200. In some examples, when the model 200 includes multiple encoders and multiple decoders, the model trainer 150 quantizes only the encoders while not quantizing the decoders because the memory requirements of the decoders are minimal in some scenarios. After training and quantization, the model trainer 150 may provide the quantized and trained model 200 to the user device 10.

[0026] Referring now to FIG. 2, exemplary models 200, 200a include a recurrent neural network - transducer (RNN - T) model architecture that adheres to latency constraints associated with an interactive application. The use of the RNN - T model architecture is exemplary, and model 200 may include other architectures such as, among others, transformer - transducer and conformer - transducer model architectures. The RNN - T model 200a provides a small computational footprint and utilizes fewer memory requirements than conventional automatic speech recognition (ASR) architectures, so the RNN - T model architecture is suitable for performing speech recognition entirely on user device 10 (e.g., communication with a remote server is not required). In this example, the RNN - T model 200a includes an encoder network 210, a prediction network 300, and a joint network 230. The encoder network 210, which is roughly similar to an acoustic model (AM) in a conventional automatic speech recognition (ASR) system, includes a stack of self - attention (self - attention) layers (e.g., conformer or transformer layers) or a recurrent network of stacked long short - term memory (LSTM) layers. For example, the encoder reads out a sequence x=(x1,x2,···,x T ) where

[0027]

Number

[0028] and generates a high - order feature representation at each output step. This high - order feature representation is

[0029]

Number

[0030] represented as. Similarly, the prediction network 300 may also be an LSTM network. This, like a language model (LM), takes the sequence of non-blank symbols y0, ···, y ui-1 and processes it into a dense representation

[0031]

Number

[0032] . Finally, by using the RNN-T model architecture, the representations generated by the encoder and prediction / decoder networks 210, 300 are combined by the joint network 230. The prediction network 300 can be replaced by an embedding lookup table to improve latency by outputting a sparse embedding lookup instead of processing a dense representation. Next, the joint network outputs the distribution of the next output symbol

[0033]

Number

[0034] Predict it. In other words, at each output step (e.g., time step), the joint network 230 generates a probability distribution of possible speech recognition hypotheses. Here, "possible speech recognition hypotheses" corresponds to a set of output labels each representing a symbol / character of the specified natural language. For example, when the natural language is English, the set of output labels may include twenty-seven (27) symbols. For example, there is one label for each of the 26 letters of the English alphabet and one label representing a space (blank). Thus, the joint network 230 may output a set of values indicating the likelihood of occurrence of each of the predetermined output label sets. Since this set of values can be a vector, it can represent a probability distribution over the set of output labels. In some cases, the output labels are graphemes (individual characters, potential punctuation marks, and other symbols, etc.), but the set of output labels is not limited thereto. For example, the set of output labels can include parts of words and / or whole words in addition to or instead of graphemes. The output distribution of the joint network 230 may include posterior probability values for each of the different output labels. Thus, if there are 100 different output labels representing different graphemes or other symbols, the output y i of the joint network 230 can have 100 different probability values, one for each output label. Next, by using the probability distribution, candidate correct writing elements (e.g., graphemes, word fragments, and / or words) can be selected in a beam search process (e.g., by the softmax layer 240), and scores can be assigned to determine the transcription 120.

[0035] The softmax layer 240 may use any technique to select the output label / symbol having the highest probability within the distribution as the next output symbol predicted by the RNN-T model 200 at the corresponding output step. In this way, the RNN-T model 200 does not make the assumption of conditional independence. Rather, the prediction of each symbol is conditioned not only on the acoustics but also on the sequence of labels output so far. The RNN-T model 200 can be used in a streaming manner by assuming that the output symbol is independent of future acoustic frames 110.

[0036] In some examples, the encoder network (i.e., audio encoder) 210 of the RNN-T model 200 comprises a stack of one or more conformable blocks / layers and / or one or more transformer blocks / layers, such as multi-head attention layers or self-attention layers / blocks. Optionally, the encoder 210 (i.e., audio encoder) comprises a first-pass causal (causal) encoder and a second-pass non-causal (non-causal) encoder for a multi-pass architecture. This multi-pass model integrates streaming and non-streaming automatic speech recognition ASR. The causal encoder generates partial results with minimal latency by using only the left context. The non-causal encoder is enabled to provide more accurate hypotheses by using both left and right contexts. In this example, each conformable block comprises a series of layers of multi-head self-attention, depthwise convolution, and feed-forward. The prediction network 300 may have two 2048-dimensional LSTM layers, each followed by a 640-dimensional projection layer. Alternatively, the prediction network 300 may include a stack of transformers or conformable blocks, or an embedding lookup table, instead of the LSTM layers. Finally, the joint network 230 may also have 640 hidden units. The softmax layer 240 may be configured with an integrated set of word pieces or graphemes generated by using all unique word pieces or graphemes within a plurality of training data sets.

[0037] FIG. 3 shows a sequence y of non-blank symbols, restricted to the N previous non-blank symbols 301a - 301n output by the final softmax layer 240 ui-n , ···, y ui-1FIG. 300 shows an exemplary prediction network 300 of an RNN-T model 200a that receives [[ID=]] as input. In some examples, N is equal to 2. In other examples, N is equal to 5, but the disclosure is non-limiting and N may be equal to any integer. The sequence of non-blank symbols 301a~301n represents the initial speech recognition result 120a (FIG. 1). In some embodiments, the prediction network 300 includes a multi-head attention mechanism 302 that shares an embedding matrix 304 across each head 302A~302H of the multi-head attention mechanism. In one example, the multi-head attention mechanism 302 includes four heads. However, any number of heads can be used in the multi-head attention mechanism 302. In particular, the multi-head attention mechanism significantly improves performance while minimizing the increase in model size. As will be described in detail below, each head 302A~302H includes a row of unique position vectors 308, and instead of increasing the model size by concatenating the outputs 318A~318H from all heads, the outputs 318A~318H are averaged by a head average module 322.

[0038] Referring to the first head 302A of the multi-head attention mechanism 302, the first head 302A uses the shared embedding matrix 304 to obtain, as input at the corresponding time step from a plurality of time steps, the sequence y of non-blank symbols that have already been received ui-n , ···, y ui-1 For each non-blank symbol 301 within, the corresponding embeddings 306, 306a~306n (e.g.,

[0039]

Number

[0040] generates it. In particular, since the shared embedding matrix 304 is shared among all the heads of the multi-head attention mechanism 302, all the other heads 302B to 302H all generate the same corresponding embedding 306 for each non-blank symbol. The first head 302A also corresponds to each non-blank symbol in the sequence y of non-blank symbols ui-n , ···, y ui-1 and assigns each respective position vector PV Aa~An 308, 308Aa to 308An (for example

[0041]

Number

[0042] ). Each position vector PV308 assigned to each non-blank symbol indicates the position within the history of the sequence of non-blank symbols (for example, the N previous non-blank symbols output by the final softmax layer 240). For example, the first position vector PV Aa is assigned to the latest position in the history, while the last position vector PV An is assigned to the last position in the history of the N previous non-blank symbols output by the final softmax layer 240. In particular, each of the embeddings 306 may include the same dimension (i.e., dimension size) as each of the position vectors PV308.

[0043] For each non-blank symbol 301 in the sequence y of non-blank symbols 301a to 301n ui-n , ···, y ui-1 , the corresponding embedding generated by the shared embedding matrix 304 is the same for all the heads 302A to 302H of the multi-head attention mechanism 302, but each head 302A to 302H defines a different set / row of the position vectors 308. For example, the first head 302A defines the rows 308Aa to 308An of the position vector PV Aa~An , the second head 302B defines different rows 308Ba to 308Bn of the position vector PV Ba~Bn , ···, the H-th head 302H defines the position vector PV Ha~HnDefine other different lines 308Ha to 308Hn.

[0044] For each non - blank symbol in the sequence of received non - blank symbols 301a to 301n, the first head 302A weights the corresponding embedding 306 in proportion to the similarity between the corresponding embedding and each of the position vectors PV308 assigned to it via the weight layer 310. In some examples, the similarity may include cosine similarity (e.g., cosine distance). In the example shown, the weight layer 310 outputs a sequence of weighted embeddings 312, 312Aa to 312An, and each weighted embedding is associated with the corresponding embedding 306 weighted in proportion to the respective position vector PV308 assigned to it. In other words, the weighted embedding 312 output by the weight layer 310 for each embedding 306 may correspond to the dot product between the embedding 306 and each position vector PV308. The weighted embedding 312 can be interpreted as being added to the embedding in proportion to the degree to which the embedding is similar to the positioning associated with each position vector PV308. To increase the calculation speed, the prediction network 300 includes a non - recurrent layer, and thus, the sequence of weighted embeddings 312Aa to 312An is not concatenated but is instead averaged by the weighted average module 316. Thus, as the output from the first head 302A, a weighted average 318A of the weighted embeddings 312Aa to 312An represented by the following equation is generated.

[0045]

Equation

[0046] In Equation 1, h represents the index of the head 302, n represents the position within the context, and e represents the embedding dimension. Further, in Equation 1, H, N, and d e, has a size corresponding to the dimension. The position vector PV308 does not need to be trainable and may include random values. In particular, even if the weighted embedding 312 is averaged, the position vector PV308 can potentially store position history information, thus reducing the need to provide regression connections in each layer of the prediction network 300.

[0047] The operations described above for the first head 302A are similarly executed for each of the other heads 302B - 302H of the multi - head attention mechanism 302. By different sets of the positioned vectors PV308 defined by each head 302, the weight layer 310 outputs sequences of weighted embeddings 312Ba - 312Bn, 312Ha - 312Hn in each of the other heads 302B - 302H, which are different from the sequence of weighted embeddings 312Aa - 312Aa in the first head 302A. Then, the weighted average module 316 generates weighted averages 318B - 318H of the corresponding weighted embeddings 312 of the sequence of non - blank symbols as outputs from each of the other corresponding heads 302B - 302H.

[0048] In the example shown, the prediction network 300 includes a head average module 322 that averages the weighted averages 318A - 318H output from the corresponding heads 302A - 302H. The projection layer 326 with SWISH may receive the output 324 from the head average module 322 corresponding to the average of the weighted averages 318A - 318H as input and generate a projection output 328 as output. The final layer normalization 330 normalizes the projection output 328 to provide a single embedding vector P ui 350. The prediction network 300 generates only a single embedding vector P ui 350 at each of the plurality of time steps following the initial time step.

[0049] In some configurations, the prediction network 300 does not implement the multi-head attention mechanism 302 and only performs the above operations with respect to the first head 302A. In these configurations, the weighted average 318A of the weighted embeddings 312Aa to 312An only passes through the projection layer 326 and the layer normalization 330 to obtain a single embedding vector P ui 350 is provided.

[0050] In some embodiments, to further reduce the sizes of the RNN-T decoder, i.e., the prediction network 300 and the joint network 230, parameter tying is applied between the prediction network 300 and the joint network 230. Specifically, for the vocabulary size |V| and the embedding dimension d e in the case of, the shared embedding matrix 304 in the prediction network is

[0051]

Number

[0052] is. On the other hand, since the last hidden layer has the dimension size d in the joint network 230 h the weights of the feed-forward projection from the hidden layer to the output logic are

[0053]

Number

[0054] resulting in the vocabulary containing an extra blank token. Thus, the feed-forward layer corresponding to the last layer of the joint network 230 has the weight matrix [d h , |V|]. The prediction network 300 has the embedding dimension d e with the size of the last hidden layer of the joint network 230 having the dimension d hBy associating, the feed-forward projection weights of the joint network 230 and the shared embedding matrix 304 of the prediction network 300 can share their weights for all non-blank symbols by a simple transpose transformation. Since the two matrices share all their values, instead of storing two individual matrices, the RNN-T decoder only needs to store those values once in memory. Embedding dimension d e Set the size of to be equal to the size of the hidden layer dimension d h By doing so, the RNN-T decoder reduces the number of parameters equal to the product of the embedding dimension d e and the vocabulary size |V|. This weight tying (weight sharing) corresponds to a regularization technique.

[0055] Referring now to FIG. 4, algorithm 400 shows native quantization of 8-bit integers (int8) in TensorFlow. By using algorithm 400, the model trainer 150 determines the scale factor 160 during training by first estimating the maximum value for the axis (axis) to be quantized (since algorithm 400 supports channel-wise quantization). Next, the model trainer 150 determines the scale factor 160 by dividing the maximum value by the integer representation value 410 (i.e., 127.0 in this example). The integer representation value 410 is based on the desired integer fixed bit width 162. That is, the scale factor 160 is based on the estimated maximum value of the axis to be quantized and the target fixed bit width 162. For example, for 8-bit quantization (i.e., an integer fixed bit width of 8), the integer representation value 410 is 127.0. As another example, for 4-bit quantization (i.e., an integer fixed bit width of 4), the integer representation value is 4.0. After determining the scale factor 160, the model trainer 150 quantizes the input tensor by dividing by the scale factor 160 and casting to an integer. De-quantization (de-quantization) can be used by multiplying the tensor by the scale factor 160.

[0056] FIG. 5 is a flowchart of an exemplary arrangement of operations of a method 500 for training an automatic speech recognition (ASR) model 200. The method 500 includes, in operation 502, obtaining a plurality of training samples 152. Each of the plurality of training samples 152 includes each respective speech utterance 154 and each respective text utterance 156 representing a transcription of each respective speech utterance 154. The method 500 includes, in operation 504, training an automatic speech recognition (ASR) model 200 with respect to the plurality of training samples 152 by using “quantization-aware training” (QAT) with native integer arithmetic. In operation 506, the method 500 includes quantizing the trained automatic speech recognition (ASR) model 200 to an integer target fixed bitwidth 162. The quantized trained automatic speech recognition (ASR) model 200 includes a plurality of weights 202. Each of the plurality of weights 202 includes an integer having the target fixed bitwidth 162. In operation 508, the method 500 includes providing the quantized trained automatic speech recognition (ASR) model 200 to a user device 10.

[0057] FIG. 6 is a schematic diagram of an exemplary computing device 600 that can be used to implement the systems and methods described in this document. The computing device 600 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The components shown here, their connections and relationships, and their functions are intended only as examples and are not intended to limit the embodiments of the invention described and / or claimed in this document.

[0058] The computing device 600 includes a processor 610, a memory 620, a storage device 630, a high-speed interface / controller 640 connected to the memory 620 and the high-speed expansion port 650, and a low-speed interface / controller 660 connected to the low-speed bus 670 and the storage device 630. Each of the components (610, 620, 630, 640, 650, and 660) is interconnected by using various buses and can be installed on a common motherboard or exist in other ways as needed. The processor 610 processes instructions for execution within the computing device 600, which includes instructions stored in the memory 620 or the storage device 630, and is enabled to display graphical information of a graphical user interface (GUI) on an external input / output device such as a display 680 coupled to the high-speed interface 640. In other embodiments, multiple memories and multiple types of memories may be used, along with multiple processors and / or multiple buses as needed. Also, multiple computing devices 600 may be connected, and each device may perform a part of the required operations (e.g., as a server bank, a group of blade servers, or a multiprocessor system).

[0059] Memory 620 stores non-temporary information within computing device 600. Memory 620 may be a computer-readable medium, a volatile memory unit(s), or a non-volatile memory unit(s). The non-temporary memory 620 may be a physical device used to temporarily or permanently store programs (e.g., sequences of instructions) or data (e.g., program state information) for use by computing device 600. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as a boot program). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.

[0060] Storage device 630 can provide large-capacity storage for computing device 600. In some embodiments, storage device 630 is a computer-readable medium. In various different embodiments, storage device 630 may be a floppy (registered trademark) disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid-state memory device, or an array of devices comprising a storage area network or other configuration of devices. In additional embodiments, a computer program product is tangibly embodied in an information carrier. The computer program product comprises instructions that perform one or more of the methods as described above when executed. The information carrier is a computer-readable medium or a machine-readable medium such as memory 620, storage device 630, or memory on processor 610.

[0061] The high-speed controller 640 manages the bandwidth-intensive operations of the computing device 600 more, and the low-speed controller 660 manages the bandwidth-intensive operations less. Such role assignments are merely examples. In some embodiments, the high-speed controller 640 is coupled to a high-speed expansion port 650 that can accept the memory 620, the display 680 (e.g., via a graphics processor or accelerator), and various expansion cards (not shown). In some embodiments, the low-speed controller 660 is coupled to the storage device 630 and the low-speed expansion port 690. The low-speed expansion port 690 may include various communication ports (such as USB, Bluetooth®, Ethernet®, wireless Ethernet®, etc.) and may be coupled to one or more input / output devices such as a keyboard, a pointing device, a scanner, or a network device such as a switch or router via a network adapter or the like.

[0062] As shown in the figure, the computing device 600 can be implemented in many different forms. For example, it can be implemented as a standard server 600a, or multiple times as a group of such servers 600a, as a laptop computer 600b, or as part of a rack server system 600c.

[0063] The various embodiments of the systems and techniques described herein can be implemented in digital electronics and / or optical circuits, integrated circuits, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can be special or general purpose, and include at least one programmable processor, at least one input device, and at least one output device, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, and can include embodiments in one or more computer programs executable and / or interpretable in a programmable system.

[0064] A software application (i.e., a software resource) may refer to computer software that causes a computing device to perform tasks. In some examples, a software application may be referred to as an "application", an "app", or a "program". Exemplary applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.

[0065] These computer programs (also known as programs, software, software applications, or code) comprise machine instructions for a programmable processor and can be implemented in high-level procedural and / or object-oriented programming languages, and / or assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" are used to provide machine instructions and / or data to a programmable processor comprising a machine-readable medium that receives the machine instructions as a machine-readable signal, and refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., magnetic disks, optical disks, memory, programmable logic circuits (PLDs)). The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0066] The processes and logical flows described in this specification can be performed by one or more programmable processors executing one or more computer programs that act on input data and generate output, also referred to as data processing hardware. The processes and logical flows can also be performed by special-purpose logic circuitry, such as an FPGA (Field Programmable Gate Array) or ASIC (Application Specific Integrated Circuit). Processors suitable for the execution of a computer program include, by way of example, both general and special purpose processors, as well as any one or more processors of any kind of digital computer. In general, a processor receives instructions and data from a read-only memory, a random access memory, or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. In general, a computer also includes, or is operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks. However, a computer need not have such devices. A computer-readable medium suitable for storing computer program instructions and data includes all forms of non-volatile memory, media, and memory devices, by way of example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices, magnetic disks, such as internal hard disks or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and memory can be supplemented by, or incorporated in, special-purpose logic circuitry.

[0067] To interact with a user, one or more aspects of the present disclosure can be implemented on a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube), an LCD (liquid crystal display) monitor, or a touch screen, and optionally, a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can input to the computer. Other types of devices can also be used to provide interaction with the user. For example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and the input from the user can be received in any form that includes acoustic, speech language, or tactile input. Further, the computer can interact with the user by sending and receiving documents to and from the devices used by the user, such as by sending a web page to a web browser on the user's client device in response to a received request from a web browser.

[0068] Some embodiments have been described. Nevertheless, it is understood that various modifications can be made without departing from the spirit and scope of the present disclosure. Accordingly, other embodiments are within the scope of the following claims.

Claims

1. A method as a computer-implemented method (500), wherein when the computer-implemented method (500) is executed on data processing hardware (62), the data processing hardware (62) is caused to perform operations, and the operations are as follows: A step of obtaining a plurality of training samples (152), wherein each of the plurality of training samples (152) among the plurality of training samples (152) includes: Each voice utterance (154), and Each text utterance (156) representing a transcription of each of the voice utterances (154), A step of obtaining a plurality of the training samples (152); A step of training an automatic speech recognition ASR model (200) related to the plurality of training samples (152) by using quantization-aware training by native integer operations, wherein the number of bits allocated to the native integer operations is an integer fixed bit width, and the native integer operations are used to represent weights of the automatic speech recognition ASR model (200), a step of training the automatic speech recognition ASR model (200); A step of quantizing the trained automatic speech recognition ASR model (200) to an integer target fixed bit width (162), wherein a plurality of weights (202) of the automatic speech recognition ASR model (200) are not only learned during training but also converted to integer values based on the target fixed bit width (162) during quantization, and the converted integer values are included in the quantized automatic speech recognition ASR model (200), a step of quantizing the trained automatic speech recognition ASR model (200) to the target fixed bit width (162); A step of providing the quantized trained automatic speech recognition ASR model (200) to a user device (10); A method (500) comprising the above steps.

2. The target fixed bit width (162) is "4", The method (500) according to claim 1.

3. The automatic speech recognition ASR model (200) further comprises a plurality of activation functions (204), Each of the plurality of activation functions (204) among the plurality of activation functions (204) comprises an integer having the target fixed bit width (162), The method (500) according to claim 1.

4. The automatic speech recognition ASR model (200) further includes a plurality of activation functions (204), each of the plurality of activation functions (204) includes an integer having a fixed bit width greater than the target fixed bit width (162), The method (500) according to claim 1.

5. The automatic speech recognition ASR model (200) further includes a plurality of activation functions (204), each of the plurality of activation functions (204) includes a floating point value, The method (500) according to claim 1.

6. The step of quantizing the trained automatic speech recognition ASR model (200) includes a step of determining a scale factor (160) based on an estimated maximum value of the axis to be quantized and the target fixed bit width (162), The method (500) according to any one of claims 1 to 5.

7. The automatic speech recognition ASR model (200) includes one or more multi-head attention layers (302), The method (500) according to any one of claims 1 to 5.

8. One or more of the multi-head attention layers (302) include one or more conformer layers or one or more transformer layers, The method (500) according to claim 7.

9. The automatic speech recognition ASR model (200) includes a plurality of encoders and a plurality of decoders, The step of quantizing the automatic speech recognition ASR model (200) includes a step of quantizing the plurality of encoders and a step of not quantizing the plurality of decoders, The method (500) according to any one of claims 1 to 5.

10. The automatic speech recognition ASR model (200) includes an audio encoder (210), The audio encoder (210) includes a cascaded encoder including a first causal encoder and a second non-causal encoder, The method (500) according to any one of claims 1 to 5.

11. A system (100), the system (100) includes data processing hardware (62), and memory hardware (64) communicating with the data processing hardware (62), and includes The memory hardware (64) stores instructions which, when executed on the data processing hardware (62), cause the data processing hardware (62) to perform operations, the operations being a step of obtaining a plurality of training samples (152), each of the plurality of training samples (152) comprising each voice utterance (154), and each text utterance (156) representing a transcription of each of the voice utterances (154), the step of obtaining a plurality of the training samples (152); a step of training an automatic speech recognition ASR model (200) for the plurality of training samples (152) using quantization-aware training with native integer arithmetic, wherein the number of bits allocated to the native integer arithmetic is an integer fixed bit width and the native integer arithmetic is used to represent weights of the automatic speech recognition ASR model (200), the step of training the automatic speech recognition ASR model (200); a step of quantizing the trained automatic speech recognition ASR model (200) to an integer target fixed bit width (162), wherein a plurality of weights (202) of the automatic speech recognition ASR model (200) are not only learned during training but are also converted to integer values based on the target fixed bit width (162) during quantization, and the converted integer values are included in the quantized automatic speech recognition ASR model (200), the step of quantizing the trained automatic speech recognition ASR model (200) to the target fixed bit width (162); a step of providing the quantized trained automatic speech recognition ASR model (200) to a user device (10); A system (100) comprising.

12. The target fixed bit width (162) is "4", The system (100) according to claim 11.

13. The automatic speech recognition ASR model (200) further comprises a plurality of activation functions (204), each of the plurality of activation functions (204) comprising an integer having the target fixed bit width (162), The system (100) according to claim 11.

14. The automatic speech recognition ASR model (200) further includes a plurality of activation functions (204), each of the plurality of activation functions (204) among the plurality of activation functions (204) includes an integer having a fixed bit width greater than the target fixed bit width (162), The system (100) according to claim 11.

15. The automatic speech recognition ASR model (200) further includes a plurality of activation functions (204), each of the plurality of activation functions (204) among the plurality of activation functions (204) includes a floating-point value, The system (100) according to claim 11.

16. The step of quantizing the trained automatic speech recognition ASR model (200) includes a step of determining a scale factor (160) based on the estimated maximum value of the axis to be quantized and the target fixed bit width (162), The system (100) according to any one of claims 11 to 15.

17. The automatic speech recognition ASR model (200) includes one or more multi-head attention layers (302), The system (100) according to any one of claims 11 to 15.

18. One or more of the multi-head attention layers (302) include one or more conformer layers or one or more transformer layers, The system (100) according to claim 17.

19. The automatic speech recognition ASR model (200) includes a plurality of encoders and a plurality of decoders, The step of quantizing the automatic speech recognition ASR model (200) includes a step of quantizing the plurality of encoders and a step of not quantizing the plurality of decoders, The system (100) according to any one of claims 11 to 15.

20. The automatic speech recognition ASR model (200) includes an audio encoder (210), The audio encoder (210) includes a cascaded encoder including a first causal encoder and a second non-causal encoder, The system (100) according to any one of claims 11 to 15.

Citation Information

Patent Citations

  • Speech feature classification method, related equipment and readable storage medium

    CN113593538A

  • Error tolerant neural network model compression

    US10229356B1

  • Fine-grained per-vector scaling for neural network quantization

    US20220067512A1

  • 4-bit quantization method and system for neural network

    WO2021258752A1