Semi-supervised training schemes for speech recognition
Patent Information
- Application Number
- JP2025532583
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-14
- Filing Date
- 2023-12-12
- Publication Date
- 2025-12-23
Smart Images

Figure 2025541793000001_ABST
Abstract
Description
[Technical Field]
[0001] This disclosure relates to semi-supervised training schemes for speech recognition. [Background technology]
[0002] Automatic speech recognition (ASR) systems attempt to provide accurate transcriptions of what a person says by taking audio input and transcribing the audio input into text. In many instances, supervised learning is used to train ASR systems using large amounts of labeled training data that include audio data and corresponding transcriptions. However, obtaining the large amounts of labeled training data required to train ASR systems is often challenging due to the time required, cost, and / or privacy concerns associated with collecting large labeled training datasets. Training an ASR system with unlabeled training data that includes only audio data can alleviate some of the difficulties associated with collecting large amounts of labeled training data. Summary of the Invention
[0003] One aspect of the present disclosure provides a cross-training network for training a speech recognition model. The cross-training network includes an unsupervised subnetwork trained with a plurality of unlabeled audio samples corresponding to speech utterances that are not paired with corresponding transcriptions. The unsupervised subnetwork includes a target branch configured to receive a sequence of acoustic frames extracted from the unlabeled audio samples as input to a supervised audio encoder of the speech recognition model and generate, at each of a plurality of output steps, a target high-order feature representation of a corresponding acoustic frame in the sequence of acoustic frames input to the supervised audio encoder at the corresponding output step. The unsupervised subnetwork also includes an augmentation branch configured to augment the sequence of acoustic frames extracted from the unlabeled audio samples by masking one or more acoustic frames in the sequence of acoustic frames and generate, at each of a plurality of output steps, a predicted high-order feature representation of a corresponding augmented acoustic frame in the sequence of augmented acoustic frames as output from the unsupervised audio encoder of the speech recognition model. The unsupervised sub-network is configured to determine, at each of the plurality of output steps, an unsupervised loss term based on the target high-level feature representation generated by the target branch at the corresponding output step and the predicted high-level feature representation generated by the augmentation branch at the corresponding output step, and to update parameters of the speech recognition model based on the unsupervised loss term determined at each of the plurality of output steps.
[0004] Implementations of the present disclosure may include one or more of the following optional features. In some implementations, the unsupervised loss term includes a contrastive loss term. In some examples, the unsupervised subnetwork is further configured to determine, at each of the plurality of output steps, a distance-based loss term between parameters of the unsupervised audio encoder and parameters of the supervised audio encoder, and updating the parameters of the speech recognition model is further based on the distance-based loss term determined at each of the plurality of output steps. Here, the distance-based loss term may be an L2 loss. In these examples, updating the parameters of the speech recognition model based on the unsupervised loss term occurs jointly with updating the parameters of the speech recognition model based on the distance-based loss term.
[0005] In some implementations, the cross-training network further includes a supervised subnetwork trained with a plurality of labeled audio samples corresponding to speech utterances paired with corresponding transcriptions. In these implementations, at each of a plurality of output steps for each labeled sample, the supervised subnetwork is configured to generate a corresponding speech recognition result for the labeled audio sample using a speech recognition model and determine a supervised loss term based on the target high-order feature representation generated by the target branch at the corresponding output step and the predicted high-order feature representation generated by the augmentation branch at the corresponding output step. Here, the supervised subnetwork updates parameters of the speech recognition model based on the supervised loss term determined at each of the plurality of output steps for each labeled audio sample in the plurality of labeled audio samples. In these implementations, the corresponding speech recognition result generated for the labeled audio sample using the speech recognition model includes a probability distribution over possible speech recognition hypotheses for the labeled audio sample at the corresponding output step. The supervised subnetwork may be further configured to update parameters of the speech recognition model based on the supervised loss term jointly with the unsupervised network updating parameters of the speech recognition model based on the unsupervised loss term and the distance-based loss term.
[0006] The target branch may be further configured to apply a stop gradient operation to the predicted high-order feature representation of the corresponding extended audio frame. In some examples, the parameters of the unsupervised audio encoder and the parameters of the supervised audio encoder are initialized with the same initial parameters. In other examples, the parameters of the unsupervised audio encoder and the parameters of the supervised audio encoder are initialized with different initial parameters. Each of the unsupervised audio encoder and the supervised audio encoder includes at least one of a respective full-context encoder or a respective cascade encoder.
[0007] Another aspect of the present disclosure provides a computer-implemented method that, when executed on data processing hardware, causes the data processing hardware to perform operations for training a speech recognition model using a cross-training network. The operations include receiving a sequence of acoustic frames extracted from unlabeled audio samples corresponding to speech utterances not paired with any transcription. In a target branch of the cross-training network, the operations include generating, at multiple output steps, target high-order feature representations of corresponding acoustic frames in the sequence of acoustic frames using a supervised audio encoder of the speech recognition model. In an augmentation branch of the cross-training network, the operations include augmenting the sequence of acoustic frames extracted from the unlabeled audio samples by masking one or more acoustic frames in the sequence of acoustic frames, and generating, at each of the multiple output steps, a predicted high-order feature representation of the corresponding augmented acoustic frame in the sequence of augmented acoustic frames as output from the unsupervised audio encoder of the speech recognition model. The operations also include, at each of the multiple output steps, determining an unsupervised loss term based on the target high-order feature representation generated by the target branch at the corresponding output step and the predicted high-order feature representation generated by the augmentation branch at the corresponding output step. The operations also include updating parameters of the speech recognition model based on the unsupervised loss terms determined at each of the plurality of output steps.
[0008] Implementations of the present disclosure may include one or more of the following optional features. In some implementations, the unsupervised loss term includes a contrastive loss term. In some examples, the operations include, at each of the plurality of output steps, determining a distance-based loss term between parameters of the unsupervised audio encoder and parameters of the supervised audio encoder, and updating the parameters of the speech recognition model is further based on the distance-based loss term determined at each of the plurality of output steps, where the distance-based loss term may include an L2 loss. In these examples, updating the parameters of the speech recognition model based on the unsupervised loss term occurs jointly with updating the parameters of the speech recognition model based on the distance-based loss term.
[0009] In some implementations, the operations further include receiving a plurality of labeled audio samples corresponding to the speech utterance paired with a corresponding transcription. In these implementations, at each of the plurality of output steps for each labeled audio sample, the operations also include generating a corresponding speech recognition result for the labeled audio sample using a speech recognition model and determining a supervised loss term based on the corresponding speech recognition result for the labeled audio sample and the corresponding transcription of the labeled audio sample. Here, the operations also include updating parameters of the speech recognition model based on the supervised loss term determined at each of the plurality of output steps for each labeled audio sample in the plurality of labeled audio samples. The corresponding speech recognition result generated for the labeled audio sample using the speech recognition model may include a probability distribution over possible speech recognition hypotheses for the labeled audio sample at the corresponding output step. Updating the parameters of the speech recognition model based on the supervised loss term may occur jointly with updating the parameters of the speech recognition model based on the unsupervised loss term and the distance-based loss term.
[0010] In some examples, the operations further include applying a gradient stopping operation to the predicted high-order feature representation of the corresponding extended audio frame. In some implementations, the parameters of the unsupervised audio encoder and the parameters of the supervised audio encoder are initialized with the same initial parameters. In other implementations, the parameters of the unsupervised audio encoder and the parameters of the supervised audio encoder are initialized with different initial parameters. Each of the unsupervised audio encoder and the supervised audio encoder may include at least one of a respective full-context encoder or a respective cascade encoder.
[0011] The details of one or more embodiments of the disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will become apparent from the description and drawings, and from the claims. [Brief explanation of the drawings]
[0012] [Figure 1] FIG. 1 is a schematic diagram of an audio environment in which an exemplary speech recognition model is implemented. [Figure 2] FIG. 2 is a schematic diagram of the exemplary speech recognition model of FIG. 1. [Figure 3A] FIG. 1 is a schematic diagram of the supervised portion of a cross-training network that performs a semi-supervised training process for a speech recognition model. [Figure 3B] FIG. 1 is a schematic diagram of the unsupervised portion of a cross-training network that performs a semi-supervised training process for a speech recognition model. [Figure 4] 1 is a flowchart of an exemplary arrangement of operations for a method of training a speech recognition model using a cross-training model. [Figure 5] FIG. 1 is a schematic diagram of an example computing device that can be used to implement the systems and methods described herein. DETAILED DESCRIPTION OF THE INVENTION
[0013] Like reference symbols in the various drawings indicate like elements.
[0014] Automatic speech recognition (ASR) systems are often trained using supervised training techniques that utilize labeled training data. The labeled training data includes speech audio data and corresponding speech transcriptions. Collecting large amounts of labeled training data is often difficult due to the associated costs, the time required to collect the training data, and user privacy concerns. In some examples, ASR systems train using unlabeled training data that includes only speech audio data without any corresponding speech transcriptions. In these examples, the ASR system may utilize only unlabeled training data to train the speech recognition system (i.e., self-supervised training), or unlabeled training data may be used in addition to labeled training data to train the speech recognition system (i.e., semi-supervised training).
[0015] However, training speech recognition models using self-supervised training often results in instability during execution. That is, although these self-supervised training systems use computationally expensive operations of the predictive network, time corrections (e.g., speed changes), and joint training, the trained ASR systems are still unstable during execution. In particular, semi-supervised training performs poorly in training scenarios including when the labeled training dataset is relatively small, when the ASR system operates in a streaming manner using cascaded encoders, and / or when there is a mismatch between the labeled and unlabeled training data. Therefore, there is a need for semi-supervised training of ASR systems that produces stable outputs during execution.
[0016] Accordingly, embodiments herein are directed to a cross-training network that uses semi-supervised training techniques to train a speech recognition model. The cross-training network includes an unsupervised subnetwork trained with a plurality of unlabeled audio samples corresponding to speech utterances that are not paired with any corresponding transcription. The unsupervised subnetwork includes a target branch configured to receive a sequence of acoustics extracted from the unlabeled audio samples as input to a supervised audio encoder. The target branch is further configured to generate a target high-order feature representation of a corresponding acoustic frame in the sequence of acoustic frames.
[0017] The unsupervised subnetwork also includes an augmentation branch configured to augment the sequence of acoustic frames extracted from the unlabeled audio samples by masking one or more acoustic frames in the sequence of acoustic frames. The augmentation branch is further configured to generate, as output from the unsupervised audio encoder, a predicted high-dimensional feature representation of a corresponding augmented acoustic frame in the sequence of augmented acoustic frames. Here, the unsupervised subnetwork determines an unsupervised loss term based on the target high-dimensional feature representation and the predicted high-dimensional feature representation, and updates parameters of the speech recognition model based on the unsupervised loss term. As will become apparent, the unsupervised subnetwork may also determine a distance-based loss between parameters of the unsupervised audio encoder and the supervised audio encoder. Here, the unsupervised subnetwork jointly updates parameters of the speech recognition model based on the unsupervised loss term and the distance-based loss term.
[0018] 1 is an example of an audio environment 100. In the audio environment 100, a way in which a user 104 interacts with a computing device, such as a user device 10, may be through voice input. The user device 10 is configured to obtain sounds (e.g., streaming audio data) from one or more users 104 within the audio environment 100. Here, the streaming audio data may refer to voice utterances 106 by the users 104 that function as audible queries, commands to the user device 10, or audible communications captured by the user device 10. A voice-enabled system of the user device 10 may process the queries or commands by responding to the queries and / or causing the commands to be executed / performed by one or more downstream applications.
[0019] The user device 10 may correspond to any computing device associated with a user 104 and capable of receiving audio data. Some examples of the user device 10 include, but are not limited to, a mobile device (e.g., a mobile phone, a tablet, a laptop, etc.), a computer, a wearable device (e.g., a smart watch), a smart appliance, an Internet of Things (IoT) device, an in-vehicle infotainment system, a smart display, a smart speaker, etc. The user device 10 includes data processing hardware 12 and memory hardware 14 in communication with the data processing hardware 12, storing instructions that, when executed by the data processing hardware 12, cause the data processing hardware 12 to perform one or more operations. The user device 10 further includes an audio system 16 having audio capture devices (e.g., microphones) 16, 16a for capturing and converting audio utterances 106 in the audio environment 100 into electrical signals, and audio output devices (e.g., speakers) 16, 16b for communicating audible audio signals (e.g., as output audio data from the user device 10). In the illustrated example, the user device 10 implements a single audio capture device 16a, however, the user device 10 may implement an array of audio capture devices 16a without departing from the scope of this disclosure, such that one or more capture devices 16a of the array may not be physically present on the user device 10 but may communicate with the audio system 16.
[0020] In the voice environment 100, an automatic speech recognition (ASR) system 118 implementing the speech recognition model 200 resides on the user device 10 of the user 104 and / or on a remote computing device 60 (e.g., one or more remote servers of a distributed system executing in a cloud computing environment) that communicates with the user device 10 via the network 40. The user device 10 and / or the remote computing device (i.e., remote server) 60 also includes an audio subsystem 108 configured to receive utterances 106 spoken by the user 104 and captured by the audio capture device 16a and convert the utterances 106 into a corresponding digital format associated with input acoustic frames 110 that can be processed by the ASR system 118. In the illustrated example, the user speaks each utterance 106, and the audio subsystem 108 converts the utterances 106 into corresponding audio data (e.g., acoustic frames) 110 for input to the ASR system 118. The speech recognition model 200 then receives, as input, audio data 110 corresponding to the utterance 106 and generates / predicts, as output, a corresponding transcription 120 (eg, a speech recognition result / hypothesis) of the utterance 106 .
[0021] The digital assistant application 50 running on the user device 10 may require streaming speech recognition so that words, word pieces, and / or individual characters are displayed on the screen as soon as they are spoken. Furthermore, the user 104 of the user device 10 may have low tolerance for latency when issuing queries for the digital assistant application 50 to execute. In such scenarios, when minimizing speech recognition latency is preferred, the speech recognition model 200 may apply zero or minimal look-ahead audio context (also referred to as "correct context") to provide real-time streaming transcription functionality as the user 104 speaks the utterance 106. On the other hand, when the user's tolerance for speech recognition latency is high and / or the utterance 106 being recognized is associated with long-form speech, the same speech recognition model 200 may apply a duration of look-ahead audio context sufficient to provide an accurate transcription 120, but may incur increased latency based on the duration of the look-ahead audio context. Thus, the ASR system 118 may implement only a single speech recognition model 200 for many different speech recognition tasks, providing both streaming and non-streaming transcription capabilities without having to utilize separate ASR models for each task.
[0022] In some implementations, the speech recognition model 200 simultaneously performs both streaming and non-streaming speech recognition on the audio data 110. For example, in the illustrated example, the speech recognition model 200 simultaneously performs streaming speech recognition on the audio data 110 to generate partial speech recognition results 120, 120a and non-streaming speech recognition on the same audio data 110 to generate final speech recognition results 120, 120b. In particular, the speech recognition model 200 may use a first look-ahead audio context, which may be set to zero (or approximately 240 milliseconds), to generate the partial speech recognition result 120a, and may use a second look-ahead audio context, which has a longer duration than the first look-ahead audio context, to generate the final recognition result 120b. Thus, the final speech recognition result 120b for the input utterance 106 may be delayed from the partial speech recognition result 120a for the input utterance by a duration based on the difference between the second look-ahead audio context and the first look-ahead audio context.
[0023] The user device 10 and / or the remote computing device 60 also execute a user interface generator 107 configured to present a representation of a transcription 120 of the utterance 106 to the user 104 of the user device 10. As described in more detail below, the user interface generator 107 may display partial speech recognition results 120a in a streaming manner during time 1, and then display final speech recognition results 120b during time 2. In some configurations, the transcription 120 output from the ASR system 118 is processed by, for example, a natural language (NLU) module executing on the user device 10 or the remote computing device 60 to execute the user command / query specified by the utterance 106. Additionally or alternatively, a text-to-speech system (not shown) (e.g., executing on any combination of the user device 10 or the remote computing device 60) may convert the transcription into synthesized speech for audible output by the user device 10 and / or other devices.
[0024] In the illustrated example, a user 104 interacts with a program or application 50 (e.g., a digital assistant application 50) on a user device 10 that uses an ASR system 118. For example, FIG. 1 shows a user 104 communicating with a digital assistant application 50 and the digital assistant application 50 displaying a digital assistant interface 18 on the screen of the user device 10 to depict a conversation between the user 104 and the digital assistant application 50. In this example, the user 104 asks the digital assistant application 50, "What time is the concert tonight?" This question from the user 104 is a speech utterance 106 that is captured by an audio capture device 16a and processed by an audio system 16 of the user device 10. In this example, the audio system 16 receives the speech utterance 106 and converts the speech utterance 106 into acoustic frames 110 for input to the ASR system 118.
[0025] Continuing with the example, as the user 104 speaks, the speech recognition model 200 receives acoustic frames (i.e., audio data) 110 corresponding to the utterance 106, encodes the acoustic frames 110 using a first look-ahead audio context, and then decodes the encoded acoustic frames 110 into partial speech recognition results 120a using the first look-ahead audio context. During time 1, the user interface generator 107 presents a representation of the partial speech recognition results 120a of the utterance 106 to the user 104 of the user device 10 via the digital assistant interface 18 in a streaming manner, such that words, word pieces, and / or individual characters are displayed on the screen as they are spoken. In some examples, the first look-ahead audio context is equal to zero.
[0026] Simultaneously, after all of the acoustic frames 110 corresponding to the utterance 106 have been received, the speech recognition model 200 encodes all of the acoustic frames 110 corresponding to the utterance 106 using a second look-ahead audio context and then decodes the acoustic frames 110 into a final speech recognition result 120b using the second look-ahead audio context. The second look-ahead audio context may be 1.2 seconds, 2.4 seconds, or any other duration. In some examples, an indication, such as an endpoint, that the user 104 has finished speaking the utterance 106 triggers the speech recognition model 200 to encode all of the acoustic frames 110 using the second look-ahead audio context. During time 2, the user interface 107 presents a representation of the final speech recognition result 120b of the utterance 106 to the user 104 of the user device 10 via the digital assistant interface 18. In some implementations, the user interface generator 107 replaces the representation of the partial speech recognition result 120a with a representation of the final speech recognition result 120b. For example, because the final speech recognition result 120b is estimated to be more accurate than the partial speech recognition result 120a generated without utilizing look-ahead audio context, the final speech recognition result 120b, which is ultimately displayed as the transcription 120, may correct any terms that may have been misrecognized in the partial speech recognition result 120a. In this example, the streaming of the partial speech recognition result 120a output by the speech recognition model 200 and displayed on the screen of the user device 10 at time 1 is associated with low latency, providing the user 104 with a sense of responsiveness that their query is being processed, while the final speech recognition result 120b output by the speech recognition model 200 and displayed on the screen at time 2 utilizes look-ahead audio context to improve the quality of the speech recognition in terms of accuracy but increase latency. However, because the partial speech recognition results 120a are displayed as the user speaks the utterance 106, the higher latency associated with generating and ultimately displaying the final recognition results is not noticeable to the user 104.
[0027] In the example shown in FIG. 1 , the digital assistant application 50 can respond to a question posed by the user 104 using natural language processing. Natural language processing generally refers to the process of interpreting written language (e.g., partial speech recognition results 120a and / or final speech recognition results 120b) and determining whether the written language prompts some action. In this example, the digital assistant application 50 uses natural language processing to recognize that the question from the user 104 is about the user's schedule and, more specifically, about a concert on the user's schedule. By recognizing these details using natural language processing, the automated assistant returns a response 19 to the user's query, where the response 19 states, "The venue doors open at 6:30 PM, and the concert starts at 8:00 PM." In some configurations, the natural language processing occurs on a remote server 60 that communicates with the data processing hardware 12 of the user device 10.
[0028] Referring to FIG. 2 , the speech recognition model 200 can provide end-to-end (E2E) speech recognition by integrating an acoustic model, a pronunciation model, and a language model into a single neural network, eliminating the need for a vocabulary or separate text normalization component. Various architectures and optimization mechanisms can improve accuracy and reduce model training time. In some implementations, the speech recognition model 200 includes a transformer transducer (TT) model architecture that adheres to latency constraints associated with interactive applications. The TT model 200 may include the TT model 200 described in U.S. Patent Application No. 17 / 210,465, filed March 23, 2021, the entire contents of which are incorporated herein by reference. Because the TT model 200 has a small computational footprint and utilizes lower memory requirements than conventional ASR architectures, the TT model architecture is suitable for performing speech recognition entirely on the user device 10 (e.g., without requiring communication with a remote server 60). The TT model 200 includes an audio encoder 210, a label encoder 220, and a collaborative network 230. The audio encoder 210, which is roughly similar to an acoustic model (AM) in a conventional ASR system, may include a neural network with a stack of successive convolutional and transformer layers. Furthermore, the audio encoder 210 may include a supervised audio encoder 212 (FIG. 3) and an unsupervised audio encoder 216 (FIG. 3). The audio encoder 210 receives a sequence of d-dimensional feature vectors (e.g., an acoustic frame, where x t ∈R d In the acoustic frame 110 (FIG. 1), x = (x1, x2, ..., x T ), and at each time step produces a high-dimensional representation (also called the "encoder output"). This high-dimensional representation is denoted by ah1,...,ah TEach transformer layer of the audio encoder 210 may include a normalization layer, a masked multi-head attention layer with relative position coding, a residual connection, a stacked / non-stacked layer, and a feedforward layer. Similarly, the label encoder 220 may also include a neural network or lookup table embedding model in the transform layer, which, like a language model (LM), encodes the sequence of non-blank symbols, y0,...,y, output so far by the final softmax layer 240. ui-1 is a dense representation Ih that encodes the predicted label history. u In embodiments, when the label encoder 220 includes a neural network of transformer layers, each transformer layer may include a normalization layer, a masked multi-head attention layer with relative position coding, a residual connection, a feedforward layer, and a dropout layer. In these embodiments, the label encoder 220 may include two transformer layers. In embodiments, when the label encoder 220 includes a lookup table embedding model with bigram label contexts, the embedding model is configured to learn a d-dimensional weight vector for each possible bigram label context, where d is the dimension of the output of the audio encoder and label encoder 210, 220. In some arrays, the total number of parameters in the embedding model is N 2 × d, where N is the label vocabulary size. Here, the learned weight vector is then used as the embedding of the bigram label context of the TT model 200 to generate the fast label encoder 220 runtime.
[0029] Finally, the TT model architecture allows the representations generated by the audio and label encoders 210, 220 to be processed by the dense layer J. u,t The joint network 230 then calculates P(Z u,t |x,t,y1,...,y u-1) at each output step (e.g., time step). In other words, the collaborative network 230 generates a probability distribution over possible speech recognition hypotheses at each output step (e.g., time step). Here, a "possible speech recognition hypothesis" corresponds to a set of output labels (also called "phonetic units"), each representing a grapheme (e.g., a symbol / character) or word piece in a specified natural language. For example, when the natural language is English, the set of output labels may include 27 symbols, e.g., one label for each of the 26 letters of the English alphabet, and one label specifying a space. Accordingly, the collaborative network 230 may output a set of values indicating the likelihood of occurrence of each of a given set of output labels. This set of values may be a vector (e.g., a one-hot vector) and may indicate a probability distribution over the set of output labels. In some cases, the output labels are graphemes (e.g., individual characters, potential punctuation marks and other symbols, etc.), but the set of output labels is not limited thereto. For example, the set of output labels may include parts of words and / or entire words in addition to or instead of graphemes. The output distribution of the collaborative network 230 may include a posterior probability value for each of the different output labels. Thus, if there are 100 different output labels representing different graphemes or other symbols, the output z of the collaborative network 230 may be u,t may contain 100 different probability values, one for each output label. The probability distributions can then be used to select and assign scores to candidate orthographic elements (e.g., graphemes, wordpieces, and / or words) in a beam search process (e.g., by softmax layer 240) to determine transcription 120.
[0030] The softmax layer 240 may use any technique to select the output label / symbol with the highest probability within the distribution as the next output symbol predicted by the TT model 200 at the corresponding output step. In this manner, the TT model 200 does not make a conditional independence assumption; rather, the prediction of each symbol is conditioned not only on the acoustics but also on the sequence of labels output so far. Although the acoustic recognition model 200 is described as having a TT model architecture, the speech recognition model 200 may include other types of transducer-based architectures, such as a conformal transducer (CT) model architecture or a recurrent neural network-transducer (RNN-T) model architecture.
[0031] 3A and 3B show schematic diagrams of a cross-training network 300 that performs a semi-supervised training process to train the speech recognition model 200 (FIG. 2). The cross-training network includes a supervised subnetwork training process 301 (FIG. 3A) and an unsupervised subnetwork training process 302 (FIG. 3B). The supervised subnetwork training process (i.e., the supervised subnetwork) 301 trains the speech recognition model 200 using a plurality of labeled audio samples 305 including sequences of acoustic frames 304 extracted from a spoken speech utterance 106 paired with corresponding transcriptions (i.e., labels) 308. The unsupervised subnetwork training process (i.e., the unsupervised subnetwork) 302 trains the speech recognition model 200 using a plurality of unlabeled audio samples 303 including sequences of acoustic frames 304 extracted from a speech utterance 106 without any paired transcriptions.
[0032] In some examples, the acoustic frames 306 used by the supervised sub-network (i.e., supervised portion) 301 are the same as the acoustic frames 304 used by the unsupervised sub-network (i.e., unsupervised portion) 302. That is, the supervised portion 301 and the unsupervised portion 302 may simultaneously use the same acoustic frames 304, 306 to train the speech recognition model 200. In other examples, the acoustic frames 306 used to train the supervised portion 301 are different from the acoustic frames 304 used to train the unsupervised portion 302. This scenario is particularly beneficial because unlabeled audio samples 303, without any corresponding transcriptions, are easy to obtain and can be utilized to train the speech recognition model 200. Thus, the speech recognition model 200 can be trained with any combination of labeled audio samples 305 and / or unlabeled audio samples 303. In some examples, the sequences of acoustic frames 304, 306 extracted from the unlabeled audio samples 303 and the labeled audio samples 305 comprise log-Mel filter bank energies. A greater number of acoustic frames 304 may be used to train the unsupervised portion 302 than the number of acoustic frames 306 used to train the supervised portion 301. Optionally, a greater number of acoustic frames 306 may be used to train the supervised portion 301 than the number of acoustic frames 304 used to train the unsupervised portion 302. In some examples, the number of acoustic frames 306 used to train the supervised portion 301 and the number of acoustic frames 304 used to train the unsupervised portion 302 are the same.
[0033] The supervised portion 301 includes a supervised audio encoder 212 of the speech recognition model 200 that is shared with the target branch 310 of the unsupervised portion 302. The unsupervised portion 302 further includes an unsupervised audio encoder 216 in an extension branch 320 distinct from the supervised audio encoder 212. Here, the supervised and unsupervised audio encoders 212, 216 may each include a stack of strided convolutional layers (e.g., two convolutional layers) and transformer layers (e.g., 20 bidirectional transformer layers). In some implementations, the supervised and unsupervised audio encoders 212, 216 each include a respective full-context encoder operating in a non-streaming manner. Here, the full-context encoder outputs an encoder output corresponding to the final speech recognition result 120b (FIG. 1) at each of multiple output steps. In other implementations, the supervised and unsupervised audio encoders 212, 216 each include a respective cascade encoder operating in a streaming manner. That is, the cascade encoder includes a causal encoder that does not receive any corrective context and outputs an encoder output corresponding to a partial speech recognition result 120a (FIG. 1) at each of a plurality of output steps, and a non-causal encoder that receives additional corrective context and outputs an encoder output corresponding to a final speech recognition result 120b (FIG. 1) at each of a plurality of output steps.
[0034] 3A , the supervised portion 301 of the cross-training network 300 trains the speech recognition model 200 using a plurality of labeled audio samples 305. Each labeled audio sample 305 of the plurality of labeled audio samples 305 corresponds to a speech utterance 106 paired with a corresponding transcription 308. The plurality of labeled audio samples 305 includes a sequence of acoustic frames 306 extracted from the labeled audio samples 305. The supervised portion 301 shares the same supervised audio encoder 212 from the speech recognition model 200 as the target branch 310 of the unsupervised portion 302. The supervised portion 301 also includes a label encoder 220 and a dense layer 346 with a bias vector 347 (e.g., a collaborative network 230).
[0035] Optionally, the supervised portion 301 may include a data augmentation module 365 (shown in dashed lines), which applies data augmentation to at least one acoustic frame 306 extracted from the labeled audio samples 305 to generate a sequence of augmented acoustic frames 306, 306A. The data augmentation module 365 in the supervised portion 301 may be the same as (or different from) the data augmentation module 360 in the unsupervised portion 302 ( FIG. 3B ). In some examples, the data augmentation module 365 in the supervised portion 301 applies a different data augmentation technique than the data augmentation module 360 in the unsupervised portion. Applying data augmentation to the acoustic frames 306 promotes acoustic diversity in the audio frames used to train the speech recognition model 200. The data augmentation module 360 may include a temporal masking component that masks portions of the acoustic frames 306. Other techniques applied by the data augmentation module 360 may include adding / injecting noise and / or adding reverberation to the labeled audio samples 305. One data augmentation technique involves using multi-style training (MTR) to inject various environmental noises into the labeled audio samples 305. Another data augmentation description that the data augmentation module 360 may apply in addition to or instead of MTR includes using spectral augmentation (SpecAugment) to make the acoustics of a labeled audio sample 305 approximate the inverse acoustics of another labeled audio sample 305. When combined, MTR and SpecAugment may inject noise into the labeled audio samples 305, tile a random external noise source previously inserted along time and superimposed on the representation, and filter the noise-injected labeled audio samples before training the speech recognition model 200.
[0036] In some examples, when the supervised portion 301 includes the data augmentation module 365, the supervised audio encoder 212 receives the augmented sequence of acoustic frames 306A and, at each output step, generates a high-order feature representation 341 of a corresponding augmented acoustic frame 306A in the sequence of augmented acoustic frames 306A. In other examples, when the supervised portion 301 does not include the data augmentation module 365, the supervised audio encoder 212 directly receives the sequence of acoustic frames 306 and, at each output step, generates a high-order feature representation 341 of a corresponding augmented acoustic frame 306 in the sequence of augmented acoustic frames 306. More specifically, successive convolutional layers of the supervised audio encoder 212 receive the augmented acoustic frame 306A (or the acoustic frame 306) and generate corresponding outputs. Here, transformer layers of the supervised audio encoder 212 receive the corresponding outputs and generate the high-order feature representation 341.
[0037] The label encoder 220 is a streaming transformer that does not pay attention to future labels 308. Thus, the label encoder 220 receives the labels 308 corresponding to the augmented acoustic frames 306A (or acoustic frames 306) received by the supervised audio encoder 212, and at each output step, it generates a language embedding 344 (i.e., a dense representation Ih u(FIG. 2)). The supervised portion 301 includes a dense layer 346 that processes language embeddings 344 from the label encoder 220 and high-level feature representations 341 (i.e., acoustic embeddings) from the supervised audio encoder 210 to generate, at each output step, a corresponding speech recognition result 342 for the corresponding augmented acoustic frame 306A (or acoustic frame 306) using the speech recognition model 200. The dense layer 346 includes a trainable bias vector 347 that performs a linear operation on the high-level feature representation 341 and language embeddings 344 to generate the speech recognition result 342. The speech recognition result 342 may include a probability distribution over possible speech recognition hypotheses for the labeled audio sample 305 at the corresponding output step. The loss module 351 of the supervised portion 301 determines a supervised loss term 350 at each of a plurality of output steps based on the corresponding speech recognition result 342 of the labeled audio sample 305 and the corresponding transcription 308 of the labeled audio sample 305. That is, the loss module 351 compares the speech recognition result 342 with the label (e.g., ground truth transcription) 308 to generate the supervised loss term 350. The supervised loss term (e.g., RNN-T loss) 350 can be expressed as follows:
number
[0038] In Equation 1, r t represents a logit vector specifying the probability of a grapheme containing a blank symbol, and a t represents the high-dimensional feature representation 341 from the supervised audio encoder 212, and l t represents the language embeddings 344 from the label encoder 220, and linear represents a conventional dense layer 346 with a trainable bias vector 347.
[0039] The supervised portion 301 updates parameters of the speech recognition model 200 based on supervised loss terms 350 determined at each of the plurality of output steps for each labeled audio sample 305 in the plurality of labeled audio samples 305. In some implementations, the supervised portion 301 is configured to update parameters of the speech recognition model 200 based on the supervised loss terms 350 independently of the unsupervised portion 302 updating parameters of the speech recognition model 200. In other implementations, the supervised portion 301 is configured to update parameters of the speech recognition model 200 based on the supervised loss terms 350 collaboratively with the unsupervised portion 302 updating parameters of the speech recognition model 200. Updating parameters of the speech recognition model 200 may include updating parameters of the supervised audio encoder 212.
[0040] Referring now to FIG. 3B, the unsupervised portion 302 trains the speech recognition model 200 using a plurality of unlabeled audio samples 303, which include a sequence of acoustic frames extracted from the speech utterance 106 that has not been paired with any transcription. The unsupervised portion 302 of the cross-training network 300 includes a target branch 310 and an augmentation branch 320. Here, the target branch 310 shares the same supervised audio encoder 212 of the speech recognition model 200 as the supervised portion 301 (FIG. 3A). The augmentation branch 320 includes the unsupervised audio encoder 216 of the speech recognition model 200. The unsupervised portion 302 is configured to extract linguistic information by matching the high-level feature representations (e.g., the target high-level feature representation 214 and the predicted high-level feature representation 218) of the target branch 310 and the augmentation branch 320.
[0041] The target branch 310 is configured to generate a target high-order feature representation 214 based on a sequence of acoustic frames 304 extracted from the unlabeled audio samples 303. The supervised audio encoder 212 of the target branch receives the sequence of acoustic frames 304 and, at each output step of a plurality of output steps, generates a target high-order feature representation 214 of the corresponding acoustic frame 304. In particular, successive convolutional layers of the supervised audio encoder 212 receive the acoustic frames 304 from the sequence of acoustic frames and generate outputs that the transformer layers use to generate the target high-order feature representation 214 of the corresponding acoustic frame.
[0042] The target branch 310 does not backpropagate gradients to train the supervised audio encoder 210 via the target branch 310. In particular, training both the supervised audio encoder 212 and the unsupervised audio encoder 216 with a contrastive loss may cause the encoders to learn irrelevant relationships (i.e., a shortcut learning problem) by learning to minimize the contrastive loss by transferring positional information to the encoder's output. Therefore, the target branch 310 applies a gradient stopping operation 314 to the target high-dimensional feature representation 214 to prevent backpropagation of gradients (e.g., contrastive loss) to the supervised audio encoder 212 via the target branch 310. Therefore, applying the gradient stopping operation 314 overcomes the shortcut learning problem.
[0043] The augmentation branch 320 of the unsupervised portion 302 includes a data augmentation module 360 that applies data augmentation to an acoustic sequence of acoustic frames 304 extracted from the unlabeled audio samples 303. That is, the data augmentation module 360 receives a sequence of acoustic frames and generates a sequence of augmented acoustic frames 304A. Here, the data augmentation module 360 augments the sequence of acoustic frames 304 by masking one or more acoustic frames 304 within the sequence of acoustic frames. As will become apparent, the data augmentation module 360 does not apply time correction to the sequence of acoustic frames 304 to avoid the output of the unsupervised portion 302 from "collapsed" to a constant value. Other techniques applied by the data augmentation module 360 may include adding / injecting noise and / or adding reverberation to the labeled audio samples. One data augmentation technique involves injecting various environmental noises into the unlabeled audio samples 303 using multi-style training (MTR). Other data augmentation descriptions that the data augmentation module 360 may apply in addition to or instead of MTR include using spectral augmentation (SpecAugment) to approximate the acoustics of the augmented acoustic frame 304 to the inverse acoustics of other unlabeled audio samples 303. When combined, MTR and SpecAugment may inject noise into the unlabeled audio samples 303, tile a random external noise source previously inserted along time and superimposed on the representation, and filter the noise-injected unlabeled audio samples 303 before training the speech recognition model 200.
[0044] The unsupervised audio encoder 216 of the augmentation branch 320 receives the augmented sequence of acoustic frames 304A from the data augmentation module 360 and generates, at each of a plurality of output steps, a predicted high-order feature representation 218 for the corresponding augmented acoustic frame 304A. In particular, successive convolutional layers of the unsupervised audio encoder 216 receive an augmented acoustic frame 304A from the sequence of augmented acoustic frames 304A and generate outputs that the transformer layers use to generate the predicted high-order feature representation 218 for the corresponding augmented acoustic frame 304A. Thus, the unsupervised audio encoder 216 generates the predicted high-order feature representation 218 to match the corresponding target high-order feature representation 214 generated by the supervised audio encoder 212 at the corresponding output step.
[0045] The unsupervised portion 302 determines an unsupervised loss term 330 based on the target high-level feature representation 214 generated by the target branch 310 at the corresponding output step and the predicted high-level feature representation 218 generated by the augmentation branch 320 at the corresponding output step. In some examples, the unsupervised loss term 330 includes a contrast loss term expressed by:
number
[0046] In Equation 2, M contains the set of masked frame indices, K contains the set of distractor indices, and h t is the encoder output, and c t is the convolutional neural network output. In another example, the unsupervised loss term 330 includes a reconstruction loss term L1 or a cosine distance loss term. The unsupervised portion 302 may update the parameters of the speech recognition model 200 based on the unsupervised loss term 330 in collaboration with the supervised portion 301 updating the parameters of the speech recognition model 200 based on the supervised loss term 350, which is represented by:
number
[0047] In Equation 3, p represents the parameters of the speech recognition model 200, and L s (p) represents the supervised loss term 350, and L u (p) represents the unsupervised loss term 330. In particular, using the acoustic frame 304, the target branch 310 generates an expected representation (i.e., the target high-level feature representation 214) based on the current state of the supervised audio encoder 212, and the augmentation branch 320 aims to match the expected representation using the augmented acoustic frame 304A.
[0048] In some examples, the unsupervised portion 302 determines a distance-based loss term 370 between the parameters 213 of the supervised audio encoder 212 and the parameters 217 of the unsupervised audio encoder 216 at each of multiple output steps. The distance-based loss term 370 may include an L2 loss. In scenarios where the supervised loss term 350 and the unsupervised loss term 330 are inconsistent and have uncorrelated or negatively correlated gradients, the parameters 213 of the supervised audio encoder 212 and the parameters 217 of the unsupervised audio encoder 216 are believed to stabilize the training of the speech recognition model 200. Therefore, the cross-training network 300 aims to make the parameters 213, 217 the same during training according to the following:
number
[0049] In Equation 4, p s represents the parameters 213 of the supervised audio encoder 212, and p u represent the parameters 217 of the unsupervised audio encoder 216. Using the Lagrangian undetermined coefficients method, Equation 4 can be rewritten as:
number
[0050] Equation 5 represents the total loss used to train the speech recognition model 200, including the supervised loss term 350, the unsupervised loss term 330, and the distance-based loss term 370. Here, the cross-training network 300 may jointly update the parameters of the speech recognition model 200 based on the supervised loss term 350, the unsupervised loss term 330, and the distance-based loss term 370. Notably, by jointly minimizing these three losses, the cross-training network 300 does not force the parameters of the supervised portion 301 and the unsupervised portion 302 to be identical at every training step (e.g., output step). Instead, the parameters of the supervised portion 301 and the unsupervised portion 302 (e.g., parameters 213 of the supervised audio encoder 212 and parameters 217 of the unsupervised audio encoder 216) have the flexibility to differ during training, while the distance-based loss term 370 gradually reduces the distance between them at each output step. Thus, in Equation 5, λ represents a knowledge transfer parameter, whereby λ=0 results in independent training of the supervised portion 301 and the unsupervised portion 302 such that the training is completely stable but does not exploit unlabeled data for the supervised portion. On the other hand, a large value of λ forces the supervised portion 301 and the unsupervised portion 302 to be identical, thereby making the training unstable. The transfer parameter may be set to any value. Updating the parameters of the speech recognition model 200 may include jointly updating the parameters 213 of the supervised audio encoder 212 and the parameters 217 of the unsupervised audio encoder 216.
[0051] Other implementations of semi-supervised training rely on a time correction and prediction network to prevent the speech recognition model's output from "folding" to a constant value during training. Advantageously, training the speech recognition model 200 jointly with the supervised loss term 350, unsupervised loss term 330, and distance-based loss term 370 enables the cross-training network 300 to exploit both unlabeled audio samples 303 and labeled audio samples 305 without the output "folding" to a constant value. Notably, the use of the distance-based loss term 370 avoids or mitigates the constant folding without the computationally expensive operations of the time correction and prediction network used by other semi-supervised training implementations. In particular, the cross-training network 300 improves the stability of the speech recognition model 200 when the labeled audio samples 305 are relatively small, the speech recognition model 200 uses cascaded encoders in a streaming fashion, and there is a mismatch between the unlabeled audio samples 303 and the labeled audio samples 305. Furthermore, the cross-training network 300 may initialize the parameters 213 of the supervised audio encoder and the parameters 217 of the unsupervised audio encoder 216 with the same initial parameters or with different transcription parameters without adversely affecting the resulting speech recognition performance. After the cross-training network 300 trains the speech recognition model 200, the speech recognition model 200 may be run on the supervised audio encoder 212.
[0052] 4 is a flowchart of an exemplary arrangement of operations for a computer-implemented method 400 for training a speech recognition model using the cross-training model 300. At operation 402, the method 400 includes receiving a sequence of acoustic frames 304 extracted from unlabeled audio samples 303 corresponding to a speech utterance 106 that has not been paired with any transcription. The target branch 310 of the cross-training network 300 includes the supervised audio encoder 212 of the speech recognition model 200. At operation 404, the method 400 includes generating a target high-order feature representation 214 of a corresponding acoustic frame 340 in the sequence of acoustic frames 304 using the supervised audio encoder 212 in multiple output steps. In the expansion branch 320 of the cross-training network 300, the method 400 performs operations 406 and 408. At operation 406, the method 400 also includes extending the sequence of acoustic frames 304 extracted from the unlabeled audio samples 303 by masking one or more acoustic frames 204 in the sequence of acoustic frames 304. At operation 408, the method 400 includes, at each of a plurality of output steps, generating a predicted high-order feature representation 218 of a corresponding extended acoustic frame 304A in the sequence of extended acoustic frames 304A as output from the unsupervised audio encoder 216 of the speech recognition model 200.
[0053] At operation 410, the method 400 includes determining, at each of the plurality of output steps, an unsupervised loss term 330 based on the target high-order feature representation 214 generated by the target branch 310 at the corresponding output step and the predicted high-order feature representation 218 generated by the augmentation branch 320 at the corresponding output step. At operation 412, the method 400 includes updating parameters of the speech recognition model 200 based on the unsupervised loss term 330 determined at each of the plurality of output steps. Here, updating the parameters of the speech recognition model 200 may include updating parameters of the supervised audio encoder 212 jointly with updating parameters of the unsupervised audio encoder 216.
[0054] 5 is a schematic diagram of an exemplary computing device 500 that may be used to implement the systems and methods described herein. Computing device 500 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown here, their connections and relationships, and their functionality are for illustrative purposes only and are not intended to limit the scope of the invention as described and / or claimed herein.
[0055] Computing device 500 includes a processor 510, memory 520, a storage device 530, a high-speed interface / controller 540 connecting to memory 520 and a high-speed expansion port 550, and a low-speed interface / controller 560 connecting to a low-speed bus 570 and storage device 530. Each of the components 510, 520, 530, 540, 550, and 560 is interconnected using various buses and may be mounted on a common motherboard or otherwise mounted as needed. Processor 510 processes instructions for execution within computing device 500, including instructions stored in memory 520 or storage device 530, and can display graphical information for a graphical user interface (GUI) on an external input / output device, such as a display 580 connected to high-speed interface 540. In other implementations, multiple processors and / or multiple buses may be used as needed, in conjunction with multiple memories and memory types. Also, multiple computing devices 500 may be connected, each providing a portion of the required operations (e.g., as a server bank, a group of blade servers, or a multiprocessor system).
[0056] The memory 520 stores information non-transiently within the computing device 500. The memory 520 may be a computer-readable medium, a volatile memory unit(s), or a non-volatile memory unit(s). The non-transient memory 520 may be a physical device used to temporarily or permanently store programs (e.g., instruction sequences) or data (e.g., program state information) for use by the computing device 500. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.
[0057] Storage device 530 can provide mass storage for computing device 500. In some embodiments, storage device 530 is a computer-readable medium. In various different implementations, storage device 530 may be a device array including a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid-state memory device, or a device in a storage area network or other configuration. In further embodiments, a computer program product is tangibly embodied on an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable or machine-readable medium, such as memory 520, storage device 530, or memory on processor 510.
[0058] High-speed controller 540 manages bandwidth-intensive operations for computing device 500, and low-speed controller 560 manages low-bandwidth-intensive operations. This allocation of roles is merely exemplary. In some implementations, high-speed controller 540 is coupled to memory 520 (e.g., via a graphics processor or accelerator), to display 580, and to high-speed expansion port 550, which can accept various expansion cards (not shown). In some implementations, low-speed controller 560 is coupled to storage device 530 and low-speed expansion port 590. Low-speed expansion port 590, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), may be connected, for example, via a network adapter, to one or more input / output devices, such as a keyboard, pointing device, scanner, or network devices, such as a switch or router.
[0059] The computing device 500, as shown, can be implemented in many different forms. For example, it may be implemented as a standard server 500a, or multiple times within a group of such servers 500a, as a laptop computer 500b, or as part of a rack server system 500c.
[0060] Various implementations of the systems and techniques described herein may be realized in digital electronic and / or optical circuitry, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementation in one or more computer programs executable and / or interpretable by a programmable system including at least one programmable processor, which may be special purpose or general purpose, coupled to receive data and instructions from and transmit data and instructions to a storage system, at least one input device, and at least one output device.
[0061] These computer programs (also known as programs, software, software applications, or code) include machine instructions for a programmable processor and may be implemented in a high-level procedural and / or object-oriented programming language, and / or assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0062] The processes and logic flows described herein may be implemented by one or more programmable processors, also referred to as data processing hardware, executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be implemented by special-purpose logic circuitry, such as an FPGA (field-programmable gate array) or an ASIC (application-specific integrated circuit). Processors suitable for executing computer programs include, by way of example, both general-purpose and special-purpose microprocessors and any one or more processors of any kind of digital computer. Generally, a processor receives instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer also includes one or more mass storage devices, such as magnetic, magneto-optical, or optical disks, for storing data, or is operably coupled to receive data from or transmit data to them, or both. A computer need not, however, have such devices. Computer-readable media suitable for storing computer program instructions and data include all types of non-volatile memory, media, and memory devices, including, by way of example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, magnetic disks such as internal or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0063] To interact with a user, one or more aspects of the present disclosure can be implemented in a computer having a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, or a touch screen, for displaying information to the user, and optionally a keyboard and pointing device, such as a mouse or trackball, through which the user can provide input to the computer. Other types of devices may also be used to provide interaction with the user; for example, feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input from the user may be received in any form, including acoustic input, voice input, or tactile input. Additionally, a computer can interact with a user by sending and receiving documents to a device used by the user, for example, by sending a web page to a web browser on the user's client device in response to a request received from the web browser.
[0064] Although several embodiments have been described, it will be understood that various modifications can be made without departing from the spirit and scope of the disclosure. Accordingly, other embodiments are within the scope of the following claims.
Claims
1. 1. A cross-training network (300) for training a speech recognition model (200), the cross-training network (300) comprising an unsupervised sub-network (302) trained with a plurality of unlabeled audio samples (303) corresponding to speech utterances (106) that have not been paired with a corresponding transcription (120), the unsupervised sub-network (302) comprising: A target branch (310), receiving a sequence of acoustic frames (104) extracted from the unlabeled audio samples (303) as input to a supervised audio encoder (212) of the speech recognition model (200); generating, at each of a plurality of output steps, a target high-level feature representation (214) for a corresponding acoustic frame (304) in the sequence of acoustic frames (104) input to the supervised audio encoder (212) at the corresponding output step; a target branch (310) configured to: An extension branch (320), augmenting the sequence of acoustic frames (104) extracted from the unlabeled audio samples (303) by masking one or more acoustic frames (104) in the sequence of acoustic frames (104); generating, at each of the plurality of output steps, a predicted high-level feature representation (218) for a corresponding extended acoustic frame (304) in a sequence of extended acoustic frames (104) as output from an unsupervised audio encoder (216) of the speech recognition model (200); an expansion branch (320) configured to: Equipped with The unsupervised sub-network (302) determining, at each of the plurality of output steps, an unsupervised loss term (330) based on the target high-level feature representation (214) generated by the target branch (310) at the corresponding output step and the predicted high-level feature representation (218) generated by the augmentation branch at the corresponding output step; updating parameters of the speech recognition model (200) based on the unsupervised loss terms (330) determined in each of the plurality of output steps; A cross-training network (300) configured to:
2. The cross-training network (300) of claim 1 , wherein the unsupervised loss term (330) comprises a contrastive loss term.
3. the unsupervised sub-network (302) is further configured to determine, at each of the plurality of output steps, a distance-based loss term between parameters of the unsupervised audio encoder (216) and parameters of the supervised audio encoder (212); updating the parameters of the speech recognition model (200) is further based on the distance-based loss term determined at each of the plurality of output steps. The cross-training network (300) of claim 1.
4. The cross-training network (300) of claim 3, wherein the distance-based loss term comprises an L2 loss.
5. 4. The cross-training network of claim 3, wherein updating the parameters of the speech recognition model based on the unsupervised loss term occurs jointly with updating the parameters of the speech recognition model based on the distance-based loss term.
6. The system further comprises a supervised sub-network (301) trained with a plurality of labeled audio samples (305) corresponding to the speech utterance (106) paired with a corresponding transcription (120), the supervised sub-network (301) comprising: In each of the plurality of output steps for each labeled audio sample (305), generating corresponding speech recognition results (342) for the labeled audio samples (305) using the speech recognition model (200); determining a supervised loss term (350) based on the corresponding speech recognition result (342) of the labeled audio sample (305) and the corresponding transcription (120) of the labeled audio sample (305); updating the parameters of the speech recognition model (200) based on the supervised loss term (350) determined for each labeled audio sample (305) in a plurality of labeled audio samples (305) at each of the plurality of output steps; The cross-training network (300) of claim 1 configured to:
7. 7. The cross-training network of claim 6, wherein the corresponding speech recognition result generated for the labeled audio sample using the speech recognition model comprises a probability distribution over possible speech recognition hypotheses for the labeled audio sample at the corresponding output step.
8. 7. The cross-training network of claim 6, wherein the supervised subnetwork is further configured to update the parameters of the speech recognition model based on the supervised loss term, in conjunction with the unsupervised network updating the parameters of the speech recognition model based on the unsupervised loss term and a distance-based loss term.
9. 2. The cross-training network of claim 1, wherein the target branch is further configured to apply a gradient stopping operation to the predicted high-dimensional feature representation of the corresponding extended acoustic frame.
10. The cross-training network (300) of claim 1 , wherein the parameters of the unsupervised audio encoder (216) and the parameters of the supervised audio encoder (212) are initialized with the same initial parameters.
11. The cross-training network (300) of claim 1 , wherein the parameters of the unsupervised audio encoder (216) and the parameters of the supervised audio encoder (212) are initialized with different initial parameters.
12. Each of the unsupervised audio encoder (216) and the supervised audio encoder (212) the respective full-context encoder, or Each cascade encoder, The cross-training network (300) of claim 1, comprising at least one of:
13. A computer-implemented method (400), when executed on data processing hardware (510), causing the data processing hardware (510) to: receiving a sequence of acoustic frames (104) extracted from unlabeled audio samples (303) corresponding to a speech utterance (106) that is not paired with a corresponding transcription (120); generating, in a target branch (310) of the cross-training network (300), target high-level feature representations (214) for corresponding acoustic frames (304) in the sequence of acoustic frames (104) using a supervised audio encoder (212) of the speech recognition model (200) at multiple output steps; In the extension branch (320) of the cross-training network (300), augmenting the sequence of acoustic frames (104) extracted from the unlabeled audio samples (303) by masking one or more acoustic frames (104) in the sequence of acoustic frames (104); generating, at each of the plurality of output steps, a predicted high-level feature representation (218) for a corresponding extended acoustic frame (304) in a sequence of extended acoustic frames (104) as output from an unsupervised audio encoder (216) of the speech recognition model (200); determining, at each of the plurality of output steps, an unsupervised loss term (330) based on the target high-level feature representation (214) generated by the target branch (310) at the corresponding output step and the predicted high-level feature representation (218) generated by the augmentation branch at the corresponding output step; updating parameters of the speech recognition model (200) based on the unsupervised loss terms (330) determined in each of the plurality of output steps; A computer-implemented method (400) for causing a computer to perform operations including:
14. 14. The computer-implemented method (400) of claim 13, wherein the unsupervised loss term (330) comprises a contrastive loss term.
15. The operation is At each of the plurality of outputting steps, further comprising determining a distance-based loss term between parameters of the unsupervised audio encoder (216) and parameters of the supervised audio encoder (212); updating the parameters of the speech recognition model (200) is further based on the distance-based loss term determined at each of the plurality of output steps. A computer-implemented method (400) according to claim 13 or 14.
16. 16. The computer-implemented method (400) of claim 15, wherein the distance-based loss term comprises an L2 loss.
17. 17. The computer-implemented method of claim 15 or 16, wherein updating the parameters of the speech recognition model based on the unsupervised loss term occurs jointly with updating the parameters of the speech recognition model based on the distance-based loss term.
18. The operation is receiving a plurality of labeled audio samples (305) corresponding to the speech utterance (106) paired with the corresponding transcription (120); In each of the plurality of output steps for each labeled audio sample (305), generating corresponding speech recognition results (342) for the labeled audio samples (305) using the speech recognition model (200); determining a supervised loss term (350) based on the corresponding speech recognition result (342) of the labeled audio sample (305) and the corresponding transcription (120) of the labeled audio sample (305); updating the parameters of the speech recognition model (200) based on the supervised loss term (350) determined for each labeled audio sample (305) in a plurality of labeled audio samples (305) at each of the plurality of output steps; The computer-implemented method (400) of any one of claims 13 to 17, further comprising:
19. 20. The computer-implemented method of claim 18, wherein the corresponding speech recognition result generated for the labeled audio sample using the speech recognition model comprises a probability distribution over possible speech recognition hypotheses for the labeled audio sample at the corresponding output step.
20. 20. The computer-implemented method (400) of claim 18 or 19, wherein updating the parameters of the speech recognition model (200) based on the supervised loss term (350) occurs jointly with updating the parameters of the speech recognition model (200) based on the unsupervised loss term (330) and a distance-based loss term.
21. 21. The computer-implemented method (400) of any one of claims 13 to 20, wherein the operations further comprise applying a gradient stopping operation (314) to the predicted high-dimensional feature representation (218) of the corresponding augmented acoustic frame (304).
22. 22. The computer-implemented method (400) of any one of claims 13 to 21, wherein the parameters of the unsupervised audio encoder (216) and the parameters of the supervised audio encoder (212) are initialized with the same initial parameters.
23. 23. The computer-implemented method (400) of any one of claims 13 to 22, wherein the parameters of the unsupervised audio encoder (216) and the parameters of the supervised audio encoder (212) are initialized with different initial parameters.
24. Each of the unsupervised audio encoder (216) and the supervised audio encoder (212) the respective full-context encoder, or Each cascade encoder, The computer-implemented method (400) of any one of claims 13 to 23, comprising at least one of:
Citation Information
Patent Citations
Efficient Streaming Non-Recurrent On-Device End-to-End Model
US20220310062A1