Semi-supervised training scheme for speech recognition

Through semi-supervised training technology, cross-training network combined with unsupervised and supervised subnetworks, the target and enhanced high-order feature representation are generated, which solves the instability problem of speech recognition system when training data is insufficient, and improves the stability and accuracy of the model.

CN120239884APending Publication Date: 2025-07-01GOOGLE LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202380082328.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-14
Filing Date
2023-12-12
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The instability caused by the lack of labeled training data during training is caused by the lack of labeled training data, especially when the labeled training data set is small, the system uses a cascading encoder operation, or the label does not match the unlabeled data, and the performance is degraded.

Method used

Using semi-supervised training technology, cross-training network combines unsupervised subnetwork and supervised subnetwork, trains using unlabeled audio samples to generate target high-order feature representations and enhanced predicted high-order feature representations, determine unsupervised loss terms, and jointly update the parameters of the speech recognition model.

Benefits of technology

Improves the stability and accuracy of the speech recognition model during execution, especially in the case of limited label data or mismatch in data, reducing instability during training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120239884A_ABST
    Figure CN120239884A_ABST
Patent Text Reader

Abstract

A method (400) includes receiving a sequence of acoustic frames (104) extracted from untagged audio samples (303) that correspond to spoken utterances that are not paired with any corresponding transcriptions. The method also includes generating a target high-order feature representation (214) of the corresponding acoustic frame using a supervised audio encoder (212). The method also includes enhancing the sequence of acoustic frames, and generating a predicted higher order feature representation (218) of a corresponding enhanced acoustic frame in the sequence of enhanced acoustic frames as an output from the unsupervised audio encoder (216). The method further includes determining an unsupervised loss term (330) based on the target higher order feature representation and the predicted higher order feature representation, and updating a parameter of the speech recognition model (200) based on the unsupervised loss term.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to semi-supervised training schemes for speech recognition. Background Art

[0002] An automatic speech recognition (ASR) system attempts to provide an accurate transcription of what people say by obtaining an audio input and transcribing the audio input into text. In many cases, supervised learning is used to train an ASR system with a large amount of labeled training data that includes audio data and corresponding transcriptions. However, obtaining a large amount of labeled training data required to train an ASR system is often difficult because of the amount of time, cost, and / or privacy issues associated with collecting a large labeled training data set. Training an ASR system using only unlabeled training data that includes only audio data can alleviate some of the difficulties in collecting a large amount of labeled training data. Summary of the Invention

[0003] One aspect of the present disclosure provides a cross-training network for training a speech recognition model. The cross-training network includes an unsupervised sub-network trained on a plurality of unlabeled audio samples corresponding to spoken utterances not paired with corresponding transcriptions. The unsupervised sub-network includes a target branch configured to: receive a sequence of acoustic frames extracted from an unlabeled audio sample as an input to a supervised audio encoder of the speech recognition model; and at each of a plurality of output steps, generate a target high-order feature representation of a corresponding acoustic frame in the sequence of acoustic frames input to the supervised audio encoder at the corresponding output step. The unsupervised sub-network further includes an enhancement branch configured to: enhance the sequence of acoustic frames extracted from the unlabeled audio sample by masking one or more acoustic frames in the sequence of acoustic frames; and at each of a plurality of output steps, generate a predicted high-order feature representation of a corresponding enhanced acoustic frame in the enhanced sequence of acoustic frames as an output from an unsupervised audio encoder of the speech recognition model. The unsupervised sub-network is configured to: at each of a plurality of output steps, determine an unsupervised loss term based on the target high-order feature representation generated by the target branch at the corresponding output step and the predicted high-order feature representation generated by the enhancement branch at the corresponding output step; and update parameters of the speech recognition model based on the unsupervised loss term determined at each of a plurality of output steps.

[0004] Implementations of the present disclosure may include one or more of the following optional features. In some implementations, the unsupervised loss term includes a contrastive loss term. In some examples, the unsupervised subnetwork is further configured to, at each output step among a plurality of output steps, determine a distance-based loss term between the parameters of the unsupervised audio encoder and the parameters of the supervised audio encoder, and update the parameters of the speech recognition model further based on the distance-based loss term determined at each output step among the plurality of output steps. Here, the distance-based loss term may be an L2 loss. In these examples, updating the parameters of the speech recognition model based on the unsupervised loss term and updating the parameters of the speech recognition model based on the distance-based loss term occur jointly.

[0005] In some implementations, the cross-training network further includes a supervised subnetwork trained on a plurality of labeled audio samples corresponding to spoken utterances paired with corresponding transcriptions. In these implementations, at each output step among the plurality of output steps for each labeled sample, the supervised subnetwork is configured to: use the speech recognition model to generate a corresponding speech recognition result for the labeled audio sample; and determine a supervised loss term based on the target high-order feature representation generated by the target branch at the corresponding output step and the predicted high-order feature representation generated by the enhancement branch at the corresponding output step. Here, the supervised subnetwork updates the parameters of the speech recognition model based on the supervised loss term determined at each output step among the plurality of output steps for each labeled audio sample among the plurality of labeled audio samples. In these implementations, the corresponding speech recognition result generated for the labeled audio sample using the speech recognition model includes a probability distribution of possible speech recognition hypotheses for the labeled audio sample at the corresponding output step. The supervised subnetwork may further be configured to, based on the supervised loss term, jointly update the parameters of the speech recognition with the unsupervised network updating the parameters of the speech recognition model based on the unsupervised loss term and the distance-based loss term.

[0006] The target branch may further be configured to apply a stop-gradient operation to the predicted high-order feature representation of the corresponding enhanced acoustic frame. In some examples, the parameters of the unsupervised audio encoder and the parameters of the supervised audio encoder are initialized with the same initial parameters. In other examples, the parameters of the unsupervised audio encoder and the parameters of the supervised audio encoder are initialized with different initial parameters. Each of the unsupervised audio encoder and the supervised audio encoder includes at least one of a corresponding full-context encoder or a corresponding cascaded encoder.

[0007] Another aspect of the present disclosure provides a computer-implemented method that, when executed on data processing hardware, causes the data processing hardware to perform operations for training a speech recognition model using a cross-training network. The operations include receiving a sequence of acoustic frames extracted from unlabeled audio samples that correspond to spoken utterances not paired with any transcriptions. At a target branch of the cross-training network, the operations include generating, at a plurality of output steps, target high-order feature representations of corresponding acoustic frames in the sequence of acoustic frames using a supervised audio encoder of the speech recognition model. At an augmentation branch of the cross-training network, the operations include augmenting the sequence of acoustic frames extracted from the unlabeled audio samples by masking one or more acoustic frames in the sequence of acoustic frames; and generating, at each of the plurality of output steps, predicted high-order feature representations of corresponding augmented acoustic frames in the augmented sequence of acoustic frames as outputs from an unsupervised audio encoder of the speech recognition model. The operations further include determining, at each of the plurality of output steps, an unsupervised loss term based on the target high-order feature representation generated by the target branch at the corresponding output step and the predicted high-order feature representation generated by the augmentation branch at the corresponding output step. The operations further include updating parameters of the speech recognition model based on the unsupervised loss terms determined at each of the plurality of output steps.

[0008] Implementations of the present disclosure may include one or more of the following optional features. In some implementations, the unsupervised loss term includes a contrastive loss term. In some examples, the operations further include determining, at each of the plurality of output steps, a distance-based loss term between the parameters of the unsupervised audio encoder and the parameters of the supervised audio encoder, and updating the parameters of the speech recognition model is further based on the distance-based loss terms determined at each of the plurality of output steps. Here, the distance-based loss term may include an L2 loss. In these examples, updating the parameters of the speech recognition model based on the unsupervised loss terms and updating the parameters of the speech recognition model based on the distance-based loss terms occur jointly.

[0009] In some implementations, the operation further includes receiving a plurality of tagged audio samples corresponding to an oral utterance paired with a corresponding transcription. In these implementations, for each tagged audio sample at each of the plurality of output steps, the operation further includes: using a speech recognition model to generate a corresponding speech recognition result for the tagged audio sample, and determining a supervised loss term based on the corresponding speech recognition result for the tagged audio sample and the corresponding transcription for the tagged audio sample. Here, the operation further includes: updating the parameters of the speech recognition model based on the supervised loss terms determined for each of the tagged audio samples among the plurality of tagged audio samples at each of the plurality of output steps. The corresponding speech recognition result generated for the tagged audio sample using the speech recognition model may include a probability distribution of possible speech recognition hypotheses for the tagged audio sample at the corresponding output step. Updating the parameters of the speech recognition model based on the supervised loss terms may occur jointly with updating the parameters of the speech recognition model based on unsupervised loss terms and distance-based loss terms.

[0010] In some examples, the operation further includes: applying a stop-gradient operation to a predicted higher-order feature representation of a corresponding enhanced acoustic frame. In some implementations, the parameters of the unsupervised audio encoder and the parameters of the supervised audio encoder are initialized with the same initial parameters. In other implementations, the parameters of the unsupervised audio encoder and the parameters of the supervised audio encoder are initialized with different initial parameters. Each of the unsupervised audio encoder and the supervised audio encoder may include at least one of a corresponding full-context encoder or a corresponding cascaded encoder.

[0011] Details of one or more implementations of the present disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is a schematic diagram of a speech environment for implementing an example speech recognition model.

[0013] Figure 2 is Figure 1 a schematic diagram of an example speech recognition model of

[0014] Figure 3A is a schematic diagram of a supervised portion of a cross-training network that performs a semi-supervised training process of a speech recognition model.

[0015] Figure 3B is a schematic diagram of an unsupervised portion of a cross-training network that performs a semi-supervised training process of a speech recognition model.

[0016] Figure 4A flowchart of an example operational arrangement for a method of training a speech recognition model using a cross-training model.

[0017] Figure 5 A schematic diagram of an example computing device that can be used to implement the systems and methods described herein.

[0018] In the various figures, like reference numerals indicate like elements. Detailed Description

[0019] Automatic speech recognition (ASR) systems are often trained using supervised training techniques that utilize labeled training data. The labeled training data includes speech audio data and corresponding transcripts of the speech. Because of the associated costs, the time required to collect training data, and privacy issues of users, it is often difficult to collect large amounts of labeled training data. In some cases, ASR systems are trained using unlabeled training data that includes only speech audio data without any corresponding transcripts of the speech. In these cases, the ASR system can train the speech recognition system using only the unlabeled training data (i.e., self-supervised training), or can use the unlabeled training data in addition to the labeled training data to train the speech recognition system (i.e., semi-supervised training).

[0020] However, training a speech recognition model using self-supervised training often results in instability during execution. That is, although these self-supervised training systems use computationally expensive operations such as prediction networks, time modification (e.g., tempo change), and joint training, the trained ASR system is still unstable during execution. Specifically, semi-supervised training results in particularly poor performance in training scenarios that include: when the labeled training data set is relatively small, the ASR system employs cascaded encoders to operate in a streaming fashion, and / or there is a mismatch between the labeled training data and the unlabeled training data. Therefore, there is a need for semi-supervised training of ASR systems that produces stable outputs during execution.

[0021] Accordingly, implementations herein relate to a cross-training network that uses semi-supervised training techniques to train a speech recognition model. The cross-training network includes an unsupervised sub-network trained on a plurality of unlabeled audio samples corresponding to spoken utterances not paired with any corresponding transcripts. The unsupervised sub-network includes a target branch configured to receive a sequence of acoustic frames extracted from the unlabeled audio samples as input to a supervised audio encoder. The target branch is further configured to generate a target high-order feature representation of the corresponding acoustic frames in the sequence of acoustic frames.

[0022] The unsupervised sub-network further includes an enhancement branch, which is configured to enhance a sequence of acoustic frames extracted from an unlabeled audio sample by masking one or more acoustic frames in the sequence of acoustic frames. The enhancement branch is further configured to generate a predicted high-order feature representation of corresponding enhanced acoustic frames in the sequence of enhanced acoustic frames as an output from the unsupervised audio encoder. Here, the unsupervised sub-network determines an unsupervised loss term based on the target high-order feature representation and the predicted high-order feature representation, and updates the parameters of the speech recognition model based on the unsupervised loss term. As will become apparent, the unsupervised sub-network may also determine a distance-based loss between the parameters of the unsupervised audio encoder and the parameters of the supervised audio encoder. Here, the unsupervised sub-network jointly updates the parameters of the speech recognition based on the unsupervised loss term and the distance-based loss term.

[0023] Figure 1 is an example of the speech environment 100. In the speech environment 100, the way the user 104 interacts with a computing device such as the user device 10 may be through voice input. The user device 10 is configured to capture the voices of one or more users 104 within the speech environment 100 (e.g., stream audio data). Here, streaming audio data may refer to the spoken utterance 106 of the user 104, which serves as an audible query, a command for the user device 10, or an audible communication captured by the user device 10. The voice-enabled system of the user device 10 may process the query or command by answering the query and / or causing the command to be executed / fulfilled by one or more downstream applications.

[0024] The user device 10 may correspond to any computing device associated with the user 104 and capable of receiving audio data. Some examples of the user device 10 include, but are not limited to, mobile devices (e.g., mobile phones, tablet computers, laptop computers, etc.), computers, wearable devices (e.g., smartwatches), smart home appliances, Internet of Things (IoT) devices, vehicle infotainment systems, smart displays, smart speakers, etc. The user device 10 includes data processing hardware 12 and memory hardware 14 that communicates with the data processing hardware 12 and stores instructions that, when executed by the data processing hardware 12, cause the data processing hardware 12 to perform one or more operations. The user device 10 also includes an audio system 16 having: audio capture devices (e.g., microphones) 16, 16a that are configured to capture spoken utterances 106 within the voice environment 100 and convert them into electrical signals; and voice output devices (e.g., speakers) 16, 16b that are configured to transmit audible audio signals (e.g., as output audio data from the user device 10). Although the user device 10 implements a single audio capture device 16a in the illustrated example, the user device 10 may implement an array of audio capture devices 16a without departing from the scope of the present disclosure, where one or more of the capture devices 16a in the array may not physically reside on the user device 10 but communicate with the audio system 16.

[0025] In the voice environment 100, an automatic speech recognition (ASR) system 118 that implements the speech recognition model 200 resides on the user device 10 of the user 104 and / or on a remote computing device 60 (e.g., one or more remote servers of a distributed system executed in a cloud computing environment) that communicates with the user device 10 via the network 40. The user device 10 and / or the remote computing device (i.e., remote server) 60 also includes an audio subsystem 108 that is configured to receive the utterance 106 spoken by the user 104 and captured by the audio capture device 16a, and convert the utterance 106 into a corresponding digital format associated with an input acoustic frame 110 that can be processed by the ASR system 118. In the illustrated example, the user speaks the corresponding utterance 106, and the audio subsystem 108 converts the utterance 106 into corresponding audio data (e.g., acoustic frames) 110 for input into the ASR system 118. Thereafter, the speech recognition model 200 receives the audio data 110 corresponding to the utterance 106 as input and generates / predicts a corresponding transcription 120 (e.g., speech recognition result / hypothesis) of the utterance 106 as output.

[0026] The digital assistant application 50 executed on the user device 10 may require that speech recognition be streamed such that words, word fragments, and / or individual characters appear on the screen as soon as they are spoken. Additionally, it may be possible that the user 104 of the user device 10 has a low tolerance for latency when issuing a query for the digital assistant application 50 to execute. In such scenarios, when it is preferred to minimize speech recognition latency, the speech recognition model 200 may apply zero or minimal look-ahead audio contexts (also referred to as "correct contexts") to provide streaming transcription capabilities in real time as the user 104 is speaking the utterance 106. On the other hand, when the user has a higher tolerance for speech recognition latency and / or the utterance 106 to be recognized is associated with long-form speech, the same speech recognition model 200 may apply a duration of look-ahead audio context sufficient to provide an accurate transcription 120, but incur an increased latency based on the duration of the look-ahead audio context. Thus, the ASR system 118 may implement only a single speech recognition model 200 for a number of different speech recognition tasks to provide both streaming transcription capabilities and non-streaming transcription capabilities, without having to utilize separate ASR models on a per-task basis.

[0027] In some implementations, the speech recognition model 200 performs both streaming speech recognition and non-streaming speech recognition on the audio data 110 in parallel. For example, in the illustrated example, the speech recognition model 200 performs streaming speech recognition on the audio data 110 in parallel to produce partial speech recognition results 120, 120a, and performs non-streaming speech recognition on the same audio data 110 to produce final speech recognition results 120, 120b. Notably, the speech recognition model 200 may use a first look-ahead audio context that can be set to zero (or approximately 240 milliseconds) to produce the partial speech recognition results 120a, and use a second look-ahead audio context having a longer duration than the first look-ahead audio context to produce the final speech recognition results 120b. Thus, the final speech recognition results 120b for the input utterance 106 may be delayed by a duration from the partial speech recognition results 120a for the input utterance based on the difference between the second look-ahead audio context and the first look-ahead audio context.

[0028] The user device 10 and / or the remote computing device 60 also execute a user interface generator 107, which is configured to present a representation of a transcription 120 of the utterance 106 to the user 104 of the user device 10. As described in more detail below, the user interface generator 107 may display partial speech recognition results 120a in a streaming manner during time 1 and then display the final speech recognition results 120b during time 2. In some configurations, the transcription 120 output from the ASR system 118 is processed, for example, by a natural language understanding (NLU) module executed on the user device 10 or the remote computing device 60 to execute the user command / query specified by the utterance 106. Additionally or alternatively, a text-to-speech system (not shown) (e.g., executed on any combination of the user device 10 or the remote computing device 60) may convert the transcription into synthetic speech for audible output by the user device 10 and / or another device.

[0029] In the example shown, the user 104 interacts with a program or application 50 (e.g., a digital assistant application 50) of the user device 10 that uses the ASR system 118. For example, Figure 1 depicts the user 104 communicating with the digital assistant application 50 and the digital assistant application 50 displaying a digital assistant interface 18 on the screen of the user device 10 for depicting the conversation between the user 104 and the digital assistant application 50. In this example, the user 104 asks the digital assistant application 50 "What time is the concert tonight?". This question from the user 104 is an oral utterance 106, which is captured by the audio capture device 16a and processed by the audio system 16 of the user device 10. In this example, the audio system 16 receives the oral utterance 106 and converts it into acoustic frames 110 for input into the ASR system 118.

[0030] Continuing with this example, when the speech recognition model 200 receives the acoustic frames (i.e., audio data) 110 corresponding to the utterance 106 while the user 104 is speaking, it encodes the acoustic frames 110 using a first look-ahead audio context and then decodes the encoded acoustic frames 110 into partial speech recognition results 120a using the first look-ahead audio context. During time 1, the user interface generator 107 presents a representation of the partial speech recognition results 120a of the utterance 106 to the user 104 of the user device 10 in a streaming manner via the digital assistant interface 18 such that words, word pieces, and / or individual characters appear on the screen as soon as they are spoken. In some examples, the first look-ahead audio context is equal to zero.

[0031] In parallel, and after receiving all the acoustic frames in the acoustic frame 110 corresponding to the utterance 106, the speech recognition model 200 encodes all the acoustic frames in the acoustic frame 110 corresponding to the utterance 106 using the second look-ahead audio context, and then decodes the acoustic frame 110 into the final speech recognition result 120b using the second look-ahead audio context. The duration of the second look-ahead audio context can be 1.2 seconds, 2.4 seconds, or any other duration. In some examples, an indication such as an indication of the end point indicating that the user 104 has finished speaking the utterance 106 triggers the speech recognition model 200 to encode all the acoustic frames 110 using the second look-ahead audio context. During time 2, the user interface generator 107 presents a representation of the final speech recognition result 120b of the utterance 106 to the user 104 of the user device 10 via the digital assistant interface 18. In some implementations, the user interface generator 107 replaces the representation of the partial speech recognition result 120a with the representation of the final speech recognition result 120b. For example, since it is assumed that the final speech recognition result 120b is more accurate than the partial speech recognition result 120a generated without using the look-ahead audio context, the final speech recognition result 120b, which is finally displayed as the transcription 120, can correct any terms that may have been misrecognized in the partial speech recognition result 120a. In this example, the streaming partial speech recognition result 120a output by the speech recognition model 200 and displayed on the screen of the user device 10 at time 1 is associated with low latency and provides the user 104 with a response that his / her query is being processed, while the final speech recognition result 120b output by the speech recognition model 200 and displayed on the screen at time 2 uses the look-ahead audio context to improve the speech recognition quality in terms of accuracy, but the latency increases. However, since the partial speech recognition result 120a is displayed when the user speaks the utterance 106, the higher latency associated with generating and finally displaying the final recognition result is not noticeable to the user 104.

[0032] At Figure 1In the example shown, digital assistant application 50 may use natural language processing to respond to questions posed by user 104. Natural language processing generally refers to the process of interpreting written language (e.g., partial speech recognition result 120a and / or final speech recognition result 120b) and determining whether the written language suggests any actions. In this example, digital assistant application 50 uses natural language processing to recognize that the question from user 104 pertains to the user's schedule, and more specifically to a concert on the user's schedule. By recognizing these details with natural language processing, the automated assistant returns a response 19 to the user's query, where the response 19 states "Venue doors open at 6:30 PM and concert starts at 8pm". In some configurations, natural language processing occurs on a remote server 60 that communicates with the data processing hardware 12 of user device 10.

[0033] Reference Figure 2 , the speech recognition model 200 can provide end-to-end (E2E) speech recognition by integrating acoustic, pronunciation, and language models into a single neural network and does not require a lexicon or separate text normalization components. Various architectures and optimization mechanisms can provide increased accuracy and reduced model training time. In some implementations, the speech recognition model 200 can include a Transformer-Transducer (T-T) model architecture that adheres to the latency constraints associated with interactive applications. The T-T model 200 can include the T-T model 200 described in U.S. Patent Application 17 / 210,465 filed on March 23, 2021, the content of which is incorporated herein by reference in its entirety. The T-T model 200 provides a small computational footprint and utilizes fewer memory requirements than conventional ASR architectures, making the T-T model architecture suitable for performing speech recognition entirely on user device 10 (e.g., without the need to communicate with remote server 60). The T-T model 200 includes an audio encoder 210, a label encoder 220, and a joint network 230. The audio encoder 210, which is roughly analogous to the acoustic model (AM) in a traditional ASR system, can include a neural network having a stack of strided convolutional layers and transformer layers. Additionally, the audio encoder 210 can include a supervised audio encoder 212 (FIG. 3) and an unsupervised audio encoder 216 (FIG. 3). The audio encoder 210 reads a sequence of d-dimensional feature vectors (e.g., acoustic frames 110 ( Figure 1 )) x = (x1, x2,... , x T ), where x t ∈R d, and generates a high-order feature representation (also referred to as "encoder output") at each time step. This high-order feature representation is denoted as ah1, ..., ah T . Each transformer layer of the audio encoder 210 may include a normalization layer, a masked multi-head attention layer with relative position encoding, a residual connection, a stack / unstack layer, and a feed-forward layer. Similarly, the label encoder 220 may also include a neural network of transformer layers or a lookup table embedding model, which, like a language model (LM), processes the sequence of non-empty symbols y0, . . . , y output by the final Softmax layer 240 so far ui−1 into a dense representation Ih encoding the predicted label history u . In an implementation where the label encoder 220 includes a neural network of transformer layers, each transformer layer may include a normalization layer, a masked multi-head attention layer with relative position encoding, a residual connection, a feed-forward layer, and a dropout layer. In these implementations, the label encoder 220 may include two transformer layers. In an implementation where the label encoder 220 includes a lookup table embedding model with a tuple label context, the embedding model is configured to learn a d-dimensional weight vector for each possible tuple label context, where d is the dimension of the outputs of the audio encoder 210 and the label encoder 220. In some examples, the total number of parameters in the embedding model is , where N is the vocabulary size of the labels. Here, the learned weight vectors are then used as the embedding of the tuple label context in the T-T model 200 to produce the runtime of the fast label encoder 220.

[0034] Finally, using the T-T model architecture, the joint network 230 combines the representations generated by the audio encoder 210 and the label encoder 220 using a dense layer J u,t . Then, the joint network 230 predicts , which is the distribution over the next output symbol. In other words, the joint network 230 generates a probability distribution over possible speech recognition hypotheses at each output step (e.g., time step). Here, the "possible speech recognition hypotheses" correspond to a set of output labels (also referred to as "speech units"), where each output label represents a grapheme (e.g., symbol / character) or word piece in a specified natural language. For example, when the natural language is English, the set of output labels can include twenty-seven (27) symbols, e.g., one label for each of the 26 letters in the English alphabet, and one label specifying a space. Thus, the joint network 230 can output a set of values that indicate the likelihood of each output label in a predetermined set of output labels occurring. The set of values can be a vector (e.g., one-hot vector) and can indicate a probability distribution over the set of output labels. In some cases, the output labels are graphemes (e.g., individual characters, and possibly punctuation marks and other symbols), but the set of output labels is not limited to this. For example, in addition to or instead of graphemes, the set of output labels can include word pieces and / or whole words. The output distribution of the joint network 230 can include posterior probability values for each of the different output labels. Thus, if there are 100 different output labels representing different graphemes or other symbols, the output z of the joint network 230 u,t can include 100 different probability values, one probability value for each output label. Then, the probability distribution can be used (e.g., by the Softmax layer 240) to select candidate orthographic elements (e.g., graphemes, word pieces, and / or words) during a beam search process and assign scores to them for determining the transcription 120.

[0035] The Softmax layer 240 can employ any technique to select the output label / symbol with the highest probability in the distribution as the next output symbol predicted by the T-T model 200 at the corresponding output step. In this way, the T-T model 200 does not make a conditional independence assumption, but rather the prediction of each symbol is conditioned not only on the acoustics but also on the sequence of labels output so far. Although the speech recognition model 200 is described as having a T-T model architecture, the speech recognition model 200 can include other types of transducer-based architectures, such as the Conformer-Transducer (C-T) model architecture or the Recurrent Neural Network Transducer (RNN-T) model architecture.

[0036] Figure 3A and Figure 3B shows a schematic diagram of a cross-training network 300 that performs a semi-supervised training process for training the speech recognition model 200 ( Figure 2 )). The cross-training network includes a supervised sub-network training process 301 (Figure 3A ) and the unsupervised subnetwork training process 302 ( Figure 3B ). The supervised subnetwork training process (i.e., the supervised subnetwork) 301 uses a plurality of labeled audio samples 305 to train the speech recognition model 200. The plurality of labeled audio samples includes a sequence of acoustic frames 304 extracted from spoken utterances 106 paired with corresponding transcripts (i.e., labels) 308. The unsupervised subnetwork training process (i.e., the unsupervised subnetwork) 302 uses a plurality of unlabeled audio samples 303 to train the speech recognition model 200. The plurality of unlabeled audio samples includes a sequence of acoustic frames 304 extracted from spoken utterances 106 without any paired transcripts.

[0037] In some examples, the acoustic frames 306 used by the supervised subnetwork (i.e., the supervised portion) 301 are the same as the acoustic frames 304 used by the unsupervised subnetwork (i.e., the unsupervised portion) 302. That is, the supervised portion 301 and the unsupervised portion 302 can concurrently use the same acoustic frames 304, 306 to train the speech recognition model 200. In other examples, the acoustic frames 306 used to train the supervised portion 301 are different from the acoustic frames 304 used to train the unsupervised portion 302. This scenario is particularly beneficial because unlabeled audio samples 303 without any corresponding transcripts are readily available and can be utilized to train the speech recognition model 200. Thus, the speech recognition model 200 can be trained on any combination of labeled audio samples 305 and / or unlabeled audio samples 303. In some examples, the sequences of acoustic frames 304, 306 extracted from unlabeled audio samples 303 and labeled audio samples 305 include log Mel filter bank energies. A greater number of acoustic frames 304 than the number of acoustic frames 306 used to train the supervised portion 301 can be used to train the unsupervised portion 302. Optionally, a greater number of acoustic frames 306 than the number of acoustic frames 304 used to train the unsupervised portion 302 can be used to train the supervised portion 301. In some examples, the number of acoustic frames 306 used to train the supervised portion 301 is the same as the number of acoustic frames 304 used to train the unsupervised portion 302.

[0038] The supervision part 301 includes a supervised audio encoder 212 shared by the speech recognition model 200 and the target branch 310 of the unsupervised part 302. The unsupervised part 302 also includes an unsupervised audio encoder 216 different from the supervised audio encoder 212 at the enhancement branch 320. Here, the supervised audio encoder 212 and the unsupervised audio encoder 216 can each include a stack of strided convolutional layers (e.g., two convolutional layers) and transformer layers (e.g., twenty bidirectional transformer layers). In some implementations, the supervised audio encoder 212 and the unsupervised audio encoder 216 each include a corresponding full-context encoder that operates in a non-streaming manner. Here, the full-context encoder outputs an encoder output corresponding to the final speech recognition result 120b ( Figure 1 ) at each output step among multiple output steps. In other implementations, the supervised audio encoder 212 and the unsupervised audio encoder 216 each include a corresponding cascaded encoder that operates in a streaming manner. That is, the cascaded encoder includes a causal encoder that does not receive any correct context and outputs an encoder output corresponding to the partial speech recognition result 120a ( Figure 1 ) at each output step among multiple output steps. In addition, the cascaded encoder includes a non-causal encoder that receives additional correct context and outputs an encoder output corresponding to the final speech recognition result 120b ( Figure 1 ) at each output step among multiple output steps.

[0039] Now referring to Figure 3A , the supervision part 301 of the cross-training network 300 uses multiple labeled audio samples 305 to train the speech recognition model 200. Each of the multiple labeled audio samples 305 in the multiple labeled audio samples 305 corresponds to an oral utterance 106 paired with a corresponding transcription 308. The multiple labeled audio samples 305 include a sequence of acoustic frames 306 extracted from the labeled audio samples 305. The supervision part 301 shares the same supervised audio encoder 212 from the speech recognition model 200 with the target branch 310 of the unsupervised part 302. The supervision part 301 also includes a label encoder 220 and a dense layer 346 (e.g., the joint network 230) with a bias vector 347.

[0040] Optionally, the supervision part 301 can include a data augmentation module 365 (represented by a dashed line) that applies data augmentation to at least one acoustic frame 306 extracted from the labeled audio samples 305 to generate a sequence of enhanced acoustic frames 306, 306A. The data augmentation module 365 of the supervision part 301 can be the same as the data augmentation module 360 of the unsupervised part 302 ( Figure 3B) are the same (or different). In some examples, the data augmentation module 365 of the supervised portion 301 applies a data augmentation technique different from that of the data augmentation module 360 of the unsupervised portion. Applying data augmentation to the acoustic frames 306 promotes the acoustic diversity of the audio frames for training the speech recognition model 200. The data augmentation module 360 may include a temporal masking component that masks portions of the acoustic frames 306. Other techniques applied by the data augmentation module 360 may include adding / injecting noise and / or adding reverberation to the labeled audio samples 305. One data augmentation technique includes injecting various environmental noises into the labeled audio samples 305 using multi-style training (MTR). In addition to or instead of MTR, another data augmentation technique that the data augmentation module 360 may apply also includes using spectral augmentation (SpecAugment) to make the acoustics of the labeled audio samples 305 closer to the adverse acoustics of other labeled audio samples 305. Combinatorially, MTR and SpecAugment can inject noise into the labeled audio samples 305, tile random external noise sources along time and inserted and overlapped onto the representation before the representation, and filter the noise-injected labeled audio samples before training the speech recognition model 200.

[0041] In some examples, when the supervised portion 301 includes the data augmentation module 365, the supervised audio encoder 212 receives a sequence of augmented acoustic frames 306A and generates, at each output step, a higher-order feature representation 341 of the corresponding augmented acoustic frame 306A in the sequence of augmented acoustic frames 306A. In other examples, when the supervised portion 301 does not include the data augmentation module 365, the supervised audio encoder 212 directly receives the sequence of acoustic frames 306 and generates, at each output step, a higher-order feature representation 341 of the corresponding acoustic frame 306 in the sequence of acoustic frames 306. More specifically, the strided convolutional layer of the supervised audio encoder 212 receives the augmented acoustic frames 306A (or acoustic frames 306) and generates a corresponding output. Here, the transformer layer of the supervised audio encoder 212 receives the corresponding output and generates the higher-order feature representation 341.

[0042] The label encoder 220 is a streaming transformer that does not attend to future labels 308. Thus, the label encoder 220 receives the labels 308 corresponding to the augmented acoustic frames 306A (or acoustic frames 306) received by the supervised audio encoder 212 and generates, at each output step, a linguistic embedding 344 (i.e., a dense representation Ih u ( Figure 2))。The supervision part 301 includes a dense layer 346 that processes the linguistic embeddings 344 from the label encoder 220 and the high-order feature representations 341 (i.e., acoustic embeddings) from the supervised audio encoder 210 to generate corresponding speech recognition results 342 for the corresponding enhanced acoustic frames 306A (or acoustic frames 306) at each output step using the speech recognition model 200. The dense layer 346 includes a trainable bias vector 347 that performs a linear operation on the high-order feature representations 341 and the linguistic embeddings 344 to generate the speech recognition results 342. The speech recognition results 342 may include a probability distribution of possible speech recognition hypotheses for the labeled audio sample 305 at the corresponding output step. The loss module 351 of the supervision part 301 determines a supervision loss term 350 at each output step among a plurality of output steps based on the corresponding speech recognition results 342 of the labeled audio sample 305 and the corresponding transcription 308 of the labeled audio sample 305. That is, the loss module 351 compares the speech recognition results 342 with the labels (e.g., ground-truth transcriptions) 308 to generate the supervision loss term 350. The supervision loss term (e.g., RNN-T loss) 350 may be represented as follows:

[0043] In Equation 1, r t represents a log-odds vector specifying the probabilities of graphemes including the blank symbol, a t represents the high-order feature representations 341 from the supervised audio encoder 212, l t represents the linguistic embeddings 344 from the label encoder 220, and linear represents the regular dense layer 346 with the trainable bias vector 347.

[0044] The supervision part 301 updates the parameters of the speech recognition model 200 based on the supervision loss terms 350 determined for each of the plurality of labeled audio samples 305 at each output step among a plurality of output steps. In some implementations, the supervision part 301 is configured to update the parameters of the speech recognition model 200 independently of the unsupervised part 302 based on the supervision loss terms 350. In other implementations, the supervision part 301 is configured to update the parameters of the speech recognition model 200 jointly with the unsupervised part 302 based on the supervision loss terms 350. Updating the parameters of the speech recognition model 200 may include updating the parameters of the supervised audio encoder 212.

[0045] Now refer to Figure 3B, the unsupervised portion 302 uses a plurality of unlabeled audio samples 303 to train the speech recognition model 200, the plurality of unlabeled audio samples including sequences of acoustic frames extracted from spoken utterances 106 that have never been paired with any transcriptions. The unsupervised portion 302 of the cross-training network 300 includes a target branch 310 and an augmentation branch 320. Here, the target branch 310 shares the same supervised audio encoder 212 of the speech recognition model 200 with the supervised portion 301 ( Figure 3A ). The augmentation branch 320 includes an unsupervised audio encoder 216 of the speech recognition model 200. The unsupervised portion 302 is configured to extract linguistic information by matching high-order feature representations (e.g., target high-order feature representation 214 and predicted high-order feature representation 218) of the target branch 310 and the augmentation branch 320.

[0046] The target branch 310 is configured to generate a target high-order feature representation 214 based on a sequence of acoustic frames 304 extracted from the unlabeled audio samples 303. The supervised audio encoder 212 of the target branch receives the sequence of acoustic frames 304 and generates a corresponding target high-order feature representation 214 of the acoustic frames 304 at each of a plurality of output steps. Specifically, the strided convolutional layer of the supervised audio encoder 212 receives the acoustic frames 304 from the sequence of acoustic frames and generates an output that the transformer layer uses to generate the corresponding target high-order feature representation 214 of the acoustic frames.

[0047] The target branch 310 does not backpropagate gradients to train the supervised audio encoder 210 through the target branch 310. Specifically, training both the supervised audio encoder 212 and the unsupervised audio encoder 216 with a contrastive loss can cause the encoders to learn unrelated relationships by learning to minimize the contrastive loss by shifting positional information to the output of the encoders (i.e., the shortcut learning problem). Accordingly, the target branch 310 applies a stop gradient operation 314 to the target high-order feature representation 214 to prevent gradients (e.g., the contrastive loss) from backpropagating through the target branch 310 to the supervised audio encoder 212. Accordingly, applying the stop gradient operation 314 overcomes the shortcut learning problem.

[0048] The augmentation branch 320 of the unsupervised portion 302 includes a data augmentation module 360 that applies data augmentation to an acoustic sequence of acoustic frames 304 extracted from unlabeled audio samples 303. That is, the data augmentation module 360 receives a sequence of acoustic frames and generates a sequence of augmented acoustic frames 304A. Here, the data augmentation module 360 augments the sequence of acoustic frames 304 by masking one or more acoustic frames 304 in the sequence of acoustic frames. It will become apparent that the data augmentation module 360 does not apply temporal modifications to the sequence of acoustic frames 304 to avoid the output of the unsupervised portion 302 from "collapsing" to a constant value. Other techniques applied by the data augmentation module 360 can include adding / injecting noise and / or adding reverberation of labeled audio samples. One data augmentation technique includes injecting various environmental noises into the unlabeled audio samples 303 using multi-style training (MTR). In addition to or instead of MTR, another data augmentation technique that the data augmentation module 360 can apply also includes using spectral augmentation (SpecAugment) to make the acoustics of the augmented acoustic frames 304 closer to the adverse acoustics of other unlabeled audio samples 303. Combinatorially, MTR and SpecAugment can inject noise into the unlabeled audio samples 303, tile random external noise sources that are inserted and overlapped onto the representation along time and before the representation, and filter the noise-injected unlabeled audio samples 303 before training the speech recognition model 200.

[0049] The unsupervised audio encoder 216 of the augmentation branch 320 receives the sequence of augmented acoustic frames 304A from the data augmentation module 360 and generates a corresponding predicted high-order feature representation 218 of the augmented acoustic frames 304A at each output step among a plurality of output steps. Specifically, the strided convolutional layer of the unsupervised audio encoder 216 receives the augmented acoustic frames 304A from the sequence of augmented acoustic frames 304A and generates an output that the transformer layer uses to generate a corresponding predicted high-order feature representation 218 of the augmented acoustic frames 304A. Thus, the unsupervised audio encoder 216 generates a predicted high-order feature representation 218 to match the corresponding target high-order feature representation 214 generated by the supervised audio encoder 212 at the corresponding output step.

[0050] The unsupervised portion 302 determines an unsupervised loss term 330 based on the target high-order feature representation 214 generated by the target branch 310 at the corresponding output step and the predicted high-order feature representation 218 generated by the augmentation branch 320 at the corresponding output step. In some examples, the unsupervised loss term 330 includes a contrastive loss term represented by:

[0051] In Equation 2, M includes a set of masked frame indices, K includes a set of interference term indices, h t is the encoder output, and c t is the convolutional neural network output. In other examples, the unsupervised loss term 330 includes a reconstruction loss term L1 or a cosine distance loss term. The unsupervised part 302 can update the parameters of the speech recognition model 200 jointly with the supervised part 301 based on the supervised loss term 350 to update the parameters of the speech recognition model 200 based on the unsupervised loss term 330, as represented by:

[0052] In Equation 3, p represents the parameters of the speech recognition model 200, represents the supervised loss term 350, and represents the unsupervised loss term 330. It is worth noting that using the acoustic frame 304, the target branch 310 generates an expected representation (i.e., the target high-order feature representation 214) based on the current state of the supervised audio encoder 212, and the enhancement branch 320 aims to match this expected representation using the enhanced acoustic frame 304A.

[0053] In some examples, the unsupervised part 302 determines a distance-based loss term 370 between the parameters 213 of the supervised audio encoder 212 and the parameters 217 of the unsupervised audio encoder 216 at each output step among multiple output steps. The distance-based loss term 370 can include an L2 loss. In a scenario where the supervised loss term 350 and the unsupervised loss term 330 are inconsistent and have uncorrelated or negatively correlated gradients, the parameters 213 of the supervised audio encoder 212 and the parameters 217 of the unsupervised audio encoder 216 are considered to stabilize the training of the speech recognition model 200. Therefore, the cross-training network 300 aims to make the parameters 213, 217 the same during training according to:

[0054] In Equation 4, p s represents the parameters 213 of the supervised audio encoder 212, and p u represents the parameters 217 of the unsupervised audio encoder 216. Using the Lagrange multiplier method, Equation 4 can be rewritten as:

[0055] Equation 5 represents the total loss for training the speech recognition model 200, including a supervised loss term 350, an unsupervised loss term 330, and a distance-based loss term 370. Here, the cross-training network 300 can jointly update the parameters of the speech recognition model 200 based on the supervised loss term 350, the unsupervised loss term 330, and the distance-based loss term 370. It is worth noting that by jointly minimizing these three losses, the cross-training network 300 does not force the parameters of the supervised part 301 and the parameters of the unsupervised part 302 to be the same at all training steps (e.g., output steps). Instead, the parameters of the supervised part 301 and the parameters of the unsupervised part 302 (e.g., the parameters 213 of the supervised audio encoder 212 and the parameters 217 of the unsupervised audio encoder 216) have the flexibility to be different parameters during training, but the distance-based loss term 370 gradually reduces the distance between them at each output step. Thus, in Equation 5, represents the knowledge transfer parameter, whereby results in independent training of the supervised part 301 and the unsupervised part 302, making the training completely stable, but without using unlabeled data for the supervised part. On the other hand, a large value of forces the supervised part 301 and the unsupervised part 302 to be the same, thus making the training unstable. The transfer parameter can be set to any value. Updating the parameters of the speech recognition model 200 can include jointly updating the parameters 213 of the supervised audio encoder 212 and the parameters 217 of the unsupervised audio encoder 216.

[0056] Other implementations of semi-supervised training rely on temporal modifications and prediction networks to avoid the output of the speech recognition model "collapsing" to a constant value during training. Advantageously, jointly training the speech recognition model 200 with a supervised loss term 350, an unsupervised loss term 330, and a distance-based loss term 370 allows the cross-training network 300 to utilize both unlabeled audio samples 303 and labeled audio samples 305 without the output "collapsing" to a constant value. Notably, using the distance-based loss term 370 alleviates the avoidance of collapsing to a constant value without using the computationally expensive operations of temporal modifications and prediction networks used by other semi-supervised training implementations. Specifically, even when the labeled audio samples 305 are relatively small, the speech recognition model 200 employs a cascaded encoder in a streaming manner, and when there is a mismatch between the unlabeled audio samples 303 and the labeled audio samples 305, the cross-training network 300 also improves the stability of the speech recognition model 200. Additionally, the cross-training network 300 can initialize the parameters 213 of the supervised audio encoder and the parameters 217 of the unsupervised audio encoder 216 with the same initial parameters or with different initial parameters without negatively affecting the resulting speech recognition performance. After the cross-training network 300 trains the speech recognition model 200, the speech recognition model 200 can be executed with the supervised audio encoder 212.

[0057] Figure 4 is a flowchart of an exemplary operational arrangement of a computer-implemented method 400 for training a speech recognition model using a cross-training network 300. At operation 402, the method 400 includes receiving a sequence of acoustic frames 304 extracted from unlabeled audio samples 303 that correspond to spoken utterances 106 not paired with any corresponding transcriptions. The target branch 310 of the cross-training network 300 includes the supervised audio encoder 212 of the speech recognition model 200. At operation 404, the method 400 includes generating, at a plurality of output steps, a target high-order feature representation 214 of corresponding acoustic frames 340 in the sequence of acoustic frames 304 using the supervised audio encoder 212. At the augmentation branch 320 of the cross-training network 300, the method 400 performs operations 406 and 408. At operation 406, the method 400 includes augmenting the sequence of acoustic frames 304 extracted from unlabeled audio samples 303 by masking one or more acoustic frames 204 in the sequence of acoustic frames 304. At operation 408, the method 400 includes generating, at each of the plurality of output steps, a predicted high-order feature representation 218 of corresponding augmented acoustic frames 304A in the sequence of augmented acoustic frames 304A as an output from the unsupervised audio encoder 216 of the speech recognition model 200.

[0058] At operation 410, method 400 includes: at each of a plurality of output steps, determining an unsupervised loss term 330 based on a target high-order feature representation 214 generated by target branch 310 at the corresponding output step and a predicted high-order feature representation 218 generated by enhancement branch 320 at the corresponding output step. At operation 412, method 400 includes updating the parameters of speech recognition model 200 based on the unsupervised loss terms 330 determined at each of the plurality of output steps. Here, updating the parameters of speech recognition model 200 may include jointly updating the parameters of supervised audio encoder 212 and the parameters of unsupervised audio encoder 216.

[0059] Figure 5 FIG. is a schematic diagram of an example computing device 500 that can be used to implement the systems and methods described in this document. Computing device 500 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown here, their connections and relationships, and their functions are only intended to be exemplary and are not intended to limit the implementation of the invention described and / or claimed in this document.

[0060] Computing device 500 includes a processor 510, a memory 520, a storage device 530, a high-speed interface / controller 540 connected to memory 520 and high-speed expansion port 550, and a low-speed interface / controller 560 connected to low-speed bus 570 and storage device 530. Each of the components 510, 520, 530, 540, 550, and 560 is interconnected using various buses and can be mounted on a common motherboard or otherwise as appropriate. Processor 510 can process instructions for execution within computing device 500, including instructions stored in memory 520 or storage device 530, to display graphical information of a graphical user interface (GUI) on an external input / output device such as a display 580 coupled to high-speed interface 540. In other implementations, multiple processors and / or multiple buses and multiple memories and multiple types of memory may be used as appropriate. Moreover, multiple computing devices 500 can be connected, where each device provides a portion of the necessary operations (e.g., as a server group, blade server cluster, or multi-processor system).

[0061] Memory 520 stores information non - transitorily within computing device 500. Memory 520 can be a computer - readable medium, a volatile memory unit, or a non - volatile memory unit. The non - transitory memory 520 can be a physical device for temporarily or permanently storing programs (e.g., sequences of instructions) or data (e.g., program state information) for use by computing device 500. Examples of non - volatile memory include, but are not limited to, flash memory and read - only memory (ROM) / programmable read - only memory (PROM) / erasable programmable read - only memory (EPROM) / electrically erasable programmable read - only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase - change memory (PCM), and magnetic disks or tapes.

[0062] Storage device 530 is capable of providing mass storage for computing device 500. In some implementations, storage device 530 is a computer - readable medium. In various different implementations, storage device 530 can be a floppy disk device, a hard disk device, an optical disk device, or a magnetic tape device, a flash memory, or other similar solid - state memory devices, or an array of devices (including devices in a storage area network or other configurations). In additional implementations, a computer program product is tangibly embodied in an information carrier. The computer program product contains instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer - readable medium or a machine - readable medium, such as memory 520, storage device 530, or memory on processor 510.

[0063] High - speed controller 540 manages bandwidth - intensive operations of computing device 500, while low - speed controller 560 manages lower - bandwidth - intensive operations. Such division of responsibilities is merely exemplary. In some implementations, high - speed controller 540 is coupled to memory 520, display 580 (e.g., via a graphics processor or accelerator), and a high - speed expansion port 550 that can accept various expansion cards (not shown). In some implementations, low - speed controller 560 is coupled to storage device 530 and a low - speed expansion port 590. The low - speed expansion port 590, which can include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), can be coupled to one or more input / output devices such as a keyboard, a pointing device, a scanner, or a networking device such as a switch or a router, for example, via a network adapter.

[0064] As shown in the figure, computing device 500 can be implemented in a variety of different forms. For example, the computing device can be implemented as a standard server 500a or implemented multiple times in a group of such servers 500a, implemented as a laptop computer 500b, or implemented as part of a rack server system 500c.

[0065] Various implementations of the systems and techniques described herein can be implemented in digital electronic and / or optical circuitry, integrated circuit systems, specially designed ASICs (Application Specific Integrated Circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementations in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor which may be special purpose or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0066] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented using high-level procedural and / or object-oriented programming languages and / or in assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, device, and / or apparatus (e.g., a magnetic disk, optical disk, memory, programmable logic device (PLD)) that provides machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal that provides machine instructions and / or data to a programmable processor.

[0067] The processes and logical flows described in this specification can be performed by one or more programmable processors (also referred to as data processing hardware) that execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by special purpose logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). By way of example, processors suitable for the execution of a computer program include both general and special purpose microprocessors, and any one or more processors of any type of digital computer. In general, a processor will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. In general, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or operatively coupled to receive data from one or more mass storage devices or to transfer data to one or more mass storage devices or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including by way of example semiconductor memory devices (such as EPROM, EEPROM, and flash memory devices), magnetic disks (such as internal hard disks or removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0068] To provide for interaction with a user, one or more aspects of the present disclosure can be implemented on a computer having a display device (e.g., a CRT (Cathode Ray Tube), an LCD (Liquid Crystal Display) monitor, or a touch screen) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. Additionally, a computer can interact with the user by sending documents to and receiving documents from the device used by the user; for example, by sending a web page to a web browser on a client device of the user in response to a request received from the web browser.

[0069] A variety of implementations have been described. However, it should be understood that various modifications can be made without departing from the spirit and scope of the present disclosure. Accordingly, other implementations are within the scope of the following claims.

Claims

1. A cross-training network (300) for training a speech recognition model (200), the cross-training network (300) comprising an unsupervised sub-network (302) trained on a plurality of unlabeled audio samples (303), the plurality of unlabeled audio samples corresponding to spoken utterances (106) not paired with corresponding transcripts (120), the unsupervised sub-network (302) comprising: A target branch (310) configured to: Receive a sequence of acoustic frames (104) extracted from the unlabeled audio samples (303) as an input to a supervised audio encoder (212) of the speech recognition model (200); And At each of a plurality of output steps, generate a target high-order feature representation (214) of a corresponding acoustic frame (304) in the sequence of acoustic frames (104) input to the supervised audio encoder (212) at the corresponding output step; And An enhancement branch (320) configured to: Enhance the sequence of acoustic frames (104) extracted from the unlabeled audio samples (303) by masking one or more acoustic frames (104) in the sequence of acoustic frames (104); And At each of the plurality of output steps, generate a predicted high-order feature representation (218) of a corresponding enhanced acoustic frame (304) in the sequence of enhanced acoustic frames (104) as an output from an unsupervised audio encoder (216) of the speech recognition model (200), Wherein the unsupervised sub-network (302) is configured to: At each of the plurality of output steps, determine an unsupervised loss term (330) based on the target high-order feature representation (214) generated by the target branch (310) at the corresponding output step and the predicted high-order feature representation (218) generated by the enhancement branch at the corresponding output step; and Update the parameters of the speech recognition model (200) based on the unsupervised loss term (330) determined at each of the plurality of output steps.

2. The cross-training network (300) according to claim 1, wherein the unsupervised loss term (330) comprises a contrastive loss term.

3. The cross-training network (300) according to claim 1, wherein: The unsupervised sub-network (302) is further configured to: at each of the plurality of output steps, determine a distance-based loss term between the parameters of the unsupervised audio encoder (216) and the parameters of the supervised audio encoder (212); and Updating the parameters of the speech recognition model (200) is further based on the distance-based loss term determined at each of the plurality of output steps.

4. The cross-training network (300) according to claim 3, wherein the distance-based loss term comprises an L2 loss.

5. The cross-training network (300) according to claim 3, wherein updating the parameters of the speech recognition model (200) based on the unsupervised loss term (330) occurs jointly with updating the parameters of the speech recognition model (200) based on the distance-based loss term.

6. The cross-training network (300) according to claim 1, further comprising a supervised sub-network (301) trained on a plurality of labeled audio samples (305), the plurality of labeled audio samples corresponding to spoken utterances (106) paired with corresponding transcripts (120), the supervised sub-network (301) being configured to: At each output step of the plurality of output steps for each labeled audio sample (305): Use the voice recognition model (200) to generate a corresponding voice recognition result (342) of the marked audio sample (305); And Determine a supervised loss term (350) based on the corresponding speech recognition result (342) of the labeled audio sample (305) and the corresponding transcript (120) of the labeled audio sample (305); And Update the parameters of the speech recognition model (200) based on the supervised loss term (350) determined for each labeled audio sample (305) among the plurality of labeled audio samples (305) at each output step of the plurality of output steps.

7. The cross-training network (300) according to claim 6, wherein the corresponding speech recognition result (342) generated using the speech recognition model (200) for the labeled audio sample (305) comprises a probability distribution of possible speech recognition hypotheses for the labeled audio sample (305) at the corresponding output step.

8. The cross-training network (300) according to claim 6, wherein the supervised sub-network (301) is further configured to: based on the supervised loss term (350), jointly update the parameters of the speech recognition model (200) with the unsupervised network updating the parameters of the speech recognition model (200) based on the unsupervised loss term (330) and the distance-based loss term.

9. The cross-training network (300) according to claim 1, wherein the target branch (310) is further configured to apply a stop-gradient operation (314) to the predicted high-order feature representation (218) of the corresponding enhanced acoustic frame (304).

10. The cross-training network (300) according to claim 1, wherein the parameters of the unsupervised audio encoder (216) and the parameters of the supervised audio encoder (212) are initialized with the same initial parameters.

11. The cross-training network (300) according to claim 1, wherein the parameters of the unsupervised audio encoder (216) and the parameters of the supervised audio encoder (212) are initialized with different initial parameters.

12. The cross-training network (300) according to claim 1, wherein each of the unsupervised audio encoder (216) and the supervised audio encoder (212) comprises at least one of the following: a corresponding full-context encoder; or a corresponding cascaded encoder.

13. A computer-implemented method (400) that, when executed on data processing hardware (510), causes the data processing hardware (510) to perform operations including: receiving a sequence of acoustic frames (104) extracted from unlabeled audio samples (303) corresponding to spoken utterances (106) not paired with a corresponding transcription (120); at a target branch (310) of a cross-training network (300), at multiple output steps, using a supervised audio encoder (212) of a speech recognition model (200) to generate target high-order feature representations (214) of corresponding acoustic frames (304) in the sequence of acoustic frames (104); at an augmentation branch (320) of the cross-training network (300): augmenting the sequence of acoustic frames (104) extracted from the unlabeled audio samples (303) by masking one or more acoustic frames (104) in the sequence of acoustic frames (104); and and at each of the multiple output steps, generating predicted high-order feature representations (218) of corresponding augmented acoustic frames (304) in the sequence of augmented acoustic frames (104) as outputs from an unsupervised audio encoder (216) of the speech recognition model (200); at each of the multiple output steps, determining an unsupervised loss term (330) based on the target high-order feature representations (214) generated by the target branch (310) at the corresponding output step and the predicted high-order feature representations (218) generated by the augmentation branch at the corresponding output step; and updating parameters of the speech recognition model (200) based on the unsupervised loss terms (330) determined at each of the multiple output steps.

14. The computer-implemented method (400) according to claim 13, wherein the unsupervised loss term (330) includes a contrastive loss term.

15. The computer-implemented method (400) according to claim 13 or 14, wherein the operations further include: at each of the multiple output steps, determining a distance-based loss term between parameters of the unsupervised audio encoder (216) and parameters of the supervised audio encoder (212); and updating the parameters of the speech recognition model (200) is further based on the distance-based loss term determined at each of the multiple output steps.

16. The computer-implemented method (400) according to claim 15, wherein the distance-based loss term includes an L2 loss.

17. The computer-implemented method (400) according to claim 15 or 16, wherein updating the parameters of the speech recognition model (200) based on the unsupervised loss term (330) occurs jointly with updating the parameters of the speech recognition model (200) based on the distance-based loss term.

18. The computer-implemented method (400) according to any one of claims 13 to 17, wherein the operations further comprise: receiving a plurality of labeled audio samples (305) corresponding to an oral utterance (106) paired with a corresponding transcription (120); at each output step of the plurality of output steps for each labeled audio sample (305): using the speech recognition model (200) to generate a corresponding speech recognition result (342) for the labeled audio sample (305); and determining a supervised loss term (350) based on the corresponding speech recognition result (342) of the labeled audio sample (305) and the corresponding transcription (120) of the labeled audio sample (305); and updating the parameters of the speech recognition model (200) based on the supervised loss term (350) determined for each labeled audio sample (305) among the plurality of labeled audio samples (305) at each output step of the plurality of output steps.

19. The computer-implemented method (400) according to claim 18, wherein the corresponding speech recognition result (342) generated for the labeled audio sample (305) using the speech recognition model (200) comprises a probability distribution of possible speech recognition hypotheses for the labeled audio sample (305) at the corresponding output step.

20. The computer-implemented method (400) according to claim 18 or 19, wherein updating the parameters of the speech recognition model (200) based on the supervised loss term (350) occurs jointly with updating the parameters of the speech recognition model (200) based on the unsupervised loss term (330) and the distance-based loss term.

21. The computer-implemented method (400) according to any one of claims 13 to 20, wherein the operations further comprise applying a stop-gradient operation (314) to the predicted high-order feature representation (218) of the corresponding enhanced acoustic frame (304).

22. The computer-implemented method (400) according to any one of claims 13 to 21, wherein the parameters of the unsupervised audio encoder (216) and the parameters of the supervised audio encoder (212) are initialized with the same initial parameters.

23. The computer-implemented method (400) according to any one of claims 13 to 22, wherein the parameters of the unsupervised audio encoder (216) and the parameters of the supervised audio encoder (212) are initialized with different initial parameters.

24. The computer-implemented method (400) according to any one of claims 13 to 23, wherein each of the unsupervised audio encoder (216) and the supervised audio encoder (212) comprises at least one of the following: A respective full-context encoder; or A respective cascaded encoder.

Citation Information

Patent Citations

  • Transformer transducer: one model unifying streaming and non-streaming speech recognition

    US11741947B2