Alignment prediction for inputting text into automatic speech recognition training
By pre-training an audio encoder with alignment outputs generated from unspoken text utterances and non-synthetic speech utterances, the ASR model learns to jointly represent speech and text, addressing the challenge of generalization to unseen data and improving domain adaptability.
Patent Information
- Application Number
- JP2024555912
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-03-20
- Filing Date
- 2023-02-13
- Publication Date
- 2025-06-09
- Estimated Expiration
- 2043-02-13
AI Technical Summary
Current automatic speech recognition (ASR) models face challenges in generalizing to unseen data due to overfitting on training data, especially when the training data is not extensive enough. Additionally, aligning text and speech modalities during training is difficult, particularly when using synthetic speech generated from unpaired text data.
The proposed solution involves pre-training an audio encoder to jointly learn shared representations of speech and text by using training data that includes unspoken text utterances, untranscribed non-synthetic speech utterances, and transcribed non-synthetic speech utterances. This is achieved by generating alignment outputs for unspoken text utterances and pre-training the audio encoder with these outputs, along with contrastive losses applied to encoded representations.
This approach improves the ability of ASR models to generalize to unseen data by enhancing the alignment between text and speech representations, leading to better performance across different domains without requiring extensive labeled human speech data.
Smart Images

Figure 0007690137000021 
Figure 0007690137000022 
Figure 0007690137000023
Abstract
Description
Technical Field
[0001] The present disclosure relates to alignment prediction for inputting text into automatic speech recognition training.
Background Art
[0002] Automatic speech recognition (ASR), that is, the process of obtaining an audio input and transcribing it into text, is a very important technology used in mobile devices and other devices. Generally, automatic speech recognition attempts to provide an accurate transcription of what a person has said by obtaining an audio input (e.g., a spoken utterance) and transcribing that audio input into text. Current ASR models, based on the ongoing development of deep neural networks, continue to improve in both accuracy (e.g., low word error rate (WER)) and latency (e.g., the delay time from a user's utterance to transcription). However, as one issue in developing a deep learning-based ASR model, the parameters of the ASR model tend to overfit the training data, and as a result, if the training data is not extensive enough, the ASR model will have difficulty generalizing to unseen data. As a result, training the ASR model with a larger training data set will improve the accuracy of the ASR model. By incorporating synthetic speech and / or data-augmented speech, the amount of training data used to train the ASR model can be increased.
Summary of the Invention
[0003] One aspect of the present disclosure provides a computer-implemented method. The computer-implemented method causes data processing hardware to perform operations for pre-training an audio encoder to jointly learn shared representations of speech and text when executed on the data processing hardware. The operations include receiving training data including unspoken text utterances, untranscribed non-synthetic speech utterances, and transcribed non-synthetic speech utterances. Each unspoken text utterance is not paired with any corresponding speech utterance of non-synthetic speech. Each untranscribed non-synthetic speech utterance is not paired with a corresponding transcription. Each transcribed non-synthetic speech utterance is paired with a corresponding transcription. In the operations, the method includes generating, for each unspoken text utterance of the received training data, a corresponding alignment output using an alignment model. In the operations, the method includes pre-training the audio encoder with the alignment outputs generated to correspond to unspoken text utterances, untranscribed non-synthetic speech utterances, and transcribed non-synthetic speech utterances to teach the audio encoder to jointly learn shared representations of speech and text.
[0004] Embodiments of the present disclosure may include one or more of the following optional features. In some embodiments, the audio encoder includes a stack of self-attention layers each including a multi-head self-attention mechanism. In some examples, pre-training the audio encoder includes generating, for each untranscribed non-synthetic speech utterance, a corresponding encoded representation of the untranscribed non-synthetic speech utterance; pre-training the audio encoder with a contrastive loss applied to the corresponding encoded representation of the untranscribed non-synthetic speech utterance; generating, for each alignment output, a corresponding encoded representation of the alignment output; pre-training the audio encoder with a contrastive loss applied to the corresponding encoded representation of the alignment output; generating, for each transcribed non-synthetic speech utterance, a corresponding encoded representation of the transcribed non-synthetic speech utterance; and pre-training the audio encoder with a contrastive loss applied to the corresponding encoded representation of the transcribed non-synthetic speech utterance.
[0005] In some embodiments, pre-training the audio encoder involves, for each of a plurality of time steps for each alignment output, using an auxiliary decoder to generate a first probability distribution over possible synthetic speech recognition hypotheses for the corresponding alignment output, determining an alignment output loss based on the first probability distribution over possible synthetic speech recognition hypotheses and the unuttered text utterance corresponding to the alignment output, pre-training the audio encoder based on the alignment output loss term, for each of a plurality of time steps for each transcribed non-synthetic speech utterance, using an auxiliary decoder to generate a second probability distribution over possible non-synthetic speech recognition hypotheses for the corresponding transcribed non-synthetic speech utterance, determining a non-synthetic speech loss term based on the second probability distribution over possible non-synthetic speech recognition hypotheses, the transcribed non-synthetic speech utterance, and the corresponding transcription paired with the transcribed non-synthetic speech utterance, and pre-training the audio encoder based on the non-synthetic speech loss term. In these embodiments, the auxiliary decoder can include one of a connectionist temporal classification (CTC) decoder, a listen attend spell (LAS) decoder, or a recurrent neural network transducer (RNN-T) decoder. Here, the first probability distribution over possible synthetic speech recognition hypotheses can include one of possible phoneme labels or possible word piece labels, and the second probability distribution over possible non-synthetic speech recognition hypotheses includes one of possible phoneme labels or possible word piece labels.
[0006] In some examples, the audio encoder includes a text encoder, an audio encoder, and a shared encoder. In these examples, the operations further include, for each alignment output, using the text encoder to determine an encoded text representation of the alignment output, using the shared encoder to generate a first encoded shared representation of the alignment output in a shared latent representation space, for each transcribed non-synthetic speech utterance, using the audio encoder to determine an encoded audio representation of the transcribed non-synthetic speech utterance, and using the shared encoder to generate a second encoded shared representation of the transcribed non-synthetic speech utterance in the shared latent representation space.
[0007] Generating a corresponding alignment output for each unspoken text utterance of the received training data can include extracting an initial text representation from the unspoken text utterance, predicting the text chunk duration of each text chunk in the unspoken text utterance, and upsampling the initial text representation using the predicted text chunk duration for each text chunk in the unspoken text utterance. In some embodiments, the operations further include training the alignment model by using the audio encoder to generate an encoded audio representation for a transcribed non-synthetic speech utterance, using the alignment model to determine an alignment output for a transcription corresponding to the transcribed non-synthetic speech utterance, generating an encoded text representation for the alignment output, and updating the parameters of the alignment model based on a comparison between the encoded audio representation for the transcribed non-synthetic speech utterance and the encoded text representation for the alignment output.
[0008] Other aspects of the present disclosure provide a system. The system includes data processing hardware and memory hardware that stores instructions, which, when executed on the data processing hardware, cause the data processing hardware to perform operations. The operations include receiving training data that includes unspoken text utterances, untranscribed non-synthetic speech utterances, and transcribed non-synthetic speech utterances. Each unspoken text utterance is not paired with any corresponding speech utterance of non-synthetic speech. Each untranscribed non-synthetic speech utterance is not paired with a corresponding transcription. Each transcribed non-synthetic speech utterance is paired with a corresponding transcription. In the operations, the method includes generating, for each unspoken text utterance of the received training data, a corresponding alignment output using an alignment model. In the operations, the method includes pre-training an audio encoder with the alignment outputs generated to correspond to unspoken text utterances, untranscribed non-synthetic speech utterances, and transcribed non-synthetic speech utterances to teach the audio encoder to jointly learn a shared representation of audio and text.
[0009] Embodiments of the present disclosure may include one or more of the following optional features. In some embodiments, the audio encoder includes a stack of self-attention layers each including a multi-head self-attention mechanism. In some examples, pre-training the audio encoder includes generating, for each untranscribed non-synthetic speech utterance, a corresponding encoded representation of the untranscribed non-synthetic speech utterance, pre-training the audio encoder with a contrastive loss applied to the corresponding encoded representation of the untranscribed non-synthetic speech utterance, generating, for each alignment output, a corresponding encoded representation of the alignment output, pre-training the audio encoder with a contrastive loss applied to the corresponding encoded representation of the alignment output, generating, for each transcribed non-synthetic speech utterance, a corresponding encoded representation of the transcribed non-synthetic speech utterance, and pre-training the audio encoder with a contrastive loss applied to the corresponding encoded representation of the transcribed non-synthetic speech utterance.
[0010] In some embodiments, pre-training the audio encoder involves, for each of a plurality of time steps for each alignment output, using an auxiliary decoder to generate a first probability distribution over possible synthetic speech recognition hypotheses for the corresponding alignment output, determining an alignment output loss based on the first probability distribution over possible synthetic speech recognition hypotheses and the unuttered text utterance corresponding to the alignment output, pre-training the audio encoder based on the alignment output loss term, for each of a plurality of time steps for each transcribed non-synthetic speech utterance, using an auxiliary decoder to generate a second probability distribution over possible non-synthetic speech recognition hypotheses for the corresponding transcribed non-synthetic speech utterance, determining a non-synthetic speech loss term based on the second probability distribution over possible non-synthetic speech recognition hypotheses, the transcribed non-synthetic speech utterance, and the corresponding transcription paired therewith, and pre-training the audio encoder based on the non-synthetic speech loss term. In these embodiments, the auxiliary decoder can include one of a connectionist temporal classification (CTC) decoder, a listen attend spell (LAS) decoder, or a recurrent neural network transducer (RNN-T) decoder. Here, the first probability distribution over possible synthetic speech recognition hypotheses can include one of possible phoneme labels or possible word piece labels, and the second probability distribution over possible non-synthetic speech recognition hypotheses includes one of possible phoneme labels or possible word piece labels.
[0011] In some examples, the audio encoder includes a text encoder, an audio encoder, and a shared encoder. In these examples, the operations further include, for each alignment output, using the text encoder to determine an encoded text representation of the alignment output, using the shared encoder to generate a first encoded shared representation of the alignment output in a shared latent representation space, for each transcribed non-synthetic speech utterance, using the audio encoder to determine an encoded audio representation of the transcribed non-synthetic speech utterance, and using the shared encoder to generate a second encoded shared representation of the transcribed non-synthetic speech utterance in the shared latent representation space.
[0012] Generating a corresponding alignment output for each unuttered text utterance of the received training data may include extracting an initial text representation from the unuttered text utterance, predicting the text chunk duration of each text chunk in the unuttered text utterance, and upsampling the initial text representation using the predicted text chunk duration for each text chunk in the unuttered text utterance. In some embodiments, the operations further include training the alignment model by using the audio encoder to generate an encoded audio representation for a transcribed non-synthetic speech utterance, using the alignment model to determine an alignment output for a transcription corresponding to the transcribed non-synthetic speech utterance, generating an encoded text representation for the alignment output, and updating the parameters of the alignment model based on a comparison between the encoded audio representation for the transcribed non-synthetic speech utterance and the encoded text representation for the alignment output.
[0013] Details of one or more embodiments of the present disclosure are set forth in the accompanying drawings and the following description. Other aspects, features, and advantages will be apparent from the following description and drawings, as well as from the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0014]
Figure 1
Figure 2
Figure 3A
Figure 3B
Figure 3C
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
DETAILED DESCRIPTION OF THE INVENTION
[0015] In various drawings, the same reference numerals indicate the same elements. With the introduction of the sequence-to-sequence (Seq2Seq) model that maps from an audio signal to a character string, automatic speech recognition has made great progress. At the same time, a text-to-speech (TTS) system or a speech synthesis system has successfully applied the Seq2Seq model to provide a state-of-the-art natural and realistic synthetic speech that is indistinguishable from human speech to the human ear.
[0016] One issue in developing an ASR model based on deep learning is that the parameters of the ASR model tend to overfit the training data. As a result, if the training data is not extensive enough, the ASR model will have difficulty generalizing unseen data. As a result, training the ASR model with a larger training data set will improve the accuracy of the ASR model. For example, the ASR model can be trained using machine learning or other statistical methods with a training data set that includes more than 10,000 hours of transcribed speech. However, the performance of the ASR model will degrade if the domain related to the training data is different from the domain in which the ASR model is deployed during inference. For example, training the ASR model with transcribed speech in a domain related to video conferencing is not very effective for the recognition of speech related to a voice search query, and vice versa.
[0017] Unpaired text data can significantly limit the amount of labeled human speech required to train an ASR model, while also providing flexibility in enabling the ASR model to function across different domains. However, using text data (i.e., unpaired text data) in addition to speech data to train an ASR model creates problems in aligning the speech and text modalities of the training data. One current approach uses multi-task training to train a single model with different objectives for each modality. This approach poses problems of interference and capacity limitations when each modality of the training data has different properties and objectives. Another current approach involves TTS systems that synthesize unpaired text data to generate synthetic speech (i.e., modality conversion). Additionally, training an ASR model using synthetic speech based on text data has been found to have a different impact on ASR training than human speech, despite examples of state-of-the-art synthetic speech being indistinguishable from human speech. This gap between synthetic and human speech results from a mismatch between human speech data and synthetic speech data, arising from the difficult one-to-many mapping problem that TTS systems are trying to solve. That is, while the overall quality of available synthetic speech is very high, synthetic speech has far less variability and minimal voice breaks compared to human speech. As a result, training an ASR model using synthetic speech based on unpaired text data makes it difficult to generalize to actual speech utterances during inference.
[0018] The embodiments described herein relate to aligning text representations used to generate synthetic speech with corresponding non-synthetic speech representations in a latent representation space for training an ASR model. That is, the alignment model can generate alignment outputs for unuttered text utterances when there is little or no available transcribed speech (e.g., non-synthetic speech) in the target domain and / or target language for training the ASR model, or when it is not very common. More specifically, the embodiments are directed to pre-training the audio encoder of the ASR model with training data including untranscribed non-synthetic speech utterances, unuttered text utterances for generating corresponding alignment outputs, and transcribed non-synthetic speech utterances such that the speech and text representations are jointly learned, and then fine-tuning (e.g., warm-start training) the pre-trained ASR model using the available transcribed non-synthetic speech utterances. In particular, generating alignment outputs for unuttered text utterances (e.g., without converting the text utterance to speech) facilitates computationally efficient alignment of representations between the text and speech modalities of the training data. As will become apparent below, pre-training the audio encoder includes updating the parameters of the audio encoder based on a combination of self-supervised contrastive loss, supervised loss, and consistency loss derived from the training data. The ASR model can include a single-language ASR model or a multilingual ASR model. Further, the learned representations between the text and speech modalities may be used in a speech translation model.
[0019] FIG. 1 shows an automatic speech recognition (ASR) system 100 that implements an ASR model 200 present in a user device 102 of a user 104 and / or a remote computing device 201 that communicates with the user device 102 (e.g., one or more servers of a distributed system running in a cloud computing environment). The user device 102 is shown as a mobile computing device (e.g., a smartphone), but the user device 102 can correspond to any type of computing device, without limitation, such as a tablet device, a laptop / desktop computer, a wearable device, a digital assistant device, a smart speaker / display, a smart appliance, an in-vehicle infotainment system, or an Internet of Things (IoT) device, and includes data processing hardware 111 and memory hardware 113.
[0020] User device 102 includes an audio subsystem 108. The audio subsystem 108 is configured to receive utterance 106 spoken by user 104 (e.g., user device 102 may include one or more microphones that record the spoken utterance 106) and convert utterance 106 into a corresponding digital format associated with input acoustic frame 110 processable by ASR system 100. In the illustrated example, the user makes respective utterances 106 in natural English language for the phrase "What's the weather like in New York City?", and audio subsystem 108 converts utterance 106 into corresponding acoustic frame 110 for input to ASR system 100. Thereafter, ASR model 200 receives acoustic frame 110 corresponding to utterance 106 as input and generates / predicts corresponding transcription 120 (e.g., recognition result / hypothesis) of utterance 106 as output. In the illustrated example, user device 102 and / or remote computing device 201 also execute a user interface generator 107 configured to present an expression of transcription 120 of utterance 106 to user 104 of user device 102. In some configurations, transcription 120 output from ASR system 100 is processed by a natural language understanding (NLU) module executed, for example, on user device 102 or remote computing device 201 to execute a user command. Additionally or alternatively, a text-to-speech system (e.g., executed on any combination of user device 102 or remote computing device 201) may convert the transcription into a synthetic voice for audible output by other devices. For example, original utterance 106 corresponds to a message that user 104 is sending to a friend, where transcription 120 is converted into a synthetic voice for audible output to the friend who hears the message conveyed by original utterance 106.
[0021] Referring to FIG. 2, an exemplary transducer model 200 based on frame alignment includes a recurrent neural network transducer (RNN-T) model architecture that complies with latency constraints related to interactive applications. The use of the RNN-T model architecture is exemplary, and the transducer model 200 based on frame alignment may include other architectures such as, among others, a transformer transducer model architecture and a conformer transducer model architecture. The RNN-T model 200 has a small computational footprint and fewer memory requirements to utilize than conventional ASR architectures, so the RNN-T model architecture is suitable for performing all speech recognition on the user device 102 (e.g., without requiring communication with a remote server). The RNN-T model 200 includes an encoder network 210, a prediction network 220, and a joint network 230. The encoder network 210 is substantially similar to the acoustic model (AM) of a conventional ASR system and includes a stack of self-attention layers (e.g., conformer layers or transformer layers), or a recurrent network of stacked long short-term memory (LSTM) layers. For example, the encoder reads a sequence of d-dimensional feature vectors (e.g., acoustic frames 110 (FIG. 1)) x=(x 1 ,x 2 ,...,x T )(where
[0022]
Number
[0023] ) and generates a high-order feature representation at each output step. This high-order feature representation is represented as
[0024]
Number
[0025] . Similarly, the prediction network 220 is also an LSTM network, which, like the language model (LM), takes as input the sequence of non-blank symbols, y 0 ,...,y ui-1 , processed so far by the final softmax layer 240, and outputs a dense representation
[0026]
Number
[0027] . Finally, the representations generated by the encoder network 210 and the prediction / decoder network 220 are combined by the joint network 230. The latency can be improved by replacing the prediction network 220 with an embedding lookup table that outputs lookup sparse embedding values instead of processing the dense representation. Next, the joint network
[0028]
Number
[0029] Predict this. This is the distribution for the following output symbols. In other words, the joint network 230 generates a probability distribution for possible speech recognition hypotheses at each output step (e.g., time step). Here, "possible speech recognition hypotheses" corresponds to a set of output labels that each represent a symbol / character of a specific natural language. For example, when the natural language is English, the set of output labels can include 27 symbols (e.g., one label for each of the 26 letters of the English alphabet and one label representing a space). Thus, the joint network 230 can output a set of values indicating the likelihood of occurrence of each of a predetermined set of output labels. This set of values can be a vector and can represent a probability distribution for the set of output labels. In some cases, the output labels are graphemes (e.g., individual characters, and optionally punctuation marks and other symbols), but the set of output labels is not limited thereto. For example, the set of output labels can include, in addition to graphemes or instead of graphemes, word pieces, phonemes, and / or whole words. The output distribution of the joint network 230 can include posterior probability values for each of different output labels. Thus, if there are 100 different output labels representing different graphemes or other symbols, the output y i of the joint network 230 can include 100 different probability values, one for each output label. Next, the probability distribution can be used to select scores and assign scores to candidate orthographic elements (e.g., graphemes, word pieces, and / or words) in a beam search process (e.g., by the softmax layer 240) for determining the transcription 120.
[0030] The softmax layer 240 can use any technique to select the output label / symbol with the highest probability within the distribution as the next output symbol predicted by the RNN-T model 200 at the corresponding output step. In this way, the RNN-T model 200 does not make the conditional independence assumption. Rather, the prediction of each symbol is conditioned not only on the acoustics but also on the sequence of labels output so far. The RNN-T model 200 assumes that the output symbol is independent of future acoustic frames 110. This enables the RNN-T model to be used in a streaming manner.
[0031] In some examples, the encoder network (i.e., audio encoder) 210 of the RNN-T model 200 includes a stack of self-attention layers / blocks such as conformer blocks. Here, each conformer block includes a series of multi-head self-attention layers, depthwise convolutional layers, and feed-forward layers. The prediction network 220 can have two 2,048-dimensional LSTM layers, each followed by a 640-dimensional projection layer. Alternatively, the prediction network 220 can include a stack of transformer blocks or conformer blocks, or an embedding lookup table, instead of the LSTM layers. Finally, the joint network 230 can also have 640 hidden units. The softmax layer 240 can be composed of a set of integrated word pieces or graphemes generated using all unique word pieces or graphemes of a plurality of training datasets.
[0032] Figures 3A - 3C show an exemplary training process 300 for pre-training the audio encoder 210 of the ASR model 200 (Figure 2). The training process 300 can pre-train the audio encoder 210 using available training data. The available training data is a set of unuttered text utterances (X text ) 320, a set of transcribed non-synthetic speech utterances (Xsup )A set of 304 and / or untranscribed non-synthetic speech utterances (X unsup )including 306. Each unuttered text utterance 320 contains only text data (i.e., unpaired data), and thus each unuttered text utterance 320 is not paired with any corresponding spoken audio representation (i.e., speech) of the utterance. The unuttered text utterance 320 can include any sequence text chunk including words, word pieces, phonemes, and / or graphemes. Each untranscribed non-synthetic speech utterance 306 (also simply referred to as "untranscribed speech utterance 306") contains only audio data (i.e., unpaired data), and thus the untranscribed speech utterance 306 is not paired with any corresponding transcription. On the other hand, each transcribed non-synthetic speech utterance 304 (also simply referred to as "transcribed speech utterance 304") includes a corresponding transcription 302 paired with the corresponding non-synthetic audio representation of the corresponding transcribed speech utterance 304.
[0033] For clarity, the training process 300 includes a self-supervised contrastive loss part 300a (Figure 3A), a supervised loss part 300b (Figure 3B), and a consistency regularization part 300c (Figure 3C). The training process 300 is as follows, i.e., the unuttered training text utterance (X text )320, the corpus of transcribed non-synthetic speech utterances (X sup )304, and the untranscribed non-synthetic speech utterances (X unsup )306, the contrastive loss (L w2v )316 derived using the self-supervised contrastive loss part 300a, the supervised losses (L text )342, 344 derived using the supervised loss part 300b from the unuttered training text utterance (X sup )320 and the transcribed non-synthetic speech utterances (X aux )304, and the consistency loss (J cons (θ))352 derived using the consistency regularization part 300c, based on the total loss (L tts4pretrain2 ) to pre-train the audio encoder 210.
[0034] Referring to FIG. 3A, the self-supervised contrastive loss part 300a of the training process 300 can use the alignment model 600. The alignment model 600 is configured to generate an alignment output (i.e., text representation) 602 for each of the plurality of unspoken training text utterances 320 at each of the plurality of output steps. The unspoken text utterance 320 (also simply referred to as "unspoken text utterance 320 (singular)") includes unspoken text that is text-only data (i.e., unpaired data), and thus, each unspoken text utterance (X text ) 320 is not paired with any synthetic or non-synthetic voice. Thus, the alignment model 600 generates a corresponding alignment output 602 for each of the unspoken text utterances 320.
[0035] Referring to FIG. 6, in some examples, the alignment model 600 includes an embedding extractor 610, a duration predictor 620, and an upsampler 630. The embedding extractor 610 receives an unspoken text utterance 320 including a sequence of text chunks including words, word pieces, phonemes, and / or graphemes, and a corresponding initial text representation (e t)Extract 612. The initial text representation 612 includes embedded vocabulary information from the unuttered text utterance 320. Additionally or alternatively, the embedding extractor 610 may receive a transcription 302 corresponding to the transcribed non-synthetic speech utterance 304 (FIG. 3C). The duration predictor 620 receives the initial text representation 612 from the embedding extractor 610 and predicts the corresponding text chunk duration (i.e., the duration of words, word pieces, phonemes, and / or graphemes) 622. The text chunk duration 622 indicates the duration for which the corresponding text chunk would be spoken if a person (or text-to-speech system) uttered the unuttered text utterance 320. For example, the unuttered text utterance 320 can include a sequence of phonemes, and the duration predictor 620 predicts the phoneme duration 622 for each phoneme in the sequence of phonemes. In this example, the duration predictor 620 predicts the probability of a non-zero duration for each phoneme and predicts the phoneme duration 622 by predicting the probability of the continuous phoneme duration for each phoneme. Since the sequence of phonemes includes normal phonemes, silence between word boundaries, and punctuation, only the normal phonemes are associated with non-zero durations, while silence and punctuation are typically associated with continuous phoneme durations. Thus, the duration predictor 620 can use a sigmoid activation value following the first of two independent activation values to predict the probability of a non-zero duration and a softplus activation value following the second of two independent projection values to predict the continuous text chunk duration 622 for each text chunk. The duration predictor 620 determines, for each text chunk, whether the probability of a non-zero duration is less than a threshold, and if the probability of a non-zero duration is less than the threshold, the multiplier may zero out the continuous text chunk duration 622 predicted by the softplus activation for the corresponding text chunk. On the other hand, if the probability of a non-zero duration is greater than or equal to the threshold, the predicted text chunk duration 622 may be set equal to the continuous phoneme duration predicted by the softplus activation.
[0036] The upsampler 630 receives, for each unuttered text utterance 320, a corresponding initial text representation 612 and a predicted text chunk duration 622, and generates an alignment output having a significant number of frames by upsampling the initial text representation 612 using the corresponding predicted text chunk duration 622.
[0037]
Number
[0038] The alignment model 600 generates 602. In some examples, the alignment model 600 sends the alignment output 602 to the text encoder 202 (FIGS. 3B and 3C) of the audio encoder 210. In other examples (not shown), the alignment model 600 sends the alignment output 602 to a shared encoder 250 of the audio encoder 210 (FIGS. 3B and 3C) (e.g., bypassing the text encoder 202). In these other examples, the alignment output 602 functions as an encoded text representation 312, such that the shared encoder 250 can receive the alignment output 602 directly from the alignment model 600 (FIGS. 3B and 3C). In some additional examples, paired training data is available, and the upsampler 630 generates the alignment output 602 as follows.
[0039]
Number
[0040] Here, the upsampler includes a resampler layer and a refiner layer that directly aligns the initial text embedding value 612 with the corresponding encoded audio representation 314 (FIGS. 3B and 3C). In other examples, paired training data is not available, and the upsampler 630 generates the alignment output 602 as follows.
[0041]
Number
[0042] In particular, the number of frames of the alignment output 602 indicates the predicted speech duration of the unspoken text utterance 320. In other words, the number of frames of the alignment output 602 maps (i.e., aligns) the sequence of text chunks of the unspoken text utterance 320 to audio frames. Here, the upsampler 630 includes a resampler layer and a refiner layer that replicate the initial text embedding value 612 to match the predicted text chunk duration 622 (i.e., the speech duration). Thus, the alignment output 602 includes a text representation of the unspoken text utterance 320 having a timing component that aligns with how a person would speak the unspoken text utterance 320.
[0043] In particular, in most cases, a text-to-speech (TTS) system generates an audible output to impart a human voice timing component to the unspoken text utterance 320, whereby the training process may use the audible output (i.e., the synthesized speech) to train the audio encoder 210. Thus, since the alignment model 600 generates an alignment output 602 that directly maps the sequence of text chunks to audio frames, the training process 300 does not require a TTS system to train the audio encoder 210 using the unspoken text utterance 320. That is, the alignment model 600 does not convert the unspoken text utterance 320 to generate synthesized speech.
[0044] FIG. 7 shows an exemplary training process 700 for training an alignment model 600 using the transcribed non-synthetic speech utterance 304 having the corresponding transcription 302 (i.e., paired training data). In the illustrated example, the speech encoder 204 receives, as input, each transcribed non-synthetic speech utterance 304 as a sequence of features / vectors (e.g., a mel-frequency spectrogram such as the acoustic frame 110 of FIG. 1), and as output, for each of a plurality of time steps, an encoded audio representation (e s ) 314 corresponding to the transcribed non-synthetic speech utterance 304 at the corresponding time step. In parallel, the alignment model 600 receives the transcription 302 corresponding to the same non-synthetic speech utterance 304 and generates an alignment output 602 according to Equation 1. The text encoder 202 receives the alignment output 602 as input and, for each of a plurality of time steps, generates, as output, an encoded text representation
[0045] [Number]
[0046] 312 corresponding to the transcription 302 at the corresponding time step. The modality loss module 750 receives the encoded text representation 312 and the encoded audio representation 314, and generates a modality loss 752 based on comparing the encoded text representation 312 with the encoded audio representation as follows.
[0047] [Number]
[0048] Equation 3 is the encoded text representation
[0049] [Number]
[0050] The mean squared error (MSE) between the audio representation (e s ) 314 encoded as 312 and the predicted text target is counted against the RNN-T model alignment value between the predicted text target and the audio representation (e s ) 314 encoded as 312 to determine the modality loss (L MM ) 752. Here, the audio representation 314 encoded as 312 functions as a ground truth label to train the alignment model 600 to generate an alignment output 602 that aligns with the corresponding non-synthetic speech utterance 304. The training process 700 may use the modality loss 752 to update the parameters of the alignment model 600. For example, the training process 700 may update the parameters of the duration predictor 620 and / or the upsampler 630 (FIG. 6).
[0051] FIG. 8 shows an exemplary training process 800 for training the alignment model 600 using paired training data and unpaired training data. That is, the training process 800 learns the speech alignment output 602 using the transcribed non-synthetic speech utterance 304 (i.e., paired training data) having the corresponding transcription 302 and the unspoken text utterance 320 (i.e., unpaired training data). In the illustrated example, the speech encoder 204 receives, as input, each transcribed non-synthetic speech utterance 304 as a sequence of features / vectors (e.g., a mel-frequency spectrogram such as the acoustic frame 110 of FIG. 1), and for each of a plurality of time steps, the encoded audio representation (e s)Generate 314 as the output. In parallel, the alignment model 600 receives the transcription 302 corresponding to the same non-synthesized speech utterance 304 and generates an alignment output 602 according to Equation 1. Additionally or alternatively, the alignment model 600 can receive an unspoken text utterance 320 and generate an alignment output 602 according to Equation 2. The text encoder 202 receives the alignment output 602 as input and, for each of a plurality of time steps, an encoded text representation
[0052] [Number]
[0053] Generate 314 as the output. The audio encoder 210 can include a shared encoder 250 that receives the encoded text representation 312 as input and generates a first encoded shared representation 322 as output. The shared encoder 250 may also receive the encoded audio representation 314 as input and generate a second encoded shared representation 324 as output. The auxiliary decoder 390 receives the first and second encoded shared representations 322, 324 as input and generates first and second probability distributions 392, 294 corresponding to possible speech recognition hypotheses as output.
[0054] The alignment masked loss module 850 receives a first probability distribution 392 corresponding to the encoded text representation 312 and a second probability distribution 394 corresponding to the encoded audio representation 314 and generates an alignment loss 852 as follows.
[0055] [Number]
[0056] The alignment loss 852 from Equation 4 can be applied to the masked, sampled, and encoded text representation 312 in both the frequency and time domains. In particular, the alignment loss 852 can be used as a training objective for both paired and unpaired training data. The training process 800 can update the parameters of the alignment model 600 using the alignment loss 852. For example, the training process 800 can update the parameters of the duration predictor 620 and / or the upsampler 630 (Figure 6).
[0057] Referring back to FIG. 3A, in some embodiments, audio encoder 210 includes audio encoder 204 and text encoder 202, which will be described in more detail with reference to FIGS. 3B and 3C. In the illustrated example, audio encoder 210 (alternatively, audio encoder 204 or text encoder 202 (FIGS. 3B and 3C)) includes a conformer encoder that includes a stack of multiple conformer blocks. Each of the multiple conformer blocks includes a series of layers that are multi-head self-attention, depthwise convolution, and feed-forward. Alternatively, audio encoder 210 may include other types of encoders (e.g., transformer encoders) having a stack of self-attention layers / blocks. Conformer encoder 210 can be naturally divided into a feature encoder that includes convolutional subsampling block 212, and a context network that includes a stack of linear layer 214 and conformer blocks 216. In some embodiments, convolutional subsampling block 212 has two two-dimensional convolutional layers, both of which have a stride of (2, 2), and thus the feature sequence length is reduced to one-fourth. Convolutional subsampling block 212 receives, as input, a sequence of input features / vectors (e.g., a mel-frequency spectrogram such as acoustic frame 110 of FIG. 1) associated with each transcribed non-synthesized speech utterance 304 and each untranscribed non-synthesized speech utterance 306, and for each of the multiple output steps, generates, as output, encoded audio features 211 corresponding to each one of the transcribed non-synthesized speech utterances 304 or each one of the untranscribed non-synthesized speech utterances 306. Convolutional subsampling block 212 may receive each alignment output 602 as input, and for each of the multiple output steps, may generate, as output, encoded text features 213 corresponding to each one of the alignment outputs 602.
[0058] The encoded audio features and text features 211, 213 output from the convolutional subsampling block 212 (i.e., interchangeably referred to as "encoded features 211, 213") can be sent to the masking module 218, where a portion of the encoded features 211, 213 are randomly selected and replaced with a trained feature vector shared among all masked time steps to provide corresponding masked and encoded audio features 211, 211m and masked and encoded text features 213, 213m. In some examples, the masking module 218 masks by randomly sampling a specific proportion p of all time steps that serve as start indices without replacement, masking the randomly selected encoded features 211, 213, and then masking the subsequent M consecutive time steps from all sample indices, such that some spans may overlap. After masking is applied, the linear layer 214 and conformer block 216 of the context network receive the masked encoded features 211m (or the encoded features 211, 213 not selected by the masking module 218) and output a corresponding contrast context vector (i.e., encoded representation) 215 from the masked encoded features 211m, 213m. Further, the quantizer 217 receives the encoded features 211, 213 as input and generates a quantization vector (i.e., target context vector) 219 as output. Then, the contrast loss module 315 derives a contrast loss (L w2v ) 316 between the contrast context vector (i.e., encoded representation) 215 and the target context vector 219 at the masked positions as follows.
[0059]
Number
[0060] where c tis the contrast context vector 215 centered around the masked time step t, q t is q t and represents the target context vector 219 at time step t in a set of K + 1 candidate target context vectors 219 that includes q and K distractors. The distractors can be uniformly sampled from other masked time steps of the same utterance.
[0061] The contrast loss 316 is optimized between the contrast context vector 215 and the target context vector 219 at the masked positions. After the audio encoder 210 converges on the untranscribed non-synthetic speech utterance 306, the pre-training procedure is repeated for both the alignment output 602 corresponding to the unuttered text utterance 320 and the transcribed non-synthetic speech utterance 304. Thus, the contrast loss (L w2v ) is optimized for both the real / human (non-synthetic) and the unuttered text utterance 320 represented by the alignment output 602, and the additional auxiliary loss by the transcribed non-synthetic speech utterance 304 and the alignment output 602 is as will be described in more detail below with reference to FIG. 3B. Thus, the training process 300a pre-trains the audio encoder 210 with the derived contrast loss 316. The derived contrast loss 316 is applied to the corresponding encoded features 211, 213 associated with each alignment output 602, each transcribed non-synthetic speech utterance 304, and each untranscribed non-synthetic speech utterance 306 supplied as input to the audio encoder 210. Pre-training the audio encoder 210 may include updating the parameters of the audio encoder 210 based on the contrast loss 316.
[0062] Referring to FIG. 3B, the supervised loss part 300b of the training process 300 is configured to input vocabulary information into the audio encoder 210 during pre-training, based on the supervised loss terms 342, 344 derived from the transcribed non-synthetic speech utterance 304 and the alignment output 602 corresponding to the unuttered text utterance 320 output by the alignment model 600. In particular, the supervised loss part 300b utilizes one or more auxiliary decoders 390 for generating the supervised loss terms 342, 344. The auxiliary decoder 390 may include a Connectionist Temporal Classification (CTC) decoder, a Listen-Attend-Spell (LAS) decoder, or an RNN-T decoder. These auxiliary decoders 390 may include at least one of a phoneme decoder configured to decode a sequence of phonemes or a word-piece decoder configured to decode a sequence of word-pieces. The auxiliary decoder 390 may also include a grapheme decoder configured to decode a sequence of graphemes.
[0063] In the supervised loss part 300b, the text encoder 202 of the audio encoder 210 is configured to receive the alignment output 602 (i.e., the text embedding value) from the alignment model 600, and the speech encoder 204 is configured to receive the transcribed non-synthetic speech utterance 304. That is, the text encoder 202 of the audio encoder 210 generates a text representation 312 encoded for the alignment output 602 (e.g., corresponding to the unuttered text utterance 320), and the speech encoder 204 of the audio encoder 210 generates an audio representation 314 encoded for the speech input (i.e., the transcribed non-synthetic speech utterance 304). Here, the encoded text representation 312 and the encoded audio representation 314 may not always match the auxiliary decoder 390. Therefore, the audio encoder 210 may also include a shared encoder 250. The shared encoder 250 receives the encoded text representation 312 as an input and the first encoded shared representation 322(e text) is generated as the output. Further, the shared encoder 250 receives the encoded audio representation 314 as input and generates a second encoded shared representation (e sup ) 324 as the output. Thus, the shared encoder 250 generates the first and second encoded shared representations 322, 324 in a shared latent representation space that is adapted to the auxiliary decoder 390.
[0064] In particular, the shared encoder 250 receives, as input, each encoded text representation 312 corresponding to the alignment output 602 generated from the unspoken text utterance 320, and for each of a plurality of time steps, a first encoded shared representation (e text ) 322 corresponding to the alignment output 602 at the corresponding time step is generated as the output. The auxiliary decoder 390, which includes a phoneme decoder or a word piece decoder, receives each first encoded shared representation 322 output from the shared encoder 250 as input and generates, as output, a first probability distribution 392 over possible speech recognition hypotheses for the corresponding alignment output 602 at the corresponding time step. In some examples, the first probability distribution 392 over possible speech recognition hypotheses includes one of possible phoneme labels, possible word piece labels, or possible grapheme labels. Thereafter, the supervised loss module 340 can determine an alignment output loss term 342 based on the first probability distribution 392 over possible speech recognition hypotheses for the alignment output 602 corresponding to the unspoken text utterance 320. Here, the corresponding unspoken text utterance 320 from which the alignment output 602 is generated also functions as the ground truth transcription 302. The supervised loss section 300b may pre-train the audio encoder 210 with the alignment output loss term 342 by updating the parameters of the audio encoder 210 using the alignment output loss term 342.
[0065] Similarly, in the supervised loss section 300b, the shared encoder 250 receives, as input, each transcribed and encoded audio representation 314 corresponding to the non-synthesized speech utterance 304, and for each of a plurality of time steps, generates, as output, a second encoded shared representation (e sup ) 324 corresponding to the transcribed non-synthesized speech utterance 304 at the corresponding time step. An auxiliary decoder 390 including a phoneme decoder or a word piece decoder receives, as input, each second encoded shared representation 324 output from the shared encoder 250, and generates, as output, a second probability distribution 394 over possible non-synthesized speech recognition hypotheses for the corresponding transcribed non-synthesized speech utterance 304 at the corresponding time step. In some examples, the second probability distribution 394 over possible non-synthesized speech recognition hypotheses includes one of possible phoneme labels, possible word piece labels, or possible grapheme labels. The supervised loss module 340 can then determine a non-synthesized speech loss term 344 based on the second probability distribution 394 over possible non-synthesized speech recognition hypotheses and the corresponding transcription 302 paired with the transcribed non-synthesized speech utterance 304. Here, the corresponding transcription 302 functions as a ground truth transcription and can include a sequence of target phonemes, target word pieces, and / or target graphemes. The supervised loss section 300b may pre-train the audio encoder 210 with the non-synthesized speech loss term 344 by updating the parameters of the audio encoder 210 using the non-synthesized speech loss term 344.
[0066] In some embodiments, the supervised loss section 300b of the training process 300 uses another auxiliary decoder 390 to generate a first encoded shared representation (e text)Based on 322, a third probability distribution 393 for possible speech recognition hypotheses is generated, whereby the supervised loss module 340 can determine another alignment output loss term 342 based on the third probability distribution 393 and the unspoken text utterance 320 corresponding to the alignment output 602. Here, the other auxiliary decoder 390 includes the other of a phoneme decoder, a word piece decoder, or a grapheme decoder, and the third probability distribution 393 for possible speech recognition hypotheses includes the other of possible phoneme labels, possible word piece labels, or possible grapheme labels. In these embodiments, the other auxiliary decoder 290 also generates a fourth probability distribution 395 for possible non-composite speech recognition hypotheses for the corresponding second encoded shared representation 324 at the corresponding time step, whereby the supervised loss module 340 can determine another non-composite speech loss term 344 based on the fourth probability distribution 395 and the corresponding transcription 302 paired with the transcribed non-composite speech representation 304. Here, the fourth probability distribution 395 for possible non-composite speech recognition hypotheses includes the other of possible phoneme labels, possible word piece labels, or possible grapheme labels. The supervised loss part 300b of the training process 300 can similarly pre-train the audio encoder 210 with the other alignment output loss term 342 and the other non-composite speech loss term 344.
[0067] The untranscribed non-composite speech utterance 306 and the unspoken text utterance 320 each correspond to "unpaired" training data, and thus, the contrastive loss (L text )316 derived from the unspoken text utterance (X w2v )320 is combined with the supervised loss J aux associated with the alignment output loss term 342 to obtain the unspoken text loss function J text .
[0068]
Number
[0069] Similarly, the untranscribed non-synthetic speech utterance (X unsup )-derived contrastive loss (L w2v )316 can be used as follows to represent the teacherless speech loss function J unsup_speech .
[0070] J unsup_speech = J w2v (x * | θ e ) (7) During the pre-training of the audio encoder 210, the alignment output 602 and the untranscribed non-synthetic utterance 306 can be separated or mixed within each batch. To have the audio encoder 210 learn valid representations for both alignment outputs 602 corresponding to the unspoken text utterance 320 and the non-synthetic (human / reality) speech, the loss function J text combined with those of equations 5 and 6, when obtaining the unpaired data loss function J unpaired as follows, applies a loss mask σ
[0071] J unpaired = σJ text + (1 - σ)J speech (8) The transcribed non-synthetic speech utterance 304 corresponds to "paired" and "teacher-present" training data, and thus the derived contrastive loss L w2v and the derived teacher-present loss J aux can be combined as follows to obtain the paired data loss function J paired .
[0072] J paired = L w2v (x | θ e ) + L aux (y | x, θ e , θ d ) (9) Referring to FIG. 3C, the consistency regularization part (i.e., the modality matching part) 300c of the training process 300 takes the transcribed non-synthetic speech utterance (Xsup ) a consistency loss term (J cons (θ)) 352 between a corresponding one of 304 and the alignment output 604 paired with the corresponding transcribed non-synthetic speech utterance 304 of the same utterance, to encourage the audio encoder 210 to learn a consistent prediction between the non-synthetic speech (e.g., real / human speech) and the alignment output 602 corresponding to the unuttered text utterance 320. Thus, the non-synthetic speech utterance 304 and the paired alignment output 604 of each training utterance pair 301 are associated with the same ground truth transcription. In short, the consistency loss term 352 between the alignment output 604 paired with the same training utterance as the transcribed non-synthetic speech utterance 304, regardless of whether the training utterance belongs to the non-synthetic speech utterance (i.e., audio training data) or the alignment output (i.e., text training data), is independent of the supervised loss term between the ground truth transcription 302 and each of the non-synthetic speech recognition hypothesis output by the auxiliary decoder 390 and the speech recognition hypothesis output by the auxiliary decoder 390, and by encouraging the audio encoder 210 to behave consistently, a self-supervised training manner is brought about.
[0073] Similar to the alignment output 602 generated from the unuttered text utterance 320 of FIG. 3B, the alignment model 600 may generate each paired alignment output 604 using the corresponding transcription 302 paired with the transcribed non-synthetic speech utterance 304. Here, the non-synthetic speech representation 304 is associated with the paired alignment output 604 generated by the alignment model 600 that maps the unuttered text utterance 320 to audio frames.
[0074] In the consistency regularization unit 300c, the text encoder 202 receives, as input, each paired alignment output 604, and for each of a plurality of time steps, generates, as output, an encoded text representation 313 corresponding to the paired alignment output 604 at the corresponding time step. The shared encoder 250 receives the encoded text representation 313 as input and generates, as output, a first encoded shared representation (e * sup ) 323. The auxiliary decoder 390 including a phoneme decoder or a word piece decoder receives each first encoded shared representation 323 output from the shared encoder 250 and generates, as output, a first probability distribution 311 for possible speech recognition hypotheses for the corresponding paired alignment output 604 at the corresponding time step. In some examples, the first probability distribution 311 for possible speech recognition hypotheses includes one of possible phoneme labels or possible word piece labels.
[0075] Similarly, the audio encoder 204 receives, as input, each transcribed non-synthetic speech utterance 304 as a sequence of features / vectors (e.g., a mel-frequency spectrogram such as the acoustic frame 110 in FIG. 1), and for each of a plurality of time steps, generates, as output, an encoded audio representation 314 corresponding to the transcribed non-synthetic speech utterance 304 at the corresponding time step. The shared encoder 250 receives the encoded audio representation 314 as input and generates, as output, a second encoded shared representation 324 (e sup ) 324. The auxiliary decoder 390 including a phoneme decoder or a word piece decoder receives each second encoded shared representation 324 output from the shared encoder 250 and generates, as output, a second probability distribution 394 for possible non-synthetic speech recognition hypotheses for the corresponding transcribed non-synthetic speech utterance 304 at the corresponding time step. In some examples, the second probability distribution 394 for possible non-synthetic speech recognition hypotheses includes one of possible phoneme labels or possible word piece labels.
[0076] Continuing to refer to FIG. 3C, the consistency regularization unit 300c of the training process 300 further, at each of a plurality of time steps for each training utterance pair 301, based on a first probability distribution 311 for possible speech recognition hypotheses and a second probability distribution 394 for possible non-synthetic speech recognition hypotheses, determines a consistency loss term (J cons (θ)) 352 for the corresponding training utterance pair 301. For example, the training process 300 may use a consistency loss term module 350, and the consistency loss term module 350 is configured to receive the corresponding non-synthetic speech and speech recognition results 311, 394 output by the auxiliary decoder 390 at each time step and determine the consistency loss term 352 for the corresponding training utterance pair 301 at the time step.
[0077] In some examples, the consistency regularization unit 300c of the training process 300 determines the consistency loss term 352 based on the Kullback-Leibler divergence (D KL ) between the first probability distribution 311 for possible speech recognition hypotheses and the second probability distribution 394 for possible non-synthetic speech recognition hypotheses. The consistency loss term 352 based on D KL can be expressed by the following formula.
[0078]
Equation
[0079] Here, the consistency loss term 352 determined for the training utterance pair 301 at each time step gives a "teacherless" loss term that is independent of the accuracy of the auxiliary decoder 390 (e.g., independent of the teacher loss terms 342, 344 in FIG. 3B), and thus can be used to update the parameters of the audio encoder 210 to enhance the consistency between the alignment output and the non-synthesized audio representation of the same utterance. In batch training, the consistency loss term 352 may correspond to the average loss term obtained for the batch. In other words, the consistency loss term 352 enables learning such that the audio encoder 210 behaves similarly regardless of whether the training utterance belongs to the non-synthesized audio or the alignment output, e.g., enabling the audio encoder 210 to learn to predict a consistent encoded representation for both non-synthesized audio (e.g., real / human voice) and the alignment output of the same training utterance.
[0080] Finally, the training process 300 can combine the unpaired data loss function (J unpaired ), the paired data loss function (J paired ), and the consistency loss term (J cons ) to obtain the total loss term J tts4pretrain2 . This can be expressed as follows. J tts4pretrain2 =J unpaired +λ 1 J paired +λ 2 J cons (11) where λ 1 may be equal to 1.0 and λ 2 is equal to 0.1. The training process 300 minimizes the total loss term J tts4pretrain2By using it to update the parameters of the audio encoder 210 and effectively teaching the audio encoder 210 to learn the shared representation between speech and text, the audio encoder 210 may be pre-trained. After pre-training the audio encoder 210, the training process 300 may fine-tune the pre-trained audio encoder with the transcribed speech utterances. The transcribed speech utterances may include supervised training samples of both the alignment output corresponding to the unuttered text utterance 320 and the non-synthetic speech (e.g., human speech).
[0081] In some embodiments, the training process 300 for pre-training the audio encoder 210 applies encoder consistency regularization. Unlike decoder consistency regularization applied to the auxiliary decoder(s) in the consistency regularization unit 300c that requires assumed labels (e.g., the transcription 302 and the unuttered text utterance 320), encoder consistency regularization has the advantage of being applicable to all of the training data 304, 306, 320 because it does not require assumed labels. Encoder consistency regularization can be applied by the Hierarchical Contrastive Consistency Regularization (HCCR) method. In hierarchical contrastive consistency regularization, the encoder activation values e, e* from the original / non-augmented speech and the augmented speech are projected through an auxiliary network to generate z and z*. Then, the positive and negative pairs are formed, and the contrastive loss l t,z,z* is calculated as follows.
[0082]
Equation
[0083] The convolutional neural network (CNN) projection network, specific to HCCR, calculates projection values for increasing length segments (30, 50, 120 ms) of the encoder activation value e, resulting in three views (V), and can produce negative examples from the same utterance of short segments and other utterances in a batch of 120 ms segments. Therefore, the HCCR loss can be calculated as follows for the alignment output 602 generated from the transcribed non-synthetic speech utterances 304 (paired speech), untranscribed non-synthetic speech utterances 306 (unpaired speech), and unuttered text utterances 320.
[0084]
Number
[0085] The HCCR loss calculated by Equation 13 can be added to Equation 11 using a coefficient of 1e-3 as part of the total loss term J used for pre-training the audio encoder 210. tts4pretrain2 of 1e-3 as part of the total loss term J used for pre-training the audio encoder 210.
[0086] The above embodiments describe the training process 300 for pre-training the audio encoder 210. However, naturally, the training process 300 can also be used to train / pre-train the single-language ASR model 200 or the multi-language ASR model 200. In some examples, the training process 300 can be used to train (i.e., non-pre-train) an end-to-end ASR model having a decoder structure, or to fine-tune the ASR model for performing downstream tasks such as speech translation or natural language understanding. Further, the training process 300 can be used with each of the unuttered text utterances 320, the transcribed non-synthetic speech utterances 304, and the untranscribed non-synthetic speech utterances 306, or with some combinations thereof.
[0087] Referring to FIG. 4, the contrast unspoken text selection process 400 can select unspoken text utterances 320 for pre-training of the audio encoder 210 from a large unspoken text corpus 402, whereby the selected unspoken text utterances 320 will be most similar to a particular domain that the audio encoder 210 is pre-trained and learning. That is, the text selection process 400 can identify unspoken text within and near the domain from the unspoken text corpus 402 for inclusion in the unspoken text utterances 320 used for pre-training of the audio encoder 210. In particular, the unspoken text utterances 320 selected by the selection process 400 enable the on-the-fly synthesis of separate utterances during batch construction, and thus, a new speaker embedding value z and latent variable Z can be sampled each time the unspoken text utterances 320 are in one batch.
[0088] The corpus of unspoken text 402 includes a number of unspoken text utterances 320, 320a - n across a wide range of domains, having a much greater language diversity than the specific domain for which the audio encoder 210 is being trained. As described above, the set of transcribed non - synthetic speech utterances 304 can be domain - specific in that they relate to a specific domain and each non - synthetic speech utterance 304 is paired with a corresponding transcription 302. The corpus of unspoken text 402 can be stored in the same or a different data store 401 as the spoken training utterances 304. The corpus of unspoken text 402 may change dynamically to incorporate new unspoken text utterances 320. Simply using all of the unspoken text utterances 320 in the unspoken text corpus 402 is not feasible for the following reasons. That is, i) for each sentence, the audio modality requires much more memory for encoding than text, and thus it is not practical to convert all of the text in corpus 402, ii) there is a vast amount of difference between the transcriptions 302 paired with the transcribed non - synthetic speech utterances 304 and the unspoken text utterances 320 in the unspoken text corpus 402, and thus an intelligent strategy is needed to balance their contributions.
[0089] The text selection process 400 aims to select a subset of the available unuttered text utterances 320 from the unuttered text corpus 402 as data for TTS synthesis. TTS synthesis results in an alignment output that is generated to pre-train the audio encoder 210 in the contrastive loss part and the supervised loss part 300a, 300b of the training process 300 described above with reference to FIGS. 3A and 3B. In other words, the text selection process 400 aims to improve the match between the selected subset of the available unuttered text utterances 320 and the specific domain of interest, thereby reducing the computational resources required to utilize a large amount of non-domain-specific data. Therefore, the text selection process 400 reduces the computational cost and memory cost by selecting the unuttered text utterances 320 that best match the specific domain that the audio encoder 210 learns through training.
[0090] In some examples, the text selection process 400 selects a subset of the available unspoken text utterances 320 that best match a particular domain by simply providing a domain identifier (not shown) associated with the particular domain as input to the background LM 406 that has already been trained on the entire unspoken text corpus 402. As described above, the unspoken text corpus 402 spans a number of different domains. In these examples, the background LM 406 may include a maximum entropy (MaxEnt LM) that can accept a domain identifier as input as needed, as described in U.S. Patent No. 9,842,592, filed Feb. 12, 2014, the entire contents of which are incorporated herein by reference. Here, the domain identifier associated with a particular domain may enable the MaxEnt LM to output from the corpus 402 a subset of the available unspoken text utterances 320 that are likely to include words and / or phrases related to the particular domain. In some configurations, rather than evaluating word likelihoods, a statistical language model operates in reverse mode to randomly generate text phrases that match the statistical distribution of words related to a particular domain.
[0091] In additional examples, as shown in FIG. 4, the text selection process 400 uses the transcription 302 paired with the transcribed non-synthetic speech utterance 304 spoken by a human speaker to select a subset of the available unspoken text utterances 320 that best match a particular domain from the corpus 402. Here, the transcribed non-synthetic speech utterance 304 includes words, phrases, and / or other terms related to a particular domain. Optionally, in addition to or instead of the transcription 304 paired with the transcribed non-synthetic speech utterance 304, a different set of transcribed utterances related to the particular domain can be used to select the unspoken text utterance 320. This provides the advantage that the transcribed non-synthetic speech utterance 304 need not all belong to a particular domain.
[0092] During the first stage (stage A), the unspoken text selection process 400 constructs two language models 404, 406 to enable the contrastive selection of unspoken text utterances 320. Here, the domain-specific LM410 is trained on each transcription 302 in the set of transcribed non-synthetic speech utterances 304. The set of transcribed non-synthetic speech utterances 304 is assumed to belong to a specific region that the audio encoder 210 is trained for learning. On the other hand, the background LM406 is trained on each unspoken text utterance 320 in the entire unspoken text corpus 402. As described above, the unspoken text corpus 402 spans a number of different domains. In some examples, the first stage constructs the two language models 404, 406 using n-gram language model training. In some examples, the first stage constructs the two language models 404, 406 using neural network language model training.
[0093] During the second state (stage B), the unspoken text selection process 400 uses the two contrastive LMs 404, 406 to determine a first probability
[0094]
Number
[0095] associated with each word in the unspoken text utterance 320 that appears in the domain-specific LM404, and a second probability
[0096]
Number
[0097] By determining, each text utterance 320 of the unspoken text corpus 402 is evaluated. Thereafter, for each text utterance 320 of the unspoken text corpus 402, process 400 determines a score S at scorer 408 based on a first probability, a second probability, and the number #(w) of words appearing in the corresponding text utterance 320 of the unspoken text. For example, the score S for each text utterance 320 of the unspoken text can be calculated as follows.
[0098]
Number
[0099] After determining the scores, the unspoken text selection process 400 selects the text utterances 320 of the unspoken text having the N best scores S because those text utterances 320 of the unspoken text best match a particular domain. The text corpus 402 can include billions of text utterances 320 of the unspoken text. The text utterances 320 of the unspoken text selected by the selection process 400 can include millions of utterances and thus can far exceed the number of untranscribed non-synthetic speech utterances 304 spoken by a human speaker. As described above, the content of the text utterances 320 of the unspoken text increases the language diversity of a particular domain that the audio encoder 210 learns by training, while when the acoustic encoder 210 is integrated within the ASR model 200, the corresponding alignment output 602 generated from the text utterances 320 of the unspoken text increases the acoustic / language diversity of the speech that the acoustic encoder 210 encodes as part of the speech recognition process.
[0100] FIG. 5 shows an exemplary projection space 500 of the alignment output and the encoder representations of non-synthesized (real / human) speech utterances. After introducing consistency regularization via the consistency regularization unit 300c shown in FIG. 3C for pre-training the audio encoder, the resulting learned audio and text encoder representations are much closer to each other compared to the audio and text encoder representations without applying consistency regularization. Thus, as is clear from the projection space 500, improved shared audio and text representations are effectively generated by using supervised training data (i.e., transcribed non-synthesized speech utterances) for pre-training the audio encoder 210.
[0101] FIG. 9 is a flowchart showing an exemplary flow of operations for a method 900 of pre-training the audio encoder 210 to jointly learn shared representations of audio and text. The method 900 can be executed by data processing hardware 1010 (FIG. 10) using instructions stored in memory hardware 1020 (FIG. 10). The data processing hardware 1010 and the memory hardware 1020 may be present in the remote computer / server 201 of FIG. 1 corresponding to the computing device 1000 (FIG. 10).
[0102] In operation 902, method 900 includes receiving training data including unspoken text utterances 320, untranscribed non-synthetic speech utterances 306, and transcribed non-synthetic speech utterances 304. Each unspoken text utterance 320 is not paired with any corresponding speech utterance of non-synthetic speech. Each untranscribed non-synthetic speech utterance 306 is not paired with a corresponding transcription. Each transcribed non-synthetic speech utterance 304 is paired with a corresponding transcription 302. In operation 904, method 900 includes using alignment model 600 to generate a corresponding alignment output 602 for each unspoken text utterance 320 of the received training data. In operation 906, method 900 includes pre-training audio encoder 210 with alignment output 602 generated to correspond to unspoken text utterances 320, untranscribed non-synthetic speech utterances 306, and transcribed non-synthetic speech utterances 304 in order to teach the audio encoder 210 to jointly learn a shared representation of audio and text.
[0103] A software application (i.e., a software resource) can refer to computer software that causes a computing device to execute tasks. In some examples, a software application may be referred to as an "application", an "app", or a "program". Exemplary applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.
[0104] A non-transitory memory can be a physical device used to store a program (e.g., an instruction sequence) or data (e.g., program state information) in temporary or permanent basic elements used by a computing device. The non-transitory memory can be a volatile and / or non-volatile addressable semiconductor memory. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., used for firmware such as a boot program normally). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disks or tapes.
[0105] FIG. 10 is a schematic diagram of an exemplary computing device 1000 that can be used to implement the systems and methods described herein. Computing device 1000 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The components shown here, their connections and relationships, and their functions are intended to be exemplary only and are not intended to limit the embodiments of the invention described herein and / or the scope of the claims.
[0106] Computing device 1000 includes processor 1010, memory 1020, storage device 1030, high-speed interface / controller 1040 connected to memory 1020 and high-speed expansion port 1050, and low-speed interface / controller 1060 connected to low-speed bus 1070 and storage device 1030. Each of components 1010, 1020, 1030, 1040, 1050, and 1060 is interconnected using various buses and may be mounted on a shared motherboard or in other manners as required. Processor 1010 can process instructions for execution within computing device 1000. The instructions include those stored in memory 1020 or storage device 1030 for displaying graphical information of a graphical user interface (GUI) on an external input / output device such as display 1080 connected to high-speed interface 1040. In other embodiments, multiple processors and / or multiple buses may be used along with multiple memories and multiple types of memory as required. Also, multiple computing devices 1000 may be connected to each device performing a part of the required operations (e.g., as a server bank, a group of blade servers, or a multiprocessor system).
[0107] Memory 1020 stores information non - transiently within computing device 1000. Memory 1020 may be a computer - readable medium, a volatile memory unit(s), or a non - volatile memory unit(s). The non - transient memory 1020 may be a physical device used to store temporarily or persistently programs (e.g., instruction sequences) or data (e.g., program state information) used by computing device 1000. Examples of non - volatile memory include, but are not limited to, flash memory and read - only memory (ROM) / programmable read - only memory (PROM) / erasable programmable read - only memory (EPROM) / electrically erasable programmable read - only memory (EEPROM) (e.g., typically used for firmware such as a boot program). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase - change memory (PCM), and disks or tapes.
[0108] Storage device 1030 can provide large - capacity storage to computing device 1000. In some embodiments, storage device 1030 is a computer - readable medium. In various different embodiments, storage device 1030 may be an array of devices including a floppy disk (registered trademark) device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid - state memory device, or a device in the form of a storage area network or other configuration. In additional embodiments, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more of the methods as described above. The information carrier is a computer - readable or machine - readable medium such as memory 1020, storage device 1030, or memory on processor 1010.
[0109] The high-speed controller 1040 manages operations that use a large amount of the bandwidth of the computing device 1000, while the low-speed controller 1060 manages operations that use less bandwidth. Such role assignments are merely examples. In some embodiments, the high-speed controller 1040 is connected to a high-speed expansion port 1050 that can accept the memory 1020, the display 1080 (e.g., via a graphics processor or accelerator), and various expansion cards (not shown). In some embodiments, the low-speed controller 1060 is connected to the storage device 1030 and the low-speed expansion port 1090. The low-speed expansion port 1090, which may include various communication ports (e.g., USB, Bluetooth®, Ethernet®, wireless Ethernet), may be connected to one or more input / output devices such as a keyboard, a pointing device, a scanner, or may be connected to a network device such as a switch or a router, e.g., via a network adapter.
[0110] As shown in the figure, the computing device 1000 can be implemented in many different forms. For example, it may be implemented as a standard server 1000a, or as a repeat within a group of such servers 1000a, as a laptop computer 1000b, or as part of a rack server system 1000c.
[0111] The various embodiments of the systems and techniques described herein can be implemented in digital electronic circuitry and / or optical circuitry, integrated circuits, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments include embodiments that can be executed in a programmable system including at least one programmable processor and / or that can be translated into machine language by one or more computer programs. The at least one programmable processor may be special purpose or general purpose and can be connected to communicate with a storage system, at least one input device, and at least one output device for sending and receiving data and instructions.
[0112] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented in high-level procedural programming languages and / or object-oriented programming languages and / or assembly / machine languages. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., magnetic disks, optical disks, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to send machine instructions and / or data to a programmable processor.
[0113] The processes and logical flows described in this specification can be performed by one or more programmable processors (also called data processing hardware) that execute one or more computer programs to act on input data and generate output. The processes and logical flows may also be performed by special purpose logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, as well as any one or more processors of any type of digital computer. Generally, a processor receives instructions and data from a read only memory, a random access memory, or both. The basic elements of a computer are a processor for performing instructions and one or more memory devices for storing instructions and data. Generally, a computer also includes one or more mass storage devices (such as magnetic disks, magneto-optical disks, or optical disks) for storing data, or the computer is also operatively coupled to such a mass storage device so as to receive data from such a mass storage device, to transmit data to such a mass storage device, or both. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including by way of example semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices, magnetic disks such as internal hard disks or removable disks, magneto-optical disks, and CD ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, special purpose logic circuitry.
[0114] To interact with a user, one or more aspects of the present disclosure can be implemented on a computer having a display device, such as a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touch screen, for displaying information to the user, and optionally a keyboard and a pointing device (such as a mouse or trackball) by which the user can input to the computer. Other types of devices can also be used to provide interaction with the user. For example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and the input from the user can be received in any form, such as acoustic input, voice input, or tactile input. Further, the computer can interact with the user by transmitting and receiving documents to and from the devices used by the user (for example, by sending a web page to a web browser on the user's client device in response to a request received from a web browser).
[0115] Some embodiments have been described. However, it will be apparent that various modifications can be made without departing from the spirit and scope of the present disclosure. Accordingly, other embodiments are within the scope of the following claims.
Claims
1. A computer-implemented method (900) that causes the data processing hardware (1010) to operate when executed by the data processing hardware (1010), the operation comprising: Receiving training data (304, 306, 320), the training data (304, 306, 320) comprising: A plurality of unspoken text utterances (320), each unspoken text utterance (320) not paired with any corresponding spoken utterance of non-synthetic speech; A plurality of untranscribed non-synthetic speech utterances (306), each untranscribed non-synthetic speech utterance (306) not paired with a corresponding transcription; A plurality of transcribed non-synthetic speech utterances (304), each transcribed non-synthetic speech utterance (304) paired with a corresponding transcription (302); receiving the training data (304, 306, 320) comprising: Using an alignment model (600) to generate, for each unspoken text utterance (320) of the received training data (304, 306, 320), a corresponding alignment output (602), the corresponding alignment output (602) being a text representation that aligns a sequence of text chunks of the corresponding unspoken text utterance (320) to audio frames; Pre-training the audio encoder (210) with the plurality of unspoken text utterances (320), the plurality of untranscribed non-synthetic speech utterances (306), and the alignment outputs (602) generated corresponding to the plurality of transcribed non-synthetic speech utterances (304) to teach the audio encoder (210) to jointly learn a shared representation of audio and text.
2. The computer-implemented method (900) of claim 1, wherein the audio encoder (210) comprises a stack of self-attention layers each comprising a multi-head self-attention mechanism.
3. Pre-training the audio encoder (210) comprises, for each untranscribed non-synthetic speech utterance (306): For each untranscribed non-synthetic speech utterance (306): Generating a corresponding encoded representation (215) of the untranscribed non-synthesized speech utterance (306); Pre-training the audio encoder (210) with a contrastive loss (316) applied to the corresponding encoded representation (215) of the untranscribed non-synthesized speech utterance (306); For each alignment output (602), Generating a corresponding encoded representation (215) of the alignment output (602); Pre-training the audio encoder (210) with a contrastive loss (316) applied to the corresponding encoded representation (215) of the alignment output (602); For each transcribed non-synthesized speech utterance (304), Generating a corresponding encoded representation (215) of the transcribed non-synthesized speech utterance (304); Pre-training the audio encoder (210) with a contrastive loss (316) applied to the corresponding encoded representation (215) of the transcribed non-synthesized speech utterance (304), the computer-implemented method (900) according to claim 1 or 2.
4. Pre-training the audio encoder (210) comprises, At each of a plurality of time steps of each alignment output (602), Using an auxiliary decoder (390) to generate a first probability distribution (392) over possible synthetic speech recognition hypotheses for the corresponding alignment output (602); Determining an alignment output loss term (342) based on the first probability distribution (392) over possible synthetic speech recognition hypotheses and the unspoken text utterance (320) corresponding to the alignment output (602); Pre-training the audio encoder (210) based on the alignment output loss term (342); At each of a plurality of time steps of each transcribed non-synthesized speech utterance (304), Using the auxiliary decoder (390) to generate a second probability distribution (394) over possible non-synthetic speech recognition hypotheses for the corresponding transcribed non-synthesized speech utterance (304); Determining a non-synthetic speech loss term (344) based on the second probability distribution (394) for possible non-synthetic speech recognition hypotheses, the corresponding transcription (302) paired with the transcribed non-synthetic speech utterance (304); Pre-training the audio encoder (210) based on the non-synthetic speech loss term (344), the computer-implemented method (900) according to claim 1 or 2. **Claim 5** The computer-implemented method (900) according to claim 4, wherein the auxiliary decoder (390) includes one of a connectionist temporal classification (CTC) decoder, a listen attend spell (LAS) decoder, or a recurrent neural network transducer (RNN-T) decoder. **Claim 6** The first probability distribution (392) for possible synthetic speech recognition hypotheses includes one of possible phoneme labels or possible word piece labels, The computer-implemented method (900) according to claim 4, wherein the second probability distribution (394) for possible non-synthetic speech recognition hypotheses includes the one of the possible phoneme labels or the possible word piece labels. **Claim 7** The computer-implemented method (900) according to claim 1 or 2, wherein the audio encoder (210) includes a text encoder (202), an audio encoder (204), and a shared encoder (250). **Claim 8** The operation further includes, For each alignment output (602), Determining an encoded text representation (312) of the alignment output (202) using the text encoder (202); Generating a first encoded shared representation (322) of the alignment output (602) in a shared latent representation space using the shared encoder (250); For each transcribed non-synthetic speech utterance (304), Determining an encoded audio representation (314) of the transcribed non-synthetic speech utterance (304) using the audio encoder (204); Generating a second encoded shared representation (324) of the transcribed non-synthetic speech utterance (304) in a shared latent representation space using the shared encoder (250), the computer-implemented method (900) according to claim 7. **Claim 9** Generating the corresponding alignment output (602) for each unspoken text utterance (320) of the received training data (304, 306, 320) comprises: extracting an initial text representation (612) from the unspoken text utterance (320); predicting a text chunk duration (622) for each text chunk in the unspoken text utterance (320); upsampling the initial text representation (612) using the predicted text chunk duration (622) for each text chunk in the unspoken text utterance (320), the computer-implemented method (900) according to claim 1 or 2.
10. The operation further comprises training the alignment model (600), the training comprising: generating an encoded audio representation (314) for a transcribed non-synthetic speech utterance (304) using an audio encoder (204); determining an alignment output (602) for a transcription (302) corresponding to the transcribed non-synthetic speech utterance (304) using the alignment model (600); generating an encoded text representation (312) for the alignment output (602) using a text encoder (202); updating parameters of the alignment model (600) based on a comparison between the encoded audio representation (314) for the transcribed non-synthetic speech utterance (304) and the encoded text representation (312) for the alignment output (612), the computer-implemented method (900) according to claim 1 or 2.
11. A system comprising data processing hardware (1010) and memory hardware (1020) communicating with the data processing hardware (1010), the memory hardware (1020) storing instructions that, when executed on the data processing hardware (1010), cause the data processing hardware (1010) to perform operations, the operations comprising: receiving training data (304, 306, 320), the training data (304, 306, 320) comprising: receiving training data (304, 306, 320), the training data (304, 306, 320) comprising: A plurality of unuttered text utterances (320), wherein each unuttered text utterance (320) is not paired with any corresponding spoken utterance of non-synthetic speech, and the plurality of unuttered text utterances (320); A plurality of untranscribed non-synthetic speech utterances (306), wherein each untranscribed non-synthetic speech utterance (306) is not paired with a corresponding transcription, and the plurality of untranscribed non-synthetic speech utterances (306); A plurality of transcribed non-synthetic speech utterances (304), wherein each transcribed non-synthetic speech utterance (304) is paired with a corresponding transcription (302), and the plurality of transcribed non-synthetic speech utterances (304), including receiving the training data (304, 306, 320); Using an alignment model (600) to generate a corresponding alignment output (602) for each unuttered text utterance (320) of the received training data (304, 306, 320), wherein the corresponding alignment output (602) is a text representation that aligns a sequence of text chunks of the corresponding unuttered text utterance (320) with audio frames; A system including pre-training the audio encoder (210) with the alignment outputs (602) generated to correspond to the plurality of unuttered text utterances (320), the plurality of untranscribed non-synthetic speech utterances (306), and the plurality of transcribed non-synthetic speech utterances (304) in order to teach the audio encoder (210) to jointly learn a shared representation of audio and text.
12. The system (100) according to claim 11, wherein the audio encoder (210) comprises a stack of self-attention layers each including a multi-head self-attention mechanism.
13. Pre-training the audio encoder (210) comprises: For each untranscribed non-synthetic speech utterance (306), Generating a corresponding encoded representation (215) of the untranscribed non-synthetic speech utterance (306); Pre-training the audio encoder (210) with a contrastive loss (316) applied to the corresponding encoded representation (215) of the untranscribed non-synthetic speech utterance (306); For each alignment output (602), Generating a corresponding encoded representation (215) of the alignment output (602); Pre-training the audio encoder (210) with a contrastive loss (316) applied to the corresponding encoded representation (215) of the alignment output (602); For each transcribed non-synthetic speech utterance (304), Generating a corresponding encoded representation (215) of the transcribed non-synthetic speech utterance (304); Pre-training the audio encoder (210) with a contrastive loss (316) applied to the corresponding encoded representation (215) of the transcribed non-synthetic speech utterance (304), the system (100) according to claim 11 or 12.
14. Pre-training the audio encoder (210) comprises At each of a plurality of time steps of each alignment output (602), Using an auxiliary decoder (390) to generate a first probability distribution (392) over possible synthetic speech recognition hypotheses for the corresponding alignment output (602); Determining an alignment output loss term (342) based on the first probability distribution (392) over possible synthetic speech recognition hypotheses and the unspoken text utterance (320) corresponding to the alignment output (602); Pre-training the audio encoder (210) based on the alignment output loss term (342); At each of a plurality of time steps of each transcribed non-synthetic speech utterance (304), Using the auxiliary decoder (390) to generate a second probability distribution (394) over possible non-synthetic speech recognition hypotheses for the corresponding transcribed non-synthetic speech utterance (304); Determining a non-synthetic speech loss term (344) based on the second probability distribution (394) over possible non-synthetic speech recognition hypotheses and the corresponding transcription (302) paired with the transcribed non-synthetic speech utterance (304); Pre-training the audio encoder (210) based on the non-synthetic speech loss term (344), the system (100) according to claim 11 or 12.
15. The auxiliary decoder (390) includes one of a connection time classification (CTC) decoder, a listen attend spell (LAS) decoder, or a recurrent neural network transducer (RNN-T) decoder, and the system (100) according to claim 14.
16. The first probability distribution (392) for possible synthetic speech recognition hypotheses includes one of possible phoneme labels or possible word piece labels, The second probability distribution (394) for possible non-synthetic speech recognition hypotheses includes the one of the possible phoneme labels or the possible word piece labels, and the system (100) according to claim 14.
17. The audio encoder (210) includes a text encoder (202), an audio encoder (204), and a shared encoder (250), and the system (100) according to claim 11 or 12.
18. The operation further includes For each alignment output (602), Using the text encoder (202) to determine an encoded text representation (312) of the alignment output (202); Using the shared encoder (250) to generate a first encoded shared representation (322) of the alignment output (602) in a shared latent representation space; For each transcribed non-synthetic speech utterance (304), Using the audio encoder (204) to determine an encoded audio representation (314) of the transcribed non-synthetic speech utterance (304); Using the shared encoder (250) to generate a second encoded shared representation (324) of the transcribed non-synthetic speech utterance (304) in a shared latent representation space, and the system (100) according to claim 17.
19. Generating the corresponding alignment output (602) for each unspoken text utterance (320) of the received training data (304, 306, 320) includes Extracting an initial text representation (612) from the unspoken text utterance (320); Predicting a text chunk duration (622) for each text chunk in the unspoken text utterance (320); upsampling the initial text representation (612) using the predicted text chunk duration (622) for each text chunk in the unspoken text utterance (320); The system (100) according to claim 11 or 12, comprising:
20. The operation further includes training the alignment model (600), and the training includes: generating an encoded audio representation (314) for the transcribed non-synthetic speech utterance (304) using an audio encoder (204); determining an alignment output (602) for a transcription (302) corresponding to the transcribed non-synthetic speech utterance (304) using the alignment model (600); generating an encoded text representation (312) for the alignment output (602) using a text encoder (202); updating parameters of the alignment model (600) based on a comparison between the encoded audio representation (314) for the transcribed non-synthetic speech utterance (304) and the encoded text representation (312) for the alignment output (612); The system (100) according to claim 11 or 12, which is performed by:
Citation Information
Cited By
Using Aligned Text and Speech Representations to Train Automatic Speech Recognition Models Without Transcribed Speech Data
JP2025525617A