Using Aligned Text and Speech Representations to Train Automatic Speech Recognition Models Without Transcribed Speech Data

By using aligned text and speech representations, the method trains ASR models on unspoken text utterances with an alignment model, addressing the challenge of lacking labeled data in low-resource languages and achieving effective speech recognition without transcribed data.

JP7799899B2Active Publication Date: 2026-01-15GOOGLE LLC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025503139
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-07-22
Filing Date
2023-07-20
Publication Date
2026-01-15
Estimated Expiration
2043-07-20

AI Technical Summary

Technical Problem

Training automatic speech recognition (ASR) models requires substantial labeled training data, which is costly and often unavailable for low-resource languages, making it challenging to develop effective ASR models without transcribed speech data.

Method used

Utilizing aligned text and speech representations, the method trains ASR models using unspoken text utterances in a target language, generating aligned outputs with an alignment model trained on transcribed speech utterances in different languages, and encoding these outputs to teach the model to recognize speech without supervised learning from transcribed data.

Benefits of technology

Enables effective ASR model training in low-resource languages by leveraging unlabeled data, reducing the need for costly labeled data and enabling speech recognition in target languages without transcribed speech data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007799899000017
    Figure 0007799899000017
  • Figure 0007799899000018
    Figure 0007799899000018
  • Figure 0007799899000019
    Figure 0007799899000019
Patent Text Reader

Abstract

The method (700) includes receiving training data including unspoken text utterances (320) in a target language. Each unspoken text utterance is not paired with any corresponding spoken utterance of unsynthesized speech. The method also includes generating a corresponding aligned output (402) for each unspoken text utterance using an alignment model (400) trained based on transcribed speech utterances (304) in one or more training languages different from the target language. The method also includes generating a corresponding coded text representation (312) for each aligned output using a text encoder (202) and training a speech recognition model (200) with the coded text representations generated for the aligned outputs. Training the speech recognition model teaches the speech recognition model to learn how to recognize speech in the target language.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to the use of aligned text and speech representations to train automatic speech recognition models without transcribed speech data. [Background technology]

[0002] Automatic speech recognition (ASR), the process of taking audio input and transcribing it into text, has become an increasingly important technology used in mobile and other devices. Generally, automatic speech recognition attempts to provide an accurate transcription of what a person said by taking audio input (e.g., a spoken utterance) and transcribing that audio input into text. Based on ongoing developments in deep neural networks, modern ASR models continue to improve in both accuracy (e.g., low word error rate (WER)) and latency (e.g., the delay between a user's utterance and transcription). However, one of the challenges in developing deep learning-based ASR models is the need for a significant amount of transcribed audio during training. In some instances, unspoken text data is used to train the ASR model to supplement a small set of transcribed audio training data. However, this challenge becomes even more complex when training an ASR model in a low-resource language for which no transcribed audio training data is available. Summary of the Invention

[0003] One aspect of the present disclosure provides a computer-implemented method that, when executed on data processing hardware, causes the data processing hardware to perform operations for using aligned text and speech representations to train an automatic speech recognition model without transcribed speech data. The operations include receiving training data including unspoken text utterances in a target language, where each unspoken text utterance is not paired with any corresponding spoken utterance in unsynthesized speech. The operations also include generating a corresponding aligned output for each unspoken text utterance in the received training data using an alignment model trained based on transcribed speech utterances in one or more training languages ​​different from the target language. The operations also include generating a corresponding encoded text representation for each aligned output using a text encoder. The operations also include training a speech recognition model with the encoded text representations generated for the aligned output corresponding to the unspoken text utterances in the target language to teach the speech recognition model to learn how to recognize speech in the target language.

[0004] Implementations of the present disclosure may include one or more of the following optional features. In some implementations, training the speech recognition model includes training the speech recognition model without using any transcribed speech utterances in the target language for supervised learning. In some examples, the speech recognition model includes an audio encoder and a decoder. The decoder may include a recurrent neural network transducer (RNN-T) architecture. In these examples, the audio encoder includes a text encoder, a speech encoder, and a shared encoder. Here, the encoder may include multiple multi-head self-attention layers. The audio encoder may include a conformer encoder. In these examples, the operations further include conditioning at least one of the audio encoder or decoder based on a language identifier that uniquely identifies the target language. Conditioning at least one of the audio encoder or decoder includes conditioning at least one of the audio encoder or decoder based on the language identifier using a residual adapter layer. The alignment model may be trained based on transcribed speech utterances in one or more training languages ​​different from the target language.

[0005] In some implementations, training the speech recognition model includes: for each alignment output, generating a first shared coded representation of the alignment output in a shared latent representation space using a shared encoder; for each transcribed speech utterance in one or more training languages, determining a coded audio representation of the transcribed speech utterance using a speech encoder; and generating a second shared coded representation of the transcribed speech utterance in the shared latent representation space. Here, training the speech recognition model includes training the speech recognition model with the first shared coded representation generated for the alignment output corresponding to an unspoken text utterance in the target language and the second shared coded representation generated for the transcribed speech utterance in one or more training languages. Each unspoken text utterance may include a sequence of words, word pieces, graphemes, and / or phonemes. In some implementations, the operations further include converting a script of the unspoken text utterance in the target language into a phoneme representation shared across multiple languages ​​using a pronunciation model.

[0006] Another aspect of the present disclosure provides a system including data processing hardware and memory hardware storing instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations. The operations include receiving training data including unspoken text utterances in a target language, where each unspoken text utterance is not paired with any corresponding spoken utterance of unsynthesized speech. The operations also include generating corresponding alignment outputs for each unspoken text utterance of the received training data using alignment models each trained based on transcribed speech utterances in one or more training languages ​​different from the target language. The operations also include generating corresponding encoded text representations for each alignment output using a text encoder. The operations also include training a speech recognition model with the generated encoded text representations for the alignment outputs corresponding to the unspoken text utterances in the target language to teach the speech recognition model to learn how to recognize speech in the target language.

[0007] Implementations of the present disclosure may include one or more of the following optional features. In some implementations, training the speech recognition model includes training the speech recognition model without using any transcribed speech utterances in the target language for supervised learning. In some examples, the speech recognition model includes an audio encoder and a decoder. The decoder may include a recurrent neural network transducer (RNN-T) architecture. In these examples, the audio encoder includes a text encoder, a speech encoder, and a shared encoder. Here, the encoder may include multiple multi-head self-attention layers. The audio encoder may include a conformer encoder. In these examples, the operations further include conditioning at least one of the audio encoder or decoder based on a language identifier that uniquely identifies the target language. Conditioning at least one of the audio encoder or decoder includes conditioning at least one of the audio encoder or decoder based on the language identifier using a residual adapter layer. The alignment model may be trained based on transcribed speech utterances in one or more training languages ​​different from the target language.

[0008] In some implementations, training the speech recognition model includes: for each alignment output, generating a first shared coded representation of the alignment output in a shared latent representation space using a shared encoder; for each transcribed speech utterance in one or more training languages, determining a coded audio representation of the transcribed speech utterance using a speech encoder; and generating a second shared coded representation of the transcribed speech utterance in the shared latent representation space. Here, training the speech recognition model includes training the speech recognition model with the first shared coded representation generated for the alignment output corresponding to an unspoken text utterance in the target language and the second shared coded representation generated for the transcribed speech utterance in one or more training languages. Each unspoken text utterance may include a sequence of words, word pieces, graphemes, and / or phonemes. In some implementations, the operations further include converting a script of the unspoken text utterance in the target language into a phoneme representation shared across multiple languages ​​using a pronunciation model.

[0009] The details of one or more embodiments of the disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will become apparent from the description and drawings, and from the claims. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 is a schematic diagram of an exemplary speech recognition system. [Figure 2] FIG. 1 is a schematic diagram of an exemplary speech recognition model. [Figure 3A] FIG. 3 is a schematic diagram of an example training process for training an audio encoder of the speech recognition model of FIG. [Figure 3B] FIG. 3 is a schematic diagram of an example training process for training an audio encoder of the speech recognition model of FIG. [Figure 3C]FIG. 3 is a schematic diagram of an example training process for training an audio encoder of the speech recognition model of FIG. [Figure 4] FIG. 4 is a schematic diagram of an alignment model used in an exemplary training process for training an audio encoder of the speech recognition model of FIGS. 3A-3C. [Figure 5] FIG. 1 is a schematic diagram of an exemplary training process for training a duration predictor of an alignment model. [Figure 6] FIG. 1 is a schematic diagram of an exemplary training process for training an upsampler for an alignment model. [Figure 7] 1 is a flowchart of an exemplary arrangement of operations of a method for using aligned text and speech representations to train a speech recognition model without transcribed speech data. [Figure 8] FIG. 1 is a schematic diagram of an exemplary computing device that can be used to implement the systems and methods described herein. DETAILED DESCRIPTION OF THE INVENTION

[0011] Like reference symbols in the various drawings indicate like elements.

[0012] Training state-of-the-art automatic speech recognition (ASR) models typically requires a substantial amount of labeled training data, including each speech utterance paired with a corresponding transcription (i.e., ground truth label). Obtaining a substantial amount of labeled training data can be prohibitively costly, especially for low-resource languages. Therefore, recent approaches to training ASR models include self-supervised training, which uses a large amount of unlabeled training data (i.e., speech not paired with any corresponding transcription and / or unspoken text not paired with any corresponding speech) to supplement a relatively small amount of labeled training data. However, even these approaches still require some small amount of labeled training data in the particular language the ASR model is trained to recognize. Furthermore, in some cases, no labeled training data is available in a particular low-resource language. Therefore, in these instances, there is no labeled training data to supplement the large amount of unlabeled training data.

[0013] Accordingly, embodiments herein are directed to methods and systems that use aligned text and speech representations to train an ASR model without transcribed speech training data. In particular, training the ASR model includes receiving training data including unspoken text utterances in a target language, where each unspoken text utterance is not paired with any corresponding utterance in unsynthesized (or synthetic) speech. The alignment model generates a corresponding aligned output (i.e., an aligned text representation) for each unspoken text utterance in the received training data. Specifically, the alignment model is trained based on transcribed (i.e., labeled) speech utterances in one or more training languages, each different from the target language that the ASR model is trained to recognize. Furthermore, a text encoder generates a corresponding coded text representation for each aligned output. The ASR model is then trained based on the coded text representations generated for the aligned outputs corresponding to the unspoken text utterances in the target language to teach the ASR model to learn how to recognize speech in the target language when labeled training data in the target language is not available. That is, the ASR model is trained without using any transcribed speech utterances in the target language for supervised learning.

[0014] 1 illustrates an ASR system 100 executing an ASR model 200 residing on a user device 102 of a user 104 and / or on a remote computing device 201 (e.g., one or more servers of a distributed system running in a cloud computing environment) that communicates with the user device 102. While the user device 102 is shown as a mobile computing device (e.g., a smartphone), the user device 102 can correspond to any type of computing device, such as, without limitation, a tablet device, a laptop / desktop computer, a wearable device, a digital assistant device, a smart speaker / display, a smart appliance, an automobile infotainment system, or an Internet of Things (IoT) device, and includes data processing hardware and memory hardware 113.

[0015] The user device 102 includes an audio subsystem 108 configured to receive utterances 106 spoken by a user 104 (e.g., the user device 102 may include one or more microphones for recording the spoken utterances 106) and convert the utterances 106 into a corresponding digital format associated with input acoustic frames 110 processable by the ASR system 100. In the example shown, the user makes each utterance 106 in the natural language of English for the phrase "What's the weather like in New York City?", and the audio subsystem 108 converts the utterance 106 into the corresponding acoustic frames 110 for input to the ASR system 100. The ASR model 200 then receives the acoustic frames 110 corresponding to the utterances 106 as input and generates / predicts a corresponding transcription 120 (e.g., a recognition result / hypothesis) of the utterance 106 as output. In the illustrated example, the user device 102 and / or the remote computing device 201 also execute a user interface generator 107 configured to present a representation of a transcription 120 of the utterance 106 to the user 104 of the user device 102. In some configurations, the transcription 120 output from the ASR system 100 is processed by a natural language understanding (NLU) module executing, for example, on the user device 102 or the remote computing device 201 to execute user commands. Additionally or alternatively, a text-to-speech system (e.g., executing on any combination of the user device 102 or the remote computing device 201) may convert the transcription into synthesized speech for audible output by another device. For example, the original utterance 106 may correspond to a message the user 104 is sending to a friend, in which case the transcription 120 is converted into synthesized speech for audible output so that the friend can hear the message conveyed in the original utterance 106.

[0016] The ASR model 200 may operate in a streaming manner, a non-streaming manner, or some combination thereof. The ASR model 200 operates in a streaming manner by receiving a sequence of acoustic frames 110, encoding the sequence of acoustic frames 110, and then decoding the encoded sequence of acoustic frames 110 into an initial transcription (e.g., a speech recognition result / hypothesis) 120. Thus, the initial transcription 120 may correspond to words, word pieces, and / or individual characters that are generated by the ASR model 200 as they are spoken. Alternatively, the ASR model 200 operates in a non-streaming manner by receiving and processing additional right context to improve the initial transcription 120 and thereby generate the final transcription 120. That is, the ASR model 200 processes additional input audio data or encoded acoustic frames (e.g., right context) to improve the transcription 120 output by the ASR model 200, but at the expense of increased latency.

[0017] Referring to FIG. 2, an exemplary ASR model 200 may include a recurrent neural network transducer (RNN-T) model architecture that complies with latency constraints in interactive applications. The use of an RNN-T model architecture is exemplary only, and the ASR model 200 may include other architectures, such as a transformer-transducer model architecture and a conformer-transducer model architecture, among others. Because the RNN-T model 200 has a small computational footprint and utilizes fewer memory requirements than traditional ASR architectures, the RNN-T model architecture is suitable for performing all speech recognition on the user device 102 (e.g., without requiring communication with a remote server). The RNN-T model 200 includes an encoder network 210, a prediction network 220, and a joint network 230. The encoder network 210 is generally similar to an acoustic model (AM) in a traditional ASR system and includes a stack of self-attention layers (e.g., conformer or transformer layers) or a recurrent network of stacked long short-term memory (LSTM) layers. For example, an encoder network (e.g., an audio encoder) 210 may generate a sequence of d-dimensional feature vectors (e.g., audio frames 110 (FIG. 1)) x=(x1, x2, . . . , x T ), in the formula,

number

number

[0018] Similarly, the prediction network 220 is also an LSTM network, and like a language model (LM), it computes the sequence of non-empty symbols y0,...,y0 output so far by the final softmax layer 240. ui-1 , a dense representation

number

number

[0019] The softmax layer 240 may use any technique to select the output label / symbol with the highest probability within the distribution as the next output symbol predicted by the RNN-T model 200 at the corresponding output step. In this way, the RNN-T model 200 does not make a conditional independence assumption; rather, the prediction of each symbol is conditioned not only on the acoustics but also on the sequence of labels output so far. The RNN-T model 200 assumes that the output symbols are independent of future acoustic frames 110, which allows the RNN-T model to be used in a streaming manner, a non-streaming manner, or some combination thereof.

[0020] In some examples, the audio encoder 210 of the RNN-T model includes multiple multi-head (e.g., eight-head) self-attention layers. For example, the multiple multi-head self-attention layers may include conformer layers (e.g., conformer encoders), transformer layers, performer layers, convolutional layers (including lightweight convolutional layers), or any other type of multi-head self-attention layer. The multiple multi-head self-attention layers may include any number of layers, for example, 16 layers. Furthermore, the audio encoder 210 may operate in a streaming manner (e.g., the audio encoder 210 outputs initial higher-level feature representations as soon as they are generated), a non-streaming manner (e.g., the audio encoder 210 outputs subsequent higher-level feature representations by processing additional right context to improve the initial higher-level feature representation), or a combination of both streaming and non-streaming manners.

[0021] 3A-3C illustrate an exemplary training process 300 for training the ASR model 200 (FIG. 2). While the training process 300 described herein describes training the audio encoder 210 of the ASR model 200, it is understood that the training process 300 may also include pre-training and / or fine-tuning training of the audio encoder 210. Furthermore, the embodiments described herein contemplate the training process 300 training the audio encoder 210 of the ASR model 200 without training the decoder (e.g., the prediction network 220 and the joint network 230) of the ASR model 200. However, it is understood that the training process 300 may additionally or alternatively train other components of the ASR model 200 (e.g., the prediction network 220 and / or the joint network 230) along with the audio encoder 210.

[0022] The training process 300 involves generating unspoken text utterances (X text ) a set of 320 transcribed non-synthesized speech utterances (X sup ) 304 set, and / or untranscribed non-synthesized speech utterances (X unsup) 306 to train the audio encoder 210. In particular, each unspoken text utterance 320 in the set of unspoken text utterances 320 includes text-only data (i.e., unpaired data) in the target language, where each unspoken text utterance 320 is not paired with a corresponding spoken audio representation (i.e., speech) of the utterance. Here, the target language is any language that the training process 300 uses to train the audio encoder 210 to recognize locations without using any transcribed (i.e., labeled) training data during training. The unspoken text utterances 320 may include any sequence of text chunks, including words, word pieces, phonemes, and / or graphemes. Optionally, the available training data may also include untranscribed, unsynthesized speech utterances 306 (also referred to simply as "untranscribed speech utterances 306"), each comprising audio-only data (i.e., unpaired data) in the target language, where the untranscribed speech utterances 306 are not paired with any corresponding transcription. In particular, if the training data includes untranscribed speech utterances 306 in addition to unspoken text utterances 306, the training process 300 cannot simply pair the untranscribed text utterances 320 with the untranscribed speech utterances 306 to generate labeled training data, because the untranscribed speech utterances 306 represent training utterances that are different from the unspoken text utterances 320.

[0023] Meanwhile, each transcribed non-synthesized speech utterance 304 (also referred to simply as a "transcribed speech utterance 304") includes a corresponding transcription 302 paired with a corresponding non-synthesized speech representation of the corresponding transcribed speech utterance 304 in one or more training languages. Each of the one or more training languages ​​is different from the target language. For example, the one or more training languages ​​may include 52 languages ​​with transcribed speech utterances, and the target language may include 50 other languages ​​(e.g., each different from the 52 training languages) with text-only training data. As will become apparent, the transcribed speech utterances 304 are used to train an alignment model 400 to generate alignment outputs 402 of the transcribed speech utterances 304 in one or more training languages. The trained alignment model 400 is then configured to receive as input unspoken text utterances 320 in a target language (e.g., different from each of the one or more training languages ​​used to train the alignment model 400) and to produce as output an aligned output 402 in the target language that was used to train the audio encoder 210. Thus, training the alignment model 400 using transcribed speech utterances 304 in one or more training languages ​​enables the alignment model 400 to generate an aligned output 402 in the target language even though the alignment model 400 was not trained with any training data in the target language.

[0024] For clarity, the training process 300 includes a symmetric self-supervised loss section 300a (FIG. 3A), a supervised loss section 300b (FIG. 3B), and a consistency regularizer section 300c (FIG. 3C). The training process 300 generates training text utterances (X text ) 320 (e.g., in the target language), transcribed non-synthesized speech utterances (X sup) 304 corpora (e.g., in one or more training languages), and untranscribed, non-synthesized speech utterances (X unsup ) 306 (e.g., in the target language) to a contrastive loss (L w2v ) 316 and the unspoken training text utterances (X text )320 and transcribed non-synthesized speech (X sup ) 304 using the supervised loss unit 300b. aux ) 342, 344 and the consistency loss (J cons (θ)) 352 and the total loss based on (L tts4pretrain2 ) to train the audio encoder 210.

[0025] 3A, the contrasting self-supervised loss portion 300a of the training process 300 may employ an alignment model 400 configured to generate, at each of a plurality of output steps, an aligned output (i.e., text representation) 402 for each of a plurality of unspoken training text utterances 320. The unspoken text utterances 320 include unspoken text that is text-only data, i.e., unpaired data, in the target language, and each unspoken text utterance (X text ) 320 is not paired with any synthesized or unsynthesized speech. Thus, for each unspoken text utterance 320, the alignment model 400 generates a corresponding alignment output 402.

[0026] 4, in some examples, the alignment model 400 includes an embedding extractor 410, a duration predictor 420, and an upsampler 430. The embedding extractor 410 receives an unspoken text utterance 320 that includes a sequence of text chunks that include words, word pieces, phonemes, and / or graphemes, and extracts a corresponding initial text representation (e t) 412. The initial text representation 412 includes embedded lexical information from the unspoken text utterance 320. Additionally or alternatively, the embedding extractor 410 may receive a transcription 302 corresponding to the transcribed non-synthesized speech utterance 304 (FIG. 3C). The duration predictor 420 receives the initial text representation 412 from the embedding extractor 410 and predicts corresponding text chunk durations (i.e., word, wordpiece, phoneme, and / or grapheme durations) 422. The text chunk durations 422 indicate the durations that the corresponding text chunks would be spoken if a human (or text-to-speech system) spoke the unspoken text utterance 320. For example, the unspoken text utterance 320 may include a sequence of phonemes, and the duration predictor 420 predicts the phoneme durations 422 for each phoneme in the sequence of phonemes. In this example, the duration predictor 420 predicts phoneme durations 422 by predicting the probability of a non-zero duration for each phoneme and the probability of consecutive phoneme durations for each phoneme. Because a sequence of phonemes includes regular phonemes, silence between word boundaries, and punctuation, only regular phonemes are associated with non-zero durations, while silence and punctuation are associated with regular consecutive phoneme durations. Thus, the duration predictor 420 can predict the probability of a non-zero duration using a sigmoid activation following the first of two independent activations and predict the consecutive text chunk duration 422 for each text chunk using a soft-plus activation following the second of two independent activations. For each text chunk, the duration predictor 420 determines whether the probability of a non-zero duration is less than a threshold, and if the probability of a non-zero duration is less than the threshold, a multiplier can zero out the consecutive text chunk duration 422 predicted by the soft-plus activation for the corresponding text chunk. On the other hand, if the probability of a non-zero duration is greater than or equal to a threshold, the predicted text chunk duration 422 can be set equal to the consecutive phoneme duration predicted by soft-plus activation.

[0027] The upsampler 430 receives, for each unspoken text utterance 320, the corresponding initial text representation 412 and the predicted text chunk duration 422, and upsamples the initial text representation 412 using the corresponding predicted text chunk duration 422 to generate an alignment output having a frame number.

number

number

number

[0028] In particular, the frame numbers of the alignment output 402 indicate the predicted audio duration of the unspoken text utterance 320. Stated another way, the frame numbers of the alignment output 402 map (i.e., align) the sequence of text chunks of the unspoken text utterance 320 to audio frames. Here, the upsampler 430 includes a resampler layer and a refiner layer that replicate the initial text embeddings 412 to match the predicted text chunk durations 422 (i.e., audio durations). Thus, the alignment output 402 includes a text representation of the unspoken text utterance 320 with timing components that align with how a human would speak the unspoken text utterance 320. In some examples, the embedding extractor 410 receives a language identifier 321 that uniquely identifies the language of one or more training and / or target languages ​​for conditioning the alignment model 400.

[0029] In some implementations, training an alignment model 400 using training data in one or more training languages ​​and then generating an alignment output 402 in a target language (e.g., different from each of the training languages) leads to a lower-quality alignment output 402 because the script of the target language (e.g., Brahmi) does not overlap with the script of the one or more training languages. To that end, in some examples, the alignment model 400 includes a pronunciation model that converts the script of the unspoken text utterance 320 in the target language into a representation (e.g., a phonemic representation) that is shared across multiple languages. In other examples, the alignment model 400 may convert the script of the unspoken text utterance 320 into a different script; that is, the different script can be aligned with the script of the transcribed speech utterance 304.

[0030] In particular, in most cases, text-to-speech (TTS) systems generate an audible output that imparts the timing components of human speech to an unspoken text utterance 320, allowing the training process to train the audio encoder 210 using the audible output (i.e., synthetic speech) from the TTS system. However, the alignment model 400 advantageously generates an alignment output 402, thereby directly mapping a sequence of text chunks to speech frames. Thus, the training process 300 does not require any TTS system to generate synthetic speech from the unspoken text utterance 320 in order to train the audio encoder 210. That is, neither the training process 300 nor the alignment model 400 generates the alignment output 402 (i.e., text alignment) rather than converting the unspoken text utterance 320 into synthetic speech.

[0031] FIG. 5 illustrates an exemplary training process 500 for training alignment model 400 using transcribed non-synthesized speech utterances 304 in one or more training languages. That is, each of the one or more training languages ​​is a different language than the target language that audio encoder 210 (FIG. 3) is trained to recognize. Furthermore, each transcribed non-synthesized speech utterance 304 has a corresponding transcription 302. In the illustrated example, speech encoder 204 receives as input each transcribed non-synthesized speech utterance 304 as a sequence of features / vectors (e.g., Mel-frequency spectrograms, such as acoustic frames 110 in FIG. 1 ) and as output, for each of a plurality of output steps, an encoded audio representation (e.g., a .epsilon.) corresponding to the transcribed non-synthesized speech utterance 304 at the corresponding output step. s) 314. In parallel, the alignment model 400 receives the transcription 302 corresponding to the same unsynthesized speech utterance 304 and generates the alignment output 402 according to Equation 1. The text encoder 202 receives the alignment output 402 as input and, for each of a plurality of output steps, generates as output an encoded text representation corresponding to the transcription 302 at the corresponding output step.

number

[0032] The modality loss module 550 receives the encoded text representation 312 and the encoded audio representation 314 and generates a modality loss 552 based on comparing the encoded text representation 312 and the encoded audio representation as follows:

number

number

[0033] FIG. 6 illustrates an exemplary training process 600 for training the alignment model 400 using paired training data (i.e., transcribed speech utterances 304 in one or more training languages) and unpaired training data (i.e., unspoken text utterances 320 in a target language). That is, the training process 600 uses transcribed non-synthesized speech utterances 304 with corresponding transcriptions 302 (i.e., paired training data) and unspoken text utterances 320 (i.e., unpaired training data) to learn a speech-aligned alignment output 402. In the illustrated example, the speech encoder 204 receives as input each transcribed non-synthesized speech utterance 304 as a sequence of features / vectors (e.g., Mel-frequency spectrograms, such as acoustic frames 110 in FIG. 1 ), and as output, for each of a plurality of output steps, an encoded audio representation (e.g., a sigma-based representation) corresponding to the transcribed non-synthesized speech utterance 304 at the corresponding output step. s ) 314. In parallel, the alignment model 400 receives the transcription 302 corresponding to the same unsynthesized speech utterance 304 and generates an aligned output 402 according to Equation 1. Additionally or alternatively, the alignment model 400 can receive an unspoken text utterance 320 and generate an aligned output 402 according to Equation 2. The text encoder 202 receives the aligned output 402 as input and, for each of a plurality of output steps, generates an encoded text representation 402 as output.

number

[0034] The audio encoder 210 may include a joint encoder 250 that receives as input the encoded text representation 312 and generates as output a first encoded shared representation 322. The joint encoder 250 may also receive as input the encoded audio representation 314 and generate as output a second encoded shared representation 324. An auxiliary decoder 390 receives as input the first and second encoded shared representations 322, 324 and generates as output corresponding first and second probability distributions 392, 394 for possible speech recognition hypotheses.

[0035] The alignment mask loss module 650 receives the first probability distribution 392 corresponding to the encoded text representation 312 and the second probability distribution 394 corresponding to the encoded audio representation 314 and generates an alignment loss 652 as follows:

number

[0036] Referring again to FIG. 3A, in some implementations, the audio encoder 210 includes a speech encoder 204 and a text encoder 202, which are described in more detail with reference to FIG. 3B and FIG. 3C. In the illustrated example, the audio encoder 210 (or the speech encoder 204 or the text encoder 202 (FIG. 3B and FIG. 3C)) includes a conformer encoder including a stack of multiple conformer blocks, each of which includes a stack of multi-head self-attention, depthwise convolution, and feedforward layers. Alternatively, the audio encoder 210 may include other types of encoders with stacks of multi-head self-attention layers / blocks, such as a Transformer or Performer encoder. The conformer encoder 210 may be naturally divided into a feature encoder including a convolutional subsampling block 212 and a context network including a stack of linear layers 214 and conformer blocks 216. In some implementations, the convolutional subsampling block 212 has two two-dimensional convolutional layers, both with a stride of (2,2), thereby reducing the feature sequence length by a factor of four. The convolutional subsampling block 212 receives as input a sequence of input features / vectors (e.g., Mel-frequency spectrograms, such as those of the acoustic frames 110 in FIG. 1 ) associated with each transcribed non-synthesized speech utterance 304 and each untranscribed non-synthesized speech utterance 306, and generates as output, for each of a plurality of output steps, encoded audio features 211 corresponding to a respective one of the transcribed non-synthesized speech utterances 304 or a respective one of the untranscribed non-synthesized speech utterances 306. The convolutional subsampling block 212 may receive as input each alignment output 402 generated by the alignment model 400 from the unspoken text utterance 320, and may generate as output, for each of a plurality of output steps, encoded text features 213 corresponding to a respective one of the alignment outputs 402.

[0037] The encoded audio and text features 211, 213 (i.e., interchangeably referred to as "encoded features 211, 213") output from the convolutional sub-sampling block 212 can be sent to a masking module 218, where a portion of the encoded features 211, 213 are replaced with a randomly selected trained feature vector shared among all masked time steps to provide corresponding masked encoded audio features 211, 211m and masked encoded text features 213, 213m. In some examples, the masking module 218 masks randomly selected encoded features 211, 213 by randomly sampling without replacement a certain percentage p of all time steps starting from a starting index, and then masks the following M consecutive time steps from all sample indices, whereby some spans may overlap. After masking is applied, the linear layer 214 and conformer block 216 of the context network receive the masked encoded features 211m, 213m (or the encoded features 211, 213 not selected by the masking module 218) and output the corresponding contrast context vector (i.e., coded representation) 215 from the masked encoded features 211m, 213m. Furthermore, a quantizer 217 receives the encoded features 211, 213 as input and generates a quantized vector (i.e., target context vector) 219 as output. Thereafter, a contrast loss module 315 calculates the contrast loss (L) between the contrast context vector 215 and the target context vector 219 at the masked positions as follows: w2v )316 is derived.

number

[0038] The contrastive loss 316 is optimized between the contrastive context vector 215 at the masked position and the target context vector 219. After the audio encoder 210 has converged on the untranscribed, unsynthesized speech utterance 306, the training procedure is repeated for both the alignment output 402 corresponding to the unspoken text utterance 320 and the transcribed, unsynthesized speech utterance 304. Thus, the contrastive loss 316 (L w2v ) is optimized for both real / human (non-synthetic) speech and unspoken text utterances 320 represented by alignment outputs 402, with additional auxiliary losses derived from transcribed non-synthetic speech utterances 304 and alignment outputs 402, as described in more detail below with reference to FIG. 3B. Thus, the contrastive self-supervised loss portion 300a of the training process 300 trains the audio encoder 210 using a contrastive loss 316 derived from corresponding encoded features 211, 213 associated with each alignment output 402, each transcribed non-synthetic speech utterance 304, and each untranscribed non-synthetic speech utterance 306 provided as input to the audio encoder 210. Training the audio encoder 210 may include updating parameters of the audio encoder 210 based on the contrastive loss 316.

[0039] Referring to FIG. 3B , the supervised loss section 300b of the training process 300 is configured to inject lexical information into the audio encoder 210 during training based on supervised loss terms 342, 344 derived from the transcribed, unsynthesized speech utterances and the alignment output 402 corresponding to the unspoken text utterance 320 output by the alignment model 400. In particular, the supervised loss section 300b utilizes one or more auxiliary decoders 390 for generating the supervised loss terms 342, 344. The auxiliary decoders 390 may include a connectionist temporal classification (CTC) decoder, a listen-and-attend-spell (LAS) decoder, or an RNN-T decoder. These auxiliary decoders 390 may include at least one of a phoneme decoder configured to decode a sequence of phonemes or a wordpiece decoder configured to decode a sequence of wordpieces. The auxiliary decoder 390 may also include a grapheme decoder configured to decode a sequence of graphemes.

[0040] In the supervised loss unit 300b, the text encoder 202 of the audio encoder is configured to receive the alignment output 402 (i.e., text embeddings) from the alignment model, and the speech encoder is configured to receive the transcribed non-synthesized speech utterance 304. That is, the text encoder 202 of the audio encoder generates an encoded text representation 312 for the alignment output 402 (e.g., corresponding to the unspoken text utterance 320), and the speech encoder 204 of the audio encoder 210 generates an encoded audio representation 314 for the speech input (i.e., the transcribed non-synthesized speech utterance 304). Note that both the encoded text representation 312 and the encoded audio representation 314 may not necessarily be compatible with the auxiliary decoder 390. Therefore, the audio encoder 210 also receives the encoded text representation 312 as input and generates a first encoded shared representation 322 (e.g., text) as output. Additionally, the shared encoder 250 receives the encoded audio representation 314 as input and generates a second encoded shared representation (e sup ) 324 as output. Thus, the joint encoder 250 generates first and second encoded shared representations 322, 324 in a shared latent representation space that is compatible with the auxiliary decoder 390.

[0041] In particular, the shared encoder 250 receives as input each encoded text representation 312 corresponding to the aligned output 402 generated from the unspoken text utterance 320, and as output, for each of a plurality of output steps, a first encoded shared representation (e text ) 322. An auxiliary decoder 390, including a phoneme decoder, a wordpiece decoder, or a byte decoder, receives as input each first encoded shared representation 322 output from the shared encoder 250 and generates as output a first probability distribution 392 over possible speech recognition hypotheses for a corresponding alignment output 402 at a corresponding time step. In some examples, the first probability distribution 392 over possible speech recognition hypotheses includes one of possible phoneme labels, possible wordpiece labels, or possible grapheme labels. The supervised loss module 340 can then determine an alignment output loss term 342 based on the first probability distribution 392 over possible speech recognition hypotheses for the alignment output 402 corresponding to the unspoken text utterance 320. Here, the corresponding unspoken text utterance 320 from which the alignment output 402 is generated also serves as a ground truth transcription. The supervised loss unit 300 b may train the audio encoder 210 with the alignment output loss term 342 by updating parameters of the audio encoder 210 based on the alignment output loss term 342 .

[0042] Similarly, during the supervised loss portion 300b, the shared encoder 250 receives as input each transcribed coded audio representation 314 corresponding to the non-synthesized speech utterance 304, and as output, for each of a plurality of output steps, a second coded shared representation (e sup ) 324. An auxiliary decoder 390, including a phoneme decoder, a wordpiece decoder, or a byte decoder, receives as input each second encoded shared representation 324 output from the shared encoder 250 and generates as output a second probability distribution 394 over possible non-synthesis speech recognition hypotheses for the corresponding transcribed non-synthesized speech utterance 304 in the corresponding output step. In some examples, the second probability distribution 394 over possible non-synthesis speech recognition hypotheses includes one of possible phoneme labels, possible wordpiece labels, or possible grapheme labels. The supervised loss module 340 can then determine a non-synthesis speech loss term 344 based on the second probability distribution 394 over possible non-synthesis speech recognition hypotheses and the corresponding transcription 302 paired with the transcribed non-synthesized speech utterance 304. Here, the corresponding transcription 302 serves as a ground truth transcription and may include sequences of target phonemes, target wordpieces, and / or target graphemes. The supervised loss unit 300 b may train the audio encoder 210 with the non-synthesized speech loss term 344 by updating parameters of the audio encoder 210 based on the non-synthesized speech loss term 344 .

[0043] In some implementations, the supervised loss portion 300b of the training process 300 uses another auxiliary decoder 390 to generate a first encoded shared representation (e text) 322 to generate a third probability distribution 393 over the possible speech recognition hypotheses, which enables the supervised loss module 340 to determine a further alignment output loss term 342 based on the third probability distribution 393 and the unspoken text utterance 320 corresponding to the alignment output 402. Here, the further auxiliary decoder 390 includes another one of a phoneme decoder, a wordpiece decoder, or a grapheme decoder, and the third probability distribution 393 over the possible speech recognition hypotheses includes another one of the possible phoneme labels, the possible wordpiece labels, or the possible grapheme labels. In these embodiments, the other auxiliary decoder 390 also generates a fourth probability distribution 395 over possible non-synthesis speech recognition hypotheses for the corresponding second-encoded shared representation 324 in the corresponding output step, allowing the supervised loss module 340 to determine a further non-synthesis speech loss term 344 based on the fourth probability distribution 395 and the corresponding transcription 302 paired with the transcribed non-synthesis speech representation 304. Here, the fourth probability distribution 395 over possible non-synthesis speech recognition hypotheses includes another one of possible phoneme labels, possible wordpiece labels, or possible grapheme labels. The supervised loss portion 300b of the training process 300 can similarly train the audio encoder 210 with the further alignment output loss term 342 and the further non-synthesis speech loss term 344.

[0044] The untranscribed, unsynthesized speech utterance 306 and the unspoken text utterance 320 each correspond to "unpaired" training data and therefore represent the unspoken text utterance (X text )320 derived from the control loss (L w2v ) 316 into the supervised loss J associated with the alignment output loss term 342 as follows: aux Together with this, the unspoken text loss function J text can be obtained. J text =L w2v (x│θ e )+L aux (y│x,θ e,θ d )(6) Similarly, untranscribed, non-synthesized speech utterances (X unsup )306 derived from the control loss (L w2v )316 is the unsupervised speech loss function J unsup_speech can be used to express as follows: J unsup_speech =J w2v (x * │θ e )(7)

[0045] During training of the audio encoder 210, the alignment outputs 402 and the untranscribed, non-synthesized speech 306 can be separated or mixed within each batch. To force the audio encoder 210 to learn a valid representation for both the alignment outputs 402 corresponding to unspoken text utterances 320 and non-synthesized (human / real) speech, a loss function J text and those in Equation 5 and Equation 6, we obtain the unpaired data loss function J unpaired A loss mask σ is applied when determining: J unpaired =σJ text +(1-σ)J speech (8)

[0046] The transcribed non-synthesized speech utterances 304 correspond to the “paired” and “supervised” training data and are therefore subjected to the derived contrastive loss L associated with the non-synthesized speech loss term 344. w2v and the derived supervised loss J aux , and combine them into a pair of data loss functions J paired can be obtained. J paired =L w2v (x│θ e )+L aux (y│x,θ e ,θ d )(9)

[0047] In some scenarios, after training the audio encoder 210, the ASR model recognizes audio from the target language using graphemes from the training language during inference. Thus, in some implementations, the supervised portion 300b of the training process 300 uses a residual adapter layer 330 that conditions at least one of the audio encoder 210 or the decoder (e.g., the prediction network 220 and the joint network 230 (FIG. 2)) based on a language identifier 321 that uniquely identifies the target language. Each residual adapter layer 330 includes a small feedforward network that includes a self-attention layer (e.g., two self-attention layers) with a bottleneck dimension. Thus, the residual identifier output 332 is fed to the shared encoder 250, so that the audio encoder 210, conditioned based on the language identifier 321, does not recognize audio from the target language during inference using graphemes from the training language.

[0048] Referring to FIG. 3C, the consistency regularizer (i.e., modality matching) 300c of the training process 300 regularizes each transcribed non-synthesized speech utterance (X sup ) 304 and the training utterance pairs 301 containing the corresponding transcribed non-synthesized speech utterances 304 and the paired alignment outputs 404 of the same utterances. cons(θ)) 352, thereby encouraging the audio encoder 210 to learn consistent predictions between the non-synthesized speech (e.g., real / human speech) and the alignment output 402 corresponding to the unspoken text utterance 320. Thus, the non-synthesized speech utterance 304 and paired alignment output 404 of each training utterance pair 301 are associated with the same ground truth transcription. In essence, the consistency loss term 352 between the transcribed non-synthesized speech utterance 304 and the paired alignment output 404 of the same training utterance provides an aspect of unsupervised training by encouraging the audio encoder 210 to behave consistently, regardless of the supervised loss term between the ground truth transcription 302 and each of the non-synthesized speech recognition hypotheses output by the auxiliary decoder 390 and the speech recognition hypotheses output by the auxiliary decoder 390, regardless of whether the training utterance belongs to the non-synthesized speech utterance (i.e., audio training data) or the alignment output (i.e., text training data).

[0049] 3B , the alignment model 400 may generate each paired alignment output 404 using a corresponding transcription 302 paired with a transcribed non-synthesized speech utterance 304. Here, the non-synthesized speech representation 304 is associated with the paired alignment output 404 generated by the alignment model 400 that maps the unspoken text utterance 320 to speech frames.

[0050] In the consistency regularizer 300c, the text encoder 202 receives as input each paired alignment output 404 and generates as output, for each of a plurality of output steps, an encoded text representation 313 corresponding to the paired alignment output 404 in the corresponding output step. The shared encoder 250 receives as input the encoded text representation 313 and generates as output a first encoded shared representation (e* sup ) 323. An auxiliary decoder 390, which may include a phoneme decoder or a wordpiece decoder, receives as input each first encoded shared representation 323 output from the shared encoder 250 and generates as output a first probability distribution 311 over possible speech recognition hypotheses for the corresponding paired alignment output 404 in the corresponding output step. In some examples, the first probability distribution 311 over possible speech recognition hypotheses includes one of the possible phoneme labels or possible wordpiece labels.

[0051] Similarly, the speech encoder 204 receives as input each transcribed non-synthesized speech utterance 304 as a sequence of features / vectors (e.g., Mel-frequency spectrograms such as acoustic frames 110 in FIG. 1 ) and generates as output, for each of a plurality of output steps, an encoded audio representation 314 corresponding to the transcribed non-synthesized speech utterance 304 at the corresponding output step. The shared encoder 250 receives as input the encoded audio representation 314 and generates as output a second encoded shared representation (e sup ) 324. An auxiliary decoder 390, which may include a phoneme decoder or a wordpiece decoder, receives as input each second encoded shared representation 324 output from the shared encoder 250 and generates as output a second probability distribution 394 over possible non-synthesized speech recognition hypotheses for the corresponding transcribed non-synthesized speech utterance 304 at the corresponding time step. In some examples, the second probability distribution 394 over possible non-synthesized speech recognition hypotheses includes one of the possible phoneme labels or possible wordpiece labels.

[0052] Continuing with FIG. 3C , the consistency regularizer 300 c of the training process 300 calculates, at each of a plurality of time steps for each training utterance pair 301, a consistency loss term (J ) for the corresponding training utterance pair 301 based on a first probability distribution 311 for possible speech recognition hypotheses and a second probability distribution 394 for possible non-synthetic speech recognition hypotheses. consFor example, the training process 300 may use a consistency loss term module 350 configured to receive, at each time step, the corresponding unsynthesized speech output by the auxiliary decoder 390 and the speech recognition results 311, 394, and to determine a consistency loss term 352 for the corresponding training utterance pair 301 at the time step.

[0053] In some examples, the consistency regularizer portion 300c of the training process 300 calculates the Kullback-Leibler divergence (D) between a first probability distribution 311 for possible speech recognition hypotheses and a second probability distribution 394 for possible non-synthetic speech recognition hypotheses. KL ) to determine the consistency loss term 352. KL The consistency loss term 352 based on can be expressed as:

number

[0054] Finally, the training process 300 calculates the unpaired data loss function (J unpaired), paired data loss function (J paired ), and the consistency loss term (J cons ) and the total loss term J tts4pretrain2 can be obtained, which can be expressed as follows: J tts4pretrain2 =J unpaired +λ1J paired +λ2J cons (11) where λ may be equal to 1.0 and λ is equal to 0.1. The training process 300 calculates the total loss term J tts4pretrain2 The audio encoder 210 may be pre-trained by using the above algorithm to update its parameters, effectively teaching it to learn shared representations between speech and text in the target language, even when labeled training data in the target language is unavailable. After training the audio encoder 210, the training process 300 may fine-tune the pre-trained audio encoder based on transcribed speech utterances, which may include supervised training samples of both alignment outputs and non-synthetic (e.g., human speech) corresponding to unspoken text utterances 320.

[0055] In some implementations, the training process 300 for training the audio encoder 210 applies encoder consistency regularization. Unlike decoder consistency regularization, which is applied to the auxiliary decoder(s) in the consistency regularizer 300c, which requires assumed labels (e.g., the transcription 302 and the unspoken text utterance 320), encoder consistency regularization has the advantage of being applicable to all of the training data 304, 306, 320, since it does not require assumed labels. Encoder consistency regularization can be applied by a hierarchical contrastive consistency regularization (HCCR) method, in which the encoder activations e, e* from the original / unaugmented and augmented speech are projected through an auxiliary network to generate z and z*. The positive and negative pairs are then constitutively compared, and a contrastive loss l is calculated. t,z,z* is calculated as follows:

number

[0056] Specific to HCCR, a convolutional neural network (CNN) projection network computes projections for increasing length segments (30, 50, 120 ms) of the encoder activations e, resulting in three views (V), from which negative examples can be extracted: from the same utterance for the shorter segments, and from other utterances in the batch for the 120 ms segments. Thus, the HCCR loss can be computed for the alignment output 402 generated from the transcribed non-synthesized speech utterance 304 (paired speech), the untranscribed non-synthesized speech utterance 306 (unpaired speech), and the unspoken text utterance 320, as follows:

number

[0057] While the above embodiments describe a training process 300 for training the audio encoder 210 for a target language, it is understood that the training process 300 may also be used to train an audio encoder for each of multiple target languages ​​that are different from the one or more training languages. Thus, the audio encoder 210 for the multilingual ASR model 200. In some examples, the training process 300 may be used to train an end-to-end ASR model with a decoder structure (i.e., without pre-training) or to fine-tune an ASR model to perform downstream tasks such as speech translation or natural language understanding. Additionally, the above embodiments describe a training process using each of the training sections 300a-c of the training process 300. Furthermore, it is understood that any combination of the training sections 300a-c may be used to independently train the audio encoder 210 using any combination of unspoken text utterances 320, transcribed non-synthesized speech utterances 304, and / or untranscribed non-synthesized speech utterances 306.

[0058] For example, transcribed unsynthesized speech utterances 304 in one or more training languages ​​can be initially used to train an alignment model 400 (FIGS. 5 and 6). Then, using the alignment model 400 trained using labeled training data in one or more training languages, the training process 300 can train an audio encoder using alignment output loss terms 342 derived from unspoken text utterances 320 in the target language(s) during the supervised loss portion 300b. Advantageously, the training process 300 can initially train the alignment model 400 utilizing high-resource languages ​​(e.g., languages ​​for which abundant labeled training data already exists) and then use the trained alignment model 400 to train the audio encoder 210 in low-resource target languages ​​(e.g., for which little or no labeled training data exists). Simply put, by training the alignment model using a large amount of labeled training data in one or more high-resource languages, the alignment model 400 learns to produce alignment output in the target language even though the alignment model 400 was not trained on any data in the target language (or no labeled training data exists).

[0059] In some scenarios, after training the audio encoder 210, the ASR model recognizes audio from the target language using graphemes from the training language during inference. Accordingly, in some implementations, the consistency regularizer 300c of the training process 300 uses a residual adapter layer 330 that conditions at least one of the audio encoder 210 or the decoder (e.g., the prediction network 220 and the joint network 230 (FIG. 2)) based on a language identifier 321 that uniquely identifies the target language. Each residual adapter layer 330 includes a small feedforward network that includes a self-attention layer (e.g., two self-attention layers) with a bottleneck dimension. Thus, the residual identifier output 332 is fed to the shared encoder 250, so that the audio encoder 210, conditioned on the language identifier 321, does not recognize audio from the target language during inference using graphemes from the training language.

[0060] 7 is a flowchart of an exemplary configuration of operations for a computer-implemented method 700 using aligned text and speech representations to train an automatic speech recognition model without transcribed speech data. Method 700 can be executed on data processing hardware 810 (FIG. 8) using instructions stored in memory hardware 820 (FIG. 8), which may reside on user device 102 and / or remote computer / server 201 of FIG. 1 corresponding to computing device 800 (FIG. 8).

[0061] At operation 702, the method 700 includes receiving training data including unspoken text utterances 320 in a target language. Each unspoken text utterance 320 is not paired with any corresponding spoken utterance in unsynthesized speech (or synthetic speech). At operation 704, the method 700 includes using an alignment model 400 to generate a corresponding aligned output 402 for each unspoken text utterance 320 in the received training data. The alignment model 400 is trained based on transcribed speech utterances 304 in each of one or more training languages ​​different from the target language. That is, the alignment model 400 is trained in the training language and generates aligned outputs 402 for unspoken text utterances 320 in the target language that were not seen by the alignment model 400 during training. At operation 706, the method 700 includes using a text encoder 202 to generate a corresponding encoded text representation 312 for each aligned output 402. At operation 708, the method 700 includes training the speech recognition model 200 with the encoded text representations 312 generated for the alignment outputs 402 corresponding to the unspoken text utterances 320 in the target language to teach the speech recognition model 200 to learn how to recognize speech in the target language. Notably, the only transcribed (i.e., paired) training data used to train the speech recognition model 200 to learn how to recognize speech in the target language are the transcribed speech utterances 304 in one or more training languages ​​used to train the alignment model 400, each of which is different from the target language.

[0062] 8 is a schematic diagram of an exemplary computing device 800 that can be used to implement the systems and methods described herein. Computing device 800 is intended to represent various types of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown here, their connections and relationships, and their functions are for illustrative purposes only and are not intended to limit the implementation of the invention as described and / or claimed in this document.

[0063] Computing device 800 includes a processor 810, memory 820, a storage device 830, a high-speed interface / controller 840 that connects to memory 820 and a high-speed expansion port 850, and a low-speed interface / controller 860 that connects to a low-speed bus 870 and storage device 830. Each of components 810, 820, 830, 840, 850, and 860 are interconnected using various buses and may be mounted on a common motherboard or reside otherwise as desired. Processor 810 can process instructions for execution within computing device 800, including instructions stored in memory 820 or storage device 830, and display graphical information for a graphical user interface (GUI) on an external input / output device, such as a display 880, coupled to high-speed interface 840. In other implementations, multiple processors and / or multiple buses may be used, along with multiple memories and multiple types of memories, as desired. Additionally, multiple computing devices 800 may be connected together (eg, as a bank of servers, a group of blade servers, or a multi-processor system) with each device providing multiple parts of multiple required operations.

[0064] The memory 820 stores information non-transiently within the computing device 800. The memory 820 may be a computer-readable medium, a volatile memory unit(s), or a non-volatile memory unit(s). The non-transient memory 820 may be a physical device used to temporarily or permanently store programs (e.g., sequences of instructions) or data (e.g., program state information) for use by the computing device 800. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.

[0065] The storage device 830 can provide mass storage for the computing device 800. In some embodiments, the storage device 830 is a computer-readable medium. In various different implementations, the storage device 830 can be a device array including a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid-state memory device, or a storage area network or other configuration of devices. In additional embodiments, a computer program product is tangibly embodied on an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable or machine-readable medium, such as memory 820, the storage device 830, or memory on the processor 810.

[0066] The high-speed controller 840 manages bandwidth-intensive operations of the computing device 800, while the low-speed controller 860 manages less bandwidth-intensive operations. This allocation of roles is merely exemplary. In some implementations, the high-speed controller 840 is coupled to memory 820, a display 880 (e.g., via a graphics processor or accelerator), and a high-speed expansion port 850 that can accept various expansion cards (not shown). In some implementations, the low-speed controller 860 is coupled to a storage device 830 and a low-speed expansion port 890. The low-speed expansion port 890 may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet, etc.) and may be coupled to one or more input / output devices such as a keyboard, pointing device, scanner, etc., or may be coupled to a network device such as a switch or router, for example, via a network adapter.

[0067] The computing device 800, as shown, may be implemented in many different forms. For example, it may be implemented as a standard server 800a, or as multiple iterations in a group of such servers 800a, as a laptop computer 800b, or as part of a rack server system 800c.

[0068] Various implementations of the systems and techniques described herein may be realized in digital electronic and / or optical circuitry, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may be specialized or general-purpose, and may include implementation in one or more computer programs executable and / or interpretable by a programmable system including at least one programmable processor, at least one input device, and at least one output device, coupled to receive data and instructions from, and transmit data and instructions to, the storage system.

[0069] These computer programs (also known as programs, software, software applications, or code) include machine instructions for a programmable processor and may be implemented in a high-level procedural and / or object-oriented programming language, and / or assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0070] The processes and logic flows described herein may be implemented by one or more programmable processors, also referred to as data processing hardware, executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be implemented by special-purpose logic circuitry, such as an FPGA (field-programmable gate array) or an ASIC (application-specific integrated circuit). Processors suitable for executing computer programs include, for example, both general-purpose and special-purpose processors, as well as any one or more processors of any kind of digital computer. Generally, a processor receives instructions and data from a read-only memory, a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer also includes one or more mass storage devices, such as magnetic, magneto-optical, or optical disks, for storing data, or is operably connected to receive data from or transmit data to them, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, magnetic disks such as internal or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0071] To interact with a user, one or more aspects of the present disclosure can be implemented in a computer that has a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, or a touch screen for displaying information to the user, and optionally a keyboard and pointing device, such as a mouse or trackball, through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user. For example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input from the user can be received in any form, such as acoustic input, speech input, or tactile input. Additionally, a computer can interact with a user by sending and receiving documents to a device used by the user, for example, by sending a web page to a web browser on the user's client device in response to a request received from the web browser.

[0072] Although several embodiments have been described, it will be understood that various modifications can be made without departing from the spirit and scope of the present disclosure. Accordingly, other embodiments are within the scope of the following claims.

Claims

1. A computer-implemented method (700) that, when executed on data processing hardware (810), causes the data processing hardware (810) to perform operations, the operations comprising: receiving training data including unspoken text utterances (320) in a target language, each unspoken text utterance (320) not paired with any corresponding spoken utterance of unsynthesized speech; generating a corresponding alignment output (402) for each unspoken text utterance (320) in the received training data using an alignment model (400), the alignment model (400) being trained based on transcribed speech utterances (304) in each of one or more training languages ​​different from the target language; generating a corresponding encoded text representation (312) for each alignment output (402) using a text encoder (202); training a speech recognition model (200) with the encoded text representation (312) generated for the alignment output (402) corresponding to the unspoken text utterance (320) in the target language to teach the speech recognition model (200) to learn how to recognize speech in the target language; A computer-implemented method (700).

2. 2. The computer-implemented method of claim 1, wherein training the speech recognition model includes training the speech recognition model without using any transcribed speech utterances in the target language for supervised learning.

3. The computer-implemented method (700) of claim 1 , wherein the speech recognition model (200) comprises an audio encoder (210) and a decoder (220, 230).

4. 4. The computer-implemented method of claim 3, wherein the decoder comprises a recurrent neural network transducer (RNN-T) architecture.

5. The audio encoder (210) the text encoder (202); an audio encoder (204); The computer-implemented method (700) of claim 3, comprising: a shared encoder (250).

6. 4. The computer-implemented method of claim 3, wherein the audio encoder comprises multiple multi-head self-attention layers.

7. The computer-implemented method of claim 3 , wherein the audio encoder comprises a conformer encoder.

8. 4. The computer-implemented method of claim 3, wherein the operations further comprise conditioning at least one of the audio encoder or the decoder based on a language identifier that uniquely identifies the target language.

9. 9. The computer-implemented method of claim 8, wherein conditioning the at least one of the audio encoder (210) or the decoder (220, 230) comprises conditioning the at least one of the audio encoder (210) or the decoder (220, 230) based on the language identifier (321) using a residual adapter layer (330).

10. training the speech recognition model (200), for each aligned output (402), generating, using a shared encoder (250), a first encoded shared representation (322) of said aligned output (402) in a shared latent representation space; For each transcribed speech utterance (304) in the one or more training languages: determining an encoded audio representation (314) of the transcribed speech utterance (304) using a speech encoder (204); generating a second encoded shared representation (324) of the transcribed speech utterance (304) in the shared latent representation space using the shared encoder (250); 2. The computer-implemented method of claim 1, wherein training the speech recognition model includes training the speech recognition model with the first coded shared representation generated for the alignment output corresponding to the unspoken text utterance in the target language and the second coded shared representation generated for the transcribed speech utterance in the one or more training languages.

11. 10. The computer-implemented method of claim 1, wherein each unspoken text utterance comprises a sequence of words, wordpieces, graphemes, and / or phonemes.

12. 12. The computer-implemented method (700) of any one of claims 1 to 11, wherein the operations further comprise: converting a script of the unspoken text utterance (320) in the target language into a phoneme representation shared across multiple languages ​​using a pronunciation model.

13. A system (100), comprising: Data processing hardware (810); memory hardware (820) in communication with the data processing hardware (810), the memory hardware (820) storing instructions that, when executed by the data processing hardware (810), cause the data processing hardware (810) to perform operations, the operations including: receiving training data including unspoken text utterances (320) in a target language, each unspoken text utterance (320) not paired with any corresponding spoken utterance of unsynthesized speech; generating a corresponding alignment output (402) for each unspoken text utterance (320) in the received training data using an alignment model (400), the alignment model (400) being trained based on transcribed speech utterances (304) in each of one or more training languages ​​different from the target language; generating a corresponding encoded text representation (312) for each alignment output (402) using a text encoder (202); training a speech recognition model (200) with the encoded text representation (312) generated for the alignment output (402) corresponding to the unspoken text utterance (320) in the target language to teach the speech recognition model (200) to learn how to recognize speech in the target language; System (100).

14. 14. The system of claim 13, wherein training the speech recognition model includes training the speech recognition model without using any transcribed speech utterances in the target language for supervised learning.

15. The system (100) of claim 13, wherein the speech recognition model (200) comprises an audio encoder (210) and a decoder (220, 230).

16. The system (100) of claim 15, wherein the decoder (220, 230) comprises a recurrent neural network transducer (RNN-T) architecture.

17. The audio encoder (210) the text encoder (202); an audio encoder (204); The system (100) of claim 15, comprising: a shared encoder (250).

18. The system (100) of claim 15, wherein the audio encoder (210) comprises multiple multi-head self-attention layers.

19. The system (100) of claim 15, wherein the audio encoder (210) comprises a conformer encoder.

20. 16. The system (100) of claim 15, wherein the operations further comprise conditioning at least one of the audio encoder (210) or the decoder (220, 230) based on a language identifier (321) that uniquely identifies the target language.

21. 21. The system (100) of claim 20, wherein conditioning the at least one of the audio encoder (210) or the decoder (220, 230) comprises conditioning the at least one of the audio encoder (210) or the decoder (220, 230) based on the language identifier (321) using a residual adapter layer (330).

22. training the speech recognition model (200), for each aligned output (402), generating, using a shared encoder (250), a first encoded shared representation (322) of said aligned output (402) in a shared latent representation space; For each transcribed speech utterance (304) in the one or more training languages: determining an encoded audio representation (314) of the transcribed speech utterance (304) using a speech encoder (204); generating a second encoded shared representation (324) of the transcribed speech utterance (304) in the shared latent representation space using the shared encoder (250); 14. The system of claim 13, wherein training the speech recognition model includes training the speech recognition model with the first coded shared representation generated for the alignment output corresponding to the unspoken text utterance in the target language and the second coded shared representation generated for the transcribed speech utterance in the one or more training languages.

23. The system (100) of claim 13, wherein each unspoken text utterance (320) comprises a sequence of words, wordpieces, graphemes, and / or phonemes.

24. 24. The system (100) of any one of claims 13 to 23, wherein the operations further comprise: converting a script of the unspoken text utterance (320) in the target language into a phoneme representation shared across multiple languages ​​using a pronunciation model.

Citation Information

Patent Citations

  • Method and system for modeling common-language speech recognition in computer with background of a plurality of dialects

    JP2010107982A

  • Method of training speech recognition model of extended language by speech in source language

    JP2022092568A

  • System and method for a multilingual speech recognition framework

    JP2023544336A

  • Using speech recognition to improve cross-language speech synthesis

    JP2024119883A

  • JPP7690137B