Using speech recognition to improve interlingual speech synthesis

A multilingual TTS model generates synthetic speech representations to enhance acoustic diversity, addressing the challenge of limited training data in low-resource languages and improving ASR model performance by leveraging consistency and adversarial loss terms.

JP7791934B2Active Publication Date: 2025-12-24GOOGLE LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024089100
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-10-21
Filing Date
2024-05-31
Publication Date
2025-12-24
Estimated Expiration
2041-10-20

AI Technical Summary

Technical Problem

Existing automatic speech recognition (ASR) systems face challenges in low-resource languages due to limited training data and the need for separate models for each language, which requires significant memory and lacks acoustic diversity.

Method used

A multilingual text-to-speech (TTS) model generates native and cross-language synthetic speech representations, conditioned on speaker characteristics, using consistency and adversarial loss terms to update the ASR model, and applies data augmentation to enhance acoustic diversity.

Benefits of technology

This approach increases acoustic and lexical diversity in training data, improving ASR model performance in low-resource languages without requiring separate models for each language, reducing memory usage and enhancing recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007791934000005
    Figure 0007791934000005
  • Figure 0007791934000006
    Figure 0007791934000006
  • Figure 0007791934000007
    Figure 0007791934000007
Patent Text Reader

Abstract

To improve cross-language speech synthesis.SOLUTION: A method (800) includes a step of obtaining a multilingual text-to-speech (TTS) model (310). This method also includes a step of generating a native synthesized speech representation (306) for an input text sequence (302) in a first language, conditioned on speaker characteristics (304) of a native speaker of the first language. The method also includes a step of generating a cross-linguistic synthesized speech representation for an input text sequence in the first language that is conditioned on the speaker characteristics of a native speaker of a different second language. The method also includes a step of generating first and second speech recognition results (312) for native and cross-linguistic synthetic speech representations. The method also includes a step of determining a consistent loss term (352) based on the first and second speech recognition results and updating parameters of a speech recognition model based on the consistent loss term.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure relates to using speech recognition to improve interlingual speech synthesis. [Background technology]

[0002] Automatic speech recognition (ASR) attempts to provide an accurate transcription of what a person said by receiving audio input and transcribing the audio input into text. Languages ​​that are little used today or have limited amounts of speech and text resources present a challenge for training an ASR system because only a limited amount of labeled training data exists. Training an ASR model using self-supervised learning can reduce the amount of labeled training data required to train the ASR model. Often, even when the ASR model has sufficient labeled training data, a unique ASR model is required for each language. Storing a separate ASR model for each language requires a significant amount of memory. Summary of the Invention [Means for solving the problem]

[0003] One aspect of the present disclosure provides a computer-implemented method that, when executed on data processing hardware, causes the data processing hardware to perform operations for training a speech recognition model. The operations include obtaining a multilingual text-to-speech (TTS) model. The operations also include using the multilingual TTS model to generate a native synthetic speech representation for an input text sequence in a first language, the native synthetic speech representation being conditioned on speaker characteristics of a native speaker of the first language. The operations also include using the speech recognition model to generate a first speech recognition result for the native synthetic speech representation and a second speech recognition result for the cross-language synthetic speech representation. The operations also include determining a consistency loss term based on the first speech recognition result and the second speech recognition result, and updating parameters of the speech recognition model based on the consistency loss term.

[0004] Implementations of the present disclosure may include one or more of the following features. In some implementations, the operations further include generating a first cross-entropy loss term based on the first speech recognition result and the input text sequence in the first language, determining a second cross-entropy loss based on the second speech recognition result and the input text sequence in the first language, and updating parameters of a speech recognition model based on the first and second cross-entropy loss terms. In some examples, the parameters of the speech recognition model are updated based on a consistency loss term independent of the first and second cross-entropy loss terms. The operations may further include back-propagating the first and second cross-entropy losses through a multilingual TTS model. Optionally, the operations may further include applying data augmentation to at least one of the native synthetic speech representation or the cross-language synthetic speech representation.

[0005] In some implementations, the multilingual TTS model includes an encoder portion that shares language embeddings across a first and a second language, and a decoder portion that shares language embeddings across the first and second languages ​​and shares speaker embeddings for both native speakers of the first language and native speakers of the second language. In these implementations, the number of speaker embeddings for native speakers of the first language may be less than the number of speaker embeddings for native speakers of the second language. The decoder portion may be further conditioned with prosodic information extracted from the synthetic speech representation using a variational autoencoder. Here, the prosodic information extracted from the synthetic speech representation using the variational autoencoder is separated from the speaker information by applying an adversarial loss for speaker classification.

[0006] In some examples, prior to generating the native and cross-language synthetic speech representations, the operations further include transliterating the input text sequence in the first language into a native script, tokenizing the native script into a phoneme sequence, encoding the phoneme sequence using an encoder of the multilingual TTS model, and decoding the encoded phoneme sequence using a decoder of the multilingual TTS model to generate each of the native synthetic speech representation or the cross-language synthetic speech representation. In some implementations, the operations further include generating a native audio encoder embedding for the native synthetic speech representation using a variational autoencoder, generating a cross-language audio encoder embedding for the cross-language synthetic speech representation using the variational autoencoder, determining an adversarial loss term conditioned on the first language based on the native and cross-language audio encoder embeddings, and updating parameters of the multilingual TTS model based on the adversarial loss term.

[0007] Another aspect of the present disclosure provides a system for training a speech recognition model, including data processing hardware and memory hardware storing instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations. The operations include obtaining a multilingual text-to-speech (TTS) model. The operations also include using the multilingual TTS model to generate a native synthetic speech representation for an input text sequence in a first language, the native synthetic speech representation being conditioned on speaker characteristics of a native speaker of the first language. The operations also include using the speech recognition model to generate a first speech recognition result for the native synthetic speech representation and a second speech recognition result for the cross-language synthetic speech representation. The operations also include determining a consistency loss term based on the first speech recognition result and the second speech recognition result, and updating parameters of the speech recognition model based on the consistency loss term.

[0008] Implementations of the present disclosure may include one or more of the following features. In some implementations, the operations further include generating a first cross-entropy loss term based on the first speech recognition result and the input text sequence in the first language, determining a second cross-entropy loss based on the second speech recognition result and the input text sequence in the first language, and updating parameters of a speech recognition model based on the first and second cross-entropy loss terms. In some examples, the parameters of the speech recognition model are updated based on a consistency loss term independent of the first and second cross-entropy loss terms. The operations may further include back-propagating the first and second cross-entropy losses through a multilingual TTS model. Optionally, the operations may further include applying data augmentation to at least one of the native synthetic speech representation or the cross-language synthetic speech representation.

[0009] In some implementations, the multilingual TTS model includes an encoder portion that shares language embeddings across a first and a second language, and a decoder portion that shares language embeddings across the first and second languages ​​and shares speaker embeddings for both native speakers of the first language and native speakers of the second language. In these implementations, the number of speaker embeddings for native speakers of the first language may be less than the number of speaker embeddings for native speakers of the second language. The decoder portion may be further conditioned with prosodic information extracted from the synthetic speech representation using a variational autoencoder. Here, the prosodic information extracted from the synthetic speech representation using the variational autoencoder is separated from the speaker information by applying an adversarial loss for speaker classification.

[0010] In some examples, prior to generating the native and cross-language synthetic speech representations, the operations further include transliterating the input text sequence in the first language into a native script, tokenizing the native script into a phoneme sequence, encoding the phoneme sequence using an encoder of the multilingual TTS model, and decoding the encoded phoneme sequence using a decoder of the multilingual TTS model to generate each of the native synthetic speech representation or the cross-language synthetic speech representation. In some implementations, the operations further include generating a native audio encoder embedding for the native synthetic speech representation using a variational autoencoder, generating a cross-language audio encoder embedding for the cross-language synthetic speech representation using the variational autoencoder, determining an adversarial loss term conditioned on the first language based on the native and cross-language audio encoder embeddings, and updating parameters of the multilingual TTS model based on the adversarial loss term.

[0011] The details of one or more implementations of the disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will be apparent from the description and drawings, and from the claims. [Brief explanation of the drawings]

[0012] [Figure 1] FIG. 1 is a schematic diagram of an exemplary speech recognition system including a speech recognition model. [Figure 2] FIG. 1 is a schematic diagram of a recurrent neural network transducer (RNN-T) model architecture. [Figure 3] FIG. 1 is a schematic diagram of an exemplary training process for training a speech recognition model and / or a multilingual text-to-speech model. [Figure 4] FIG. 1 is a schematic diagram of an exemplary training process for training a multilingual text-to-speech model. [Figure 5] FIG. 1 is a schematic diagram of a multilingual text-to-speech model for training multiple speech recognition models. [Figure 6] FIG. 1 is a schematic diagram of an exemplary speech recognition system. [Figure 7] 1 is a flowchart of an exemplary sequence of operations for a method of training an automated speech recognition model. [Figure 8] FIG. 1 is a schematic diagram of an exemplary computing device that can be used to implement the systems and methods described herein. DETAILED DESCRIPTION OF THE INVENTION

[0013] Like numbers in the various drawings indicate like elements.

[0014] Training an automatic speech recognition (ASR) model requires a vast amount of transcription data. That is, an ASR model requires training data pairs including audio data and a corresponding transcription of the audio data. Together, the audio data and the corresponding transcription (e.g., training data pairs) train the ASR model. An ASR model receives audio data, predicts a transcription of the audio data, and compares the predicted transcription with the corresponding transcription (i.e., ground truth labels). However, the amount of training data pairs required to train an ASR model is difficult to collect. In some cases, user permission is required to access the training data pairs. In other cases, low-resource languages ​​contain only a limited number of available speakers, creating little acoustic diversity for training an ASR model that provides only minor improvements in ASR model performance. For example, for low-resource languages ​​such as Indic languages ​​(e.g., Kannada, Telugu, Tamil, and Bengali), only a limited number of speakers with formal, constrained speaking styles are available for training Indic language ASR models. In contrast, for a high-resource language such as English, there may be thousands of speakers with a variety of speaking styles available to train an English ASR model.

[0015] Implementations herein are directed to a system and method for training an ASR model. Specifically, for an input text sequence input to a multilingual TTS model, the TTS model generates a native synthetic speech representation in a first language. The native synthetic speech representation is conditioned on speaker characteristics of a native speaker of the first language. The TTS model also generates a cross-language synthetic speech representation in the first language for the same input text sequence. Here, the cross-language synthetic speech representation is conditioned on speaker characteristics of a native speaker of a second language. That is, the cross-language synthetic speech representation is conditioned on a non-native speaker of the first language. An ASR model receives the native and cross-language synthetic speech representations and generates a first speech recognition result for the native synthetic speech representation and a second speech recognition result for the cross-language synthetic speech representation. A consistency loss term module determines a consistency loss term based on a comparison of the first and second speech recognition results, and the ASR model updates parameters based on the consistency loss term.

[0016] Implementations herein are further directed to systems and methods for separating speaker embeddings, language embeddings, and prosodic embeddings from synthetic speech representations produced by a multilingual TTS model for training the multilingual TTS model. A variational autoencoder (VAE) may receive native and cross-language synthetic speech representations and generate native audio encoder embeddings and cross-language audio encoder embeddings for the native and cross-language synthetic speech representations, respectively. A classifier may then determine an adversarial loss term conditioned on a first language based on the native and cross-language audio encoder embeddings. The adversarial loss term may be used to update parameters to prevent the multilingual TTS model from generating accented synthetic speech representations. That is, the adversarial loss term prevents the synthetic speech representations from closely resembling the speaker embeddings, language embeddings, and / or prosodic embeddings of the speaker characteristics on which the synthetic speech representations are conditioned.

[0017] 1 illustrates an automated speech recognition (ASR) system 100 that implements an ASR model 200 that resides on a user device 102 of a user 104 and / or on a remote computing device 201 (e.g., one or more servers of a distributed system executing in a cloud computing environment) that communicates with the user device 102. Although the user device 102 is shown as a mobile computing device (e.g., a smartphone), the user device 102 may correspond to any type of computing device, such as, but not limited to, a tablet device, a laptop / desktop computer, a wearable device, a digital assistant device, a smart speaker / display, a smart appliance, an in-vehicle infotainment system, or an Internet of Things (IoT) device.

[0018] The user device 102 includes an audio subsystem 108 configured to receive utterances 106 spoken by the user 104 (e.g., the user device 102 may include one or more microphones for recording the spoken utterances 106) and convert the utterances 106 into a corresponding digital format associated with input acoustic frames 110 that can be processed by the ASR system 100. In the illustrated example, the user 104 speaks each utterance 106 in natural English for the phrase "What is the weather in New York City?", and the audio subsystem 108 converts the utterance 106 into a corresponding acoustic frame 110 for input to the ASR system 100. The ASR model 200 then receives the acoustic frame 110 corresponding to the utterance 106 as input and generates / predicts a corresponding transcription (e.g., recognition result / hypothesis) 120 of the utterance 106 as output. In the illustrated example, the user device 102 and / or the remote computing device 201 also execute a user interface generator 107 configured to present a representation of a transcription 120 of the utterance 106 to the user 104 of the user device 102. In some configurations, the transcription 120 output from the ASR system 100 is processed by a natural language understanding (NLU) module executing on the user device 102 or the remote computing device 201, for example, to execute user commands. Additionally or alternatively, a text-to-speech system (e.g., executing on any combination of the user device 102 or the remote computing device 201) may convert the transcription 120 into synthesized speech for audible output by another device. For example, the original utterance 106 may correspond to a message the user 104 is sending to a friend, where the transcription 120 is converted into synthesized speech for audible output to the friend, who should hear the message conveyed in the original utterance 106.

[0019] Referring to FIG. 2 , the ASR model 200 may include an end-to-end (E2E) sequence-to-sequence model. The E2E sequence-to-sequence model may include a recurrent neural network transducer (RNN-T) model architecture that adheres to latency constraints associated with interactive applications. The RNN-T model 200 offers a small computational footprint and uses fewer memory requirements than traditional ASR architectures, making the RNN-T model architecture suitable for performing speech recognition entirely on the user device 102 (e.g., no communication with a remote server is required). The RNN-T model 200 includes an encoder network 210, a prediction network 220, and a collaboration network 230. The encoder network 210 is generally similar to an acoustic model (AM) in traditional ASR systems and includes a recurrent network consisting of stacked long short-term memory (LSTM) layers. For example, the encoder may receive a sequence of d-dimensional feature vectors (e.g., acoustic frame 110 ( FIG. 1 )) x=(x1, x2,..., x T ), where

number

number

[0020] Similarly, the prediction network 220 is also an LSTM network, which, like a language model (LM), calculates the sequence of non-empty symbols output so far by the final softmax layer 240, i.e., y 0 , ..., y ui-1 to obtain a dense representation

number

number

[0021] The softmax layer 240 may utilize any technique for selecting the output label / symbol with the highest probability in the distribution as the next output symbol predicted by the RNN-T model 200 at the corresponding output step. In this manner, the RNN-T model 200 does not make a conditional independence assumption; rather, the prediction of each symbol is conditioned not only on the acoustics but also on the sequence of labels output so far. The RNN-T model 200 assumes that the output symbols are independent of future acoustic frames 110, which allows the RNN-T model 200 to be used in a streaming manner.

[0022] In some examples, the encoder network 210 of the RNN-T model 200 consists of eight 2,048-dimensional LSTM layers, each followed by a 640-dimensional projection layer. The prediction network 220 may have two 2,048-dimensional LSTM layers, each also followed by a 640-dimensional projection layer. Finally, the collaboration network 230 may also have 640 hidden units. The softmax layer 240 may consist of a unified word piece or grapheme set generated using all unique word pieces or graphemes in multiple training text utterances.

[0023] 3 illustrates an exemplary training process 300 for training the ASR model 200 and / or the multilingual TTS model 310 (also referred to as the TTS model 310). The TTS model 310 is configured to generate a synthetic speech representation 306 for each of a plurality of unspoken training text utterances 302 (also referred to as the input text sequences 302) at each of a plurality of time steps. The input text sequences 302 include unspoken text that is text-only data, i.e., unpaired data, such that each input text sequence 302 is not paired with any synthetic or non-synthetic speech representation, i.e., the input text sequences 302 are not paired with a corresponding utterance of a human voice. Thus, the TTS model 310 generates a corresponding synthetic speech representation 306 for each input text sequence 302. That is, the TTS model 310 creates paired data through self-supervised learning by predicting the synthetic speech representation 306 for the unpaired input text sequences 302. In particular, the synthetic speech representation 306 may include mel-frequency spectrogram frames for training the ASR model 200, thereby eliminating the need for the TTS model 310 to include a vocoder and / or synthesizer for synthesizing the mel-frequency spectrogram frames into synthetic speech.

[0024] In some examples, the TTS model 310 generates multiple synthetic speech representations 306 for each of the input text sequences 302. Each of the synthetic speech representations 306 may be conditioned on speaker characteristics 304 of native speakers of a different language. In some examples, the speaker characteristics 304 for each native speaker of a respective language include a speaker embedding 305 (FIG. 6), a language embedding 303 (FIG. 6), and / or a local embedding 307 (FIG. 6) representing accent / dialect information, which are discussed in more detail below with reference to FIG. 6.

[0025] In the illustrated example, the TTS model 310 receives an input text sequence 302 in a first language. For example, the input text sequence 302 may represent input text in a low-resource language (i.e., the first language) called Kannada. The TTS model 310 then generates native synthetic speech representations 306, 306a for the input text sequence 302 in the first language, conditioned on speaker characteristics 304, 304a of native speakers of the first language. The speaker characteristics 304a of native speakers of the first language may be referred to as first conditioned input 304a. Continuing with this example, the TTS model 310 generates a native synthetic speech representation 306a in Kannada, conditioned on speaker characteristics 304 of native speakers of Kannada. Thus, the native synthetic speech representation 306a is conditioned on speaker characteristics 304a of native speakers of the corresponding language of the input text sequence 302.

[0026] In some implementations, the training process 300 generates a coarse-to-fine (C2F) loss 315 by comparing the native synthetic speech representation 306a with ground truth audio of a native speaker of the first language for the corresponding input text sequence 302. The C2F loss 315 thus represents the difference between the unsynthesized speech from the native speaker (e.g., the ground truth audio) and the native synthetic speech representation 306a. The training process 300 provides the C2F loss 315 to a TTS model 310 and updates parameters of the TTS model 310 based on the C2F loss 315.

[0027] However, in some cases, for some low-resource languages, the acoustic diversity of the native synthetic speech representation 306a is limited because there are only a limited number of native speakers of the first language. That is, an Indic language such as Kannada may only have speaker characteristics (e.g., conditioning input) 304 for one or two native speakers of Kannada on which the TTS model 310 can condition the synthetic speech representation 306. The limited acoustic diversity of the synthetic speech representation 306 provides only incremental improvements when training the ASR model. In contrast, a high-resource language such as English has thousands of native speakers with a variety of speaking styles, thereby providing significant improvements to the ASR model during training.

[0028] In some implementations, the TTS model 310 generates a cross-language synthetic speech representation 306, 306b for the same input text sequence 302 in a first language, conditioned with speaker characteristics 304, 304b of a native speaker of a second language different from the first language. That is, the TTS model 310 generates a cross-language synthetic speech representation 306b for the input text sequence 302 in the first language that conveys speaker characteristics 304b of a native speaker of the second language. The speaker characteristics 304b of the native speaker of the second language may be referred to as a second conditioned input 304a. For example, the TTS model 310 generates a cross-language synthetic speech representation 306b for the same input text sequence 302 in Kannada, conditioned with speaker characteristics 304b of a native speaker of English. In other words, the cross-language synthetic speech representation 306b represents Kannada speech spoken as a native speaker of English. Thus, the cross-language synthesized speech representation 306 b is conditioned on speaker characteristics 304 b of native speakers of a language different from the language of the input text sequence 302 .

[0029] By conditioning the synthetic speech representations 306 with speaker characteristics of native speakers of different languages, the TTS model 310 can generate multiple synthetic speech representations 306 for each input text sequence 302 in a first language to increase acoustic diversity among the synthetic speech representations 306. That is, the TTS model 310 can use one or more speakers from a second language (e.g., English) to synthesize an input text sequence 302 in a first language (e.g., Kannada) even if the speaker of the second language does not speak the first language. Thus, the TTS model 310 can leverage a high-resource language, such as English, to capture speaker characteristics for native speakers of English to increase acoustic diversity among synthetic speech representations generated for low-resource languages ​​to train the ASR model. Moreover, the TTS model 310 can generate synthetic speech representations 306 for unspoken input text sequences 302, thereby increasing the lexical diversity of the training data for the ASR model 200.

[0030] In some implementations, the training process 300 generates a spectrogram consistency loss (not shown). That is, the training process may extract latent variables from ground truth audio of a native speaker of a first language to perform teacher constraints on a cross-language synthetic speech representation 306b conditioned on speaker characteristics 304b of native speakers of a second language. Then, for a corresponding input text sequence 302, the training process 300 calculates the mean squared error (MSE) between the cross-language synthetic speech representation 306b (e.g., based on teacher constraints) and the ground truth audio of the native speaker of the first language to determine the spectrogram consistency loss. Thus, the spectrogram consistency loss promotes consistency between the cross-language synthetic speech representation 306b and the ground truth audio of the native speaker of the first language. The spectrogram consistency loss may be provided as feedback to the TTS model 310 to update its parameters.

[0031] In the illustrated example, the TTS model 310 generates a synthetic speech representation 306 that is conditioned on the speaker characteristics 304a, 304b of only two speakers 304a, 304b for clarity's sake. That is, the TTS model 310 may generate any number of synthetic speech representations 306 that are conditioned on the speaker characteristics 304 of any number of speakers. For example, the TTS model 310 may generate a third synthetic speech representation 306 in Kannada that is conditioned on the speaker characteristics of a native speaker of a third language (e.g., a native speaker of Spanish spoken in Spain). Optionally, the TTS model 310 may generate a fourth synthetic speech representation 306 in Kannada that is conditioned on the speaker characteristics of a second native speaker of the first language.

[0032] In some implementations, the training process 300 includes a data augmentation module 360 ​​that applies data augmentation to at least one of the native synthetic speech representation 306a and / or the cross-lingual synthetic speech representation 306b. Data augmentation of the synthetic speech representation 306 is configured to promote acoustic diversity in the training samples used to train the ASR model 200. In some examples, the data augmentation module 360 ​​applies a data augmentation technique that includes at least one of adding / injecting noise, adding reverberation, or manipulating the timing of the synthetic speech representation 306. Another data augmentation technique includes using multi-style training (MTR) to inject diverse environmental noise into the synthetic speech representation. Yet another data augmentation technique that the data augmentation module 360 ​​can apply in addition to or instead of MTR includes using spectral augmentation (SpecAugment) to make the synthetic speech representation sound more similar. Combined, MTR and SpecAugment inject noise into the synthetic speech representation 306, inserting a random external noise source in front of and tiling it over time, and filtering the noise-injected synthetic speech representation 306 prior to training the ASR model 200.

[0033] The exemplary training process 300 trains the ASR model 200 using multiple synthetic speech representations 306 generated by a multilingual TTS model 310. In the illustrated example, the ASR model 200 is trained to recognize speech spoken in a first language (e.g., Kannada). For each synthetic speech representation 306, the ASR model 200 generates a corresponding speech recognition result 312. The speech recognition result 312 may represent a probability distribution over possible speech recognition hypotheses. The ASR model 200 generates a first speech recognition result 312, 312a for the native synthetic speech representation 306a and a second speech recognition result 312, 312b for the cross-language synthetic speech representation 306b. Continuing with the above example, the ASR model 200 receives a mel-frequency spectrogram frame for a native synthesized speech representation 306a conditioned on a first conditioning input 306a (e.g., speaker characteristics of a native speaker of Kannada 306a) and generates a first speech recognition result 312a. The ASR model also receives a mel-frequency spectrogram frame for a cross-language synthesized speech representation 306b conditioned on a second conditioning input 304b (e.g., speaker characteristics of a native speaker of English 304b) and generates a second speech recognition result 312b.

[0034] In some examples, the training process 300 determines a consistency loss term 352 based on the first and second speech recognition results 312 a, 312 b. For example, the training process 300 may utilize a consistency loss term module 350 configured to receive corresponding speech recognition results 312 a, 312 b output by the ASR model 200 at each of a plurality of time steps and determine a consistency loss term 352 between the corresponding speech recognition results 312 a, 312 b at each of the plurality of time steps. The consistency loss term module 350 may calculate a Kullback-Leibler divergence (D ) between a first probability distribution for a first possible synthetic speech recognition result hypothesis and a second probability distribution for a second possible synthetic speech recognition result hypothesis. KL ), a consistency loss term 352 may be determined.

[0035] The consistency loss term 352 provides an “unsupervised” loss term that is unrelated to the accuracy of the ASR model 200 and may be used to update parameters of the ASR model 200 to promote consistency between a first speech recognition result 312a recognized from a native synthetic speech representation 306a and a second speech recognition result 312b recognized from a cross-language synthetic speech representation 306b. In particular, the synthetic speech representations 306a, 306b are generated from the same input text sequence 302 that serves as ground truth for the ASR model 200. In other words, the consistency loss term 352 allows the ASR model 200 to learn to behave in the same way, e.g., make consistent predictions for both the native synthetic speech representation 306a conditioned on the speaker characteristics 304a of a native speaker of a first language and the cross-language synthetic speech representation 306b conditioned on the speaker characteristics 304b of a native speaker of a second language, for the same input text sequence 302. During training, the consistency loss term module 350 may return the consistency loss term 352 to the ASR model 200 to update the parameters of the ASR model 200 based on the consistency loss term 352.

[0036] In some implementations, the training process 300 executes a supervised loss term module 340 configured to receive as input the speech recognition result 312 and generate as output a supervised loss term 342 based on the input text sequence 302 serving as ground truth. In the illustrated example, the training supervised loss term module 340 receives the input text sequence 302 (i.e., the ground truth transcription) and the first speech recognition result 312a and outputs a first supervised loss term 342, 342a (also referred to as a first cross-entropy loss term 342a). Thus, the first supervised loss term 342a is based on a comparison between the first speech recognition result 312a and the corresponding input text sequence 302 (e.g., the target speech recognition result). The first supervised loss term 342a represents the accuracy of the first speech recognition result 312a based on the native synthesized speech representation 306a.

[0037] Additionally, the supervised loss term module 340 receives the input text sequence 302 and the second speech recognition result 312b and outputs a second supervised loss term 342, 342b (also referred to as a second cross-entropy loss term 342b). The second supervised loss term 342b is based on a comparison between the second speech recognition result 312b and the corresponding input text sequence 302 (e.g., the target speech recognition result). Thus, the second supervised loss term represents the accuracy of the second speech recognition result 312b based on the cross-language synthetic speech representation 306b.

[0038] The supervised loss term module 340 may return the first supervised loss term 342 a and the second supervised loss term 342 b to the ASR model 200, which updates its parameters based on the first supervised loss term 342 a and the second supervised loss term 342 b. In some examples, the training process 300 updates the parameters of the ASR model 200 based on the consistency loss term 352, independently of the first supervised loss term 342 a and the second supervised loss term 342 b. Optionally, the training process 300 may back-propagate the first supervised loss term 342 a and the second supervised loss term 342 b to the TTS model 310 to update the parameters of the TTS model 310. Here, the ASR model 200 is fixed such that the parameters of the ASR model 200 are static (e.g., not updated) while updating the parameters of the TTS model 310 based on the first and second supervised loss terms 342a, 342b.

[0039] 4 shows an exemplary training process 400 for training a multilingual TTS model 310. In some implementations, the TTS model 310 generates a synthetic speech representation 306 that closely corresponds to the speaker characteristics of a particular speaker. For example, as a result of the limited number of native speakers of the first language 304a, the TTS model 310 may generate native synthetic speech representations 306a that closely resemble the speaker characteristics of a limited number of speakers. Therefore, the ASR model 200 (FIG. 3) trains only with synthetic speech representations 306 that resemble the speaker characteristics of a limited number of speakers. Therefore, the training process 400 includes a hierarchical variational autoencoder (VAE) 410 configured to separate speaker, language, and / or prosodic information from the synthetic speech representation 306.

[0040] In some examples, the TTS model 310 generates multiple synthetic speech representations 306 for each of the input text sequences 302. Each of the synthetic speech representations 306 may be conditioned with conditioning inputs 304 representing different speakers having different speaker characteristics 304. In some examples, the speakers include speaker characteristics 304 that represent a speech style for a particular speaker. That is, the speaker characteristics 304 for a particular speaker may include a speaker embedding 305 (FIG. 6), a language embedding 303 (FIG. 6), and / or a local embedding 307 (FIG. 6) that represents accent / dialect information, which are discussed in more detail with reference to FIG. 6.

[0041] In the illustrated example, a TTS model 310 receives an input text sequence 302 in a first language. For example, the input text sequence 302 may represent input text in a low-resource language (i.e., the first language) called Kannada. The TTS model 310 then generates native synthetic speech representations 306, 306a for the input text sequence 302 in the first language, conditioned on speaker characteristics 304a of native speakers 304, 304a of the first language. Continuing with this example, the TTS model 310 generates a native synthetic speech representation 306a in Kannada, conditioned on speaker characteristics 304a of the native speaker 304a of Kannada. Thus, the native synthetic speech representation 306a is conditioned on speaker characteristics 304a of a native speaker of the corresponding language of the input text sequence 302.

[0042] 3, the training process 400 of FIG. 4 generates a coarse-to-fine (C2F) loss 315 by comparing the native synthesized speech representation 306a with ground truth audio of a native speaker of a first language for the corresponding input text sequence 302. The C2F loss 315 thus represents the difference between the unsynthesized speech from the native speaker (e.g., the ground truth audio) and the native synthesized speech representation 306a. The training process 400 provides the C2F loss 315 to the TTS model 310 and updates parameters of the TTS model 310 based on the C2F loss 315.

[0043] In some implementations, the TTS model 310 also generates cross-language synthetic speech representations 306, 306b for the same input text sequence 302 in a first language that are conditioned on speaker characteristics 304, 304b of a native speaker of a second language different from the first language. That is, the TTS model 310 generates a cross-language synthetic speech representation 306b for the input text sequence 302 when a native speaker of the second language speaks in the first language. For example, the TTS model 310 generates a cross-language synthetic speech representation 306b for the same input text sequence 302 in Kannada that is conditioned on speaker characteristics 304b of a native speaker of English. In other words, the cross-language synthetic speech representation 306b represents Kannada speech spoken by a native speaker of English. Therefore, the cross-language synthetic speech representation 306b is conditioned on speaker characteristics 304b of a native speaker of a language different from the language of the input text sequence 302.

[0044] By conditioning the synthetic speech representations 306 with multiple speakers 304, the TTS model 310 can generate multiple synthetic speech representations 306 for each input text sequence 302 in a first language to increase the acoustic diversity of the synthetic speech representations 306. That is, the TTS model 310 can use one or more speakers from a second language (e.g., English) to synthesize an input text sequence 302 from a first language (e.g., Kannada) even if the speaker of the second language does not speak the first language. Thus, the TTS model 310 can leverage a high-resource language, such as English, to generate synthetic speech representations for low-resource languages ​​for training an ASR model. Moreover, the TTS model 310 can generate synthetic speech representations 306 for unspoken input text sequences 302, thereby increasing the lexical diversity of the training data for the ASR model 200.

[0045] In some implementations, the training process 400 generates a spectrogram consistency loss (not shown). That is, the training process may extract latent variables from the ground truth audio of a first language native speaker to perform teacher constraints on the second language native speaker 304b. Then, for the corresponding input text sequence 302, the training process 400 calculates the mean squared error (MSE) between the cross-language synthetic speech representation 306b (e.g., based on teacher constraints) and the ground truth audio of the first language native speaker to determine the spectrogram consistency loss. Thus, the spectrogram consistency loss promotes consistency between the cross-language synthetic speech representation 306b and the ground truth audio of the first language native speaker. The spectrogram consistency loss may be provided as feedback to the TTS model 310 to update its parameters.

[0046] In the illustrated example, the TTS model 310 generates a synthetic speech representation 306 that is conditioned on only two speakers 304 a, 304 b for clarity's sake. That is, the TTS model 310 may generate any number of synthetic speech representations 306 that are conditioned on the speaker characteristics 304 of any number of speakers 304. For example, the TTS model 310 may generate a third synthetic speech representation 306 in Kannada that is conditioned on the speaker characteristics of a native speaker of a third language (e.g., a native speaker of Spanish spoken in Spain). Optionally, the TTS model 310 may generate a fourth synthetic speech representation 306 in Kannada that is conditioned on the speaker characteristics of a second native speaker of the first language.

[0047] In some implementations, the training process 400 includes a data augmentation module 360 ​​that applies data augmentation to at least one of the native synthetic speech representation 306a and / or the cross-lingual synthetic speech representation 306b. Data augmentation of the synthetic speech representation 306 is configured to promote acoustic diversity in the training samples used to train the ASR model 200. In some examples, the data augmentation module 360 ​​applies a data augmentation technique that includes at least one of adding / injecting noise, adding reverberation, or manipulating the timing of the synthetic speech representation 306. Another data augmentation technique includes using multi-style training (MTR) to inject diverse environmental noise into the synthetic speech representation. Yet another data augmentation technique that the data augmentation module 360 ​​can apply in addition to or instead of MTR includes using spectral augmentation (SpecAugment) to make the synthetic speech representation sound more similar. Combined, MTR and SpecAugment inject noise into the synthetic speech representation 306, inserting a random external noise source in front of and tiling it over time, and filtering the noise-injected synthetic speech representation 306 prior to training the ASR model 200.

[0048] The exemplary training process 400 trains the ASR model 200 using multiple synthetic speech representations 306 generated by the multilingual TTS model 310. In the illustrated example, the ASR model 200 is trained to recognize speech in a first language (e.g., Kannada). For each synthetic speech representation 306, the ASR model 200 generates a corresponding speech recognition result 312. The speech recognition result 312 may represent a probability distribution over possible speech recognition hypotheses. The ASR model 200 generates a first speech recognition result 312, 312a for the native synthetic speech representation 306a and a second speech recognition result 312, 312b for the cross-language synthetic speech representation 306b. Continuing with the above example, the ASR model 200 receives a mel-frequency spectrogram frame for a native synthesized speech representation 306a conditioned on a first conditioning input 304a (e.g., speaker characteristics of a native speaker of Kannada) and generates a first speech recognition result 312a. The ASR model also receives a mel-frequency spectrogram frame for a cross-language synthesized speech representation 306b conditioned on a second conditioning input 304b (e.g., speaker characteristics of a native speaker of English) and generates a second speech recognition result 312b.

[0049] The training process 400 also includes a hierarchical VAE (interchangeably referred to as a VAE) 410 configured to generate encoder embeddings 412 for the synthetic speech representations 306. The VAE 410 includes a local encoder configured to encode fixed, 1-second overlapping, 2-second chunks of each synthetic speech representation 306 and a global encoder configured to encode the entire synthetic speech representation 306. In the illustrated example, the VAE 410 receives the native synthetic speech representation 306a and generates native audio encoder embeddings 412, 412a. The native audio encoder embedding 412a represents latent variables extracted from the native synthetic speech representation 306a. For example, the native audio encoder embedding 412 may represent prosody / accent information extracted from the native synthetic speech representation 306a. Additionally, the VAE 410 generates cross-language audio encoder embeddings 412, 412b for the cross-language synthetic speech representation 306b. The cross-language audio encoder embedding 412b represents latent variables (eg, prosody / accent information) extracted from the cross-language synthesized speech representation 306b.

[0050] The training process 400 also executes a classifier 420 that receives the native audio encoder embedding 412 a and the cross-language audio encoder embedding 412 b. The classifier 420 may be a language classifier. The classifier 420 determines an adversarial loss term 422 based on the native audio encoder embedding 412 a and the cross-language audio encoder embedding 412 b. The TTS model 310 receives the adversarial loss term 422 from the classifier 420 and updates the parameters of the TTS model 310 based on the adversarial loss term 422. That is, the adversarial loss term 422 prevents the TTS model 310 from generating a synthesized speech representation 306 that resembles only the prosody / accent information of the speaker 304. In other words, the encoder embedding 412 extracted from the synthesized speech representation 306 using the VAE 410 is separated from the speaker information by applying the adversarial loss term 422 to the speaker classification.

[0051] 5, in some implementations, a TTS model 310 generates synthetic speech representations 306 in different languages ​​to separately train multiple monolingual ASR models 200. In the illustrated example, a first ASR model 200, 200a is trained with a synthetic speech representation 306, 306A generated by the TTS model 310 in a first language to recognize speech in the first language, a second ASR model 200, 200b is trained with a synthetic speech representation 306, 306B generated by the TTS model 310 in a second language to recognize speech in the second language, and a third ASR model 200, 200c is trained with a synthetic speech representation 306, 306C generated by the TTS model 310 in a third language to recognize speech in the third language. In another example, the TTS model 310 generates synthetic speech representations 306 in multiple languages ​​to train a single multilingual ASR model 200 to recognize speech in multiple different languages. Thus, the multilingual TTS model 310 generates synthetic speech representations from input text sequences in multiple different languages ​​for use as training audio data to train one or more monolingual and / or multilingual ASR models 200. The synthetic speech representations in each language may include both a native synthetic speech representation 306a in the respective language conditioned on speaker characteristics 304a of native speakers of the respective language and / or a cross-language speech representation 306b in the respective language conditioned on speaker characteristics 304b of native speakers of a different language ( FIG. 3 ). Although the illustrated example shows ASR model 200 being trained on a synthetic speech representation that may include a Mel-frequency spectrogram, ASR model 200 may similarly be trained on a time-domain audio waveform of synthetic speech that has been converted from the synthetic speech representation, for example, via a vocoder (not shown) or other synthesizer device (not shown).

[0052] In the illustrated example, the TTS model 310 receives as input an input text sequence 302 and one or more conditioning inputs 304, and generates as output a synthetic speech representation 306 in each language for training the ASR model 200 in the respective language. Here, the conditioning inputs 304 received by the TTS model 310 for conditioning the synthetic speech representation 306 generated from the input text sequence 302 in the respective language may include at least one of a language embedding 303 associated with the respective language, a speaker embedding 305 specifying the voice characteristics of the respective speaker, or a local embedding 307 specifying accent / dialect information. Thus, the resulting synthetic speech representation 306 may convey a speech style having the accent / dialect identified by the local embedding and in the target speaker's voice identified by the speaker embedding 305.

[0053] 6 shows a schematic diagram of an exemplary speech recognition system 600, in which a TTS model 310 generates synthetic speech representations 306, each conditioned on speaker characteristics 304 of a respective speaker for a corresponding input text sequence 302. The speaker characteristics 304 may include a language embedding 303, a speaker embedding 305, and / or a local embedding 307. That is, the language embedding 303 may specify linguistic information associated with the language of the synthetic speech representation 306 to be produced, the speaker embedding 305 may represent the vocal characteristics of the target speaker, and the local embedding 307 may specify the accent / dialect associated with the synthetic speech representation 306 to be produced.

[0054] A phoneme tokenizer 610 receives the input text sequence 302 and transliterates the language of the text sequence 302 into a native script. In some implementations, the phoneme tokenizer 610 transliterates the language of the text sequence 302 into a native script based on the language embeddings 303. The phoneme tokenizer 610 tokenizes the native script into a phoneme sequence 612. All languages ​​in the phoneme tokenizer 610 share a global Speech Assessment Methods Phonetic Alphabet (SAMPA)-derived phoneme set. The phoneme tokenizer 610 provides the phoneme sequence 612 corresponding to the input text sequence 302 as input to the TTS model 310.

[0055] The TTS model 310 includes an encoder portion 316 that shares language embeddings (i.e., language identifiers) 303 across the first and second languages. The language embeddings 303 may be input to the encoder portion 316 to improve phoneme embedding extraction, where the language embeddings may be jointly trained with the TTS model 310. The TTS model 310 also includes a decoder portion 318 that shares the language embeddings 303 across the first and second languages ​​and shares speaker embeddings 305 for different speakers. The decoder portion 318 may also share local embeddings 307.

[0056] The encoder portion 316 is configured to encode the phoneme sequence 612 to generate the coded phoneme sequence 612, 612E. In some implementations, the encoder portion 316 includes an attention network configured to receive the phoneme sequence 612 and generate the corresponding coded phoneme sequence 612E as a fixed-length context vector for each output step of the decoder portion 318. That is, the attention network in the encoder portion 316 may generate a fixed-length vector for each frame of a mel-frequency spectrogram (e.g., the synthetic speech representation 306) that the decoder portion 318 subsequently generates. The attention network may generate the fixed-length vector by determining a weight for each element of the output of the encoder portion 316 and determining a weighted sum of each element. The attention weights may vary for each decoder portion 318 time step.

[0057] Thus, the decoder portion 318 is configured to receive as input the coded phoneme sequence 612 from the encoder portion 316, the speaker embedding 305, and the language embedding 303 (and optionally the local embedding 307) to generate the synthetic speech representation 306. In some examples, the decoder portion 318 is further conditioned on prosodic information (i.e., the encoder embedding 412) extracted from the synthetic speech representation 306. The synthetic speech representation 306 is in the language identified by the language embedding 303 and represents the voice of the target speaker (e.g., which may be a native speaker of the first language 304a or a native speaker of the second language 304b) identified by the speaker embedding 303.

[0058] The decoder portion 318 decodes the coded phoneme sequence 612E to generate the native synthesized speech representation 306a or the cross-language synthesized speech representation 306b, respectively. For example, if the decoder portion 318 receives the language embedding 303 for the first language and the speaker embedding 305 for the native speaker of the first language 304a, the decoder portion 318 decodes the coded phoneme sequence 612E to generate the native synthesized speech representation 306a. In an alternative example, if the decoder portion 318 receives the language embedding 303 for the first language and the speaker embedding 305 for the native speaker of the second language 304b, the decoder portion 318 decodes the coded phoneme sequence 612E to generate the cross-language synthesized speech representation 306b.

[0059] In the illustrated example, the TTS model 310 provides the synthetic speech representation 306 as input to the VAE 410. The VAE 410 is configured to consume the synthetic speech representation 306 and output either a native audio encoder embedding 412a or a cross-language audio encoder embedding 412b. That is, the VAE 410 extracts latent variables (e.g., prosodic information) from the synthetic speech representation 306. For example, if the VAE 410 receives the native synthetic speech representation 306a, the VAE 410 extracts latent variables of the native synthetic speech representation 306a and generates the native audio encoder embedding 412a. Alternatively, if the VAE 410 receives the cross-language synthetic speech representation 306b, the VAE 410 extracts latent variables of the cross-language synthetic speech representation 306b and generates the cross-language audio encoder embedding 412b. In some implementations, the decoder portion 318 is further conditioned on prosody information extracted from the synthetic speech representation 306 by the VAE 410. That is, the decoder portion 318 receives the audio encoder embedding 412 and further conditions the synthetic speech representation 306 based on the audio encoder embedding 412. An adversarial loss 422 may be applied to the coded phoneme sequence 612E to separate prosody from speaker information.

[0060] The classifier 420 receives the audio encoder embeddings 412 and generates an adversarial loss 422. In some examples, the classifier 420 includes a language classifier. In other examples, the classifier 420 includes an adversarial or speaker classifier. The classifier 420 may be configured to separate the prosodic information extracted from the synthetic speech representation 306 by applying the adversarial loss to the speaker classification. The TTS model 310 may receive the adversarial loss 422 and update parameters based on the adversarial loss 422.

[0061] 7 is a flowchart of an example sequence of operations for a computer-implemented method 700 for training an automated speech recognition (ASR) model 200. At operation 702, the method 700 includes obtaining a multilingual text-to-speech (TTS) model 310. At operation 704, the method 700 includes using the multilingual TTS model 310 to generate native synthetic speech representations 306, 306a for an input text sequence 302 in a first language, conditioned on speaker characteristics 304a of native speakers of the first language. Here, the first language may be Kannada, a low-resource language with only a few native speakers with a constrained speech style. At operation 706, the method 700 includes using the TTS model 310 to generate cross-language synthetic speech representations 306, 306b for the input text sequence 302 in the first language, conditioned on speaker characteristics 304b of native speakers of a different, second language. The second language may be English, which has thousands of native speakers 304b with a variety of speech styles.

[0062] At operation 708, the method 700 includes generating a first speech recognition result 312, 312a for the native synthesized speech representation 306a and a second speech recognition result 312, 312b for the cross-language synthesized speech representation 306b using the ASR model 200. At operation 710, the method 700 includes determining a consistency loss term 352 based on the first speech recognition result 312a and the second speech recognition result 312b. At operation 712, the method 700 includes updating parameters of the ASR model 200 based on the consistency loss term 352. Optionally, the parameters of the ASR model 200 may be fixed while the consistency loss term 352 is back-propagated through the TTS model 310 to update the parameters of the TTS model 310 based on the consistency loss term 352.

[0063] 8 is a schematic diagram of an exemplary computing device 800 that can be used to implement the systems and methods described herein. Computing device 800 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown herein, their connections and relationships, and their functionality are for illustrative purposes only and are not intended to limit the implementation of the invention as described and / or claimed herein.

[0064] Computing device 800 includes a processor 810, a memory 820, a storage device 830, a high-speed interface / controller 840 that connects to memory 820 and a high-speed expansion port 850, and a low-speed interface / controller 860 that connects to a low-speed bus 870 and storage device 830. Each of components 810, 820, 830, 840, 850, and 860 are interconnected using various buses and may be mounted on a common motherboard or in other manners as desired. Processor 810 can process instructions for execution within computing device 800, including instructions stored in memory 820 or on storage device 830, for displaying graphical information for a graphical user interface (GUI) on an external input / output device, such as a display 880 coupled to high-speed interface 840. In other implementations, multiple processors and / or multiple buses may be used, along with multiple memories and multiple types of memory, as desired. Also, multiple computing devices 800 may be connected, each providing a portion of the required operations (eg, as a bank of servers, a group of blade servers, or a multi-processor system).

[0065] The memory 820 stores information non-transiently within the computing device 800. The memory 820 may be a computer-readable medium, a volatile memory unit, or a non-volatile memory unit. The non-transient memory 820 may be a physical device used to temporarily or permanently store programs (e.g., sequences of instructions) or data (e.g., program state information) for use by the computing device 800. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.

[0066] The storage device 830 is capable of providing mass storage for the computing device 800. In some implementations, the storage device 830 is a computer-readable medium. In various different implementations, the storage device 830 may be a floppy disk device, a hard disk device, an optical disk device, or an array of devices including a tape device, a flash memory or other similar solid-state memory device, or a device in a storage area network or other configuration. In additional implementations, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer- or machine-readable medium, such as the memory 820, the storage device 830, or memory on the processor 810.

[0067] The high-speed controller 840 manages bandwidth-intensive operations for the computing device 800, while the low-speed controller 860 manages more bandwidth-intensive operations. Such an allocation of roles is merely exemplary. In some implementations, the high-speed controller 840 is coupled to memory 820, a display 880 (e.g., through a graphics processor or accelerator), and a high-speed expansion port 850, which may receive various expansion cards (not shown). In some implementations, the low-speed controller 860 is coupled to a storage device 830 and a low-speed expansion port 890. The low-speed expansion port 890 may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), but may also be coupled to one or more input / output devices, such as a keyboard, pointing device, scanner, etc., or to a network device, such as a switch or router, for example, through a network adapter.

[0068] Computing device 800 may be implemented in a number of different ways, as shown in the figure, such as as a standard server 800a, or multiple of a group of such servers 800a, as a laptop computer 800b, or as part of a rack server system 800c.

[0069] Various implementations of the systems and techniques described herein may be realized in digital electronic and / or optical circuitry, integrated circuitry, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementation in one or more computer programs executable and / or translatable on a programmable system including at least one programmable processor, which may be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0070] A software application (i.e., a software resource) may refer to computer software that causes a computing device to perform a task. In some examples, a software application may be referred to as an "application," "app," or "program." Exemplary applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.

[0071] These computer programs (also known as programs, software, software applications, or code) include machine instructions for a programmable processor and may be implemented in a high-level procedural and / or object-oriented programming language and / or in an assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" also refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives the machine instructions as a machine-readable signal. The term "machine-readable signal" also refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0072] The processes and logic flows described herein can be implemented by one or more programmable processors, also referred to as data processing hardware, that execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be implemented by special-purpose logic circuitry, such as FPGAs (field-programmable gate arrays) and ASICs (application-specific integrated circuits). Processors suitable for executing computer programs include, by way of example, both general-purpose and special-purpose microprocessors, as well as any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer also includes one or more mass storage devices, such as magnetic, magneto-optical, or optical disks, for storing data, or is operably coupled to the mass storage devices to receive data from, transfer data to, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include, by way of example, all forms of non-volatile memory, media, and memory devices, including semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices, magnetic disks, e.g., internal hard disks or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0073] To enable interaction with a user, one or more aspects of the present disclosure can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, or a touch screen, for displaying information to the user, and optionally a keyboard and pointing device, e.g., a mouse or trackball, by which the user can provide input to the computer. Other types of devices can also be used to enable interaction with the user; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback, and input from the user can be received in any form, including acoustic, speech, or tactile input. Furthermore, the computer can interact with the user by sending documents to and receiving documents from a device used by the user, e.g., by sending a web page to a web browser on the user's client device in response to a request received from the web browser.

[0074] Although several implementations have been described, it will be understood that various modifications can be made without departing from the spirit and scope of the present disclosure. Accordingly, other implementations are within the scope of the following claims. [Explanation of symbols]

[0075] 100 Automated Speech Recognition (ASR) Systems 102 User Devices 107 User Interface Generator 108 Audio Subsystem 200 ASR model, RNN-T model 201 Remote Computing Devices 210 Encoder Network, Prediction Network 220 Prediction Network 230 Collaborative Network 240 softmax layers 310 Multilingual TTS Model, TTS Model 316 Encoder part 318 Decoder part 340 Supervised Loss Term Module 350 Consistency Loss Term Module 360 Data Augmentation Module 410 Hierarchical Variational Autoencoder (VAE), VAE 420 Classifier 600 Voice Recognition System 610 Phoneme Tokenizer 800 computing devices 800a Standard Server, Server 800b laptop computer 800c Rack Server System 810 Processor, Components 820 Memory, Components 830 Storage devices, components 840 High-Speed ​​Interface / Controller, Components 850 High-Speed ​​Expansion Port, Components 860 Low-Speed ​​Interface / Controller, Components 870 Slow Bus 880 Display 890 Low-Speed ​​Expansion Port

Claims

1. 1. A computer-implemented method that, when executed by data processing hardware, causes the data processing hardware to perform an operation, the operation comprising: Obtaining a multilingual TTS model; generating a native synthetic speech representation for an input text sequence in a first language using the multilingual TTS model, the native synthetic speech representation being conditioned on speaker characteristics of a native speaker of the first language; generating a cross-language synthetic speech representation for the input text sequence in the first language using the multilingual TTS model, the cross-language synthetic speech representation being conditioned on speaker characteristics of a native speaker of a different second language; determining an adversarial loss term conditioned on the first language based on the native synthetic speech representation and the cross-language synthetic speech representation; updating parameters of the multilingual TTS model based on the adversarial loss term; 11. A computer-implemented method comprising:

2. The operation applying data augmentation to at least one of the native synthetic speech representation or the cross-language synthetic speech representation. The computer-implemented method of claim 1 further comprising:

3. The computer-implemented method of claim 1, wherein the multilingual TTS model shares language embedding across the first and second languages.

4. Prior to generating the native synthetic speech representation and the cross-language synthetic speech representation, transliterating the input text sequence in the first language into a native script; tokenizing the native script into a sequence of phonemes using a set of global phonemes shared between the first and second languages; and The computer-implemented method of claim 1 further comprising:

5. The computer-implemented method of claim 4, wherein generating the native synthetic speech representation includes generating the native synthetic speech representation based on the phoneme sequence.

6. The computer-implemented method of claim 4, wherein generating the cross-language synthetic speech representation includes generating the cross-language synthetic speech representation based on the phoneme sequence.

7. The operation encoding the sequence of phonemes using an encoder of the multilingual TTS model; decoding the encoded phoneme sequence using a decoder of the multilingual TTS model to generate each of the native synthesized speech representation or the cross-language synthesized speech representation; The computer-implemented method of claim 4, further comprising:

8. The operation is generating a first speech recognition result for the native synthetic speech representation and a second speech recognition result for the cross-language synthetic speech representation using a speech recognition model; determining a consistency loss term associated with a difference between the first speech recognition result and the second speech recognition result; updating parameters of the speech recognition model based on the consistency loss term; 8. The computer-implemented method of claim 2, further comprising:

9. The operation is generating a first cross-entropy loss term based on the first speech recognition result and the input text sequence in the first language; determining a second cross-entropy loss term based on the second speech recognition result and the input text sequence in the first language; and updating parameters of the speech recognition model based on the first and second cross-entropy loss terms; and The computer-implemented method of claim 8, further comprising:

10. The operation is back-propagating the first and second cross-entropy loss terms to the multilingual TTS model to update parameters of the multilingual TTS model. The computer-implemented method of claim 9, further comprising:

11. 1. A system comprising: data processing hardware; memory hardware in communication with the data processing hardware; the memory hardware storing instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations, the operations comprising: Obtaining a multilingual TTS model; generating a native synthetic speech representation for an input text sequence in a first language using the multilingual TTS model, the native synthetic speech representation being conditioned on speaker characteristics of a native speaker of the first language; generating a cross-language synthetic speech representation for the input text sequence in the first language using the multilingual TTS model, the cross-language synthetic speech representation being conditioned on speaker characteristics of a native speaker of a different second language; determining an adversarial loss term conditioned on the first language based on the native synthetic speech representation and the cross-language synthetic speech representation; updating parameters of the multilingual TTS model based on the adversarial loss term; Including, the system.

12. The operation applying data augmentation to at least one of the native synthetic speech representation or the cross-language synthetic speech representation. The system of claim 11 further comprising:

13. The system described in claim 11, wherein the multilingual TTS model shares language embedding across the first and second languages.

14. Prior to generating the native synthetic speech representation and the cross-language synthetic speech representation, the operation comprises: transliterating the input text sequence in the first language into a native script; tokenizing the native script into a sequence of phonemes using a set of global phonemes shared between the first and second languages; and The system of claim 11 further comprising:

15. The system of claim 14, wherein generating the native synthetic speech representation includes generating the native synthetic speech representation based on the phoneme sequence.

16. The system of claim 14, wherein generating the cross-language synthetic speech representation includes generating the cross-language synthetic speech representation based on the phoneme sequence.

17. The operation encoding the sequence of phonemes using an encoder of the multilingual TTS model; decoding the encoded phoneme sequence using a decoder of the multilingual TTS model to generate each of the native synthesized speech representation or the cross-language synthesized speech representation; The system of claim 14 further comprising:

18. The operation is generating a first speech recognition result for the native synthetic speech representation and a second speech recognition result for the cross-language synthetic speech representation using a speech recognition model; determining a consistency loss term associated with a difference between the first speech recognition result and the second speech recognition result; updating parameters of the speech recognition model based on the consistency loss term; 18. The system of any one of claims 12 to 17, further comprising:

19. The operation is generating a first cross-entropy loss term based on the first speech recognition result and the input text sequence in the first language; determining a second cross-entropy loss term based on the second speech recognition result and the input text sequence in the first language; and updating parameters of the speech recognition model based on the first and second cross-entropy loss terms; and 20. The system of claim 18, further comprising:

20. The operation is back-propagating the first and second cross-entropy loss terms to the multilingual TTS model to update parameters of the multilingual TTS model.

20. The system of claim 19, further comprising:

Citation Information

Patent Citations

  • Audio signal generation model learning device, audio signal generation device, method, and program

    JP2019139102A

  • JPP7502561B

  • Multi-lingual speech synthesis

    US20050144003A1

  • Speech translation method and system using multilingual text-to-speech synthesis model

    WO2019139431A1