Systems and methods for neural codec language model for zero sample text-to-speech synthesis

By using a zero-sample cross-language text-to-speech model and a neural codec language model, the problem of traditional TTS systems' dependence on labeled data is solved, enabling high-quality personalized speech generation and cross-language speech cloning, and improving speech naturalness and speaker similarity.

CN121014076APending Publication Date: 2025-11-25MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380093015.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-03-02
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Traditional TTS systems require a large amount of labeled data for personalized speech generation, especially for new speakers and languages. Furthermore, speaker similarity and speech naturalness decrease in zero-shot scenarios, and existing methods require additional training data and complex feature engineering.

Method used

Employing a zero-sample cross-language text-to-speech model, utilizing large-scale multi-speaker training data and a neural codec language model, personalized speech is generated using a small amount of audio data from the target speaker. Speech features are modified by combining attribute IDs to achieve cross-language speech cloning.

Benefits of technology

It achieves high-quality personalized speech generation without the need for additional training data, improves speech naturalness and speaker similarity, supports cross-language applications, and reduces computational costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121014076A_ABST
    Figure CN121014076A_ABST
Patent Text Reader

Abstract

Systems and methods are provided for accessing a machine learning model configured as a zero-sample cross-language text-to-speech model, the machine learning model having been previously trained on a text-to-speech training dataset comprising different bilingual speech transcription pairs; obtaining a first text prompt in a first language, a second text prompt in a second language, and a voice sample including audio data from an unseen target speaker; providing the first text prompt in a first language, the second text prompt in a second language and a voice sample from the target speaker as input to the machine learning model; and finally generating a personalized speech output based on the input and at least by converting the second textual cue in a second language using a synthetic sound of the target speaker based on the speech sample from the target speaker.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] Automatic speech recognition systems and other speech processing systems are used to process and decode audio data to detect spoken utterances (e.g., words, phrases, and / or sentences). The processed audio data is then used for various downstream tasks, such as search-based queries, speech-to-text transcription, language translation, etc. Conversely, text-to-speech (TTS) systems are used to detect text-based utterances and subsequently generate simulated spoken utterances corresponding to the detected text-based utterances.

[0002] In most TTS systems, the raw text is tagged as words and / or phonological units. Each word or phonological unit is then associated with a specific phonological transcription and prosodic unit, forming a linguistic representation of the text. The phonological transcription contains information about how to pronounce the phonological units, while the prosodic unit contains information about larger phonological units, including intonation, stress, rhythm, timbre, and speech rate. Once the linguistic representation is generated, a synthesizer or vocoder is able to convert it into synthesized speech that is audible and recognizable to the human ear.

[0003] Traditional TTS systems typically require large amounts of labeled training data, initially used to train the TTS system to be speaker-independent and / or multilingual. However, even larger amounts of labeled data are still needed, especially when the TTS system has not been previously trained for new speakers and / or new languages ​​to personalize the system. Given the foregoing, there is an ongoing need for improved systems and methods to build and use personalized TTS systems to generate personalized synthesized speech from text-based input.

[0004] The subject matter claimed herein is not limited to the embodiments that address any shortcomings or operate only in environments such as those described above. Rather, this background is provided merely to illustrate an exemplary technical field in which some of the embodiments described herein may be practiced. Summary of the Invention

[0005] The disclosed embodiments include systems, methods, and apparatus for performing TTS processing and for generating and utilizing machine learning modules configured for zero-shot learning, which are personalized to facilitate the generation of personalized voices from text-based inputs that will be used to generate synthetic voices.

[0006] Some disclosed embodiments relate to systems and methods for generating cross-lingual personalized speech. For example, the system accesses a machine learning model configured as a zero-sample cross-lingual text-to-speech model, which has been previously trained on a text-to-speech training dataset comprising multiple different bilingual speech transcription pairs. The system also acquires a first text prompt in a first language, a second text prompt in a second language, and speech samples including audio data from a target speaker, wherein the target speaker is an unseen target speaker, and therefore no audio data from that target speaker is included in the text-to-speech training dataset. The system then feeds the first text prompt in the first language, the second text prompt in the second language, and the speech samples from the target speaker as input to the machine learning model, and ultimately generates a personalized speech output based on these inputs and, at least by using a synthesized voice of the target speaker based on the speech samples from the target speaker.

[0007] Some disclosed embodiments also relate to generating modified personalized speech based on different attributes. For example, systems and methods are provided to access a machine learning model configured for zero-shot cross-lingual text-to-speech modeling, which has been previously trained on a text-to-speech training dataset. The system also acquires text prompts and speech samples containing audio data from a target speaker, wherein the target speaker is an unseen target speaker, and therefore no audio data from the target speaker is included in the text-to-speech training dataset. Additionally, the system accesses an attribute ID configured to modify the personalized speech output based on specific attributes. Subsequently, the system applies the text prompts, speech samples, and attribute IDs to the machine learning model and generates personalized speech output based on the text prompts using a synthesized voice of the target speaker using the speech samples, the synthesized voice of the target speaker being modified according to the attribute ID.

[0008] Some disclosed embodiments also relate to systems and methods for accessing a machine learning model configured as a zero-shot cross-lingual speech-to-speech model, which has previously been trained on a text-to-speech training dataset comprising multiple different bilingual speech transcription pairs. The system acquires speech samples containing audio data from a target speaker, wherein the target speaker is unseen and therefore no audio data from that target speaker is included in the text-to-speech training dataset. The system then generates a first transcription of the speech samples in a first language and, based on translating the first transcription of the speech samples into a second language, generates a second transcription of the speech samples in the second language. The system then applies the first transcription, the second transcription, and the speech samples to the machine learning model and generates a personalized speech output using a synthesized voice of the target speaker based on the speech samples from the target speaker and the second transcription of the speech samples in the second language.

[0009] This overview is provided to present, in a simplified form, a selection of concepts that will be further described in the detailed description below. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to help determine the scope of the claimed subject matter.

[0010] Additional features and advantages will be set forth in the following description, and in part will be obvious from the description, or may be learned by practice of the teachings herein. The features and advantages of this disclosure may be realized and obtained by means of the instruments and combinations particularly pointed out in the appended claims. The features of this disclosure will become more apparent from the following description and the appended claims, or may be learned by practice of this disclosure as set forth below. Attached Figure Description

[0011] To describe how the above-described and other advantages and features can be obtained, a more specific description of the subject matter briefly described above will be presented with reference to various specific embodiments, which are illustrated in the accompanying drawings. It should be understood that these drawings only describe typical embodiments and are therefore not intended to limit the scope of the invention. The embodiments will be described and explained with additional specificity and detail using the drawings, in which:

[0012] Figure 1 Examples are illustrated of computing environments in which computing systems are incorporated and / or used to perform the various aspects of the disclosed embodiments.

[0013] Figure 2 An example diagram of a machine learning model is shown, including its input and output, which is configured to perform personalized text-to-speech generation.

[0014] Figure 3An example graph illustrating a machine learning model, including its input and output, is shown. This model is configured as an encoder-decoder network to generate and process acoustic tokens as part of personalized text-to-speech generation.

[0015] Figure 4 An example diagram illustrating a neural audio codec model is shown.

[0016] Figure 5 An example diagram illustrating a machine learning model, including its inputs and outputs, is shown. This model is configured to perform cross-language personalized text-to-speech generation.

[0017] Figure 6 An example diagram illustrates a neural audio codec model for generating cross-language personalized text-to-speech output.

[0018] Figure 7 Example diagrams are shown illustrating neural audio codec language models applicable to zero-shot cross-lingual text-to-speech generation and zero-shot speech-to-speech generation.

[0019] Figure 8 An example embodiment of a flowchart is illustrated, which involves multiple actions associated with generating personalized text-to-speech output using a zero-sample personalized text-to-speech model.

[0020] Figure 9 An example embodiment of a flowchart is illustrated, which involves multiple actions associated with generating modified personalized text-to-speech output using a zero-sample personalized text-to-speech model.

[0021] Figure 10 An example embodiment is illustrated with a flowchart of multiple actions associated with using a zero-sample personalized speech-to-speech model. Detailed Implementation

[0022] The disclosed embodiments are systems, methods, and frameworks for improving the training and use of machine learning models to synthesize personalized speech for unseen target speakers and cross-lingual personalized speech.

[0023] Traditional TTS systems and methods utilize cascaded TTS systems, which employ acoustic models, vocoders, and Mel spectrograms as intermediate representations. While traditional TTS systems can synthesize speech from single or multiple speakers, they still require high-quality, clean data from recording studios. Large-scale data crawled from the internet is insufficient and leads to performance degradation in TTS systems. Due to the relatively small amount of training data currently used, existing TTS systems still suffer from poor generalization ability. Furthermore, in zero-shot scenarios, speaker similarity and speech naturalness decrease significantly for unseen speakers. Current methods to improve this problem attempt to utilize speaker adaptation and speaker coding methods, which require additional training data, fine-tuning, complex pre-designed features, and / or extensive structural engineering.

[0024] First, shift your attention Figure 1 , Figure 1 The computing system 110, which is part of the computing environment 100, has been described. The computing environment 100 also includes remote systems 120 that are in communication with the computing system 110 (via network 130). The computing system communicates with the remote systems 120, which include one or more processors 122, one or more computer-readable instructions 118, and one or more hardware storage devices 124. In some cases, the remote systems 120 may further include databases containing data that can be used as training data (e.g., text data not included in local storage). Additionally or alternatively, the remote systems 120 may include machine learning systems and / or software programs or applications external to the computing system 110.

[0025] The computing system 110 includes, for example, one or more processors 112 (such as one or more hardware processors) and storage (i.e., hardware storage devices 140) storing computer-readable instructions 118, wherein the one or more hardware storage devices 140 are capable of accommodating any number of data types and any number of computer-readable instructions 118. The computing system 110 is configured to implement one or more aspects of the disclosed embodiments by means of the computer-readable instructions 118 when executed by the one or more processors 112. The computing system 110 is also shown to include user interfaces 114 and input / output (I / O) devices 116.

[0026] The computing system 110 is configured to generate, train, and use various machine learning models, including a zero-shot model 146 and a cross-lingual model 147 (which is also a zero-shot model), which can generate personalized synthetic speech for a new target speaker. Regarding the use of the term "zero-shot," as used with reference to the disclosed zero-shot models, it should be understood that the term generally means that the corresponding zero-shot model is capable of and can be configured to generate personalized speech for a new target speaker when the zero-shot model is applied to target reference speech (audio) from the new target speaker, even if the model has not previously been applied to any target reference speech or audio associated with the new target speaker.

[0027] like Figure 1 As shown, hardware storage devices 140 are represented as single storage units. However, it will be understood that hardware storage devices 140 are distributed storage systems distributed across several separate, and sometimes remote, systems 120. Computing system 110 may also include distributed systems, wherein components of one or more computing systems 110 are maintained / operated by different discrete systems that are geographically separated and each performs different tasks. In some instances, multiple distributed systems perform similar and / or shared tasks to achieve the disclosed functionality, such as in a distributed cloud environment.

[0028] The memory (e.g., hardware storage device 140) includes computer-readable instructions 118 for instantiating or executing one or more models shown in the computing system 110. These models are configured as machine learning models or machine learning-enabled models, such as deep learning models and / or algorithms and / or neural networks. In some instances, the one or more models are configured as engines or processing systems (e.g., computing systems integrated within the computing system 110), where each engine includes one or more processors (e.g., hardware processors 112) and computer-readable instructions 118 corresponding to the computing system 110. In some configurations, the model is a set of numerical weights embedded in a data structure, while the engine is a separate piece of code that, when executed, is configured to load the model and compute its output in the context of input audio.

[0029] Hardware storage device 140 is configured to store and / or cache different data types in memory storage, including training data 141, text prompts 142, acoustic prompts 143, attribute IDs 144, and synthesized speech 145 as described herein.

[0030] Training data

[0031] Training data 141 is used for initial training of the zero-shot model 146 and / or the cross-lingual model 147. For example, the zero-shot method for speaker voice cloning described herein advantageously utilizes a well-trained multi-speaker TTS source model. The machine learning model (e.g., model 100) is trained using a corpus (e.g., training data 141) consisting of thousands of hours of multi-speaker speech data. In some cases, the training data includes approximately 60,000 hours of speaker data. This corpus includes speaker data in a single language, or alternatively, speaker data in multiple languages. When the raw audio data contained in the corpus is only audio, the system employs a speech recognition model to generate transcriptions of the speaker audio data.

[0032] Compared to previous TTS training datasets that required very clean audio data as training data, the system in this paper is able to utilize noisy speech data, even including some erroneous transcriptions, because the disclosed embodiments provide a method robust to noise and errors and can accurately generalize by utilizing large datasets. Conventional systems typically use tens or hundreds of hours of speaker data for training, contrasting with the thousands of hours of training data used in the disclosed embodiments. Because the disclosed embodiments are able to model language using neural codecs, the system and method are able to perform contextual learning. Conventional systems are limited in this respect.

[0033] Text prompt

[0034] Text prompt 142 includes sequences of characters, symbols, and / or numbers extracted from various sources. For example, text prompt 142 may include text message data, email content, news articles, web pages, books, mobile application pages, etc. In some cases, the characters of text prompt 142 are identified through optical text recognition of physical or digital samples of text prompt 142. Alternatively, the characters of text prompt 142 may be identified by processing metadata of digital samples of text prompt 142. Text prompt 142 is processed by a zero-shot model 146 to generate synthesized speech 145.

[0035] Acoustic cues

[0036] Natural language audio is used for speech samples of new target speakers (e.g., acoustic cues 143). Natural language audio is extracted from previously recorded files, such as video recordings with audio or audio-only recordings. Examples of recordings include videos, podcasts, voicemails, voice memos, songs, etc. Natural language audio is also extracted from actively streaming content that is real-time, continuous speech, such as news broadcasts, telephone calls, virtual or face-to-face meetings, etc.

[0037] In some cases, previously recorded audio files are streamed. Natural audio data includes spoken speech utterances without corresponding clean speech reference signals. Natural audio data is recorded from multiple sources, including applications, conferences involving one or more speakers, background noise, and the environment surrounding the human speaker. It should be understood that natural language audio includes one or more spoken languages ​​from around the world. Therefore, the zero-shot model 146 can be trained in different languages, where the source language of the training data 141 and the source language of the acoustic cues 143 are the same.

[0038] Attribute ID

[0039] Attribute ID 144 refers to an additional guide to model parameters that modifies the personalized synthesized speech based on specific attributes during inference or runtime. For example, if the attribute ID is a language ID corresponding to a specific language, the personalized synthesized speech is modified to include the accent associated with that language.

[0040] If the attribute ID is an emotion ID corresponding to a specific emotion, the personalized synthesized speech is modified to convey that emotion, even if that emotion is different from the emotion associated with the target speaker's acoustic cues. If the attribute ID is a speech style ID corresponding to a specific speech style, the personalized synthesized speech is modified to be generated in that specific speech style, even if the speech style corresponding to the speaker style ID is different from the speaker style associated with the target speaker's acoustic cues.

[0041] Synthetic speech

[0042] Synthetic speech 145 includes synthetic audio data generated by a zero-shot model 146 and / or a cross-linguistic model 147. Synthetic speech 145 includes spoken utterances corresponding to words, phrases, and sentences identified in text prompts 142. Synthetic speech 145 uses a cloned voice of the target speaker, based on acoustic and text prompts. The spoken utterances of synthetic speech 145 can be generated with different target speaker voices (i.e., cloned voices), different languages, different speaking styles, etc. The spoken utterances of synthetic speech 145 are characterized by the speech features of the target speaker (e.g., acoustic features, linguistic features, and / or prosodic features). Generating synthetic speech 145 facilitates the imitation of natural language audio (e.g., the natural speaking voice of the target speaker).

[0043] By implementing the disclosed embodiments in this manner (e.g., by using a computing system such as computing system 110), numerous technical advantages over existing systems are achieved, including the generation and utilization of a high-quality TTS system architecture, sometimes referred to herein as a zero-shot personalized text-to-speech model. Compared to conventional systems that require additional training with new labeled training data, the zero-shot model 146 and the cross-lingual model 147 are capable of generating personalized voices for new target speakers without applying the model to new labeled training data associated with the new target speaker, and without sacrificing the quality of the synthesized speech.

[0044] Traditional zero-shot processing systems require additional training because they rely on techniques that generate speaker embeddings using speaker verification systems. These embeddings are then fed into their text-to-speech (TTS) systems without capturing the prosodic features of the target speaker, such as the fundamental frequency, energy, and duration of the target speaker, even though prosodic features play an important role in voice cloning.

[0045] By implementing the disclosed embodiments, TTS systems such as computational system 110 can generate more natural and expressive synthesized speech, thereby improving the similarity between synthesized speech and natural spoken language. The disclosed system can synthesize a personalized voice (i.e., a personal voice; a cloned voice) for a target speaker using only a few audio clips without requiring a text transcription from the speaker. After a training process, the TTS system can clone specific features of the target speaker and incorporate them into the personalized voice. The zero-shot method disclosed herein can clone a speaker's voice using only a few seconds of audio without requiring a corresponding text transcription from a new or unseen speaker as a reference. Furthermore, as previously described, the disclosed system can rapidly clone the features of a target speaker using speaker information extracted from a few seconds of reference audio.

[0046] To clone unseen voices, the system directly synthesizes speech for the new target speaker using only the speaker information input into the source model, without requiring additional training. By using a zero-shot method for voice cloning, the computational cost of training is significantly reduced in training time, and because it eliminates the need to generate new training datasets for new target speakers, the system achieves this.

[0047] It should be understood that this is another advantage of the disclosed embodiments over conventional zero-shot TTS systems, which focus on monolingual TTS scenarios, meaning that their synthesized speech is generated in the same language as the reference speech. Unlike these conventional systems, the disclosed embodiments advantageously provide a framework for cross-lingual TTS voice cloning, meaning that synthesized speech can be generated in a language different from the language corresponding to the reference audio.

[0048] Furthermore, when examining the experimental results of the zero-shot model 145, the disclosed embodiments significantly outperform conventional models in terms of speech naturalness and speaker similarity. For example, machine learning model 200 achieves a +0.12 improvement in Comparative Mean Option Score (CMOS) and a +0.93 improvement in Similarity Mean Option Score (SMOS) compared to conventional TTS systems. It also achieves a +0.04 CMOS score relative to ground reality, indicating that the synthesized speech of an unseen speaker is as natural as human recordings. Moreover, qualitative analysis shows that the disclosed embodiments are capable of synthesizing diverse outputs using the same text and target speaker, which is beneficial for creating pseudo-data to generate training data for speech recognition tasks. The speech output of machine learning model 200 also preserves the acoustic environment (e.g., reverberation or other environmental features) and emotion of the acoustic cues.

[0049] These advantages are particularly evident in real-time applications of voice cloning and speech synthesis, as well as in cross-language applications. Some examples of real-time applications include Skype translators and other voice translators in IoT devices.

[0050] Personalized TTS generation

[0051] Now let's turn our attention to... Figure 2 It shows an example diagram of a machine learning model (e.g., model 200), including inputs and outputs configured to perform personalized text-to-speech generation. It should be understood that model 200 represents... Figure 1 The zero-sample model 146. (e.g.) Figure 2 As shown, text prompt 202 (e.g., from...) Figure 1 Text cue 142) and acoustic cue 204 (e.g., from Figure 1 The acoustic cue 143 is provided as input to model 200. The text cue 202 is the text provided to the model for synthesizing speech.

[0052] Acoustic cue 204, also referred to herein as a registration record or target speaker sample speech, comprises a limited amount of audio data from an unseen target speaker. In some cases, acoustic cue 204 includes a registration record of 3 seconds or other durations (e.g., 4 seconds, 5 seconds, 5-10 seconds, 10-15 seconds, but preferably less than 10 seconds). Acoustic cue 204 can also be defined as audio data comprising a single spoken utterance of short duration (e.g., less than 10 seconds or even less than 5 seconds).

[0053] The text prompt 202 (via phoneme conversion 206) is converted into a phoneme sequence 208. The acoustic prompt 204 is converted into a source acoustic token 212 using an audio codec encoder 210.

[0054] Based on phoneme sequence 208 and source acoustic token 212, model 200 performs neural codec language modeling 214 and generates target acoustic token 216. Target acoustic token 216 is provided as input to audio codec decoder 218, which converts target acoustic token 216 into personalized speech waveform 220. It should be understood that audio codec encoder 210 and audio codec decoder 218 are part of the same audio codec model referenced herein.

[0055] As described above, conventional systems and methods for generating personalized speech involve feeding phoneme input to a machine learning model, converting the phoneme input into a Mel spectrogram, and then converting the Mel spectrogram into a waveform. In contrast, the novel embodiments described herein convert phoneme input (e.g., phoneme sequence 208) into discrete codes (e.g., target acoustic token 216), which can be processed to output a final waveform (e.g., personalized speech waveform 220) using a machine learning model.

[0056] As will be understood from the foregoing and following descriptions, the disclosed embodiments provide a technological improvement over conventional systems by configuring the training and use of a neural codec language model to generate text-to-speech output as a conditional language modeling task, rather than performing continuous signal regression as in conventional systems that use intermediate outputs of Mel spectrograms. This enables the system to use a wider range of initial training data and requires less training data to personalize the TTS model.

[0057] For example, the machine learning model described in this paper (e.g., Model 200) generates discrete audio codec codes based on phonemes and acoustic cues, corresponding to the target content and the speaker's voice. Through methods such as... Figure 2 The configuration shown, featuring machine learning models like Model 200, advantageously supports a variety of speech synthesis applications, such as zero-shot TTS, speech editing, and content creation combined with other generative AI models (e.g., GPT-3). This allows for the use of cue-based, high-level large-scale model techniques, such as those used in GPT, for TTS tasks. Acoustic tokens also allow the system to generate diverse synthesized results in TTS by using different sampling strategies during inference and to utilize a wider range of training data.

[0058] Now let's turn our attention to... Figure 3 It shows an example diagram of a machine learning model, including the inputs and outputs of a neural audio codec model (e.g., model 300). Figure 3 As shown, model 300 includes encoder 302 (representing...). Figure 2The audio codec encoder 210), multiple vector quantizers (VQ) (e.g., VQ1, VQ2, ..., VQ8) and decoder 304 (representing... Figure 2 (audio codec decoder).

[0059] Acoustic cue 306 from the target speaker (representing) Figure 2 Acoustic cues 202 are provided as input to encoder 302. A first quantization layer (e.g., VQ1) generates a first quantization acoustic token set (tokens 12, 43, 8, ..., 59). This first quantization acoustic token set corresponds to the codebook of stage 1 included in quantization token 308. This token set is provided as residual 1 as input to the next quantization layer (e.g., VQ2).

[0060] Based on the first set of quantized acoustic tokens as input, a second quantization layer generates a second set of quantized acoustic tokens (e.g., tokens 71, 21, 38, ..., 67), corresponding to the codebook of stage 2, and is provided as residual 2 as input to a subsequent quantization layer (not shown). The acoustic tokens are processed by the remaining quantizers until residual 7 is provided as input to a final quantization layer (e.g., VQ8), which generates a final set of quantized acoustic tokens (e.g., tokens 9, 16, 52, ..., 84). This final set corresponds to the codebook of stage 8. Each of the different sets of quantized acoustic tokens is summed and provided as input to decoder 304. Decoder 304 then converts the cascaded sets of acoustic tokens into the final waveform 310 (i.e., synthesized speech).

[0061] Compared to traditional systems, Model 300 achieves several technical advantages. For example, since audio data is typically stored as a sequence of 16-bit integer values, conventional generative models are configured to output 2^16 = 65536 probabilities at each time step to synthesize the original audio. Furthermore, audio sampling rates exceeding 10,000 result in exceptionally long sequence lengths, making original audio synthesis even more difficult to process. To address this, by implementing speech quantization as in Model 300, and as included in the embodiments disclosed herein, the TTS system is able to compress both integer values ​​and sequence length. For example, the neural audio codec model 300 is able to represent speech using discrete tokens. To compress audio for network transmission, Model 300 is able to encode waveforms into discrete acoustic codes and reconstruct high-quality waveforms even if the target speaker does not appear in the training data.

[0062] Compared to traditional audio codec methods, neural codec models (e.g., Model 300) perform significantly better at low bit rates. Furthermore, the quantization token 308 contains sufficient information about the speaker and recording conditions to generate high-quality synthesized speech from the target speaker's voice. Compared to other quantization methods, the neural audio codec model 300 offers the following technical advantages: 1) It contains rich speaker and acoustic information, helping to preserve speaker identity during reconstruction. 2) The disclosed system and method can utilize readily available codec decoders to convert discrete tokens into waveforms without additional effort in vocoder training, which requires processing continuous spectra. 3) It also reduces the time step and sequence length, improving system / model efficiency.

[0063] like Figure 3 The illustrated neural audio codec model 300 is a tokenizer configured as an encoder-decoder model, such as a convolutional encoder-decoder model. The encoder 306 generates embeddings at specific frequencies (e.g., 75 Hz) for input waveforms of different frequencies (e.g., 24 kHz), with a sampling rate reduction of 320 times. Each embedding is modeled by residual vector quantization (RVQ), where the system selects multiple quantizers. In some cases, such as... Figure 3 As shown, an eight-level quantizer with multiple entries (e.g., 1024 entries) is used. In some cases, the neural audio codec model 300 is configured at a 6k bit rate for 24kHz audio reconstruction. In such an example configuration, given a 10-second waveform, the discrete representation consists of a matrix of 750 × 8 entries, where 750 = (24,000 × 10) / 320 is the downsampling time step, and 8 is the number of quantizers (e.g., VQ1, VQ2, ..., VQ8).

[0064] It should be understood that, although Figure 3 Eight quantizers are illustrated, but the Neural Audio Codec Model 300 can be configured according to different bitrate settings, resulting in different time steps and different numbers of quantizers. For example, the higher the bitrate, the more quantizers are used, and therefore the better the reconstruction quality. As another example, if the bitrate is set to 12kHz and 16 quantizers are used, then a 10-second waveform will correspond to a 750×16 entry matrix, further improving the quality of the reconstructed audio.

[0065] In summary, by utilizing the discrete code from all quantizers, Model 300 is able to generate real-valued embeddings and reconstruct the output waveform from acoustic tokens at the desired frequency (e.g., 24 Hz).

[0066] Now let's turn our attention to... Figure 4 It illustrates an example diagram of a neural audio codec model (e.g., Model 400), which represents Figure 3 Model 300. Model 400 has an autoregressive (AR) converter decoder (e.g., AR model 402) and a non-autoregressive (NAR) converter decoder (e.g., NAR model 404), which are configured to perform conditional codec language modeling.

[0067] In this example, text 406 uses a grapheme-to-phoneme (G2P) model (representing...). Figure 1 The phoneme conversion) is converted to the phoneme sequence "x". The acoustic cue 408 is converted to the source acoustic token (e.g., token) using the audio codec encoder 410 (representing encoder 302). ).

[0068] Note that, given a dataset containing audio samples and their corresponding phoneme transcriptions, the token notation above is as follows: for example, “c” represents a two-dimensional acoustic code matrix, and T is the downsampled utterance length. The row vectors of each acoustic code matrix (e.g., c) t,: ) represents the code of frame t (e.g., 8 codes), and the column vector of each acoustic code matrix (e.g., c) :,j Let represent the code sequence from the j-th codebook, where j∈{1,...,8}. In this case, the set reaches 8 because the system uses 8 quantizers. However, as mentioned earlier, the system can also use other numbers of quantizers for different bit rate settings.

[0069] Back Figure 4 Phoneme sequences and acoustic tokens are provided as input to AR model 402. In AR model 402, for each source acoustic token, only the token to the left (i.e., the next subsequent token) is considered, see matrix 410. For example, token c 0,1 Pay attention to token c 1,1 Token c 1,1 Pay attention to token c 2,1 And so on, until the final token is processed (e.g., c). T,1 In some cases, a special <eos>The token is attached to the source acoustic token.

[0070] NAR model 404 also receives the phoneme sequence "x" generated by the G2P model based on text cue 406 as input. Similarly, acoustic cue 408 is converted into a source acoustic token set. All tokens in this set can follow all other tokens in the set, see matrix 412. The output from AR model 402 (associated with the first quantization layer, e.g.) Figure 3 VQ1 is provided as input to NAR model 404, which is associated with all subsequent quantization layers (e.g., VQ1). Figure 3 VQ2, ..., VQ8).

[0071] Compared to traditional systems and methods used for quantization in TTS applications, Model 400 offers several technical advantages. For example, the neural audio codec model (also known as the neural speech codec model or Model 400) allows the system to operate on discrete audio representations. Figure 3 As described, tokens have a hierarchical structure due to residual quantization in the neural codec model. Notably, tokens from previous quantizers recover acoustic properties (such as speaker identity), while subsequent quantizers learn finer acoustic details. Furthermore, each quantizer is trained to model the residuals from previous quantizers.

[0072] To utilize this configuration for data processing, the disclosed embodiments employ a neural audio codec model that includes multiple language models situated within the described hierarchical structure. For example, as... Figure 4 As shown, the neural audio codec model includes two conditional language models: an autoregressive (AR) decoder-only language model and a non-autoregressive (NAR) decoder-only language model.

[0073] For discrete tokens from the first quantizer, the system trains an AR model conditioned on a phoneme sequence and acoustic cues. For discrete tokens from the second to the final quantizer, the system trains a NAR model.

[0074] Since tokens can be accessed by each other in a NAR manner, an acoustic cue matrix is ​​used as an acoustic cue to constrain speaker identity. Therefore, the NAR model is conditioned on the phoneme sequence, the acoustic cue, and the predicted acoustic tokens belonging to the previous codebook (e.g., the set of tokens output from each quantization layer).

[0075] The combination of AR and NAR models offers a good trade-off between speech quality and inference speed. On the one hand, the rate of speech generation should be consistent with the registration records. In some cases, training a length predictor for different speakers is difficult because each person's speaking speed can vary significantly. In such cases, the AR model compensates for this difficulty because of its flexibility in predicting the acoustic sequence length. On the other hand, for consecutive stages, the NAR model advantageously reduces the time complexity of token processing because the number of output slots follows the sequence length of the first stage.

[0076] Many of the disclosed embodiments include AR networks configured with an AR-NAR architecture containing a specific number of quantization layers. However, it should be understood that the disclosed functionality and benefits can also be obtained by utilizing other types of AR networks containing the same or different number of quantization layers as described. For example, in some alternative embodiments, the disclosed text-to-speech synthesis is provided and performed using a single AR network that is larger than the disclosed dual AR-NAR architecture and / or may include more quantization layers. It is noteworthy that, regardless of the specific architecture used, the AR network is configured to process input tokens and generate a final set of acoustic tokens, which can then be reconstructed into the final waveform. Additional details regarding AR model configuration and functionality will now be provided.

[0077] AR model

[0078] The AR model generates tokens from a first quantizer. It includes phoneme embedding, acoustic embedding, a converter decoder, and a prediction layer. To generate speech with specific content, the system uses a phoneme sequence as a phoneme cue for the language model. Therefore, the model input is a concatenation of the phoneme sequence and the acoustic cue.

[0079] In some cases, two special appendices are made after each of the above inputs. <eos>Tokens. The system calculates sinusoidal position embeddings for both the prompt and input tokens. For the converter model, each token can focus on the token to its left, as shown in Figure 410. The model is optimized to maximize the probability of the next token in the first codebook. The system shares the parameters of the output projection layer with the parameters of the acoustic embedding.

[0080] In the AR model, the system does not explicitly extract audio segments as cues during training. Instead, the training process is a purely causal language model training. In this way, any prefix sequence is treated as a cue for the latter half of the sequence. During inference, given a registration record, the system concatenates the phoneme sequence of the registration record (i.e., the transcription of the acoustic cue) with the phoneme sequence used for synthesis (i.e., the text cue). Simultaneously, the acoustic token sequence of the registration record is used as a prefix in AR decoding.

[0081] NAR model

[0082] When the system obtains the first quantizer code through the AR model, it then uses the NAR model to generate the code for subsequent quantizers. The NAR model has a similar architecture to the AR model, but it contains multiple separate acoustic embedding layers. In some cases, the NAR model includes eight separate acoustic embedding layers. In each training step, the system randomly samples a training phase, where the model is trained to maximize acoustic tokens from the quantizer codebook corresponding to the sampled training phase. The NAR model is trained to maximize acoustic tokens from that corresponding quantizer codebook. Acoustic tokens from the first phase to the sampling phase are embedded and summed as model input. Phoneme sequences are also considered as cues for the language model. Furthermore, to clone the unique voice of a given speaker, the system also uses acoustic tokens from registered speech as acoustic cues.

[0083] Specifically, the system first tokenizes the registered speech using a neural codec model. Embedsion representations from all different codebooks are summed as acoustic cues. To predict acoustic tokens from the codebook corresponding to the sampling training phase, the transducer input is a concatenation of the inputs. Positional embeddings are also computed separately for the cues and the acoustic sequence. The current phase is injected into the network, for example using an adaptive layer normalization operator. Unlike AR, the NAR model allows each token to attend to all input tokens in a self-attention layer, as illustrated in matrix 412. The system shares parameters between the acoustic embedding layer and the output prediction layer, meaning that the weights of a particular prediction layer are the same as the weights of subsequent prediction layers.

[0084] Some of the technical benefits of the disclosed embodiments include the ability of the language model to predict labels for unseen input without additional parameter updates. In other words, the model can synthesize high-quality speech for unseen speakers without fine-tuning (i.e., performing contextual learning). This is an improvement over conventional systems and methods for personalized text-to-speech generation, which require additional fine-tuning or experience significant quality degradation for unseen speakers.

[0085] For the language model, the cue enables context learning in the zero-shot scenario described in this paper. In summary, the system converts text into phoneme sequences and encodes registration records into acoustic matrices. This forms phoneme cues and acoustic cues, respectively. Both types of cues are used in both AR and NAR models. For example, in the AR model, the system uses sample-based decoding conditioned on the cue. This is an improvement over beam search techniques, which can cause language models to enter infinite loops, never achieving a predicted output. Furthermore, the sample-based approach significantly increases the diversity of the output. For the NAR model, the system uses greedy decoding to select the token with the highest probability. Finally, the system uses a neural codec decoder to generate waveforms conditioned on different code sequences. The acoustic cue can be semantically related to the synthesized speech or not.

[0086] For example, in some embodiments, when the acoustic cue and the generated speech are semantically unrelated, the machine learning model is configured to generate content without seeing the speaker. The model is given a text sentence, a registered speech, and its corresponding transcription. The system adds the transcribed phonemes of the registered speech to the beginning of the phoneme sequence of the given sentence as a phoneme cue, and uses the first-level acoustic token of the registered speech as an acoustic prefix. Using the phoneme cue and the acoustic prefix, the model generates acoustic tokens for the given text using a cloned voice of the speaker.

[0087] Furthermore, or alternatively, in some embodiments, when the acoustic cue and the generated speech are semantically related, the system uses the first few seconds (e.g., 3 seconds) of the complete transcription and spoken utterance as the phoneme cue and acoustic cue, respectively. The system then instructs the model to generate continuous speech for the remaining transcription. The inference process is the same as in the previous embodiments, except that the registered speech and the generated speech are semantically continuous.

[0088] Cross-language personalized text-to-speech generation

[0089] End-to-end TTS synthesis has made progress over the years. However, traditional models cannot generate high-quality speech for cross-language applications. The quality of traditional models degrades due to the inherent scarcity of data (speakers typically only speak the source language, making it difficult to obtain source language pairs from the same speaker in the training data), and because of the limited capacity of the models (e.g., traditional TTS models are not powerful enough to transfer speaker voice, speech context, and speaker emotion from the source language speech to the target language speech).

[0090] Traditional solutions to these problems include adding specific subnetworks of the speaker network for different speakers, and adding additional language networks for language control. However, even with multiple encoders introduced for each language, these models still suffer from a loss in preserving speaker identity, even if the models are explicitly trained using speech data from the target speaker. This quality degradation is further exacerbated in zero-shot scenarios. When such models attempt to synthesize target speech from a source speaker they have never seen before, they encounter problems with low speaker similarity and foreign language accents.

[0091] As an improvement to such models, the disclosed embodiments also relate to systems and methods for cross-lingual personalized speech synthesis, which transfer a speaker's voice from one language to another. To further improve upon the aforementioned shortcomings of traditional TTS models, the following disclosed embodiments relate to systems and methods for utilizing cross-lingual neural codec language models, where robust contextual learning capabilities provide a solution for achieving zero-shot cross-lingual speech synthesis. Since the cross-lingual neural codec language model is initially trained on thousands of hours of speech training data, it is able to accurately transfer the target speaker's speech features (e.g., voice, prosody, emotion, speaking context) into the synthesized cross-lingual speech, while also mitigating accent mismatch between the source and target languages.

[0092] Cross-language TTS model training

[0093] For example, systems and methods for training cross-lingual neural codec language models are provided. For instance, the system acquires multilingual transcription data by directly using ASR data or by identifying large amounts of unlabeled speech data as pseudo-transcriptions with an offline ASR model. The system then uses a rule-based transducer to convert the transcription into phoneme sequences and an offline neural codec encoder to convert the speech data into acoustic tokens. Utilizing the multilingual acoustic tokens and phoneme sequences, the system trains a multilingual neural codec language model that predicts acoustic codec sequences of target language speech from target language text, using source language speech as cues. The predicted acoustic token sequences are then converted into the final speech of the target language using an offline audio codec decoder.

[0094] The cross-lingual neural codec language model is trained using two large-scale datasets, totaling thousands of hours (e.g., over 70,000 hours) of speech data. The first large-scale dataset includes thousands of hours (e.g., over 60,000 hours) of unlabeled speech data, such as audiobooks in a first language (e.g., English). The second large-scale dataset includes thousands of hours (e.g., over 10,000 hours) of multi-domain, multi-speaker speech data, such as ASR data in a second language. In some cases, the second dataset contains less speech data than the first dataset—for example, half, a quarter, an eighth, or less. The combination of these two datasets constitutes a large-scale, multi-lingual, multi-speaker, multi-domain, non-clean speech dataset, which significantly improves coverage of different speakers. Furthermore, it enhances the generalization ability of the cross-lingual neural codec language model.

[0095] Regarding the foregoing, it should be understood that while specific examples of large training dataset sizes have been described above, the datasets cited can actually contain more training data than specified. In particular, any number of hours can be used to train a machine learning model, including but not limited to datasets containing more than 70,000 hours, 100,000 hours, or even 200,000 hours of training data. Training datasets of less than 70,000 hours can also be used. However, there is an inverse relationship between the size of the training dataset required to train the model and the size of the runtime samples required for the personalized model. For example, the larger the training dataset used, the smaller the runtime samples required for the personalized model. During testing, it was found that a training dataset of approximately 70,000 hours could support samples of approximately 3 seconds. If the training dataset is larger than 70,000 hours, then samples can sometimes be smaller than 3 seconds while still obtaining the desired results. If the training dataset is smaller than 70,000 hours, then samples are best longer than 3 seconds.

[0096] Based on experimental results, the cross-lingual neural codec language model achieves higher speaker similarity scores than previous models for tasks where the speaker is not present. Through large-scale training (i.e., thousands of hours, rather than hundreds of hours as in traditional TTS models), the published cross-lingual model significantly reduces the word error rate from 8.53 to 40.7 in the cross-lingual English TTS task, achieves a 3.17 BLEU score improvement compared to the baseline model in the S2S translation task, and achieves better speech naturalness. Furthermore, human evaluation shows that the published cross-lingual model outperforms the robust baseline model in both SMOS (4.00 vs. 2.88 in cross-lingual TTS, 4.12 vs. 3.06 in S2ST) and MOS (3.87 vs. 3.81 in S2ST).

[0097] Now let's turn our attention to... Figure 5 It shows an example diagram of a machine learning model (e.g., model 500), including inputs and outputs configured to perform cross-language personalized text-to-speech generation. For example, in some cases, model 200 is modified / extended to model 500, such as... Figure 5 As shown, it now includes a multilingual conditional codec language model to predict acoustic codec sequences of target language speech using acoustic cues containing source language speech and text cues containing both source and target language text. It should also be understood that Model 500 represents... Figure 1 Cross-language model 147.

[0098] like Figure 5 As shown, multiple text prompts (e.g., from...) Figure 1 The text prompts 142 include: language 1 text (e.g., source text prompt 502) and language 2 text (e.g., target text prompt 504) and acoustic prompts 506 (e.g., language 1 speech) (e.g., from...). Figure 1 The acoustic cue 143, along with the language ID 508, is provided as input to model 500.

[0099] Acoustic cues 506 (e.g., Language 1 speech), also known as registration records or target speaker sample speech, contain a limited amount of audio data from an unseen target speaker. In some cases, acoustic cues may include a 3-second registration record or audio data containing a single spoken utterance.

[0100] Language IDs are used to guide speech generation for a specific language. Without language IDs, the model may become confused when selecting appropriate acoustic tokens for speech in a particular language, as it is trained on multilingual data and the input text is converted into phonemes. Furthermore, some languages ​​have very different characteristics. For example, Chinese is a tonal language, while English is a non-tonal language, which increases the difficulty of speech synthesis.

[0101] By adding language IDs to the input of a cross-language neural codec model, the model is guided to generate speech with correct speaking features and mitigate foreign language accent issues. In other words, for example, even if the target speaker's acoustic cues are English, and the generated target speech is Chinese, the generated target speech will simultaneously possess (i) the target speaker's voice based on acoustic tokens derived from the target speaker's acoustic cues, and (ii) a Chinese accent guided by the language ID, even if the acoustic cues have an English accent. It should be understood that... Figure 5 The language ID can also be configured as an attribute ID, such as an emotion ID, a speaking style ID, or other attribute IDs, to modify personalized speech in monolingual or cross-lingual speech synthesis.

[0102] The text prompt is converted into a phoneme sequence using a multilingual G2P model (e.g., multilingual G2P 510). For example, language 1 text is converted into language 1 phonemes, and language 2 text is converted into language 2 phonemes. The acoustic prompt 506 is converted into a source acoustic token (e.g., a language 1 acoustic token) using an audio codec encoder 512, which in some cases represents... Figure 4 Audio codec encoder.

[0103] In some alternative embodiments, the system is configured to use a text tokenizer instead of a phoneme tokenizer (e.g., a G2P model) to convert text cues into word tokens. These word tokens, along with the source acoustic tokens, are provided as input to subsequent machine learning model layers to generate the final target acoustic token. Other types of tokenizers configured to convert characters or text strings into tokens can also be used in a similar process.

[0104] Back Figure 5 Based on phoneme sequences and source acoustic tokens, model 500 performs neural codec language modeling 514 and generates target acoustic tokens 516 (e.g., language-2 acoustic tokens). The target acoustic tokens 516 are provided as input to audio codec decoder 518, which converts the target acoustic tokens 516 into personalized speech waveforms 520 (e.g., personalized language-2 speech). It should be understood that audio codec encoder 512 and audio codec decoder 518 are part of the same audio codec model referenced herein.

[0105] Machine Learning Model 500 achieves many of the same technical advantages as Machine Learning Model 200, including additional cross-lingual advantages. For example, Model 500 inherits strong context learning capabilities, enabling zero-shot cross-lingual TTS synthesis. It can even be used for zero-shot speech-to-speech (STS) translation tasks. Model 500 uses a limited set of speech from the source language (e.g., only one utterance) as cues to generate high-quality speech in the target language while preserving the speaker's voice, emotion, and acoustic environment, even without seeing the actual speaker. Furthermore, Machine Learning Model 500 mitigates the foreign language accent problem that often occurs in cross-lingual speech generation. For example, in some cases, while phonemes may be correctly generated for the new language in the speaker's voice, the speaker's source language accent may be used instead of the correct target language accent. This problem is mitigated by using language IDs to control speech synthesis.

[0106] Furthermore, it should be noted that the disclosed embodiments for cross-language speech generation do not require cross-language speech data from the same speaker for model training. This significantly reduces training data costs and training time compared to traditional systems that require target language acoustic cues and source language acoustic cues from the target speaker to generate cross-language speech.

[0107] Using phoneme sequences derived from source and target texts, and acoustic tokens (generated by a neural audio codec encoder based on acoustic cues) as cues, Model 500 can generate acoustic tokens for the target language, which are then converted into target speech by the corresponding audio codec decoder. Model 500 is suitable for various cross-lingual speech generation tasks, such as cross-lingual TTS generation and STS translation.

[0108] Compared to traditional TTS models that treat TTS as a continuous regression task with Mel spectrograms as intermediate products, this model treats TTS as a conditional language modeling task with audio codec codes as intermediate representations. For example, in... Figure 6 As described in more detail, the cross-lingual neural codec language model employs a two-stage modeling approach. The model first uses an AR language model to generate acoustic tokens for the first layer from phoneme sequences, and then uses a NAR transducer model to generate codec codes for subsequent layers. To learn cross-lingual acoustic transduction information for cross-lingual TTS and S2S translation tasks, the system utilizes a bilingual speech-transcription corpus to train the multilingual AR and NAR codec models. After training on a first large-scale speech-transcription dataset, the cross-lingual neural codec language model acquires strong context learning capabilities, generating personalized speech using limited speech segments (e.g., 3-second registration records) as acoustic cues.

[0109] Cross-language neural codec language sub-model

[0110] Now let's turn our attention to... Figure 6 It illustrates an example diagram of a cross-lingual neural codec language model (e.g., model 600), which includes a multilingual AR codec language model (multilingual AR model 602) and a multilingual NAR codec language model (multilingual NAR model 604) to generate acoustic tokens at different granularities. Figure 6 The training data, as illustrated, comprises multilingual speech-transcription pairs 606. On one hand, the transcription is phonemicized using a G2P model to generate a phoneme token set 608 (e.g., S = {HH, AH, L, OW, ..., D}). On the other hand, the speech data is quantized using an audio codec encoder to generate a superset of quantization tokens (e.g., token set 610), which includes the quantization token set for each quantization layer. For example, the token set for layer 1 is A... :,1 ={731,284,78,32,...,669}. This token set is generated by the multilingual AR model 602. Subsequent layers' token sets include token set A. :,2:l Layer 2 is associated with tokens {325,71,435,90,...,7}, and layer 1 is associated with tokens {12,504,32,8,...,743}. (Layers 3 through 1-1 are not shown). These token layers after layer 1 are generated by the multilingual NAR model 604.

[0111] As mentioned above, the audio codec model is used as an acoustic tokenizer; it is an encoder-decoder model concatenated with multiple quantization layers, as shown in the reference. Figure 3 In more detail, each quantization layer generates a quantization token, which includes multiple entries for a specific frequency of the input waveform. For example... Figure 6 The illustrated bit rate setting is configured to use 8 quantization layers (i.e., quantizers), where the quantization tokens comprise 1024 entries at 75Hz. However, it should be understood that the model can be configured for different bit rate settings, resulting in different numbers of quantization layers.

[0112] Multilingual AR Model

[0113] The multilingual AR codec model is a one-way converter decoder that generates acoustic tokens autoregressively based on semantic tokens. To improve the storage memory of the computing system, the multilingual AR codec model is used to predict the acoustic tokens of the first layer. This advantageously prevents the AR model from predicting acoustic tokens of multiple layers simultaneously, which would otherwise result in sequences that are too long to train and infer the model effectively.

[0114] like Figure 6 As shown, S represents the transcribed phoneme sequence, and A represents the first-layer acoustic token extracted from the corresponding speech X. As an autoregressive decoder, the model is trained to predict the first-layer A token-by-token, given its prefixes. This is optimized by maximizing the log-likelihood of the speech-transcriptional cue data.

[0115] Multilingual NAR Model

[0116] The multilingual NAR model does not use an autoregressive generative model, but rather a non-autoregressive transducer language model configured to use phoneme sequences (e.g., the phoneme sequence "S") and acoustic tokens from the preceding sentence (e.g., ...). As a cue, the remaining acoustic tokens are generated iteratively. Here, the previous sentence is expected to have the same acoustic features as the current sentence (speaker identity, speed, background, etc.) and is used to provide additional reference information for cloning the target voice. Similar to machine learning model 200, at each layer of the cross-language extension, the embeddings of the quantization layers after the first quantization layer are summed layer by layer as input.

[0117] Cross-language reasoning

[0118] Now let's turn our attention to... Figure 7 It shows an example diagram of a cross-lingual neural codec language model (e.g., model 600) applicable to zero-shot cross-lingual TTS and zero-shot S2ST. As previously described, for zero-shot cross-lingual TTS, the source text 702 and the target text 704 are converted into phoneme sequences S using the G2P tool 706, respectively. s and S t The source speech (i.e., acoustic cue 708) is converted into source acoustic token A using audio codec encoder 710. s .

[0119] For example, after training, the cross-language neural codec language model can perform cross-language speech synthesis inference. The model first sets the source phoneme S... s and target phoneme S t Connect them as input, and use the first layer acoustic token A s:,1 As a decoding prefix for the multilingual AR model 602, to generate the first-layer target acoustic token A t:,1 As mentioned earlier, Model 600 adds source language embeddings and target language embeddings to S. s S t A s:,1 and A t<i,1 Each token is embedded to control the intonation of the final generated speech. After obtaining the first layer of target acoustic tokens from the multilingual AR model 602, the multilingual NAR model 604 is used to predict the acoustic tokens of the remaining layers (e.g., {, A}) through a greedy search (i.e., selecting the token with the highest probability). t:,l |l=2,...,8). Finally, the audio codec decoder is used to synthesize the target speech from the complete target acoustic token set.

[0120] Cross-language speech-to-speech translation

[0121] like Figure 7 The illustrated cross-lingual neural codec language model can be applied to zero-shot cross-lingual TTS and zero-shot S2ST. Zero-shot S2ST is achieved by using an additional speech recognition and translation model, which is responsible for recognizing the source speech and translating it into a source phoneme sequence and a target phoneme sequence.

[0122] In some cases, additional speech recognition and translation models are unified modal speech-to-text pre-trained frameworks that use hidden units at the modal bridge between speech and text. This supports a variety of speech-to-text tasks, including ASR and speech-to-text translation.

[0123] In some cases, the intermediate hidden units are replaced with phonemes, enabling additional speech recognition and translation models to predict both source and target phonemes simultaneously. Specifically, this additional model includes a speech encoder 712, a semantic encoder 714, and a semantic decoder 716. All of these components are pre-trained on an ASR corpus containing source speech and source phonemes, as well as a machine translation (MT) corpus containing source and target phonemes, where the phoneme sequences are derived from text.

[0124] After pre-training, the system uses triplet data—including source speech data, source phoneme data, and target phoneme data—to fine-tune components. The system performs multi-task learning to transcribe source phonemes using CTC loss on the semantic encoder and multi-task learning to translate target phonemes using cross-entropy loss on the semantic decoder.

[0125] Figure 7 The reasoning process for speech-to-speech translation is illustrated. For example, given a source speech sample (e.g., acoustic cue 718), the speech recognition and translation model (ASR model) first generates the source phoneme S from the semantic encoder 714. s The system then generates the target phoneme St from the semantic decoder 716. Next, the system uses the audio codec encoder 710 to compress the source speech into the source acoustic token A. s Then, the system will use the source phoneme S s Target phoneme S t Heyuan Acoustics Token A s These are concatenated and used as input to the model to generate a codec sequence for the target speech. The generated codec tokens are then converted into the final target speech by the decoder of the audio codec model.

[0126] Example Method

[0127] Now let's turn our attention to... Figure 8 The example illustrates a flowchart with multiple actions (e.g., actions 810, 820, 830, 840, 850, and 860) that are associated with a method implemented by a computing system (e.g., computing system 110) using a zero-shot personalized text-to-speech model (e.g., zero-shot model 146).

[0128] The first action illustrated includes the action of the system accessing a machine learning model (e.g., a cross-language model 147) configured as a zero-shot cross-language text-to-speech model, which has been previously trained on a text-to-speech training dataset (e.g., training data 141) that includes multiple different bilingual speech transcription pairs.

[0129] The system also acquires a first text prompt in a first language (e.g., text prompt 142) (action 820), a second text prompt in a second language (e.g., text prompt 142) (action 830), and a speech sample including audio data from the target speaker (e.g., acoustic prompt 143), wherein the target speaker is an unseen target speaker, and therefore no audio data from that target speaker is included in the text-to-speech training dataset (action 840).

[0130] The system provides or applies a first text prompt in the first language, a second text prompt in the second language, and a speech sample from the target speaker as input to a machine learning model (action 850), and uses the model to generate personalized speech output (e.g., synthesized speech 145) based on these inputs and at least by converting the second text prompt in the second language into the synthesized voice of the target speaker based on the speech sample from the target speaker (action 860).

[0131] In some embodiments, the system also obtains a language ID (e.g., attribute ID144) associated with a second language, which is configured to guide the generation of personalized speech output using the second language based on accents or other speaking styles / attributes associated with the second language.

[0132] Understandably, speech samples include audio data obtained from the target speaker, which may consist of a single spoken utterance, or alternatively, a limited-duration recording from the target speaker (e.g., 3-second, 4-second, or 5-second segments extracted from a single spoken utterance). The duration of the samples can be predetermined based on the known training quality of the model.

[0133] As in Figure 2-7 As described above, by providing the aforementioned inputs to the machine learning model, the system is able to convert text cues into phoneme sequences and speech samples into discrete acoustic tokens. For example, the system converts a first text cue into a first phoneme sequence, a second text cue into a second phoneme sequence, and speech samples from a target speaker into a set of source acoustic tokens. By converting speech samples or acoustic cues into source tokens, the system can leverage the contextual learning capabilities of the machine learning model and avoid processing the input using continuous spectrum methods. This configuration and initial training can help reduce the amount of new sample data required for personalized models.

[0134] As described, the system uses a transformed input to generate an acoustic token set for the target language based on a first phoneme sequence, a second phoneme sequence, and a source acoustic token set, such that personalized speech output is generated based on the target language acoustic token set. In this case, the personalized speech output is generated by applying the target language acoustic token set to the audio codec decoder. By implementing the system in this way, the audio codec decoder can directly reconstruct the target acoustic tokens while maintaining high speaker similarity, emotion, and acoustic context of the speech samples.

[0135] As referenced above Figure 2-7 (in particular Figure 3 In more detail, the machine learning model is configured as a neural codec model that generates acoustic tokens at various quantization layers according to different granularities. For example, the machine learning model includes a multilingual autoregressive codec language model and a multilingual non-autoregressive codec language model. The multilingual autoregressive codec language model generates acoustic tokens at a first quantization layer, and the multilingual non-autoregressive codec language model generates acoustic tokens at multiple subsequent quantization layers based on the acoustic tokens generated at the first quantization layer. By utilizing both AR and NAR models, the system achieves a beneficial balance between accuracy and time reduction in processing different token sets.

[0136] By configuring the AR and NAR models in a hierarchical manner, tokens from previous quantization layers advantageously recover coarse acoustic properties, while subsequent quantization layers advantageously learn fine acoustic properties.

[0137] Now let's turn our attention to... Figure 9 , Figure 9 Flowchart 900 is explained. Flowchart 800 includes various actions (actions 910, 920, 930, 940, 950, and 960) associated with exemplary methods that can be implemented by computing system 110 to generate modified personalized speech using the aforementioned zero-sample personalized text-to-speech model and configured as a new target speaker.

[0138] The first action illustrated includes the action of the system accessing a machine learning model (e.g., a cross-language model 147) configured as a zero-shot cross-language text-to-speech model that has been previously trained on a text-to-speech training dataset (action 910).

[0139] The next action involves the system acquiring a text prompt (e.g., text prompt 142) (action 920) and a speech sample containing audio data from the target speaker (e.g., acoustic prompt 143), where the target speaker is an unseen target speaker and therefore no audio data from the target speaker is included in the text-to-speech training dataset (action 930).

[0140] The system also accesses attribute IDs (e.g., attribute IDs) that are configured to modify personalized speech output based on specific attributes (Action 940). For example, the attribute ID can be identified and accessed based on user input that specifies a particular speaking accent, language, and / or other speaking style to be used, which is mapped to the attribute ID and attribute ID features used by the model.

[0141] Subsequently, the system applies the text prompt, speech sample, and attribute ID to the machine learning model (action 950), and generates personalized speech output (e.g., synthetic speech 145) based on the text prompt using the synthesized voice of the target speaker using the speech sample, which is modified according to the attribute ID (action 960).

[0142] Understandably, the attribute ID can be configured based on different attributes, such as different languages, different emotions, and / or different speaking styles. For example, in some cases, the attribute ID is a language ID corresponding to a specific language, causing the synthesized voice of the target speaker to be modified to include an accent associated with that specific language. Additionally, or alternatively, in some cases, the attribute ID is an emotion ID corresponding to a specific emotion, causing the synthesized voice of the target speaker to be modified to convey that specific emotion. Additionally, or alternatively, the attribute ID is a speaking style ID corresponding to a specific speaking style, causing the synthesized voice of the target speaker to be modified according to that specific speaking style.

[0143] Now let's turn our attention to... Figure 10 The flowchart 1000 is illustrated, which includes various actions (actions 1010, 1020, 1030, 1040, 1050, and 1060) associated with exemplary methods that may be implemented by computing system 110.

[0144] The illustrated first action includes the action of the system accessing a machine learning model (e.g., a cross-language model 147) configured as a zero-shot cross-language text-to-speech model that has previously been trained on a text-to-speech training dataset (e.g., training data 141) that includes multiple different bilingual speech transcription pairs (action 1010).

[0145] The system also acquires speech samples containing audio data from the target speaker (e.g., acoustic cue 143), where the target speaker is an unseen target speaker, and therefore no audio data from the target speaker is included in the text-to-speech training dataset (action 1020).

[0146] Then, the system generates a first transcription of a speech sample in the first language (action 1030), and based on translating the first transcription of the speech sample into the second language, generates a second transcription of a speech sample in the second language (action 1040).

[0147] Subsequently, the system applies the first transcription of the speech sample, the second transcription of the speech sample, and the speech sample to the machine learning model (action 1050), thereby generating personalized speech output (e.g., synthesized speech 145) based on the second transcription of the speech sample using the second language (action 1060) using the synthesized voice of the target speaker based on the speech sample from the target speaker.

[0148] In some cases, the system also obtains a language ID associated with the second language (e.g., attribute ID 144), which is configured to guide the generation of personalized speech output in the second language based on the accent or other speaking style associated with it. This language ID can be mapped to and based on user input specifying the desired accent or other style for the second language.

[0149] In light of the foregoing, it will be appreciated that the disclosed embodiments offer numerous technical advantages over conventional systems and methods for generating personalized voices for new target speakers using personalized text-to-speech models with zero-shot learning. By implementing the disclosed embodiments in this manner, numerous technical advantages relative to existing systems are achieved. For example, the disclosed embodiments provide a TTS framework with robust context learning capabilities, replacing traditional Mel spectrograms with language modeling mediated by audio codec codes. The disclosed embodiments support a cue-based approach for zero-shot TTS that does not require additional structural engineering, pre-designed acoustic features, or fine-tuning as in previous systems.

[0150] The disclosed embodiments provide a general TTS system at the speaker dimension by utilizing thousands of hours of semi-supervised data, and are able to produce diverse outputs for the same input text while maintaining the context and speaker emotion of the acoustic cues. The generated synthesized speech exhibits high speaker similarity by cuing zero-sample scenarios. The disclosed embodiments also relate to cross-language TTS and S2ST, which achieve the aforementioned technical advantages, as well as cross-language specific advantages, such as high-quality cross-language speech synthesis, by maintaining speaker similarity (in terms of identity and emotion), providing accurate translation, and generating synthesized speech that sounds like natural speech.

[0151] Example computing system

[0152] Various embodiments of this disclosure may include or utilize a dedicated or general-purpose computer (e.g., computing system 110) comprising computer hardware and software. Embodiments within the scope of this disclosure include, for example, physical storage media and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. These computer-readable media may be any available media accessible to general-purpose or dedicated computer systems.

[0153] Stores computer-executable instructions (e.g., Figure 1 The computer-readable medium (e.g., computer-readable instructions 118) Figure 1 The (various) hardware storage devices 140) are physical hardware storage media or devices excluding transmission media. Physical computer storage media / devices are hardware and include RAM, ROM, EEPROM, CD-ROM or other optical disc storage (such as CD, DVD, etc.), disk storage or other magnetic storage devices, or any other hardware that can be used to store program code in the form of computer-executable instructions or data structures and can be accessed by a general-purpose or special-purpose computer.

[0154] A computer-readable medium that carries computer-executable instructions or computer-readable instructions (e.g., computer-readable instruction 118) in one or more carrier waves or signals is a transmission medium. A transmission medium may include a network and / or data link that can be used to carry desired program code in the form of computer-executable instructions or data structures and is accessible to a general-purpose or special-purpose computer.

[0155] Therefore, by way of example and not limitation, various embodiments of this disclosure may include at least two completely different types of computer-readable media: physical computer-readable storage media / devices and transmission computer-readable media.

[0156] The aforementioned combinations also fall within the scope of computer-readable media, particularly when computer-executable instructions or data structures can be automatically transferred from a transmission computer-readable medium to a physical computer-readable storage medium (or vice versa). For example, computer-executable instructions or data structures received via a network or data link may be cached in RAM within a network interface module (e.g., a "NIC") and then ultimately transferred to the computer system RAM and / or a less lossy computer-readable physical storage medium at the computer system. Therefore, computer-readable physical storage media may be included in computer system components that also (or even primarily) utilize transmission media.

[0157] The computer-executable instructions referred to herein are instructions and data that cause a general-purpose computer, special-purpose computer, or special-purpose processing device to perform a specific function or set of functions (such as the functions disclosed above). Computer-executable instructions may be constructed in the form of binary code, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in language specific to structural features and / or methodological actions, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the features or actions described above. Rather, the features and actions described above are disclosed as exemplary forms of implementing the claims.

[0158] Those skilled in the art will understand that this disclosure can be practiced in networked computing environments with many types of computer system configurations, including personal computers, desktop computers, laptop computers, message processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics devices, network PCs, minicomputers, mainframes, mobile phones, PDAs, pagers, routers, switches, and so on. This disclosure can also be implemented in distributed system environments in which both local and remote computer systems perform tasks via network links (or via hardwired data links, wireless data links, or a combination of hardwired and wireless data links). In a distributed system environment, program modules can reside on both local and remote memory storage devices.

[0159] Alternatively or additionally, the functionality described herein may be performed at least in part by one or more hardware logic components. For example, but not limited to, illustrative types of hardware logic components that may be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0160] This disclosure may be embodied in other specific forms without departing from its essential characteristics. The described embodiments should be considered in all respects as illustrative rather than restrictive. Thus, the scope of the invention is indicated by the appended claims rather than the foregoing description. All modifications falling within the meaning and scope of the equivalents of the claims are covered by the scope of the claims.< / eos> < / eos>

Claims

1. A method for generating cross-lingual personalized speech, the method comprising: Access a machine learning model configured as a zero-sample cross-lingual text-to-speech model, which has previously been trained on a text-to-speech training dataset that includes multiple different bilingual speech transcription pairs. Get the first text prompt written in the first language; Get a second text prompt in a second language; Acquire speech samples containing audio data from a target speaker, wherein the target speaker is an unseen target speaker and therefore no audio data from the target speaker is included in the text-to-speech training dataset; The first text prompt in the first language, the second text prompt in the second language, and the speech sample from the target speaker are provided as input to the machine learning model. as well as Personalized speech output is generated based on these inputs and at least by using a second text prompt in a second language based on a synthesized voice conversion of the target speaker from a speech sample of the target speaker.

2. The method as described in claim 1, characterized in that, Further includes: Obtain a language ID associated with a second language, which is configured to guide the generation of personalized speech output using the second language based on the accent associated with the second language.

3. The method as described in claim 1, characterized in that, The speech sample includes audio data acquired from the target speaker, comprising approximately one spoken utterance.

4. The method as described in claim 1, characterized in that, Further includes: Convert the first text prompt into a first phoneme sequence; Convert the second text prompt into a second phoneme sequence; as well as The speech samples from the target speaker are converted into a source acoustic token set.

5. The method as described in claim 4, characterized in that, Further includes: An acoustic token set for the target language is generated based on the first phoneme sequence, the second phoneme sequence, and the source acoustic token set, such that the personalized speech output is generated based on the acoustic token set using the target language.

6. The method as described in claim 5, characterized in that, The personalized voice output is generated by applying an acoustic token set using the target language to the audio codec decoder.

7. The method as described in claim 1, characterized in that, The machine learning model is configured as a neural codec model, which generates acoustic tokens at each quantization layer according to different granularities.

8. The method as described in claim 7, characterized in that, The machine learning model includes a multilingual autoregressive codec language model and a multilingual non-autoregressive codec language model. The multilingual autoregressive codec language model generates acoustic tokens at a first quantization layer, and the multilingual non-autoregressive codec language model generates acoustic tokens at multiple subsequent quantization layers based on the acoustic tokens generated at the first quantization layer.

9. The method as described in claim 7, characterized in that, Tokens from previous quantization layers recover coarse acoustic properties, while subsequent quantization layers learn fine acoustic properties.

10. A method for generating modified personalized speech, the method comprising: Access a machine learning model configured as a zero-shot cross-language text-to-speech model, which has previously been trained on a text-to-speech training dataset; Get text hints; Acquire speech samples containing audio data from a target speaker, wherein the target speaker is an unseen target speaker and therefore no audio data from the target speaker is included in the text-to-speech training dataset; Access is configured to modify the attribute ID of personalized voice output based on specific attributes; The text prompts, the voice samples, and the attribute IDs are applied to the machine learning model; as well as Personalized speech output is generated based on the text prompt using the synthesized voice of the target speaker from the speech sample, the synthesized voice of the target speaker being modified according to the attribute ID.

11. The method as described in claim 10, characterized in that, The attribute ID is a language ID corresponding to a specific language, which causes the synthesized voice of the target speaker to be modified to include an accent associated with the specific language.

12. The method as described in claim 10, characterized in that, The attribute ID is an emotion ID corresponding to a specific emotion, which modifies the synthesized voice of the target speaker to convey the specific emotion.

13. The method as described in claim 10, characterized in that, The attribute ID is a speaking style ID corresponding to a specific speaking style, which allows the synthesized voice of the target speaker to be modified according to the specific speaking style.

14. A method for generating personalized speech output, the method comprising: Access a machine learning model configured as a zero-shot cross-lingual speech-to-speech model, which has previously been trained on a text-to-speech training dataset comprising multiple different bilingual speech transcription pairs. Acquire speech samples containing audio data from a target speaker, wherein the target speaker is an unseen target speaker and therefore no audio data from the target speaker is included in the text-to-speech training dataset; Generate a first transcription of the speech sample using the first language; Based on translating the first transcription of the speech sample into a second language, a second transcription of the speech sample using the second language is generated; The first transcription of the speech sample, the second transcription of the speech sample, and the speech sample are applied to the machine learning model; as well as Personalized speech output is generated using a synthesized voice of the target speaker based on a speech sample from the target speaker, and a second transcription of the speech sample in a second language.

15. The method as described in claim 1, characterized in that, Further includes: Obtain a language ID associated with a second language, which is configured to guide the generation of personalized speech output using the second language based on the accent associated with the second language.