Residual adapter for less sample text-to-speech speaker adaptation
By using the residual adapter stack to enhance the pre-trained model in the text-to-speech model, the problem of high model training cost when the few-sample speakers are adapted is solved, and the improvement of the adaptability of the multi-speaker and the reduction of the training cost is achieved.
Patent Information
- Application Number
- CN202380074492.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-25
- Filing Date
- 2023-10-24
- Publication Date
- 2025-06-03
AI Technical Summary
When the prior art realizes adaptive adaptation of small sample text to speech speakers, it faces the problem of high cost of model training and difficulty in expanding to thousands of speakers.
The residual adapter stack is used to enhance the pre-trained multispeaker TTS model. By optimizing the parameters of the residual adapter stack, we learn how to synthesize speech in the target speaker's voice, freezing the parameters of the backbone model to control the training cost.
The ability to adapt to multiple speakers without adding backbone model parameters is achieved, reducing service costs and avoiding the problems of overfitting and catastrophic forgetting.
Smart Images

Figure CN120092286A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a residual adapter for few-shot text-to-speech speaker adaptation. Background Art
[0002] A speech synthesis system uses a text-to-speech (TTS) model to generate speech from a text input. The generated / synthesized speech should accurately convey the message (intelligibility), while sounding like human speech (naturalness) with an expected prosody (expressiveness). While traditional speech synthesis models are able to provide intelligible speech, recent advances in speech neural modeling have significantly improved the naturalness and fidelity of synthesized speech. These neural models are large, having millions of parameters that can be fine-tuned during training. Due to the size of these models, training can be a long and computationally expensive process. Summary of the Invention
[0003] One aspect of the present disclosure provides a computer-implemented method for implementing a residual adapter for few-shot text-to-speech speaker adaptation. The computer-implemented method is executed by data processing hardware such that the data processing hardware performs operations including obtaining a text-to-speech (TTS) model configured to convert text into a representation of synthesized speech, the TTS model being pre-trained on an initial training dataset. The operations further include enhancing the TTS model with a stack of residual adapters. The operations include receiving an adaptation training dataset including one or more spoken utterances spoken by a target speaker, each spoken utterance in the adaptation training dataset being paired with a corresponding input text associated with a transcription of the spoken utterance. The operations include adapting the TTS model enhanced with the stack of residual adapters using the adaptation training dataset to learn how to synthesize speech in the voice of the target speaker while freezing the parameters of the TTS model by optimizing the stack of residual adapters.
[0004] Implementations of the present disclosure may include one or more of the following optional features. In some implementations, the TTS model includes a decoder that includes a stack of corresponding multi-head self-attention layers. In some implementations, enhancing the TTS model with a stack of residual adapters includes inserting the stack of residual adapters into the decoder, each respective residual adapter in the stack of residual adapters being disposed between each multi-head self-attention layer in the stack of corresponding multi-head self-attention layers of the decoder. In these implementations, the selected decoder may include a Conformer decoder that includes a first feed-forward module, a multi-head self-attention module, a convolutional module, a second feed-forward module, and a layer normalization module.
[0005] In some implementations, each verbal utterance in the adaptive training dataset includes a corresponding non-synthetic speech representation of the target speaker who uttered the verbal utterance. In these implementations, the residual adapter stack is optimized by: for each verbal utterance in the adaptive training dataset, using the TTS model enhanced by the residual adapter stack to generate a corresponding synthetic speech representation from the corresponding input text paired with the verbal utterance; determining a loss based on the non-synthetic speech representation of the target speaker who uttered the verbal utterance and the synthetic speech representation generated from the input text; and updating the parameters of the residual adapter stack based on the loss. In these implementations, the residual adapter stack can be further optimized by optimizing the speaker embedding associated with the target speaker.
[0006] The initial training dataset for pre-training the TTS model may not include any utterances spoken by the target speaker. The input text paired with each corresponding verbal utterance in the adaptive training dataset may include a corresponding phoneme sequence and a corresponding grapheme sequence. In some implementations, each respective residual adapter in the residual adapter stack includes a layer normalization layer, a down-projection layer, a rectified linear unit (ReLU) layer, an up-projection layer, and a residual connection.
[0007] In some implementations, the TTS model includes a bidirectional encoder representation from Transformer (BERT) encoder, which is configured to receive the corresponding input text of each verbal utterance in the adaptive training dataset and generate an encoded text representation for the corresponding input text utterance. In these implementations, the TTS model may further include a variance adapter, which is configured to receive a first concatenated output that includes a speaker embedding and an accent embedding concatenated with the encoded text representation generated by the BERT encoder, and generate an output corresponding to the first concatenated output. In these implementations, the TTS model may further include a duration adapter, which is configured to receive the first concatenated output and predict the phoneme duration of each phoneme in the corresponding input text based on the first concatenated output. Additionally, in these implementations, the TTS model may include an upsampler, which is configured to receive a second concatenated output that includes the first concatenated output concatenated with the output generated by the variance adapter, and upsample the second concatenated output. In these implementations, the TTS model may further include a decoder, which is configured to receive the upsampled second concatenated output and generate a synthetic speech representation for the corresponding input text.
[0008] The initial training dataset can pre-train the TTS model to synthesize speech in a first language. Here, each verbal utterance in the adaptive training dataset is spoken in a second language different from the first language.
[0009] Another aspect of the present disclosure provides a system for implementing a residual adapter for few-shot text-to-speech speaker adaptation. The system includes data processing hardware and memory hardware communicatively coupled to the data processing hardware. The memory hardware stores instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations. The operations include obtaining a text-to-speech (TTS) model configured to convert text into a representation of synthesized speech, the TTS model being pre-trained on an initial training dataset. The operations further include enhancing the TTS model via a stack of residual adapters. The operations include receiving an adaptive training dataset that includes one or more spoken utterances spoken by a target speaker, each spoken utterance in the adaptive training dataset being paired with corresponding input text associated with a transcription of the spoken utterance. The operations include adapting the TTS model enhanced via the stack of residual adapters using the adaptive training dataset to learn how to synthesize speech in the voice of the target speaker while freezing the parameters of the TTS model by optimizing the stack of residual adapters.
[0010] This aspect may include one or more of the following optional features. In some implementations, the TTS model includes a decoder that includes a corresponding stack of multi-head self-attention layers. In some implementations, enhancing the TTS model via a stack of residual adapters includes inserting the stack of residual adapters into the decoder, with each respective residual adapter in the stack of residual adapters being disposed between each multi-head self-attention layer in the corresponding stack of multi-head self-attention layers of the decoder. In these implementations, the decoder may include a Conformer decoder that includes a first feed-forward module, a multi-head self-attention module, a convolutional module, a second feed-forward module, and a layer normalization module.
[0011] In some implementations, each spoken utterance in the adaptive training dataset includes a corresponding non-synthesized speech representation of the target speaker who spoke the spoken utterance. In these implementations, the stack of residual adapters is optimized by, for each spoken utterance in the adaptive training dataset, using the TTS model enhanced via the stack of residual adapters to generate a corresponding synthesized speech representation from the corresponding input text paired with the spoken utterance; determining a loss based on the non-synthesized speech representation of the target speaker who spoke the spoken utterance and the synthesized speech representation generated from the input text; and updating the parameters of the stack of residual adapters based on the loss. In these implementations, the stack of residual adapters may be further optimized by optimizing a speaker embedding associated with the target speaker.
[0012] The initial training dataset for pre-training the TTS model may not include any utterances spoken by the target speaker. The input text paired with each corresponding oral utterance in the adaptive training dataset may include a corresponding phoneme sequence and a corresponding grapheme sequence. In some implementations, each corresponding residual adapter in the residual adapter stack includes a layer normalization layer, a down-projection layer, a rectified linear unit (ReLU) layer, an up-projection layer, and a residual connection.
[0013] In some implementations, the TTS model includes a Bidirectional Encoder Representations from Transformers (BERT) encoder configured to receive the corresponding input text for each oral utterance in the adaptive training dataset and generate an encoded text representation for the corresponding input text utterance. In these implementations, the TTS model may further include a variance adapter configured to receive a first concatenated output that includes a speaker embedding and an accent embedding concatenated with the encoded text representation generated by the BERT encoder and generate an output corresponding to the first concatenated output. In these implementations, the TTS model may further include a duration adapter configured to receive the first concatenated output and predict the phoneme duration of each phoneme in the corresponding input text based on the first concatenated output. Additionally, in these implementations, the TTS model may include an upsampler configured to receive a second concatenated output that includes the first concatenated output concatenated with the output generated by the variance adapter and upsample the second concatenated output. In these implementations, the TTS model may further include a decoder configured to receive the upsampled second concatenated output and generate a synthetic speech representation for the corresponding input text.
[0014] The initial training dataset may pre-train the TTS model to synthesize speech in a first language. Here, each oral utterance in the adaptive training dataset is spoken in a second language different from the first language.
[0015] Details of one or more implementations of the present disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a schematic diagram of an example system of a residual adapter for few-shot text-to-speech speaker adaptation.
[0017] Figure 2 is a schematic diagram of an example text-to-speech model configured with a residual adapter.
[0018] Figure 3Schematic diagram of an example Conformer decoder configured with a residual adapter.
[0019] Figure 4 Schematic diagram of an example residual adapter.
[0020] Figure 5 Schematic diagram of an example training process for a text-to-speech model configured with a residual adapter.
[0021] Figure 6 Flowchart of an example operational arrangement of a method for a residual adapter for few-shot text-to-speech speaker adaptation.
[0022] Figure 7 Schematic diagram of an example computing device that can be used to implement the systems and methods described herein.
[0023] In the various figures, like reference numerals indicate like elements. Detailed Description
[0024] Due to the growing interest in many different commercially customized voice applications (e.g., personalized voice assistants in low-resource scenarios), few-shot speaker adaptation remains an important research area in text-to-speech synthesis (TTS). Typical few-shot speaker adaptation methods involve adapting a pre-trained multi-speaker backbone acoustic model to a target speaker by fine-tuning all or part of the set of model parameters. However, it is difficult to scale such methods to thousands of speakers because the model service cost typically grows linearly with the number of target speakers supported. Additionally, in some cases, they require supplementary data (e.g., the training data used for pre-training the backbone model) during adaptation to avoid overfitting and / or catastrophic forgetting.
[0025] The implementations herein relate to few-shot speaker adaptation methods based on the parameter efficiency of residual adapters for text-to-speech. During adaptation, the pre-trained multi-speaker TTS backbone model is frozen but enhanced with a simple / lightweight neural module called a residual adapter, the parameters of which are optimized to synthesize the target speaker's voice, thereby allowing many different speakers to share the same backbone model for inference. The methods described herein are more flexible than the aforementioned fine-tuning methods because the size of the layers in the backbone model is fixed, but the size of the residual adapter can be adjusted accordingly based on the current task. Moreover, since the residual adapter is typically extremely small, adding 0.1% of the parameters to the backbone model, it can prevent the service cost from growing linearly with the number of target speakers. Additionally, since the current method freezes the backbone model during the speaker adaptation phase, the quality of the speakers included in the backbone model is not affected by the design.
[0026] Reference Figure 1 , in some implementations, the voice environment 100 includes the user 10 (also referred to herein as the target speaker 10) transmitting an oral utterance 12 to a voice-enabled device 110 (also referred to as the device 110 or the user device 110). The user 10 (i.e., the speaker of the utterance 12) may speak the utterance 12 as a query or command to solicit a response from the device 110 or to cause the device 110 to perform a task specified by the query. The device 110 is configured to capture sound from one or more users 10 within the voice environment 100. Here, the audio sound may refer to the oral utterance 12 of the user 10, which is used as an audible query, a command for the device 110, or an audible communication captured by the device 110. A voice-enabled system (e.g., a digital assistant interface) of or associated with the device 110 may process the query of the command by answering the query and / or causing the command to be executed.
[0027] Here, the device 110 captures audio data 14 corresponding to the oral utterance 12 of the user 10. The device 110 may correspond to any computing device associated with the user 10 and capable of receiving the audio data 14. Some examples of the user device 110 include, but are not limited to, mobile devices (e.g., mobile phones, tablets, laptops, e-book readers, etc.), computers, wearable devices (e.g., smartwatches), music players, broadcast devices, smart home appliances (e.g., smart TVs) and Internet of Things (IoT) devices, remote controls, smart speakers, etc. The device 110 includes data processing hardware 112 and a memory hardware 114 that communicates with the data processing hardware 112 and stores instructions that, when executed by the data processing hardware 112, cause the data processing hardware 112 to perform one or more operations related to voice and / or text processing. In some examples, the device 110 includes one or more applications (i.e., software applications), where each application may utilize one or more voice processing systems / models 140, 150, 200 associated with the device 110 to perform various functions within the application. For example, the device 110 includes an assistant application that is configured to transmit synthesized playback audio 154 (also referred to as synthesized voice 154) to the user 10 to talk to the user 10 and assist in performing various tasks.
[0028] The apparatus 110 further includes an audio subsystem having: an audio capture device (e.g., a microphone) 116 for capturing audio data 14 within the acoustic environment 100 and converting the audio data into an electrical signal; and a voice output device (e.g., a speaker) 118 for delivering an audible audio signal (e.g., a synthesized playback signal 154 from the apparatus 110). Although in the illustrated example, the apparatus 110 implements a single audio capture device 116, the apparatus 110 may implement an array of audio capture devices 116 without departing from the scope of the present disclosure, where one or more of the audio capture devices 116 in the array may not physically reside on the apparatus 110 but may communicate with the audio subsystem (e.g., a peripheral of the apparatus 110). For example, the apparatus 110 may correspond to an in-vehicle infotainment system that utilizes an array of microphones positioned throughout the vehicle. Similarly, the voice output device 118 may include one or more speakers that reside on the apparatus 110, communicate with the apparatus, or a combination thereof, where one or more speakers reside on the apparatus 110 and one or more other speakers are physically removed from the apparatus 110 but communicate with the apparatus 110.
[0029] In addition, device 110 is configured to communicate with remote system 130 via network 120. Remote system 130 may include remote resources 132, such as remote data processing hardware 134 (e.g., a remote server or CPU) and / or remote memory hardware 136 (e.g., a remote database or other storage hardware). Device 110 may utilize remote resources 132 to perform various functions related to voice processing and / or playback of synthesized communications. For example, device 110 is configured to perform speech recognition using speech recognition system 140 and / or convert text to speech using a TTS system 150 (e.g., using TTS model 200). These systems / models 140, 150, 200 may reside on device 110 (referred to as on-device systems), or remotely (e.g., reside on remote system 130) but communicate with device 110. In some examples, some of these systems 140, 150, 200 reside locally or on-device, while other systems reside remotely. In other words, any one of these systems 140, 150, 200 may be local, remote, or any combination of both. For example, when systems 140, 150, 200 are relatively large in size or processing requirements, systems 140, 150, 200 may reside in remote system 130. However, when device 110 can support the size or processing requirements of one or more systems 140, 150, 200, the one or more systems 140, 150, 200 may reside on device 110 using data processing hardware 112 and / or memory hardware 114. Optionally, one or more of systems 140, 150, 200 may reside both locally / on-device and remotely. For example, when the connection to network 120 between device 110 and remote system 130 is available, one or more of systems 140, 150, 200 may default to executing on remote system 130, but when the connection is lost or network 120 is unavailable, systems 140, 150, 200 instead execute locally on device 110. In some implementations, TTS model 200 may be a large trained model residing on a server (i.e., remote system 130) and further configured with a residual adapter 400 trained based on a target speaker (i.e., user 10 associated with device 110).
[0030] The speech recognition system 140 receives the audio data 14 as input and transcribes the audio signal into a transcription 142 as output. Generally, by converting the audio data 14 into a transcription 142, the speech recognition system 140 allows the device 110 to identify when the spoken words 12 of the user 10 correspond to a query, command, or some other form of audio communication. That is, the speech recognition system 140 may include a natural language understanding (NLU) function to perform query interpretation (e.g., semantic analysis) on the transcription 142. The transcription 142 refers to a text sequence that the device 110 can subsequently use to generate a response to a query or command. For example, if the user 10 asks the device 110 the question "what will the weather be like today", the device 110 passes the audio data 14 corresponding to the question "what will the weather be like today" to the speech recognition system 140. The speech recognition system 140 converts the audio data 14 into a transcription 142 that includes the text "what will the weather be like today?". Then, the device 110 can then use the text or a portion of the text to determine a response to the query. For example, to determine the weather for the day (i.e., today), the device 110 passes the text (e.g., "what will the weather be like today?") or an identified portion of the text (e.g., "weather" and "today") to a search engine. The search engine can then return one or more search results, and the device 110 interprets the one or more search results to generate a response for the user 10.
[0031] In some implementations, the device 110 or a system associated with the device 110 identifies text 152 (also referred to as a sequence of text 152 or input text 152) that the device 110 will transmit to the user 10 as a response to the query of the spoken utterance 12. The device 110 can then use the TTS system 150 to convert the text 152 into corresponding synthesized playback audio 154 for the device 110 to transmit to the user 10 as a response to the query of the spoken utterance 12 (e.g., audibly transmitted to the user 10). In other words, the TTS system 150 receives the text 152 as input and converts the text 152 into an output of synthesized playback audio 154 (e.g., through a series of neural networks), wherein the synthesized playback audio 154 is an audio signal that defines an audible reproduction of the text 152. For example, the playback audio 154 is a verbal expression or narration of the input text 152. In some examples, the input text 152 refers to a text or character sequence in a particular natural language (e.g., English, Spanish, or French). The character sequence may include letters, numbers, punctuation marks, and / or other special characters. When TTS system 150 generates playback audio 154 , playback audio 154 includes synthesized speech that approximates how a human would verbally express the sequence of characters defining input text 152 .
[0032] The TTS system 150 (or other speech synthesis system) includes a TTS model 200 (e.g., Figure 2 154 ). In some implementations, the TTS model 200 generates a synthesized speech representation 261 based on the input text 152, which is then further converted into the synthesized playback audio 154. In some implementations, the TTS model 200 processes the embedding as an encoded representation of speech features (e.g., features of the input text 152) to generate an audio waveform (e.g., a time-domain audio waveform defining the amplitude of an audio signal over time). Once generated, the TTS system 150 transmits the synthesized playback audio 154 to the device 110 to allow the device 110 to output the synthesized playback audio 154. For example, the device 110 audibly outputs the synthesized playback audio 154 "today is sunny" from one or more speakers 118. Here, the TTS model 200 of the TTS system 150 is configured to control the speech-related properties of the synthesized speech 154. In other words, the TTS model 200 is configured to simulate the voice of a human speaker in terms of naturalness, while also being able to generate different synthesized speech by modeling fine-grained latent features.
[0033] In some implementations, the TTS model 200 is enhanced by one or more residual adapters 400 (i.e., a stack of residual adapters 400). For example, the TTS model 200 may include a base / trunk model that is trained on a large user dataset of a large number of users. Then, the base model portion of the TTS model 200 can be frozen, and then one or more residual adapters 400 can be trained to learn how to synthesize speech in the voice of a target speaker (e.g., speaker 10 associated with the device 110). One or more residual adapters 400 may go through a training process ( Figure 5 ), where the residual adapter is optimized for the target speaker. The residual adapter 400 can be inserted at various points within the TTS model 200.
[0034] Although Figure 1 an example of the TTS system 150 is depicted in the context of an assistant application, the TTS system 150 (e.g., using the TTS model 200) can be applicable to other text-to-speech scenarios, such as voice search, navigation, or reading a document.
[0035] Figure 2 An example TTS model 200 is shown that is adapted to synthesize speech in the voice of a target speaker using a residual adapter 400. The TTS model 200 may include one or more components, including a speaker embedding and / or a lexical stress embedding 205, a bidirectional encoder representations from Transformers (BERT) encoder 210, a variance adapter 220, a duration adapter 230, an upsampler 240, and / or a decoder 300. The decoder 300 may include a plurality of multi-head self-attention layers. For example, the decoder 300 may include a Conformer decoder. Although this disclosure refers to the decoder 300 as the Conformer decoder 300, other decoders with multi-head attention layers are also contemplated without departing from the scope of this disclosure. The TTS model 200 is configured to receive input text 152 (e.g., in the form of a sequence of phonemes and graphemes) and output a synthetic speech representation 261 (e.g., a mel spectrogram) of the input text 152 using various components 205, 210, 220, 230, 240, and / or 300. In some implementations, to convert the synthetic speech representation 261 corresponding to the mel spectrogram in the frequency domain to a time-domain audio waveform corresponding to the synthesized playback audio 154, the TTS system 150 may implement a vocoder (not shown). In some examples, the TTS system 150 implements a neural network-based vocoder, such as a speaker-independent generative pre-trained WaveRNN-based neural vocoder.
[0036] The speaker embedding and / or lexical stress embedding 205 may affect the timbre or other elements of the final output synthesized speech 261. The BERT encoder 210 may receive the input text 152 and generate an encoded text representation 211 for the corresponding input text. The variance adapter 220 may receive the first concatenated output 212, including the speaker embedding and stress embedding 205 (e.g., from the speaker embedding and / or lexical stress embedding 205) concatenated with the encoded text representation 211 from the BERT encoder 210. Then, the variance adapter 220 may generate an output 221 based on the first concatenated output 212. Additionally, the duration adapter 230 may also be configured to receive the first concatenated output 212 and predict the phoneme duration 231 of each phoneme in the corresponding input text 152 based on the first concatenated output 212. The upsampler 240 may be configured to receive a second concatenated output 213 that is based on the concatenation of the output 221 generated by the variance adapter 220 and the first concatenated output 212. Based on the predicted phoneme duration 231 for each phoneme in the corresponding input text 152, the upsampler 231 may upsample the second concatenated output 231, and the Conformer decoder 300 may generate a synthesized speech representation 261 for the corresponding input text 152 based on the upsampled second concatenated output 241.
[0037] In some implementations, during the pre-training ( Figure 5 ) of the TTS model 200, the components 205, 210, 220, 230, 240, 300 (also referred to as sub-models) of the TTS model 200 are optimized. The TTS model 200 may be further adapted to synthesize speech in the voice of a target speaker by setting one or more stacks of residual adapters 400 at various components 205, 210, 220, 230, 240, 300 of the TTS model 200. Although Figure 2 the residual adapter 400 is shown as being set in the variance adapter 220, the duration adapter 230, and / or the Conformer decoder 300, the residual adapter 400 may be set in any one of the components 205, 210, 220, 230, 240, 300. In some implementations, multiple residual adapters 400 may be set throughout the TTS model 200. During the training of the residual adapter 400, the parameters of the components 205, 210, 220, 230, 240, 300 may be frozen. In other words, although the TTS model 200 and the residual adapter 400 may be used to generate an output based on the training data, only the residual adapter 400 is optimized based on the loss. The training of the residual adapter 400 will be discussed in more detail below ( Figure 5 ).
[0038] In some implementations, the TTS model 200 is a non-autoregressive variant of PnG NAT (Phoneme and Grapheme Non-Attentive Tacotron), which replaces the autoregressive LSTM-based decoder with a non-autoregressive Conformer-based decoder. In these implementations, instead of implementing a decoder with long short-term memory (LSTM), the Conformer decoder 300 with multiple multi-head attention layers is more easily adaptable by the residual adapter 400.
[0039] The TTS model 200 can receive input text 152 corresponding to both phoneme and grapheme sequences, and then output a synthetic speech representation 261 as a 128-bin mel spectrogram with a 50 ms frame window and a 12.5 ms frame step. In some implementations, the TTS model 200 includes a PnG BERT (Bidirectional Encoder Representations from Transformers) encoder (BERT encoder 210), 6 stacked Conformer decoders 300, a duration-based Gaussian upsampler 240, and convolutional variance adapters 220, 230 that predict phoneme-level log durations, log F0, and energy, such as FastPitch. In these implementations, the encoder output 211 is concatenated with the utterance-level speaker and lexical stress embeddings 205, and then the utterance-level speaker and lexical stress embeddings are broadcast to the phoneme level before the variance and duration adapters 220, 230 are applied. Before the upsampler 240, the log F0 and energy are concatenated with the speaker and lexical stress embeddings 205 to the encoder output 211.
[0040] In some implementations, the TTS model 200 is first pre-trained using a masked language model (MLM) objective on a large text corpus such as Wikipedia, and then used to initialize the encoder 210. To prevent possible overfitting, when pre-training the TTS model 200, the phoneme and grapheme embeddings and the lower 4 encoder layers 210 are frozen.
[0041] Figure 3Provide an example of a Conformer block / layer implemented by the Conformer decoder 300 of the TTS model 200. The Conformer decoder 300 includes a first half feed-forward layer 310, a second half feed-forward layer 340, and a connection operator 305, where a multi-head self-attention block 320 and a convolutional layer 330 are arranged between the first half feed-forward layer 310 and the second half feed-forward layer 340. The first half feed-forward layer 310 processes the upsampled second concatenated output 241 of the sequence for the input text 152. Subsequently, the multi-head self-attention block 320 receives the upsampled second concatenated output 241 concatenated with the output of the first half feed-forward layer 310. The convolutional layer 330 subsamples the output of the multi-head self-attention block 320 concatenated with the output of the first half feed-forward layer 310. Thereafter, the second half feed-forward layer 340 receives the concatenation of the output of the convolutional layer 330 and the multi-head self-attention block 320. The layer normalization module 350 processes the output from the second half feed-forward layer 340. Mathematically, the Conformer block 300 transforms the input feature x using the modulation feature m to produce the output feature y as follows:
[0042]
[0043] x" = x′ + MHCA(x',n')
[0044] x″′ = x'☉r(x") + h(x")
[0045] x″″ = x' + MHCA(x′,x"′)
[0046]
[0047] In some implementations, a residual adapter 400 is inserted in the Conformer decoder 300. Although the residual adapter 400 is shown as being inserted at the end of the Conformer decoder 300, the residual adapter 400 can be inserted into any suitable component of the Conformer decoder 300. For example, a stack of residual adapters 400 can be inserted into the multi-head self-attention layer 320 of the Conformer decoder 300. In this example, each residual adapter 400 in the stack of residual adapters 400 is arranged between each multi-head self-attention layer 320 in the corresponding stack of multi-head self-attention layers 320 of the Conformer decoder 300. Additionally, as described above, the residual adapter 400 can be arranged in any suitable component of the TTS model 200 such that the TTS model 200 is adapted to synthesize speech in the voice of the target speaker.
[0048] In Figure 3In the example of, the residual adapter 400 can be used to further adapt the output of the Conformer 300 so that the output is the synthetic speech representation 261 of the target speaker trained on the training data corresponding to the target speaker. Figure 4 Figure 4 shows an example residual adapter 400. The residual adapter may include a layer normalization layer 410, a down-projection layer 420, a rectified linear unit (ReLU) layer 430, an up-projection layer 440, and a residual connection 450. In some implementations, the residual adapter 400 receives an input 405 (i.e., the output from a layer of the Conformer decoder 300). The residual adapter 400 then processes the input 405 through components 410, 420, 430, 440, and / or 450 to produce an output 455. In some implementations, the output 455 is a modified version of the input 405 such that the output 455 corresponds to a synthetic utterance in the voice of the target speaker.
[0049] In some implementations, to adapt the pre-trained TTS model 200 to a new target speaker, the residual adapter 400 can be used to adapt the TTS model. The residual adapter may include lightweight neural modules inserted between the layers of the TTS model 200. As Figure 4 shown, each residual adapter 400 can include applying layer normalization to a d-dimensional input vector h ∈ R d , followed by down-projecting to a bottleneck dimension r (where Wdown ∈ R d×r ), ReLU activation, up-projection (where Wup ∈ R r×d), and a residual connection, and has the following form:
[0050] h0 = h + ReLU(Layernorm(h)Wdown)Wup (3)
[0051] Although the residual adapter 400 can be inserted anywhere in the TTS model 200, in some implementations, the residual adapter 400 is most efficient in the Conformer decoder 300. Additionally, in addition to optimizing the residual adapter 400, the training can also optimize the speaker embedding such that the variance adapters 220, 230 ( Figure 2 ) are appropriately conditioned by the learned target speaker embedding. During the training / optimization of the residual adapter 400, the TTS model 200 may be frozen, including moving mean and variance updates.
[0052] Figure 5A training process 500 is shown for adapting a text-to-speech (TTS) model 200 enhanced by a stack of residual adapters 400 to learn how to synthesize speech in the voice of a target speaker. In some implementations, the process 500 employs a two-step training technique. First, the TTS model 200 is pre-trained on an initial training data set 505 that includes spoken utterances from a large number of speakers to obtain a backbone TTS model 200 that produces synthesized speech in a generic voice. The next step of the two-step training technique involves training the stack of residual adapters 400 on the adaptive training data 510 while freezing the parameters of the backbone TTS model 200. The result is a TTS model 200 that is adapted to synthesize speech in the voice of a target speaker by means of the residual adapters 400. The training data 510 may include one or more spoken utterances of the target speaker (e.g., Figure 1 12), wherein each spoken utterance is paired with corresponding input text 512 associated with a transcription of the spoken utterance.
[0053] refer to Figure 5 , the process 500 begins by pre-training the TTS model 200 using pre-training data 505 (i.e., initial training data 505). A pre-trained model is a technique for initializing a model that can then be further fine-tuned based on additional training data 510. For the TTS model 200, pre-training can include initiating the TTS model 200 to synthesize text into speech in a generic voice. In some implementations, the pre-training data 505 does not include any utterances spoken by the target speaker (i.e., non-synthesized speech 513 from the training data 510, used to adapt the TTS model 200 to synthesize speech in the voice of the target speaker). The pre-training data 505 can include spoken utterances paired with textual representations of the spoken utterances.
[0054] Then, the process 500 can adapt the TTS model 200 to synthesize speech in the voice of the target speaker using the training data 510 to fine-tune one or more residual adapters 400 of the TTS model 200 while freezing the parameters of the TTS model 200 after pre-training. That is, while the TTS model 200 is used to generate output 515 based on the input text 512, only the residual adapter 400 is optimized based on the determined loss 540. The training process may include fine-tuning the residual adapter 400 ( Figure 4 ) components 410, 420, 430, 440. Process 500 includes feeding training input 510 to TTS model 200. In some implementations, training input 510 includes a plurality of spoken utterances from a target speaker (e.g., Figure 1the spoken utterance 12 of the target speaker 10), including a non-synthetic speech representation 513 of the target speaker 10 uttering the spoken utterance 12. Additionally, the training input 510 can include an input training text 512 that corresponds to a transcription of the spoken utterance 12 of the target user (e.g., Figure 1 a transcription 152 of the utterance 12). The training data 510 including the input training text 512 can include a corresponding sequence of phonemes and graphemes. In some implementations, the training data set 510 is completely different from the initial training data 505. For example, the initial training data 505 can be in a first language, while the training data set 510 can be in a second language. This enables the backbone TTS model 200 to be trained on a greater variety of pre-training data 505, while the adaptation using the residual adapter 400 can be trained on the target speaker with a significantly smaller set of training data 510.
[0055] After receiving the training input 510, the TTS model 200 enhanced by the residual adapter 400 can generate an output 515 (e.g., a synthetic speech representation 261 from the corresponding input training text 512). The TTS model 200 can process the input training text 512 in the manner described for any of the Figures 1 to 4 figures herein or any other suitable manner for text-to-speech conversion. For example, referring to Figure 2 , the input training text 512 can be received by the BERT encoder 210 of the TTS model 200. The BERT encoder 210 can generate an encoded text representation 211 for the corresponding input text 512. The variance adapter 220 can receive a first concatenated output 212 that includes a speaker embedding and an accent embedding 205 concatenated with the encoded representation 211 generated by the BERT encoder. Then, the variance adapter 220 can generate an output 221 corresponding to the first concatenation 212. The duration adapter 230 can receive the first concatenated output 212 and predict a phoneme duration 231 for each phoneme in the corresponding input training text 512 based on the first concatenated output 212. In some implementations, the upsampler 240 receives a second concatenated output 213 that includes the first concatenated output 212 concatenated with the output 221 generated by the variance adapter 220. Then, the upsampler 240 can upsample the second concatenation. Then, the Conformer decoder 300 can receive the upsampled second concatenated output 241 and generate a synthetic speech representation 261 for the corresponding input text 512. In this example, Figure 2 the synthetic speech representation 261 can correspond to Figure 5Output 515. Additionally, when training the residual adapter 400, one or more stacks of the residual adapter 400 can be inserted into any component of the TTS model 200 as needed to adapt the TTS model 200 to synthesize speech in the voice of the target speaker.
[0056] In some implementations, the output 515 is used by the loss function 530 to generate a loss 540. That is, the loss function 530 compares the output 515 with the spoken utterance of the target speaker to generate a loss 540, where the loss 540 indicates the difference between the non-synthesized speech 513 corresponding to the spoken utterance of the target speaker (i.e., the target output) and the synthesized speech generated by the TTS model 200 from the input text 512 (i.e., the output 515). The loss function 350 can implement any suitable technique to determine the loss, such as regression loss, mean squared error, mean squared logarithmic error, mean absolute error, binary classification, binary cross-entropy, hinge loss, multi-classification loss, etc.
[0057] The loss 540 can then be fed directly into the TTS model 200. Here, the TTS model 200 is frozen, so processing the loss 540 includes adjusting only one or more parameters of the residual adapter 400 to account for the loss 540. In some implementations, one or more of the residual adapters 400 include speaker embeddings for speaker adaptation in the TTS model 200. For example, the speaker embedding can be extracted from the reference mel spectrogram of the target speaker and / or adapted to the speaker embedding in the speaker embedding table that is closest to the timbre of the target speaker. Here, optimizing the residual adapter 400 includes optimizing the speaker embedding associated with the target speaker.
[0058] Figure 6 is a flowchart of an exemplary operational arrangement of a computer-implemented method 600 of a residual adapter for implementing few-shot text-to-speech speaker adaptation. For example, the method 600 can be performed by Figure 1 the exemplary voice environment 100 and / or Figure 7are performed by various components of computing device 700. At operation 602, method 600 includes obtaining a text-to-speech (TTS) model 200 that is configured to convert text 152 into a representation of synthesized speech 261, and TTS model 200 is pre-trained on an initial training dataset 505. At operation 604, method 600 includes enhancing TTS model 200 via a stack of residual adapters 400. At operation 606, method 600 includes receiving an adaptive training dataset 510 that includes one or more spoken utterances 12 spoken by a target speaker 10, and each spoken utterance 12 in adaptive training dataset 510 is paired with a corresponding input text 512 associated with a transcription of spoken utterance 12. At operation 608, method 600 includes adapting the TTS model 200 enhanced via the stack of residual adapters 400 using the adaptive training dataset 510 to learn how to synthesize speech in the voice of target speaker 10 while freezing the parameters of TTS model 200 by optimizing the stack of residual adapters 400.
[0059] Figure 7 is a schematic diagram of an example computing device 700 that can be used to implement the systems and methods described in this document. Computing device 700 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown herein, their connections and relationships, and their functions are only intended to be exemplary and are not intended to limit the implementation of the invention described and / or claimed in this document.
[0060] Computing device 700 includes a processor 710, a memory 720, a storage device 730, a high-speed interface / controller 740 connected to memory 720 and a high-speed expansion port 750, and a low-speed interface / controller 760 connected to a low-speed bus 770 and storage device 730. Each of the components 710, 720, 730, 740, 750, and 760 is interconnected using various buses and can be mounted on a common motherboard or otherwise as appropriate. Processor 710 can process instructions for execution within computing device 700, including instructions stored in memory 720 or on storage device 730, to display graphical information of a graphical user interface (GUI) on an external input / output device, such as a display 780 coupled to high-speed interface 740. In other implementations, multiple processors and / or multiple buses and multiple memories and multiple types of memory can be used as appropriate. Additionally, multiple computing devices 700 can be connected, where each device provides a portion of the necessary operations (e.g., as a server farm, a blade server group, or a multi-processor system).
[0061] Memory 720 stores information non - transiently within computing device 700. Memory 720 can be a computer - readable medium, a volatile memory unit, or a non - volatile memory unit. The non - transient memory 720 can be a physical device for storing programs (e.g., sequences of instructions) or data (e.g., program state information) on a temporary or permanent basis for use by computing device 700. Examples of non - volatile memory include, but are not limited to, flash memory and read - only memory (ROM) / programmable read - only memory (PROM) / erasable programmable read - only memory (EPROM) / electrically erasable programmable read - only memory (EEPROM) (e.g., commonly used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase - change memory (PCM), and magnetic disks or tapes.
[0062] Storage device 730 is capable of providing large - scale storage for computing device 700. In some implementations, storage device 730 is a computer - readable medium. In various different implementations, storage device 730 can be a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory, or other similar solid - state memory devices, or an array of devices (including devices in a storage area network or other configurations). In additional implementations, a computer program product is tangibly embodied in an information carrier. The computer program product contains instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer - readable medium or a machine - readable medium, such as memory 720, storage device 730, or memory on processor 710.
[0063] High - speed controller 740 manages bandwidth - intensive operations of computing device 700, while low - speed controller 760 manages lower - bandwidth - intensive operations. This division of responsibilities is merely exemplary. In some implementations, high - speed controller 740 is coupled to memory 720, display 780 (e.g., via a graphics processor or accelerator), and a high - speed expansion port 750 that can accept various expansion cards (not shown). In some implementations, low - speed controller 760 is coupled to storage device 730 and a low - speed expansion port 790. The low - speed expansion port 790, which can include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), can be coupled to one or more input / output devices, such as a keyboard, a pointing device, a scanner; or to a networking device, such as a switch or a router, via a network adapter.
[0064] As shown in the figure, the computing device 700 can be implemented in many different forms. For example, the computing device can be implemented as a standard server 700a or implemented multiple times in a group of such servers 700a, implemented as a laptop computer 700b, or implemented as part of a rack server system 700c.
[0065] Various implementations of the systems and techniques described herein can be implemented in digital electronic and / or optical circuitry, integrated circuit systems, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementations in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor which can be special purpose or general purpose and coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0066] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented using high-level procedural and / or object-oriented programming languages and / or in assembly / machine language. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, non-transitory computer-readable medium, device, and / or apparatus (e.g., a magnetic disk, an optical disk, a memory, a programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal for providing machine instructions and / or data to a programmable processor.
[0067] A software application (i.e., a software resource) can refer to computer software that causes a computing device to perform tasks. In some examples, a software application can be referred to as an “application,” an “app,” or a “program.” Example applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.
[0068] The processes and logical flows described in this specification can be performed by one or more programmable processors, also referred to as data processing hardware, which execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by special purpose logic circuitry, such as an FPGA (field programmable gate array) or ASIC (application specific integrated circuit). By way of example, processors suitable for the execution of a computer program include both general and special purpose microprocessors, and any one or more processors of any type of digital computer. Generally, a processor will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or be operatively coupled to receive data from one or more mass storage devices or transfer data to one or more mass storage devices or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0069] To provide for interaction with a user, one or more aspects of the present disclosure may be implemented on a computer having a display device (e.g., a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touch screen) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. Additionally, a computer can interact with the user by sending documents to and receiving documents from the device used by the user; for example, by sending a web page to a web browser on a client device of the user in response to a request received from the web browser.
[0070] A variety of implementations have been described. However, it should be understood that various modifications can be made without departing from the spirit and scope of the present disclosure. Accordingly, other implementations are within the scope of the following claims.
Claims
1. A computer-implemented method (600) that, when executed by data processing hardware (112), causes the data processing hardware (112) to perform operations, the operations comprising: obtaining a text-to-speech (TTS) model (200) configured to convert text (152) into a representation of synthesized speech (261), the TTS model (200) being pre-trained on an initial training dataset (505); enhancing the TTS model (200) via a stack of residual adapters (400); receiving an adaptive training dataset (510) comprising one or more spoken utterances (12) spoken by a target speaker (10), each spoken utterance (12) in the adaptive training dataset (510) being paired with a corresponding input text (512) associated with a transcription of the spoken utterance (12); and adapting the TTS model (200) enhanced via the stack of residual adapters (400) using the adaptive training dataset (510) to learn how to synthesize speech in the voice of the target speaker (10) while freezing the parameters of the TTS model (200) by optimizing the stack of residual adapters (400).
2. The computer-implemented method (600) according to claim 1, wherein the TTS model (200) comprises a decoder (300) comprising a stack of corresponding multi-head self-attention layers (320).
3. The computer-implemented method (600) according to claim 1 or 2, wherein enhancing the TTS model (200) via the stack of residual adapters (400) comprises inserting the stack of residual adapters (400) into the decoder (300), each respective residual adapter (400) in the stack of residual adapters (400) being disposed between each multi-head self-attention layer (320) in the stack of multi-head self-attention layers (320) of the decoder (300).
4. The computer-implemented method (600) according to any one of claims 1 to 3, wherein the decoder (300) comprises a Conformer decoder (300), the Conformer decoder comprising: a first feed-forward module (310); a multi-head self-attention module (320); a convolutional module (330); a second feed-forward module (340); and a layer normalization module (350).
5. The computer-implemented method (600) according to any one of claims 1 to 4, wherein: each spoken utterance (12) in the adaptive training dataset (510) comprises a corresponding non-synthesized speech representation (513) of the target speaker (10) who spoke the spoken utterance (12); and the stack of residual adapters (400) is optimized by: For each spoken utterance (12) in the adaptive training dataset (510), a corresponding synthetic speech representation (515) is generated from the corresponding input text (512) paired with the spoken utterance (12) using the TTS model (200) enhanced by the stack of residual adapters (400); A loss (540) is determined based on the non-synthetic speech representation (513) of the target speaker (10) uttering the spoken utterance (12) and the synthetic speech representation (515) generated from the input text (512); and The parameters of the stack of residual adapters (400) are updated based on the loss (540).
6. The computer-implemented method (600) according to any one of claims 1 to 5, wherein the stack of residual adapters (400) is further optimized by optimizing the speaker embedding associated with the target speaker.
7. The computer-implemented method (600) according to any one of claims 1 to 6, wherein the initial training dataset (505) used to pre-train the TTS model (200) does not include any utterances (12) spoken by the target speaker (10).
8. The computer-implemented method (600) according to any one of claims 1 to 7, wherein the input text (512) paired with each corresponding spoken utterance (12) in the adaptive training dataset (510) includes a corresponding phoneme sequence and a corresponding grapheme sequence.
9. The computer-implemented method (600) according to any one of claims 1 to 8, wherein each respective residual adapter (400) in the stack of residual adapters (400) comprises: a layer normalization layer (410); a down-projection layer (420); a rectified linear unit ReLU layer (430); an up-projection layer (440); and a residual connection (450).
10. The computer-implemented method (600) according to any one of claims 1 to 9, wherein the TTS model (200) comprises at least one of the following: a bidirectional encoder representation from Transformer BERT encoder (210), the BERT encoder being configured to: receive the corresponding input text (512) of each spoken utterance (12) in the adaptive training dataset (510); and generate an encoded text representation (211) for the corresponding input text (512); a variance adapter (220), the variance adapter being configured to: receive a first concatenated output (212), the first concatenated output including a speaker embedding and an accent embedding (205) concatenated with the encoded text representation (211) generated by the BERT encoder (210); and generate an output (221) corresponding to the first concatenated output (212); a duration adapter (230), the duration adapter being configured to: receive the first concatenated output (212); and Predict a phoneme duration (231) for each phoneme in the corresponding input text (512) based on the first concatenated output (212); An upsampler (240), the upsampler being configured to: Receive a second concatenated output (213), the second concatenated output including the first concatenated output (212) concatenated with the output (221) generated by the variance adapter (220); and Upsample the second concatenated output (213); And A decoder (300), the decoder being configured to: Receive the upsampled second concatenated output (241); and Generate a synthetic speech representation (261) for the corresponding input text (512).
11. The computer-implemented method (600) according to any one of claims 1 to 10, Wherein: The initial training dataset (505) pre-trains the TTS model (200) to synthesize speech in a first language; and Each spoken utterance (12) in the adaptive training dataset (510) is spoken by the target speaker (10) in a second language different from the first language.
12. A system (100), Comprising: Data processing hardware (112); And Memory hardware (114) communicating with the data processing hardware (112), the memory hardware (114) storing instructions that, when executed on the data processing hardware (112), cause the data processing hardware (112) to perform operations, the operations including: Obtain a text-to-speech TTS model (200), the TTS model being configured to convert text (152) into a representation of synthetic speech (261), the TTS model (200) being pre-trained on an initial training dataset (505); Enhance the TTS model (200) through a stack of residual adapters (400); Receive an adaptive training dataset (510), the adaptive training dataset including one or more spoken utterances (12) spoken by a target speaker (10), each spoken utterance (12) in the adaptive training dataset (510) being paired with a corresponding input text (512) associated with a transcription of the spoken utterance (12); and Use the adaptive training dataset (510) to adapt the TTS model (200) enhanced by the stack of residual adapters (400) to learn how to synthesize speech in the voice of the target speaker (10) while freezing the parameters of the TTS model (200) by optimizing the stack of residual adapters (400).
13. The system (100) according to claim 12, wherein the TTS model (200) includes a decoder (300), the decoder including a stack of corresponding multi-head self-attention layers (320).
14. The system (100) according to claim 12 or 13, wherein enhancing the TTS model (200) by the residual adapter (400) stack includes inserting the residual adapter (400) stack into the decoder (300), and each corresponding residual adapter (400) in the residual adapter (400) stack is arranged between each multi-head self-attention layer in the corresponding multi-head self-attention layer (320) stack of the decoder (300).
15. The system (100) according to any one of claims 12 to 14, wherein the decoder (300) includes a Conformer decoder (300), and the Conformer decoder comprises: a first feed-forward module (310); a multi-head self-attention module (320); a convolutional module (330); a second feed-forward module (340); and a layer normalization module (350).
16. The system (100) according to any one of claims 12 to 15, wherein: each uttered speech (12) in the adaptive training dataset (510) includes a corresponding non-synthetic speech representation (513) of the target speaker (10) who uttered the uttered speech (12); and the residual adapter (400) stack is optimized by: for each uttered speech (12) in the adaptive training dataset (510), using the TTS model (200) enhanced by the residual adapter (400) stack to generate a corresponding synthetic speech representation (515) from the corresponding input text (512) paired with the uttered speech (12); determining a loss (540) based on the non-synthetic speech representation (513) of the target speaker (10) who uttered the uttered speech (12) and the synthetic speech representation (515) generated from the input text (512); and updating the parameters of the residual adapter (400) stack based on the loss (540).
17. The system (100) according to any one of claims 12 to 16, wherein the residual adapter (400) stack is further optimized by optimizing the speaker embedding associated with the target speaker.
18. The system (100) according to any one of claims 12 to 17, wherein the initial training dataset (505) for pre-training the TTS model (200) does not include any utterances (12) spoken by the target speaker (10).
19. The system (100) according to any one of claims 12 to 18, wherein the input text (512) paired with each corresponding uttered speech (12) in the adaptive training dataset (510) includes a corresponding phoneme sequence and a corresponding grapheme sequence.
20. The system (100) according to any one of claims 12 to 19, wherein each corresponding residual adapter (400) in the residual adapter (400) stack comprises: a layer normalization layer (410); a down-projection layer (420); Rectified Linear Unit ReLU layer (430); Upward projection layer (440); And Residual connection (450).
21. The system (100) according to any one of claims 12 to 20, wherein the TTS model (200) comprises at least one of the following: Bidirectional Encoder Representations from Transformers BERT encoder (210), the BERT encoder being configured to: Receive the corresponding input text (512) of each spoken utterance (12) in the adaptive training dataset (510); and Generate an encoded text representation (211) for the corresponding input text (512); Variance adapter (220), the variance adapter being configured to: Receive a first concatenated output (212), the first concatenated output comprising a speaker embedding and an accent embedding (205) concatenated with the encoded text representation (211) generated by the BERT encoder (210); and Generate an output (221) corresponding to the first concatenated output (212); Duration adapter (230), the duration adapter being configured to: Receive the first concatenated output (212); and Predict the phoneme duration (231) of each phoneme in the corresponding input text (512) based on the first concatenated output (212); Upsampler (240), the upsampler being configured to: Receive a second concatenated output (213), the second concatenated output comprising the first concatenated output (212) concatenated with the output (221) generated by the variance adapter (220); and Upsample the second concatenated output (213); And Decoder (300), the decoder being configured to: Receive the upsampled second concatenated output (241); and Generate a synthetic speech representation (261) for the corresponding input text (512).
22. The system (100) according to any one of claims 12 to 21, Wherein: The initial training dataset (505) pre-trains the TTS model (200) to synthesize speech in a first language; and Each spoken utterance (12) in the adaptive training dataset (510) is spoken by the target speaker (10) in a second language different from the first language.