Speech-to-text prompts for speech tasks

CN122804268APending Publication Date: 2026-09-22GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202580016459.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-02-29
Filing Date
2025-02-14
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

虽然文本到语音(TTS)模型通过生成合成语音数据来提供潜在的解决方案,但是对TTS用于训练数据的依赖引入了附加的限制

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122804268A_ABST
    Figure CN122804268A_ABST
Patent Text Reader

Abstract

A method (500) includes receiving a reference utterance (302) and an input text utterance (304). The reference utterance includes multiple terms spoken by a reference speaker, and the input text sequence includes a corresponding transcription of each of the multiple terms spoken by the reference speaker (306). The method includes obtaining a speaker embedding characterizing the speaker features of the reference speaker (308). The method includes generating a replacement input text sequence (314) by replacing the corresponding transcription of a single term among the multiple terms with a replacement transcription corresponding to a different term not included in the reference utterance (316). The method includes generating resynthesized speech based on the reference speaker's voice using a text-to-speech (TTS) model (130) conditioned on the reference utterance and speaker embeddings (134).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to voice-text prompts for speech tasks. Background Technology

[0002] Automatic speech recognition (ASR) models have shown significant progress in recent years, but their performance remains heavily reliant on the availability of large, diverse, and high-quality training datasets. While text-to-speech (TTS) models offer potential solutions by generating synthetic speech data, this reliance on TTS for training data introduces additional limitations. Specifically, while current TTS techniques can produce perceptibly realistic speech, they often struggle to capture the full spectrum of natural human speech variability, including subtle differences in accent, intonation, speaking style, and background noise. Consequently, ASR models trained primarily on TTS-generated data may exhibit reduced robustness and accuracy when encountering real-world speech, particularly in challenging acoustic environments or when processing speech from diverse groups. The lack of diversity in training data and the limitation of real-world representations underscore the ongoing need for improved data augmentation methods and training strategies to enhance the performance and generalization capabilities of ASR systems. Summary of the Invention

[0003] One aspect of this disclosure provides a computer-implemented method executed on data processing hardware, which enables the hardware to perform operations of speech-text prompting for a speech task. These operations include receiving a reference utterance and an input text sequence. The reference utterance comprises a plurality of words spoken by a reference speaker, and the input text sequence comprises a corresponding transcription of each of the plurality of words spoken by the reference speaker. These operations include obtaining a speaker embedding characterizing speaker features of the reference speaker who spoke the plurality of words. These operations include generating a replacement input text sequence by replacing the corresponding transcription of a corresponding word among the plurality of words with a replacement transcription corresponding to a different word not included in the reference utterance. These operations also include generating resynthesized speech based on the reference speaker's voice using a text-to-speech (TTS) model conditioned on the reference utterance and speaker embedding.

[0004] Implementations of this disclosure may include one or more optional features from the following optional features. In some implementations, these operations further include: receiving an audio signal comprising one or more other terms spoken by different speakers; generating a reference utterance of the voice clone based on the reference utterance and the audio signal using a voice cloning model; and generating resynthesized speech of the voice clone based on a replacement input text sequence using TTS further conditioned on the reference utterance of the voice clone. The reference utterance of the voice clone comprises synthesized speech of a different speaker corresponding to the input text sequence. The resynthesized speech of the voice clone comprises synthesized speech of a different speaker corresponding to the replacement input text sequence. In some examples, a corresponding term among a plurality of terms includes a first hot word, and different terms include a second hot word different from the first hot word. In these examples, these operations may further include training a hot word model on the resynthesized speech.

[0005] These operations may further include training a speech recognition model on a reference utterance paired with an input text sequence and resynthesized speech paired with a replacement input text sequence. Distinct terms may include speech disfluency terms. In some implementations, distinct terms are sampled from a speech domain different from the speech domain associated with the reference utterance. In some examples, these operations further include post-processing the resynthesized speech based on the reference utterance to preserve audio from the reference utterance for terms included in both the resynthesized speech and the reference utterance. In these examples, post-processing of the resynthesized speech includes at least one of crossfading, forced alignment, or dynamic time warping.

[0006] These operations may further include: modifying alternative transcriptions of different terms from the alternative input text sequence; generating an updated alternative input text sequence by replacing the alternative transcriptions with the modified alternative transcriptions; and generating additional resynthesized speech based on the updated alternative input text sequence using a TTS model conditioned on a reference utterance and speaker embeddings. Here, modifying alternative transcriptions of different terms from the alternative input text sequence includes at least one of the following: modifying syllables of the alternative transcription; modifying punctuation of the alternative transcription; or inserting speech disfluency into the alternative transcription. In some implementations, these operations further include: generating an enhanced speaker embedding that enhances at least one of the speaker characteristics of a reference speaker uttering multiple terms; and generating additional resynthesized speech using a TTS model further conditioned on the enhanced speaker embedding, the additional resynthesized speech comprising synthesized speech corresponding to the alternative input text sequence as the speech of a reference speaker having the enhanced at least one of the speaker characteristics.

[0007] Another aspect of this disclosure provides a system including data processing hardware and memory hardware storing instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations. These operations include receiving a reference utterance and an input text sequence. The reference utterance includes a plurality of words spoken by a reference speaker, and the input text sequence includes a corresponding transcription of each of the plurality of words spoken by the reference speaker. These operations include obtaining a speaker embedding characterizing the speaker's speech features of the reference speaker who spoke the plurality of words. These operations include generating a replacement input text sequence by replacing the corresponding transcription of a corresponding word among the plurality of words with a replacement transcription corresponding to a different word not included in the reference utterance. These operations also include generating resynthesized speech based on the reference speaker's speech using a text-to-speech (TTS) model conditioned on the reference utterance and speaker embedding.

[0008] Implementations of this disclosure may include one or more optional features from the following optional features. In some implementations, these operations further include: receiving an audio signal comprising one or more other terms spoken by different speakers; generating a reference utterance of the voice clone based on the reference utterance and the audio signal using a voice cloning model; and generating resynthesized speech of the voice clone based on a replacement input text sequence using TTS further conditioned on the reference utterance of the voice clone. The reference utterance of the voice clone comprises synthesized speech of a different speaker corresponding to the input text sequence. The resynthesized speech of the voice clone comprises synthesized speech of a different speaker corresponding to the replacement input text sequence. In some examples, a corresponding term among a plurality of terms includes a first hot word, and different terms include a second hot word different from the first hot word. In these examples, these operations may further include training a hot word model on the resynthesized speech.

[0009] These operations may further include training a speech recognition model on a reference utterance paired with an input text sequence and resynthesized speech paired with a replacement input text sequence. Distinct terms may include speech disfluency terms. In some implementations, distinct terms are sampled from a speech domain different from the speech domain associated with the reference utterance. In some examples, these operations further include post-processing the resynthesized speech based on the reference utterance to preserve audio from the reference utterance for terms included in both the resynthesized speech and the reference utterance. In these examples, post-processing of the resynthesized speech includes at least one of crossfading, forced alignment, or dynamic time warping.

[0010] These operations may further include: modifying alternative transcriptions of different terms from the alternative input text sequence; generating an updated alternative input text sequence by replacing the alternative transcriptions with the modified alternative transcriptions; and generating additional resynthesized speech based on the updated alternative input text sequence using a TTS model conditioned on a reference utterance and speaker embeddings. Here, modifying alternative transcriptions of different terms from the alternative input text sequence includes at least one of the following: modifying syllables of the alternative transcription; modifying punctuation of the alternative transcription; or inserting speech disfluency into the alternative transcription. In some implementations, these operations further include: generating an enhanced speaker embedding that enhances at least one of the speaker characteristics of a reference speaker uttering multiple terms; and generating additional resynthesized speech using a TTS model further conditioned on the enhanced speaker embedding, the additional resynthesized speech comprising synthesized speech corresponding to the alternative input text sequence as the speech of a reference speaker having the enhanced at least one of the speaker characteristics.

[0011] Details of one or more implementations of this disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will become apparent from the specification, drawings, and claims. Attached Figure Description

[0012] Figure 1 This is a schematic diagram of an example speech recognition system.

[0013] Figure 2 This is a schematic diagram of an example automatic speech recognition model.

[0014] Figure 3 This is a schematic diagram of an example speech synthesis process.

[0015] Figure 4 This is a schematic diagram of an example training process.

[0016] Figure 5 This is a flowchart illustrating an example layout of a computer-implemented method for voice-text prompts for voice tasks.

[0017] Figure 6 This is a schematic diagram of an example computing device that can be used to implement the systems and methods described herein.

[0018] In the various figures, the same reference numerals indicate the same elements. Detailed Implementation

[0019] Automatic speech recognition (ASR) systems have become increasingly prevalent in a wide range of applications, from voice assistants and dictation apps to automated customer service and accessibility tools. However, the performance of these ASR systems is fundamentally tied to the availability of robust and representative training data. To this end, ASR models are trained on massive corpora of transcribed audio representing a broad range of speakers, accents, speaking styles, and acoustic conditions. However, creating such comprehensive datasets is a significant challenge. Collecting and transcribing real-world speech is expensive, time-consuming, and often raises privacy concerns. Furthermore, certain speech patterns, such as those from specific groups of characteristics or those including rare words or phrases, may be underrepresented in existing datasets, leading to performance bias and reduced accuracy for these groups.

[0020] Text-to-speech (TTS) systems offer a potential solution to data scarcity by generating synthesized speech data. TTS systems generate synthesized speech from text, effectively expanding the training data available for ASR models. While advances in TTS have yielded remarkable results in speech naturalness and intelligibility, relying solely on TTS-generated data also presents its own set of challenges. Current TTS models, though sophisticated, often struggle to capture the full spectrum of human speech variability. Numerous differences in articulation, intonation, speech rate, and timbre—essential for accurate ASR—may be lost or simplified in synthesized speech. Furthermore, TTS systems can introduce artifacts or biases that then propagate into ASR models, negatively impacting the performance of real-world speech.

[0021] For example, TTS systems trained primarily on formal scripted languages ​​may produce synthesized speech lacking the fluency and conversational nuances present in everyday speech. Consequently, ASR models trained primarily on such data may exhibit reduced robustness when encountering spontaneous speech in noisy or informal environments. This discrepancy between synthesized and real-world speech underscores the urgent need for improved training data strategies to bridge the gap between the capabilities of current TTS technologies and the demand for robust and accurate ASR systems.

[0022] Therefore, this paper's implementation involves a speech resynthesis process. The speech resynthesis process includes receiving a reference utterance and an input text sequence. The reference utterance includes multiple words spoken by a reference speaker, and the input text sequence includes the corresponding transcription of each of the multiple words spoken by the reference speaker. The speech resynthesis process includes obtaining a speaker embedding that characterizes the speaker features of the reference speaker who spoke the multiple words. The speech resynthesis process includes generating a replacement input text sequence by replacing the corresponding transcription of one of the multiple words with a replacement transcription corresponding to a different word not included in the reference utterance. The speech resynthesis process includes generating resynthesized speech based on the reference speaker's voice using a text-to-speech (TTS) model conditioned on the reference utterance and speaker embedding, based on the replacement input text sequence.

[0023] Advantageously, the training process can use resynthesized speech generated by a speech resynthesis process to train the ASR model. This training process addresses the limitations of current ASR training data by generating synthetic speech that more closely mimics real-world speech patterns and speaker variability. By leveraging reference utterances and corresponding transcripts, the training process extracts speaker embeddings that capture the unique voice characteristics of the reference speaker. The training process conditioned the TTS model on the reference utterance and speaker embeddings, ensuring that the newly generated speech maintains the voice and speaking style of the reference speaker. Notably, the training process does not simply generate speech from existing transcripts. That is, the training process introduces variation by replacing words within the original transcript with new words, thereby creating new sentences while preserving the speaker's voice. This process generates synthetic speech data that is diverse in content and consistent in speaker characteristics, effectively extending the ASR model's training data with realistic variations in vocabulary and sentence structure, all anchored to the voice profiles of real speakers. This allows ASR models to be trained on richer and more representative datasets, thereby improving robustness and accuracy when dealing with real-world speech (especially from different speakers and in a variety of acoustic environments).

[0024] Figure 1An example system 100 is shown that implements an Automated Speech Recognition (ASR) model 200, which resides on a user device 102 of user 104 and / or on a remote computing device 201 (e.g., one or more servers of a distributed system executing in a cloud computing environment) communicating with user device 102. Although user device 102 is depicted as a mobile computing device (e.g., a smartphone), user device 102 can correspond to any type of computing device, such as, but not limited to, tablet devices, laptop / desktop computers, wearable devices, digital assistant devices, smart speakers / displays, smart appliances, automotive infotainment systems, or Internet of Things (IoT) devices, and is equipped with data processing hardware 111 and memory hardware 113.

[0025] User device 102 includes an audio subsystem 108 configured to receive utterance 106 spoken by user 104 (e.g., user device 102 may include one or more microphones for recording spoken utterance 106) and convert utterance 106 into a corresponding digital format associated with an input acoustic frame 110 that can be processed by ASR system 100. In the example shown, user 104 speaks the phrase “What is the weather in New York City?” in natural language English, and audio subsystem 108 converts utterance 106 into a corresponding acoustic frame 110 for input to ASR system 100. ASR model 200 then receives the acoustic frame 110 corresponding to utterance 106 as input and generates / predicts a corresponding transcription 120 (e.g., identification result / hypothesis) of utterance 106 as output. In the example shown, user device 102 and / or telecomputing device 201 also execute user interface generator 107, which is configured to present a representation of the transcription 120 of utterance 106 to user 104 of user device 102. In some configurations, the transcription 120 output from ASR system 100 is processed by, for example, a natural language understanding (NLU) module executing on user device 102 or telecomputing device 201 to execute user commands. Alternatively or additionally, a TTS model (executed on any combination of user device 102 or telecomputing device 201) can convert the transcription 120 into synthesized speech for audible output by audio subsystem 108 or another device. For example, the original utterance 106 may correspond to a message that user 104 is sending to a friend, where the transcription 120 is converted into synthesized speech for audible output to the friend so that the friend can hear the message conveyed in the original utterance 106.

[0026] refer to Figure 2In some examples, ASR model 200 includes a recurrent neural network transducer (RNN-T) model architecture that follows latency constraints associated with interactive applications. The use of the RNN-T model architecture is exemplary, and ASR model 200 may include other architectures such as transformer-transducer and conformer-transducer model architectures. The RNN-T model architecture offers a small computational footprint and uses less memory than conventional ASR architectures, making it suitable for performing speech entirely on user device 102 (e.g., without needing to communicate with a remote server). The RNN-T model architecture of ASR model 200 includes an encoder network 210, a prediction network 220, and a joint network 230. The encoder network 210, broadly similar to the acoustic model (AM) in a conventional ASR system, is a recurrent network consisting of stacked or stacked long short-term memory (LSTM) layers of self-attention layers (e.g., conformer or transformer layers). For example, the encoder network 210 reads... d 110-dimensional feature vector sequences (e.g., acoustic frames 110) Figure 1 )) x = (x 1 , x 2 , ··· , x T ),in x t ∈ d Furthermore, at each output step, a higher-order feature representation is generated. This higher-order feature representation is represented as... .

[0027] Similarly, prediction network 220 is also an LSTM network, which, like a language model (LM), will output the non-whitespace symbol sequence so far by the final Softmax layer 240. y 0 , . . . , y ui-1 Processing as dense representation Finally, utilizing the RNN-T model architecture, a joint network 230 combines the representations generated by the encoder network 210 and the predictor / decoder network 220. The predictor network 220 can be replaced by an embedding lookup table to improve latency by outputting the looked-up sparse embeddings instead of processing dense representations. Then, the joint network predicts... This is a distribution of the next output symbol. In other words, the joint network 230 generates a probability distribution over possible speech recognition hypotheses at each output step (e.g., time step). Here, “possible speech recognition hypotheses” correspond to a set of output labels, each output label representing a symbol / character in a specified natural language. For example, when the natural language is English, the set of output labels may include twenty-seven (27) symbols, such as one label for each of the 24 letters of the English alphabet, and one label for a space. Thus, the joint network 230 may output a set of values ​​indicating the probability of each output label appearing in the predetermined set of output labels. This set of values ​​may be a vector and may indicate a probability distribution over the set of output labels. In some cases, the output labels are graphemes (e.g., individual characters, and possibly punctuation marks and other symbols), but the set of output labels is not limited to this. For example, in addition to graphemes or alternative graphemes, the set of output labels may also include word segments, phonemes, and / or whole words. The output distribution of the joint network 230 may include posterior probability values ​​for each of the different output labels. Therefore, if there are 100 different output labels representing different characters or other symbols, then the output y of the joint network 230... i It can include 100 different probability values, one probability value for each output label. The probability distribution can then be used (e.g., by a Softmax layer 240) to select candidate orthographic elements (e.g., graphemes, word segments, and / or words) during the beam search process and assign them scores for determining transcription 120.

[0028] The Softmax layer 240 can employ any technique to select the output label / symbol with the highest probability in the distribution as the next output symbol predicted by the ASR model 200 at the corresponding output step. In this way, the RNN-T model architecture of the ASR model 200 does not make a conditional independence assumption; instead, the prediction of each symbol is conditional not only on acoustics but also on the sequence of labels output so far. The ASR model 200 does assume that the output symbol is independent of future acoustic frames 110, which allows the RNN-T model architecture of the ASR model 200 to be used in a streaming manner.

[0029] In some examples, the encoder network (i.e., the audio encoder) 210 of the ASR model 200 includes a stack of self-attention layers / blocks, such as conformer blocks. Here, each conformer block includes a series of multi-head self-attention layers, depthwise convolutional layers, and feedforward layers. The prediction network 220 may have two 2,048-dimensional LSTM layers, each followed by a 440-dimensional projection layer. Alternatively, the prediction network 220 may include a stack of transformer blocks or conformer blocks, or an embedding lookup table instead of LSTM layers. Finally, the joint network 230 may also have 440 hidden units. The softmax layer 240 may consist of a uniform set of word pieces or pixels generated using all unique word pieces or pixels from multiple training datasets.

[0030] Figure 3 A speech resynthesis process 300 implementing enhancement model 310, post-processing module 320, modifier 330, speech cloning model 340, and TTS model 130 is illustrated. The speech resynthesis process 300 obtains multiple reference utterances 302 and multiple input text sequences 304. Each reference utterance 302 is a speech segment comprising multiple terms spoken by a reference speaker. For example, reference utterance 302 may include “the quick brown fox jumps over the lazydog” spoken by a reference speaker. Reference utterances 302 may include non-synthesized speech (e.g., human speech) or synthesized speech. Each input text sequence 304 corresponds to one of the reference utterances 302 and includes a corresponding transcription 306 of each of the multiple terms spoken by the reference speaker in reference utterance 302. That is, input text sequence 304 includes a corresponding transcription 306 of each word or term spoken in reference utterance 302. Therefore, input text sequence 304 directly corresponds to reference utterance 302. Furthermore, the speech resynthesis process 300 obtains a speaker embedding 308 representing the speaker characteristics of a reference speaker uttering multiple words. The speaker embedding 308 is a multidimensional representation that captures various attributes of the reference speaker's speech, such as intonation, language, prosody, pitch, tone, and speaking style. As will become clear below, these attributes can be used to generate synthesized speech later in the process that closely matches the natural speech patterns of the reference speaker.

[0031] The enhancement model 310 obtains each input text sequence 304 comprising a transcription 306 of each of a plurality of terms, and generates a replacement input text sequence 314 by manipulating the input text sequence 304. Here, manipulation may include replacing the corresponding transcription 306 of a particular term among the plurality of terms with a replacement transcription 316 corresponding to a different term not included in the reference discourse 302. For example, if the input text sequence 304 is “the quick brown fox jumps over the lazy dog,” the enhancement model 310 may replace the transcription corresponding to the term “fox” with a replacement transcription 316 corresponding to the term “cat.” Manipulation may also include inserting, deleting, or modifying words or phrases within the input text sequence 304. For example, the enhancement model 310 may insert the word “very” before “quick,” resulting in “the very quick brown fox jumps over the lazy dog.” Alternatively, the augmentation model can remove the word "the" before "lazy," resulting in "the quick brown fox jumps over lazy dog." The substitution, insertion, deletion, or modification process introduces variability into the training data, which is beneficial for creating more robust ASR models. For example, continuing the previous example, the augmentation model 310 can replace the transcription 306 of the term "fox" with the substitution transcription 316 of the term "cat," resulting in the substitution input text sequence 314 of "the quick brown cat jumps over the lazy dog."

[0032] In some examples, different terms include speech disfluency terms. Speech disfluency terms refer to any interruption in the normal speech flow, such as filler words (e.g., “um”, “uh”), repetitions (e.g., “III”), or corrections. Speech disfluency terms may replace one of the terms within the input text sequence 304 or be added to the input text sequence 304. Including such terms can help the ASR model 200 better handle real-world speech patterns that typically include these speech disfluencies. For example, the enhancement model 310 may insert the filler word “um” before the word “fox” to obtain “the quick brown um fox jumps over the lazy dog”. Alternatively, the enhancement model 310 may replace the word “fox” with the repetition “ff-fox” to obtain “the quick brown ff-fox jumps over the lazy dog”. As another example, the enhancement model can insert corrections, changing "the quick brown fox" to "the quick brown, I mean, the quick red fox jumps over the lazydog." In addition to replacing transcription 306 or replacing transcription 306, non-fluent words can also be added to the input text sequence 304. For example, enhancement model 310 can replace "fox" with "cat" and insert the filler word "uh" before "cat," resulting in "the quick brown uh cat jumps over the lazy dog."

[0033] In some implementations, the enhancement model 310 samples different terms from a different speech domain than the one associated with the reference utterance 302. Sampling from a different speech domain means selecting terms used in different contexts or environments. For example, if the reference utterance 302 is from a formal speech domain, different terms can be sampled from a casual speech domain, a technical slang domain, or another different context. Therefore, sampling from different speech domains ensures that the ASR model 200 is exposed to a wide range of vocabulary and speaking styles, thereby enhancing the ASR model 200's ability to generalize across various speech types. For example, if the reference utterance 302 is "Good evening, ladies and gentlemen" from a formal context, the enhancement model 310 can replace the term "ladies and gentlemen" with "folks" from a casual context, resulting in "Good evening, folks".

[0034] Then, the speech resynthesis process 300 conditions the TTS model 130 with the reference utterance 302 and the speaker embedding 308. Specifically, "conditioning" refers to the process of providing these inputs to the TTS model 130 so that the TTS model 130 considers these inputs when generating speech. Conditioning can be achieved through various techniques, such as feeding the reference utterance 302 and the speaker embedding 308 as input features to the model, or by using techniques such as attention mechanisms to allow the TTS model 130 to focus on relevant aspects of these inputs. By conditioning the TTS model 130 with these inputs, the speech resynthesis process 300 ensures that the TTS model 130 generates speech that not only matches the replaced input text sequence 314 but also preserves the unique voice characteristics of the reference speaker. That is, even if the words have been changed, the speech will sound as if it were spoken by the reference speaker. In fact, the reference speaker may never have said the changed words. Subsequently, the conditional TTS model 130 generates resynthesized speech 134 based on the replaced input text sequence 314 and the voice of the reference speaker. Continuing the example above, the conditional TTS model 130 generates resynthesized speech 134, which produces a new audio segment uttering "The quick brown cat jumps over the lazy dog" in the voice of the reference speaker. It is noteworthy that the reference speaker never uttered the entire utterance of the resynthesized speech 134 (e.g., the reference speaker never said "cat"), but the resynthesized speech 134 is in the voice of the reference speaker. Therefore, the speech resynthesis process 300 ensures that the synthesized speech data is context-diverse and consistent in speaker characteristics, which is important for training the ASR model 200 to effectively handle real-world speech variability.

[0035] In some implementations, a corresponding term among multiple terms includes a first hot word, and different terms include a second hot word that differs from the first hot word. Hot words are keywords or phrases that activate a specific function or response of the ASR model 200. The substitution process is particularly useful for training the ASR model 200 to identify and respond to a variety of hot words, thereby enhancing the flexibility and robustness of the ASR model 200. For example, reference utterance 302 may include “Hey Google.” Here, the augmented model 310 can substitute the term “Google” with the term “Pixel” to create a substitution input text sequence 312 of “Hey Pixel.” The variations produced by the substitution input text sequence 312 allow the ASR model 200 to learn and adapt to different hot words. Furthermore, the method can be extended to other hot words and phrases, enabling the creation of diverse training datasets that better represent the range of user interactions with the ASR model 200.

[0036] Optionally, the speech resynthesis process 300 may employ a post-processing module 320 to post-process the resynthesized speech 134 to enhance its quality and naturalness. The post-processing module 320 may perform several functions to refine the resynthesized speech 134 and generate a post-processed resynthesized speech 324. Alternatively, the post-processing module 320 may perform functions to refine the resynthesized speech 136 of the speech clone to generate a post-processed resynthesized speech 324. First, the post-processing module 320 analyzes the resynthesized speech 134 in conjunction with the reference utterance 302 to identify and retain audio segments common to both the reference utterance 302 and the resynthesized speech 134. Therefore, the post-processing module 320 ensures that the terms present in both the resynthesized speech 134 and the reference utterance 302 retain the acoustic properties of the reference utterance 302, thereby maintaining the naturalness and consistency of the reference speaker's voice. Post-processing of the resynthesized speech may include applying at least one of crossfading, forced alignment, or dynamic time warping to the resynthesized speech 134. In some implementations, the post-processing module 320 may employ a combination of these techniques.

[0037] Crossfading helps smooth transitions between audio segments by gradually blending the end of one segment with the beginning of the next, minimizing abrupt changes and creating a more natural flow. Forced alignment ensures precise synchronization of phonological elements by aligning the transcribed pronunciation of the resynthesized speech 134 with the reference utterance 302, thus ensuring accurate timing of each phoneme. Using forced alignment and crossfading as post-processing steps allows for the complete preservation of the original waveform samples of specific words. For example, given the reference utterance 302 of “Hey Google, what is the time?” and the resynthesized speech 134 of “Hey Gemini, what is the time?”, forced alignment can be used to obtain timing information for each word, and only the edited word (“Google”) is replaced with (“Gemini”), while crossfading is applied near word boundaries. In this approach, the audio waveforms of the word “Hey” and the query (i.e., “What’s the time?”) are fully preserved.

[0038] Dynamic Time Warping (DTW) adjusts the timing of the resynthesized speech 134 to match a natural speaking pattern by stretching or compressing the time axis to align with the reference utterance 302, thus ensuring a close match in rhythm and tempo. In some examples, DTW can be used to partially preserve the original waveform samples of specific words. For example, given the original “Hey Google” + query utterance and the newly synthesized “Hey Gemini” + query utterance, DTW can be applied to the feature space (e.g., waveform, MFCC, log-Mel filter bank energy, or F0) between the two utterances. Based on DTW alignment, some synthesized samples in the synthesized samples can be replaced with the corresponding original samples, or a weighted average of the synthesized and original samples can be used. Thus, the post-processing module 320 produces a post-processed resynthesized speech 324 that is not only diverse in content but also of high quality, thus closely matching the nuances of real-world speech. For example, if the original resynthesized speech 134 has abrupt transitions between words, crossfade-in and crossfade-out will smooth these transitions. If the timing of phonemes is slightly off, forced alignment will correct the timing. If the overall speed of the speech feels unnatural, dynamic warping will adjust the resynthesized speech to sound more natural.

[0039] In some configurations, the speech resynthesis process 300 includes a modifier 330 configured to modify the substitution transcription 316 from the substitution input text sequence 312 for different terms. That is, the modifier 330 generates an updated substitution input text sequence 334 by replacing the substitution transcription 316 with the modified substitution transcription 336. The modification process introduces additional variability into the synthesized speech data. For example, if the substitution input text sequence 312 is “Hey Pixel,” the modifier 300 can update the substitution input text sequence 332 to “Hey, Pixel!” or “Hey, Pix-el” or “Heh Peexel”, “Hey, Hey Pixel”, “Hey Pixl”, “Hey; Pixel”, or “Hey uh Pixel”. Subsequently, the TTS model 130, conditioned on the reference utterance 302 and the speaker embedding 308, generates additional resynthesized speech 134 based on the updated replacement input text sequence 332, using the reference speaker's voice. This iterative process of modifying and resynthesizing speech ensures that the speech data covers a broad vocabulary and sentence structure while preserving the unique voice characteristics of the reference speaker. In some embodiments, the resynthesized speech 134 based on certain modified replacement input text sequences 334, such as "Hey tickle" in the current example, can be labeled as negative examples for training hot word or keyword models.

[0040] The modifier 330 can also change the pronunciation of syllables, for example, changing "Hey Pixel" to "Heh Peexel". In some implementations, a pronunciation dictionary is used to identify text prompt variations, where the dictionary specifies how words are pronounced. All pronunciations of the target keyword are found, and then a dictionary search is performed to find words with similar pronunciations. The modifier 330 uses these words, including variations with different spellings and pronunciations, to create a modified replacement transcription 336. In other embodiments, a large language model is used to identify text prompt variations. The large language model can be prompted with the target keyword and required to generate a list of spelling variations, pronunciation variations and / or other text prompt variations. The large language model can also be prompted to generate negative examples, such as similar-sounding phrases that are not keyword triggers. For example, the large language model can be prompted with the target keyword "Hey Pixel" and required to generate a list of spelling variations such as "Hey Pixl", "Hey, Pixel" and "Hey; Pixel". The large language model can also be prompted to generate pronunciation variations, such as "Heh Peexel" and "Hey, um, Pixel". In addition, the large language model can be prompted to generate negative examples, such as "Hey tickle" and "Heypickle". The large language model can also be prompted to generate text prompt variations that are not spelling or pronunciation variations, such as "Hey, what's up Pixel?" and "Hey Pixel, how's it going?"

[0041] The modifier 330 can modify the replacement transcription 316 in various ways to create an updated replacement input text sequence 332. These modifications may include at least one of the following: modifying syllables of the replacement transcription; modifying punctuation of the replacement transcription; or inserting speech disfluencies into the replacement transcription. For example, the modifier 330 can insert an additional syllable into the replacement input text sequence 312 of "HeyPixel", resulting in "Hey Hey Pixel". Alternatively, the modifier 330 can delete a syllable, changing "Hey Pixel" to "Hey Pixl". The modifier 330 can also adjust the punctuation of the replacement input text sequence 312, transforming "Hey Pixel" into "Hey, Pix-el". Additionally or alternatively, the modifier 330 can insert speech disfluencies such as stutters or filled pauses to make the synthesized speech sound more natural and diversified. For example, the modifier 330 can modify the replacement input text sequence 312 of "Hey Pixel" into "Hey, um, Pixel".

[0042] In some examples, modifier 330 generates an enhanced speaker embedding 338 that enhances at least one speaker characteristic of a reference speaker uttering multiple words. The enhancement process involves altering specific attributes of the speaker embedding 308. These attributes may include pitch, tone, prosody, or other speech characteristics of the reference speaker. For example, the enhanced speaker embedding 338 may adjust intonation patterns to make the speech sound more formal or casual, or adjust prosody to emphasize different parts of the speech. The modification process may also include changing the speaker identity by utilizing the target SpeakerId, thereby producing synthesized speech with the characteristics of the target speaker. Similarly, the LanguageId may be modified to generate speech with a foreign accent, where the characteristics of the reference speaker are preserved but the pronunciation follows the rules of the target language.

[0043] Furthermore, prosodic features such as duration, broad timing, local timing, and pitch shifts can be controlled to generate speech distinct from the reference audio while maintaining correspondences with various high- or low-level attributes. Subsequently, in addition to or instead of speaker embedding 308, the speech resynthesis process 300 can condition the TTS model 130 with an enhanced speaker embedding 338. By conditioning the TTS model 130 with the enhanced speaker embedding 338, the TTS model can generate a wider range of synthesized speech outputs that still sound like the reference speaker but exhibit different speech characteristics. In some implementations, the modifier 330 simultaneously controls speaker features, prosodic features, and / or dictionary content to generate speech distinct from the reference audio in a predictable manner. This allows manipulation of spoken words, accent, speaker identity, timing, and / or pitch.

[0044] The TTS model 130 can be further conditional on the modified speaker embedding 338. Therefore, the conditionalized TTS model 130 generates additional resynthesized speech 134, which includes synthesized speech corresponding to the replaced input text sequence 314, but now incorporates the voice of a reference speaker possessing at least one enhanced speaker characteristic from the enhanced speaker characteristics of the enhanced speaker embedding 338. For example, if the original speaker embedding represents a calm and even speaking style, the enhanced speaker embedding 338 can introduce a more dynamic and expressive style, resulting in the additional resynthesized speech 134 that sounds more lively while still being recognizable as the reference speaker's voice. Another example could be that if the original speaker embedding 308 has a formal tone, the enhanced speaker embedding 338 can be modified to sound more casual and conversational. This flexibility allows the synthesized speech to be customized for different scenarios and user interactions, making the speech more universal for various applications. The conditional TTS model 130 can also generate additional resynthesized speech 134 based on the modified replacement input text sequence 334, whereby the TTS model 130 is conditional on the reference utterance 302 and the speaker embedding 308 or the enhanced speaker embedding 338.

[0045] In some implementations, the speech resynthesis process 300 employs a speech cloning model 340, which is configured to clone the speech of a different speaker having speech characteristics different from a reference speaker. Here, the speech cloning model 340 receives an audio signal 301 including one or more additional words spoken by a different speaker. The speech cloning model 340 generates a speech clone reference 342 based on the reference utterance 302 and the audio signal 301. Here, the speech clone reference utterance 342 includes synthesized speech corresponding to the input text sequence 304, representing the speech of a different speaker. Therefore, in addition to or instead of the reference utterance 302, the speech resynthesis process 300 may further condition the TTS model 130 with the speech clone reference utterance 342. Here, the conditionalized TTS model 130 generates the resynthesized speech 136 of the speech clone based on a replacement input text sequence 314. For example, if the replacement input text sequence 314 is "jumping over the lazy dog", and the TTS model 130 is conditional on the speech clone reference utterance 342, then the resynthesized speech 136 will be a synthesized version of "jumping over the lazy dog" spoken in the voice of the target speaker. The resynthesized speech 136 includes synthesized speech with a different speaker's voice corresponding to the replacement input text sequence 314.

[0046] Figure 4A training process 400 for training the ASR model 200 is illustrated. The training process 400 obtains multiple reference utterances 302 and multiple input text sequences 304. Here, each reference utterance 302 corresponds to one input text sequence in the input text sequence 304. That is, each reference utterance 302 is paired with one input text sequence in the input text sequence 304. For example, the reference utterance 302 might be an audio recording of the phrase "the quick brown fox" spoken by a specific individual, and the corresponding input text sequence 304 would be "the quick brown fox". Furthermore, the training process 400 obtains resynthesized speech 134 paired with a replacement input text sequence 314 or an updated replacement input text sequence 334. The training process 400 may also obtain resynthesized speech 136 of a speech clone paired with a replacement input text sequence 314. Alternatively, the training process 400 may obtain post-processed resynthesized speech 324 paired with a replacement input text sequence 314 or an updated replacement input text sequence 334. For each speech input (e.g., reference utterance 302, resynthesized speech 134, resynthesized speech 136 from a speech clone, and post-processed resynthesized speech 324), the ASR model 200 processes the speech input to generate a corresponding transcription 120. For example, if the speech input is reference utterance 302 “the quickbrown fox”, then the ASR model 200 will ideally generate transcription 120 “the quick brown fox”. However, during training, the transcription 120 generated by the ASR model 200 may be incorrect, such as “the quick brown focks” or “the quik brown fox”.

[0047] Subsequently, the loss module 410 determines the loss 412 by comparing the transcription 120 with the corresponding input text sequence 304, the alternative input text sequence 314, or the updated alternative input text sequence 334. For example, the loss module 410 compares the transcription 120 of "the quick brown focks" with the correct input text sequence 304, the alternative input text sequence 314, or the modified alternative input text sequence 334 corresponding to "the quick brown fox," and determines a loss value reflecting the difference. The training process 400 trains the ASR model 200 on the loss 412 by updating the parameters of the ASR model 200. Therefore, the training process 400 trains the ASR model 200 on various speech inputs such as reference utterance 302, resynthesized speech 134, resynthesized speech 136 of speech clone, and post-processed resynthesized speech 324. After the training process 400 trains the ASR model 200, the trained ASR model 200 can obtain utterance 106 and generate the transcription 120 of utterance 106. In some implementations, the ASR model 200 includes a hot word model, such that the training process trains the hot word model on the resynthesized speech 134.

[0048] Figure 5 This is a flowchart illustrating an example layout of a computer-implemented method 500 for voice-text prompts for voice tasks. Method 500 can utilize memory hardware 620 (…). Figure 6 Instructions on the data processing hardware 610 ( Figure 6 The data processing hardware 610 and memory hardware 620 can reside on [the server / system]. Figure 1 Each of the computing devices 600 ( Figure 6 On the corresponding user device 102 and / or remote computing device 201.

[0049] At operation 502, method 500 includes receiving a reference utterance 302 and an input text sequence 304. The reference utterance 302 includes multiple terms spoken by a reference speaker. The input text sequence 304 includes a corresponding transcription 306 for each of the multiple terms spoken by the reference speaker. At operation 504, method 500 includes obtaining a speaker embedding 308 characterizing the speaker features of the reference speaker who spoke the multiple terms. At operation 506, method 500 includes generating a replacement input text sequence 314 by replacing the corresponding transcription 306 of a corresponding term among the multiple terms with a replacement transcription 316 corresponding to a different term not included in the reference utterance 302. At operation 508, method 500 includes generating resynthesized speech 134 based on the reference speaker's voice using a TTS model 130 conditioned on the reference utterance 302 and the speaker embedding 308.

[0050] Figure 6 This is a schematic diagram of an example computing device 600 that can be used to implement the systems and methods described in this document. The computing device 600 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown herein, their connections and relationships, and their functions are intended to be exemplary only and are not intended to limit the implementations of the invention described and / or claimed in this document.

[0051] Computing device 600 includes a processor 610, a memory 620, a storage device 630, a high-speed interface / controller 640 connected to the memory 620 and a high-speed expansion port 650, and a low-speed interface / controller 660 connected to a low-speed bus 670 and the storage device 630. Each of the components 610, 620, 630, 640, 650, and 660 is interconnected using various buses and may be mounted on a common motherboard or otherwise. The processor 610 can process instructions for execution within the computing device 600, including instructions stored in the memory 620 or the storage device 630, to display graphical information of a graphical user interface (GUI) on an external input / output device such as a display 680 coupled to the high-speed interface 640. In other implementations, multiple processors and / or multiple buses, as well as multiple memories and various types of memory, may be used as appropriate. Furthermore, multiple computing devices 600 may be connected, with each device providing a portion of the necessary operation (e.g., as a server library, blade server group, or multiprocessor system).

[0052] Memory 620 stores information non-temporarily within computing device 600. Memory 620 may be a computer-readable medium, a volatile memory cell, or a non-volatile memory cell. Non-temporary memory 620 may be a physical means for temporarily or permanently storing programs (e.g., instruction sequences) or data (e.g., program state information) for use by computing device 600. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., commonly used in firmware, such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase-change memory (PCM), and magnetic disks or magnetic tapes.

[0053] Storage device 630 provides mass storage for computing device 600. In some implementations, storage device 630 is a computer-readable medium. In various implementations, storage device 630 may be a floppy disk device, hard disk device, optical disk device, magnetic tape device, flash memory, or other similar solid-state storage device or array of devices, including devices arranged in a storage area network or other configuration. In additional implementations, a computer program product is tangibly embodied in an information carrier. The computer program product contains instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer or machine-readable medium, such as memory 620, storage device 630, or memory on processor 610.

[0054] High-speed controller 640 manages bandwidth-intensive operations of computing device 600, while low-speed controller 660 manages lower bandwidth-intensive operations. This allocation of responsibilities is merely exemplary. In some implementations, high-speed controller 640 (e.g., via a graphics processor or accelerator) is coupled to memory 620, display 680, and high-speed expansion port 650, which can accept various expansion cards (not shown). In some implementations, low-speed controller 660 is coupled to storage device 630 and low-speed expansion port 690. Low-speed expansion port 690, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, Wireless Ethernet), may be coupled to one or more input / output devices, such as keyboards, pointing devices, scanners, or networking devices, such as switches or routers, for example, via a network adapter.

[0055] The computing device 600 can be implemented in a variety of different forms, as shown in the figure. For example, it can be implemented as a standard server 600a or multiple times in a group of such servers 600a, as a laptop computer 600b, or as part of a rack server system 600c.

[0056] Various implementations of the systems and techniques described herein can be implemented in digital electronic and / or optical circuit systems, integrated circuit systems, specially designed ASICs (Application-Specific Integrated Circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system, which includes at least one programmable processor, which may be dedicated or general-purpose and is coupled to receive data and instructions from a storage system, at least one input device, and at least one output device, and to transfer data and instructions to the storage system, at least one input device, and at least one output device.

[0057] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented using high-level procedural and / or object-oriented programming languages ​​and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0058] The processes and logic flows described in this specification can be executed by one or more programmable processors, also known as data processing hardware, which execute one or more computer programs to perform functions by manipulating input data and generating output. The processes and logic flows can also be executed by special-purpose logic circuit systems, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). For example, processors suitable for executing computer programs include both general-purpose microprocessors and special-purpose microprocessors, as well as any one or more processors of any type of digital computer. Typically, the processor receives instructions and data from read-only memory or random access memory, or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or operatively coupled to receive data from or transfer data to said mass storage device, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. The processor and memory may be supplemented by or incorporated into a dedicated logic circuit system.

[0059] To provide interaction with the user, one or more aspects of this disclosure can be implemented on a computer having a display device for displaying information to the user (e.g., a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touchscreen) and possibly a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback, such as visual, auditory, or tactile feedback; and input from the user can be received in any form, including sound, speech, or tactile input. Additionally, the computer can interact with the user by sending documents to and receiving documents from a device used by the user; for example, by sending a webpage to a web browser in response to a request received from a web browser on the user's client device.

[0060] Various implementations have been described. However, it should be understood that various modifications can be made without departing from the spirit and scope of this disclosure. Therefore, other implementations are within the scope of the appended claims.

Claims

1. A computer-implemented method (500) executed on data processing hardware (610), the computer-implemented method causing the data processing hardware (610) to perform operations, the operations including: Receive a reference utterance (302) and an input text sequence (304), the reference utterance (302) comprising a plurality of words spoken by a reference speaker, and the input text sequence (304) comprising a corresponding transcription (306) of each of the plurality of words spoken by the reference speaker; Obtain speaker embeddings (308) that represent speaker characteristics of the reference speaker who utters the plurality of terms; A replacement input text sequence (314) is generated by replacing the corresponding transcription (306) of a specific term among the plurality of terms with a replacement transcription (316) corresponding to a different term not included in the reference discourse (302); and Using a text-to-speech (TTS) model (130) conditioned on the reference utterance (302) and the speaker embedding (308), a resynthesized speech (134) is generated based on the replaced input text sequence (314) using the voice of the reference speaker.

2. The computer-implemented method (500) according to claim 1, wherein the operation further comprises: Receive audio signals including one or more other words spoken by different speakers (301); The speech cloning model (340) is used to generate a speech cloning reference utterance (342) based on the reference utterance (302) and the audio signal (301), the speech cloning reference utterance (342) comprising synthesized speech corresponding to the different speakers’ speech in the input text sequence (304); as well as Using the TTS model (130) further conditioned on the reference utterance (342) of the voice clone, a resynthesized speech (136) of the voice clone is generated based on the replacement input text sequence (314), the resynthesized speech (136) of the voice clone comprising synthesized speech of the voice of the different speaker corresponding to the replacement input text sequence (314).

3. The computer-implemented method (500) according to claim 1 or 2, wherein the corresponding term among the plurality of terms includes a first hot word, and the different terms include a second hot word that is different from the first hot word.

4. The computer-implemented method (500) according to claim 3, wherein the operation further includes training a hot word model on the resynthesized speech (134).

5. The computer-implemented method (500) according to any one of claims 1 to 4, wherein the operation further includes training a speech recognition model (200) on the reference utterance (302) paired with the input text sequence (304) and the resynthesized speech (134) paired with the replacement input text sequence (314).

6. The computer-implemented method (500) according to any one of claims 1 to 5, wherein the different terms include speech disfluency terms.

7. The computer-implemented method (500) according to any one of claims 1 to 6, wherein the different terms are sampled from a speech domain different from the speech domain associated with the reference utterance (302).

8. The computer-implemented method (500) according to any one of claims 1 to 7, wherein the operation further comprises post-processing the resynthesized speech (134) based on the reference utterance (302) to preserve audio from the reference utterance (302) for terms included in the resynthesized speech (134) and the reference utterance (302).

9. The computer-implemented method (500) according to claim 8, wherein post-processing of the resynthesized speech (134) includes at least one of the following: Cross fade-in / fade-out; Force alignment; or Dynamic time regularization.

10. The computer-implemented method (500) according to any one of claims 1 to 9, wherein the operation further comprises: Modify the substitution transcription (316) of the different terms from the substitution input text sequence (314); as well as An updated replacement input text sequence (334) is generated by replacing the replacement transcript (316) with a modified replacement transcript (336); and Using the TTS model (130) conditioned on the reference utterance (302) and the speaker embedding (308), additional resynthesized speech (134) is generated based on the updated replacement input text sequence (334) with the voice of the reference speaker.

11. The computer-implemented method (500) of claim 10, wherein the substitution transcription (316) of the different terms from the substitution input text sequence (314) comprises at least one of the following: Modify the syllables of the replacement transcription (316); Modify the punctuation of the replacement transcription (316); or The speech disfluency is inserted into the replacement transcription (316).

12. The computer-implemented method (500) according to any one of claims 1 to 11, wherein the operation further comprises: Generate an enhanced speaker embedding (338) that enhances at least one speaker characteristic among the speaker characteristics of the reference speaker uttering the plurality of terms; as well as Additional resynthesized speech (134) is generated using the TTS model (130) further conditioned on the enhanced speaker embedding (338), the additional resynthesized speech comprising synthesized speech in the form of the speech of the reference speaker corresponding to the replacement input text sequence (314), the reference speaker having at least one of the enhanced speaker features.

13. A system (100) comprising: Data processing hardware (610); as well as A memory hardware (620) communicating with the data processing hardware (610), the memory hardware (620) storing instructions that, when executed on the data processing hardware (610), cause the data processing hardware (610) to perform operations, the operations including: Receive a reference utterance (302) and an input text sequence (304), the reference utterance (302) comprising a plurality of words spoken by a reference speaker, and the input text sequence (304) comprising a corresponding transcription (306) of each of the plurality of words spoken by the reference speaker; Obtain speaker embeddings (308) that represent speaker characteristics of the reference speaker who utters the plurality of terms; A replacement input text sequence (314) is generated by replacing the corresponding transcription (306) of a specific term among the plurality of terms with a replacement transcription (316) corresponding to a different term not included in the reference discourse (302); and Using a text-to-speech (TTS) model (130) conditioned on the reference utterance (302) and the speaker embedding (308), a resynthesized speech (134) is generated based on the replaced input text sequence (314) using the voice of the reference speaker.

14. The system (100) of claim 13, wherein the operation further comprises: Receive audio signals including one or more other words spoken by different speakers (301); The speech cloning model (340) is used to generate a speech cloning reference utterance (342) based on the reference utterance (302) and the audio signal (301), the speech cloning reference utterance (342) comprising synthesized speech corresponding to the different speakers’ speech in the input text sequence (304); as well as Using the TTS model (130) further conditioned on the reference utterance (342) of the voice clone, a resynthesized speech (136) of the voice clone is generated based on the replacement input text sequence (314), the resynthesized speech (136) of the voice clone comprising synthesized speech of the voice of the different speaker corresponding to the replacement input text sequence (314).

15. The system (100) according to claim 13 or 14, wherein the corresponding term of the plurality of terms includes a first hot word, and the different terms include a second hot word that is different from the first hot word.

16. The system (100) of claim 15, wherein the operation further comprises training a hot word model on the resynthesized speech (134).

17. The system (100) according to any one of claims 13 to 16, wherein the operation further comprises training a speech recognition model (200) on the reference utterance (302) paired with the input text sequence (304) and the resynthesized speech (134) paired with the alternative input text sequence (314).

18. The system (100) according to any one of claims 13 to 17, wherein the different terms include speech disfluency terms.

19. The system (100) according to any one of claims 13 to 18, wherein the different terms are sampled from a speech domain different from the speech domain associated with the reference utterance (302).

20. The system (100) according to any one of claims 13 to 19, wherein the operation further comprises post-processing the resynthesized speech (134) based on the reference utterance (302) to retain audio from the reference utterance (302) for terms included in the resynthesized speech (134) and the reference utterance (302).

21. The system (100) of claim 20, wherein post-processing of the resynthesized speech (134) comprises at least one of the following: Cross fade-in / fade-out; Force alignment; or Dynamic time regularization.

22. The system (100) according to any one of claims 13 to 21, wherein the operation further comprises: Modify the substitution transcription (316) of the different terms from the substitution input text sequence (314); as well as An updated replacement input text sequence (334) is generated by replacing the replacement transcript (316) with a modified replacement transcript (336); and Using the TTS model (130) conditioned on the reference utterance (302) and the speaker embedding (308), additional resynthesized speech (134) is generated based on the updated replacement input text sequence (334) with the voice of the reference speaker.

23. The system (100) of claim 22, wherein the substitution transcription (316) of the different terms from the substitution input text sequence (314) comprises at least one of the following: Modify the syllables of the replacement transcription (316); Modify the punctuation of the replacement transcription (316); or The speech disfluency is inserted into the replacement transcription (316).

24. The system (100) according to any one of claims 13 to 23, wherein the operation further comprises: Generate an enhanced speaker embedding (338) that enhances at least one speaker characteristic among the speaker characteristics of the reference speaker uttering the plurality of terms; as well as Additional resynthesized speech (134) is generated using the TTS model (130) further conditioned on the enhanced speaker embedding (338), the additional resynthesized speech comprising synthesized speech in the form of the speech of the reference speaker corresponding to the replacement input text sequence (314), the reference speaker having at least one of the enhanced speaker features.