Generating speech of a given style using machine learning models

The system addresses inefficiencies in generating speech by using a language model neural network to select audio prompts based on natural language descriptions, enabling flexible and resource-efficient speech generation with improved training data variety.

WO2026060285A1PCT designated stage Publication Date: 2026-03-19GDM HOLDING LLC
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Conventional systems for generating speech in a target style require an audio signal as input, which can be difficult and inefficient for users, and existing methods for specifying a target style in natural language are cumbersome and resource-intensive.

Method used

A system that uses a language model neural network to identify an audio prompt compatible with a target style described in natural language, combined with a speech processing model to generate an output audio signal, allowing flexible control over speech style and speaker identity without requiring extensive storage or computing resources.

Benefits of technology

Enables efficient generation of speech in a specified target style and speaker voice, reducing the need for large audio databases and computing resources, while improving training data variety and model performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025046220_19032026_PF_FP_ABST
    Figure US2025046220_19032026_PF_FP_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for generating an audio signal representing speech spoken in a given style. One of the methods includes maintaining a plurality of audio prompts, wherein each audio prompt represents speech spoken by a corresponding speaker in a corresponding style; obtaining a transcript that comprises text to be spoken in speech represented by an output audio signal; obtaining a style description that comprises text specifying a target style for the speech represented by the output audio signal; providing an input comprising the style description and data characterizing the plurality of audio prompts to a language model neural network to generate a network output identifying a subset of the plurality of audio prompts that are compatible with the target style; and providing i) the transcript and ii) one or more of the identified audio prompts as input to a speech processing model to generate the output audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Attorney Docket No.: 45288-0457WO1

[0002] GENERATING SPEECH OF A GIVEN STYLE USING MACHINE LEARNING

[0003] MODELS

[0004] CROSS-REFERENCE TO RELATED APPLICATIONS

[0005] This application claims priority to U.S. Provisional Application No. 63 / 694,151 filed on September 12, 2024. The disclosure of the prior application is considered part of and is incorporated by reference in the disclosure of this application.

[0006] BACKGROUND

[0007] This specification relates to generating audio using machine learning models.

[0008] Machine learning models receive an input and generate an output, e.g., a predicted output, based on the received input. Some machine learning models are parametric models and generate the output based on the received input and on values of the parameters of the model.

[0009] Some machine learning models are deep models that employ multiple layers of models to generate an output for a received input. For example, a deep neural network is a deep machine learning model that includes an output layer and one or more hidden layers that each apply a non-linear transformation to a received input to generate an output.

[0010] SUMMARY

[0011] This specification generally describes a system implemented as one or more computer programs on one or more computers in one or more locations that generates an audio signal representing speech spoken in a given style.

[0012] According to one aspect there is provided a computer-implemented method comprising: maintaining a plurality of audio prompts, wherein each audio prompt represents speech spoken by a corresponding speaker in a corresponding style; obtaining a transcript that comprises text to be spoken in speech represented by an output audio signal; obtaining a style description that comprises text specifying a target style for the speech represented by the output audio signal; providing an input comprising the style description and data characterizing the plurality of audio prompts as input to a language model neural network to generate a network output identifying a subset of the plurality of audio prompts that are compatible with the target style; and providing i) Attorney Docket No.: 45288-0457WO1 the transcript and ii) one or more of the identified audio prompts as input to a speech processing model to generate the output audio signal.

[0013] In some implementations, the method further comprises maintaining, for each audio prompt, a corresponding prompt style description, wherein the corresponding prompt style description is a description of the corresponding style of the audio prompt.

[0014] In some implementations, the data characterizing the plurality of audio prompts comprises the corresponding prompt style descriptions for the plurality of audio prompts.

[0015] In some implementations, the data characterizing the plurality of audio prompts comprises the plurality of audio prompts, and wherein the language model neural network comprises a multimodal language model neural network.

[0016] In some implementations, the input further comprises a respective identifier for each of the plurality of audio prompts.

[0017] In some implementations, the network output comprises the respective identifier for each of the subset of the plurality of audio prompts that are compatible with the target style.

[0018] In some implementations, the plurality of audio prompts comprise one or more initial audio prompts and one or more synthetic audio prompts, and wherein the one or more synthetic audio prompts are generated by: for each of one or more of the initial audio prompts: providing at least data representing the initial audio prompt and a target speaker prompt representing speech of a target speaker as input to a voice conversion model to generate a synthetic audio prompt, wherein the synthetic audio prompt represents speech of the initial audio prompt, spoken by the target speaker.

[0019] In some implementations, the method further comprises: maintaining, for each audio prompt, a corresponding prompt style description, wherein the corresponding prompt style description is a description of the corresponding style of the audio prompt; and for each of the one or more initial audio prompts: assigning, as the corresponding prompt style description for the synthetic audio prompt, the corresponding style description of the initial audio prompt.

[0020] In some implementations, the target speaker prompt is derived from an initial audio prompt that represents speech spoken by a different corresponding speaker than the corresponding speaker of the initial audio prompt. Attorney Docket No.: 45288-0457WO1

[0021] In some implementations, the data representing the initial audio prompt comprises any one or more of: an audio signal, energy features of the audio signal, or pitch features of the audio signal.

[0022] In some implementations, providing at least data representing the initial audio prompt and a target speaker prompt as input to a voice conversion model to generate a synthetic audio prompt comprises providing the data representing the initial audio prompt, the target speaker prompt, and an initial transcript of speech represented by the initial audio prompt as input to the voice conversion model to generate the synthetic audio prompt.

[0023] In some implementations, the method further comprises: providing at least data representing the output audio signal and a target speaker prompt as input to a voice conversion model to generate a converted audio signal.

[0024] In some implementations, the data representing the output audio signal comprises any one or more of: the output audio signal, energy features of the output audio signal, or pitch features of the output audio signal.

[0025] In some implementations, providing at least data representing the output audio signal and a target speaker prompt as input to a voice conversion model to generate a converted audio signal comprises providing data representing the output audio signal, the target speaker prompt, and the transcript as input to the voice conversion model to generate the converted audio signal.

[0026] In some implementations, the method further comprises: generating a training example comprising i) the transcript, ii) the style description, iii) the target speaker prompt, and iv) the converted audio signal; and including the training example in a first set of training data.

[0027] In some implementations, the method further comprises training a speech generation model on the first set of training data.

[0028] In some implementations, the method further comprises: generating a training example comprising i) the transcript, ii) the style description, and iii) the output audio signal; and including the training example in a second set of training data.

[0029] In some implementations, the method further comprises training a speech generation model on the second set of training data.

[0030] Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods. Attorney Docket No.: 45288-0457WO1

[0031] Particular embodiments of the subject matter described in this specification can be implemented so as to realize one or more of the following advantages.

[0032] The system described in this specification can perform text to speech (otherwise known as text to speech conversion) according to a specified target style. For example, the system can generate audio signals representing speech spoken in the target style, where the target style is described in natural language.

[0033] Some conventional systems for generating speech of a target style require providing a text transcript and an audio signal as input to an audio generative model to generate speech in the style of the audio signal. In order for a user to generate speech in a target style, the user must provide an audio signal of the target style, which can be difficult and inefficient for the user for many reasons, e.g., because generating or locating an audio signal that directly matches the desired style is difficult and because listening to many different audio signals to determine which audio signals have a given style is time-consuming.

[0034] To address these issues and allow users to specify a desired target style in natural language, the system described in this specification automates the selection of an audio prompt representing speech in a target style using a language model neural network. As a result, the system can allow users to specify the target style by describing the target style in natural language text.

[0035] More specifically, the system can use a language model neural network to select an appropriate audio prompt for the target style. For example, the system can maintain multiple audio prompts. In some examples, the system can also maintain corresponding prompt style descriptions for the audio prompts. Rather than searching for or requiring an exact match between corresponding prompt style descriptions, or training a machine learning model to select an audio prompt from a set of audio prompts that matches the target style, both of which would require maintaining a large number of audio prompts and corresponding prompt style descriptions, the system can use the language model neural network to identify an audio prompt that is compatible with the target style. The system can then generate an output audio signal given the identified audio prompt and the transcript using a speech processing model. Thus the system can provide for flexible control of the style of synthesized speech, with only a small amount of annotated data. Attorney Docket No.: 45288-0457WO1

[0036] In some examples, the system described in this specification can perform text to speech according to a specified target style and spoken by a particular speaker, allowing for the enforcement of a desired voice identity and for further control over synthesized speech. For example, the system can receive a target style description and a target speaker prompt representing speech by the particular speaker. The system can generate an output audio signal using a speech processing model that represents speech in the target style as described above. The system can then provide data representing the output audio signal and the target speaker prompt as input to a voice conversion model to generate a converted audio signal. The system can thus provide for flexibility and control in the generation of audio signals representing speech. Furthermore, the system can perform text to speech according to a specified target style and spoken by a particular speaker without having to store target speaker prompts for the particular speaker, reducing the amount of computer memory required for storage of target speaker prompts for different speakers.

[0037] In some examples, the system can perform text to speech according to a specified target style and spoken by the particular speaker using fewer computing resources at runtime. For example, the system can generate an output audio signal given an identified audio prompt and a transcript using a speech processing model. The identified audio prompt can represent speech by the particular speaker. In these examples, the system generates and maintains synthetic audio prompts for the particular speaker of different styles. In some examples, the system maintains corresponding prompt style descriptions for the synthetic audio prompts. By preparing audio prompts for the particular speaker in a diverse range of styles prior to generating the output audio signal, the system does not have to perform voice conversion at runtime, reducing the amount of computing resources needed to perform text to speech according to a specified target style and spoken by the particular speaker.

[0038] As a particular example, the system can perform text to speech for dialogue agents and speech-based dialogue interfaces. By adapting the style of the responses as described in this specification, the system can generate responses for users that lead to more natural and engaging interactions. For example, the particular speaker can be a speaker associated with the dialogue agent. The system can generate speech representing a response from the dialogue agent, in the voice of the dialogue agent, and in a style that is relevant to the context of the conversation with the user. Attorney Docket No.: 45288-0457WO1

[0039] In some examples, the system described in this specification can generate training data for training a speech generation model. Training a speech generation model to perform tasks such as text to speech according to a given style description requires a large amount of training data. However, training data that includes input style descriptions and output speech in the style of the input style descriptions is difficult to obtain. By using a language model neural network and a speech processing model to generate a set of training data with audio signals representing speech of a particular style, the system increases the number of training examples available for training, resulting in improved training and performance of the speech generation model. Training the speech generation model on a larger number and greater variation of training examples allows the speech generation model to generalize better to previously unseen inputs at inference.

[0040] For example, the system can include a transcript, a style description, and an output audio signal generated using the speech processing model that represents speech of the transcript spoken in the style of the style description in a training example. The system can train the speech generation model to generate audio representing speech of a given transcript spoken in the style of a given style description directly, without requiring maintaining multiple audio prompts and corresponding prompt style descriptions, and without requiring the use of a language model neural network and a speech processing model. As another example, the system can include a transcript, a style description, a target speaker prompt, and the converted audio signal generated using the voice conversion model that represents speech of the transcript, spoken in the style of the style description, and spoken by the speaker of the target speaker prompt, in a training example. The system can train the speech generation model to generate audio representing speech of a given transcript, spoken in the style of a given style description, and spoken by a particular speaker, directly, without requiring maintaining multiple audio prompts and corresponding prompt style descriptions, and without requiring the use of a language model neural network, a speech processing model, and a voice conversion model. Thus the system can use the speech generation model to perform text to speech without requiring the use of multiple machine learning models, reducing the amount of computing resources used to maintain and run the multiple machine learning models. Furthermore, in some examples the speech generation model can learn to combine aspects from different prompts during training, Attorney Docket No.: 45288-0457WO1 rather than generating audio representing speech according to a single style prompt or a single target speaker prompt.

[0041] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.

[0042] BRIEF DESCRIPTION OF THE DRAWINGS

[0043] FIG. 1A is a block diagram of an example audio generation system.

[0044] FIG. IB is a block diagram of another example audio generation system.

[0045] FIG. 2 shows an example process for generating synthetic audio prompts.

[0046] FIG. 3 is a block diagram of another example audio generation system.

[0047] FIG. 4 is a flow diagram of an example process for generating audio.

[0048] Like reference numbers and designations in the various drawings indicate like elements.

[0049] DETAILED DESCRIPTION

[0050] FIG. 1A is a block diagram of an example audio generation system 100. The system 100 is an example of a system implemented as one or more computer programs on one or more computers in one or more locations in which the systems, components, and techniques described below are implemented.

[0051] The audio generation system 100 generates a prediction of an output audio signal 150 given a style description 102 and a transcript 104. The output audio signal 150 represents speech specified by the transcript 104, spoken in a style that is described by the style description 102.

[0052] Generally, the output audio signal is an output audio example that includes a sample of an audio wave at each of a sequence of output time steps that span a specified time window. For example, the output time steps can be arranged at regular intervals within the specified time window.

[0053] The audio sample at a given output time step can be an amplitude value of the audio wave or an amplitude value that has been compressed, companded, or both. For example, the audio sample can be a raw amplitude value or a mu-law companded representation of the amplitude value. Attorney Docket No.: 45288-0457WO1

[0054] The system 100 maintains multiple audio prompts 110, e.g., in a database. Each audio prompt represents speech spoken by a corresponding speaker in a corresponding style. The corresponding style of the speech refers to how the corresponding speaker says the content of the speech, as opposed to the identity of the speaker or the content of the speech. For example, the corresponding speaker can speak in different emotions, moods, or tones, or according to different settings, audiences, or situations. In some examples, each audio prompt is an audio signal that represents speech. In some examples, each audio prompt includes prosody features of an audio signal that represents speech, such as energy features and pitch features.

[0055] Generally, an audio signal is an audio example that includes a sample of an audio wave at each of a sequence of time steps that span a specified time window. For example, the time steps can be arranged at regular intervals within the specified time window. The audio sample at a given time step can be an amplitude value of the audio wave or an amplitude value that has been compressed, companded, or both. For example, the audio sample can be a raw amplitude value or a mu-law companded representation of the amplitude value.

[0056] In some examples, one or more of the audio prompts 110 can be synthetic audio prompts. For example, the system 100 can generate synthetic audio prompts for inclusion in the audio prompts 110 using a voice conversion model. Generating synthetic audio prompts is described in further detail below with reference to FIG. 2.

[0057] In some examples, the system 100 also maintains corresponding prompt style descriptions 120. Each corresponding prompt style description is a description of the corresponding style for one of the audio prompts 110. Each corresponding prompt style description can describe the style in which the speech of the audio prompt is spoken. For example, each corresponding prompt style description can describe properties of the style of the speech of the audio prompt, such as emotion, mood, tone, setting, audience, or situation. Each corresponding prompt style description can include a natural language description, or textual tags, conveying the corresponding style. In some examples, the corresponding prompt style descriptions can be obtained manually, e.g., from a user. In some examples, the corresponding prompt style descriptions can have been generated using a machine learning model configured to output description of the style of an input audio prompt, such as a classifier or a multimodal model. For example, the system can provide each audio prompt as input to the machine learning model to generate the corresponding prompt style description. As a particular example, the Attorney Docket No.: 45288-0457WO1 machine learning model can be a language model neural network. Some example corresponding prompt style descriptions are described below with reference to FIG. IB.

[0058] In some examples, the system 100 also maintains corresponding speaker identifiers. Each speaker identifier identifies a corresponding speaker for one of the audio prompts 110.

[0059] To generate the output audio signal 150, the system 100 receives the style description 102 and the transcript 104. The style description 102 includes a natural language text description of a target style for speech represented by the output audio signal 150. The style description 102 can describe properties of speaking style such as emotion, mood, tone, setting, audience, or situation. In the example of FIG. 1 A, the style description includes the text “soothing.” Other example style descriptions include “with a soothing voice like on a meditative podcast,” “like a person talking in a bar,” “as someone very irritated by what’s happening around them,” and “as an anchorman reading the news.”

[0060] The transcript 104 includes text to be spoken in speech represented by the output audio signal 150. In the example of FIG. 1 A, the transcript 104 includes the text “Soften your gaze, allowing the edges of your vision to blur slightly.”

[0061] In some examples, the system 100 receives the style description 102, the transcript 104, or both, from a user. For example, the system 100 can receive the style description 102 or the transcript 104 from a user through a user interface of a user device.

[0062] As a particular example, the system 100 can receive a user query from the user that includes the style description 102 and the transcript 104. In the example of FIG. 1A, the user query can include the text ‘Say this, with a soothing voice: “Soften your gaze, allowing the edges of your vision to blur slightly.’” The system 100 can obtain the style description 102 and the transcript 104 from the user query.

[0063] The system 100 processes the style description 102 to identify a subset of the audio prompts 110 that are compatible with the target style. For example, as described in more detail with reference to FIG. IB, the system 100 can use a language model neural network to identify the subset. In the example of FIG. 1A, the subset includes audio prompts that have corresponding prompt styles of “calm.”

[0064] The system 100 processes one or more audio prompts from the subset and the transcript 104 to generate the output audio signal 150. For example, as described with reference to FIG. IB, the system 100 can use a speech processing model to generate the output audio signal 150. Attorney Docket No.: 45288-0457WO1

[0065] The audio signal 150 represents speech specified by the transcript 104, spoken in a style that is described by the style description 102. In the example of FIG. 1A, the output audio signal 150 represents the speech “Soften your gaze, allowing the edges of your vision to blur slightly,” spoken in a “soothing” style.

[0066] In some examples, the output audio signal 150 represents speech spoken by one of the corresponding speakers of the subset, i.e., one of the corresponding speakers of the one or more audio prompts processed by the speech processing model. In some examples, the system generates a converted audio signal that represents speech spoken by a target speaker from the output audio signal 150, as described below with reference to FIG. 3.

[0067] In some examples, the system 100 provides the output audio signal 150 for presentation to the user. For example, the system 100 can provide data representing the audio signal 150 to the user device and cause playback of the audio signal 150. Alternatively or in addition, the system 100 may provide data representing the audio signal 150 for storage.

[0068] FIG. IB is a block diagram of the example audio generation system 100 described with reference to FIG. 1 A. In particular, the audio generation system 100 generates a prediction of the output audio signal 150 given the style description 102 and the transcript 104 using a language model neural network 130 and a speech processing model 140.

[0069] The system 100 receives the style description 102 and the transcript 104 as described above with reference to FIG. 1A.

[0070] The system 100 processes the style description 102 using the language model neural network 130 to generate a network output 132 identifying a subset of the audio prompts 110 that are compatible with the target style. An example language model neural network 130 is described in more detail further below.

[0071] To generate the network output 132, the system 100 can generate an input 112 for the language model neural network 130 that includes the style description 102 and data characterizing the audio prompts.

[0072] In some examples, the data characterizing the audio prompts can include the audio prompts 110. For example, the input 112 can include an instruction to identify the audio prompts that have compatible styles or are the most appropriate with the style description 102. In some examples, the input 112 can also include a respective identifier for each of the audio prompts Attorney Docket No.: 45288-0457WO1

[0073] In some examples, the data characterizing the audio prompts can include the corresponding prompt style descriptions 120. For example, the input 112 can include an instruction to identify the audio prompts that have compatible corresponding prompt style descriptions with the style description 102. In some examples, the input 112 can also include a respective identifier for each of the audio prompts 110.

[0074] In the example of FIG. IB, the input 112 can include:

[0075] Given this table, containing annotated audio prompts, id, speaker, style 0, A, calm

[0076] 1 , A, neutral

[0077] 2, A, apologetic

[0078] 3, B, calm

[0079] 4, B, Sport commentary, excited

[0080] 5, C, News, neutral

[0081] 6, C, firm

[0082] 7, D, Sport commentary, calm

[0083] 9, E, Someone speaking from afar

[0084] Return the ids of the audio prompts compatible with this query: “soothing”

[0085] In some examples, the system 100 also obtains an input identifying the target speaker. For example, the input can include an identifier for the speaker B. In these examples, the system 100 can generate an input 112 for the language model neural network 130 that also includes the identifier for the target speaker. For example, the input 112 can include an instruction to identify the audio prompts that have compatible corresponding prompt style descriptions with the style description 102 and are spoken by the target speaker.

[0086] The system 100 can provide the input 112 to the language model neural network 130 to generate the network output 132. In the example of FIG. IB, if the input 112 includes corresponding style descriptions, the network output 132 can identify the subset of the audio prompts 110 that are compatible with the target style as audio prompts for which the corresponding style description is “calm.” In examples where the input 112 includes audio prompts, the network output 132 can identify the subset of the audio prompts 110 that have styles that are compatible with “soothing.”

[0087] In some examples where the input 112 also includes respective identifiers for the audio prompts 110, the network output 132 can include a respective identifier for each audio prompt of Attorney Docket No.: 45288-0457WO1 the subset of the audio prompts 1 10 that are compatible with the target style. In the example of FIG. IB, the network output 132 can identify the audio prompts with identifiers 0 and 3. For example, the network output 132 can include the text “0, 3.”

[0088] The system 100 can select one or more of the identified audio prompts. For example, if the subset of identified audio prompts includes more than one identified audio prompt, the system 100 can randomly select an audio prompt. In the example of FIG. IB, the system can select the audio prompt with the identifier 3.

[0089] In some examples, if the subset of identified audio prompts includes more than one identified audio prompt, the system 100 can select more than one audio prompt. In the example of FIG. IB, the system 100 can select the audio prompts with identifiers 0 and 3. As another example, if the system 100 maintains corresponding style descriptions 120, the system 100 can select more than one audio prompt based on the corresponding prompt styles for the audio prompts. For example, the system can determine, using the language model neural network 130, that a combination of the corresponding prompt style descriptions is more compatible with the style description 102 than the individual corresponding prompt style descriptions. The system can select the audio prompts for the corresponding prompt style descriptions.

[0090] As another example, the system 100 can select more than one audio prompt in response to determining, using the language model neural network 130, that a combination of the audio prompts is more compatible with the style description 102 than the individual audio prompts.

[0091] The system 100 can provide the transcript 104 and the selected audio prompt from the audio prompts 110 as input to the speech processing model 140 to generate the output audio signal 150. The output audio signal 150 represents speech specified by the transcript 104, spoken in a style of the selected audio prompt, and spoken by the speaker of the selected audio prompt.

[0092] In examples where the system 100 selects multiple audio prompts, the system 100 can provide the selected audio prompts as input to the speech processing model 140. In these examples, the output audio signal 150 represents speech specified by the transcript 104, spoken in a style that is a combination of the styles of the selected audio prompts, and spoken by one of the corresponding speakers of the selected audio prompts. For example, the speech processing model 140 can be trained to generate the output audio signal 150 representing speech spoken by one of the corresponding speakers of the selected audio prompts. Attorney Docket No.: 45288-0457WO1

[0093] The speech processing model 140 is configured to perform audio generation tasks such as text to speech. For example, the speech processing model 140 processes a text transcript and audio prompt to generate an output audio signal representing speech of the text transcript, spoken by the speaker of the audio prompt, and in the style of the audio prompt.

[0094] More generally, the speech processing model 140 can be implemented as any appropriate type of machine learning model that can perform a speech processing task. The speech processing model 140 can include any appropriate types of neural network layers (e.g., fully connected layers, message passing layers, convolutional layers, attention layers, recurrent layers, pooling layers, and so forth), in any appropriate number (e.g., 5 layers, or 10 layers, or 50 layers), and connected in any appropriate configuration (e.g., as a directed graph of layers).

[0095] As a particular example, the speech processing model 140 can be configured to generate an audio signal conditioned on multiple types of inputs. For example, the types of inputs can include input audio signals, features of input audio signals such as pitch features, energy features, or spectral features, semantic representations of input audio signals, embeddings of input audio signals, text inputs such as text transcripts, a text input representing a sequence of one or more dialogue turns, corresponding input audio signals representing speech of the one or more dialogue turns, and features of an input video. In some examples, the speech processing model 140 can derive one or more of the types of inputs, e.g., from an input audio signal or an input video.

[0096] As a particular example, the speech processing model 140 can be configured to generate the output audio signal 150 by processing an encoded representation derived from the transcript 104 and the one or more selected audio prompts using a token decoder neural network to generate a sequence of output tokens representing the output audio signal 150. The speech processing model 140 can include multiple encoders, e g., a corresponding encoder for each type of input. Each encoder is configured to generate a respective representation of the type of input.

[0097] The speech processing model 140 can process each input using the corresponding encoder to generate a respective representation for the input. For example, the speech processing model 140 can process the transcript 104 using the corresponding encoder to generate the respective representation for the transcript 104. The speech processing model 140 can process the one or more selected audio prompts using the corresponding encoder to generate the respective representation for the selected audio prompts. Attorney Docket No.: 45288-0457WO1

[0098] In examples where the system 100 provides multiple selected audio prompts as input to the speech processing model 140, the speech processing model 140 can concatenate the audio prompts, or features derived from the audio prompts, in the time dimension prior to processing the audio prompts, or features derived from the audio prompts, using the corresponding encoder.

[0099] In some examples where the audio prompts include embeddings of audio signals, the speech processing model 140 can process the one or more selected audio prompts using the corresponding encoder for embeddings of input audio signals. In some examples where the audio prompts include audio signals, the speech processing model 140 can generate the embeddings of the one or more selected audio prompts. For example, the speech processing model 140 can generate the embedding from one or more of the selected audio prompts using an encoder. As an example, the encoder can include an encoder neural network of a neural audio codec. As a particular example, the encoder can be a SoundStream encoder of the SoundStream neural audio codec.

[0100] An "embedding" can refer to an ordered collection of numerical values, e.g., a vector or matrix of numerical values. The speech processing model 140 processes the respective representations using a shared encoder to generate the encoded representation. In some examples, the shared encoder is configured to generate a combination, e.g., a concatenation, of the respective representations. In some examples, the shared encoder includes a shared encoder neural network that processes the concatenation to generate the encoded representation.

[0101] The speech processing model 140 processes the encoded representation using the token decoder neural network to generate a sequence of output tokens representing the output audio signal. Each of the output tokens can be selected from a vocabulary of output tokens. As an example, the token decoder neural network can have a Transformer-based architecture.

[0102] In particular, the token decoder neural network can be an auto-regressive neural network that auto-regressively generates the sequence of output tokens by generating each particular output token in the sequence conditioned on the encoded representation and a current input sequence that includes any tokens that precede the particular output token in the output sequence. The token decoder neural network can apply a cross-attention mechanism over the encoded representation and the current input sequence. Attorney Docket No.: 45288-0457WO1

[0103] In some examples, the output tokens in the vocabulary include semantic tokens. Each semantic token is selected from the vocabulary and represents semantic content of the output audio signal.

[0104] In some examples, the output tokens in the vocabulary include acoustic tokens. Each acoustic token is selected from the vocabulary and represents acoustic properties of the output audio signal. Examples of acoustic properties represented by the acoustic tokens can include reverberation, distortion, speaker identity, and background noise. Any appropriate set of acoustic tokens may be used. For example, an acoustic token can represent one of a plurality of code vectors in a codebook for a quantizer, e.g., a codebook for a vector quantizer included in a residual (i.e., multi-stage) vector quantizer (RVQ). For example, the set of acoustic tokens may be provided using the codebook of an audio codec such as a Soundstream neural audio codec.

[0105] Throughout this specification, a “residual vector quantizer” (RVQ) can refer to a multistage vector quantization technique that is based on a sequence of (residual) vector quantizers. A vector quantizer can quantize an input vector, e.g., by identifying a code vector from a codebook of code vectors associated with the vector quantizer, e.g., that has a smallest distance from the input vector, e.g., according to a distance metric (e.g., based on an LI norm). The residual vector quantizer can quantize an input vector (or “signal”) by iteratively quantizing the residual errors from previous quantization stages. Thus each stage in a residual vector quantizer encodes the difference (or residual) between the original signal and the reconstructed signal from the previous stage, thereby progressively refining the approximation of the original signal with each step.

[0106] In this example, the neural audio codec can include a hierarchy of multiple vector quantizers that each generate a respective acoustic token from a corresponding codebook of token vectors for the vector quantizer. The hierarchy includes one or more coarse vector quantizers at one or more first levels in the hierarchy and one or more fine vector quantizers at one or more last levels in the hierarchy. The output tokens can include, for each vector quantizer, a respective acoustic token selected from the codebook for the vector quantizer.

[0107] For example, the hierarchy can include Q vector quantizers arranged in the order of 1...Q’, (Q' + 1) ... Q. and the vector quantizers 1...Q' can be coarse vector quantizers, and the vector quantizers (Q' -I- 1) ... Q can be fine vector quantizers. The coarse vector quantizers generate coarse acoustic tokens, or acoustic tokens for coarse vector quantizers, that can Attorney Docket No.: 45288-0457WO1 represent acoustic properties such as speaker identity and recording conditions. The fine vector quantizers generate fine acoustic tokens, or acoustic tokens for fine vector quantizers, that can represent fine acoustic details. For example, fine acoustic tokens can be used to remove lossy compression artifacts in the coarse acoustic tokens.

[0108] In some examples, the output tokens in the vocabulary include acoustic and semantic tokens. The sequence of output tokens can thus include semantic tokens, acoustic tokens, or both. For example, the sequence of output tokens can include interleaved semantic tokens and acoustic tokens.

[0109] The speech processing model 140 can process the sequence of output tokens using an audio decoder neural network to generate the output audio signal 150.

[0110] In examples where the output tokens include acoustic tokens, the audio decoder neural network is configured to reconstruct an audio signal by processing acoustic tokens representing the audio signal. For example, the audio decoder neural network can include the decoder of the Soundstream neural audio codec, as described in Zeghidour, Neil, et al. "Soundstream: An end- to-end neural audio codec." IEEE / ACM Transactions on Audio, Speech, and Language Processing 30 (2021): 495-507.

[0111] In some examples where the output tokens include acoustic tokens and semantic tokens, the speech processing model 140 is configured to extract the acoustic tokens and provide the acoustic tokens as input to the audio decoder neural network. The audio decoder neural network is configured to reconstruct an audio signal by processing the acoustic tokens as described above.

[0112] In examples where the output tokens include semantic tokens, the speech processing model 140 is configured to generate acoustic tokens from the semantic tokens and provide the acoustic tokens as input to the audio decoder neural network. For example, the speech processing model 140 can use one or more generative neural networks to convert semantic tokens to acoustic tokens. Example generative neural networks for converting semantic tokens to acoustic tokens are described in Z. Borsos et al., "AudioLM: a Language Modeling Approach to Audio Generation," arXiv:2209.03143, which is hereby incorporated by reference in its entirety. The audio decoder neural network is configured to reconstruct an audio signal by processing the acoustic tokens as described above.

[0113] In some examples, one or more of the corresponding encoder neural networks can include pre-trained components such as text tokenizers, text encoders, audio tokenizers, audio encoders, Attorney Docket No.: 45288-0457WO1 etc. In some examples, an encoder neural network can be pre-trained on unlabeled training data based on optimizing a self-supervised or unsupervised loss function to generate embeddings. In some examples, the encoder neural network can be the encoder of an autoencoder neural network corresponding to the type of input.

[0114] The shared encoder neural network and the token decoder neural network can have been trained on multiple training examples that each include a training input and a corresponding target output. In some examples, each training input can include a set of inputs, or a set of representations of inputs. In some examples, the corresponding target output can include a corresponding ground-truth sequence of output tokens representing a target audio signal, or a target audio signal.

[0115] For example, the shared encoder neural network, the token decoder neural network, or both, can be trained using a machine learning training technique, e.g., a gradient descent with backpropagation training technique that uses a suitable optimizer, e.g., stochastic gradient descent, RMSprop, Adam optimizer, or Adafactor optimizer, to optimize an objective function for a next token prediction task. For example, the system can train the shared encoder neural network, the token decoder neural network, or both, to optimize an objective function that, for each training example, measures an error (e.g., cross-entropy error) between (i) the ground-truth sequence of output tokens specified by the training example and (ii) the output tokens generated by the token decoder neural network for the training input specified by the training example.

[0116] In some implementations, the system 100 can generate a training example for training a speech generation model. For example, the training example can include the transcript 104, the style description 102, and the output audio signal 150. The system 100 can include the training example in a set of training data.

[0117] A training system of the system 100 or another training system can train a speech generation model on the set of training data. The speech generation model can be configured to process one or more inputs in accordance with current values of parameters of the speech generation model to generate an output audio signal. For example, the speech generation model can be configured to receive a transcript and a style description to generate an output audio signal.

[0118] The speech generation model can have any appropriate architecture for performing a speech generation task. More generally, the speech generation model can be implemented as any Attorney Docket No.: 45288-0457WO1 appropriate type of machine learning model that can perform a speech generation task. The speech generation model can include any appropriate types of neural network layers (e.g., fully connected layers, message passing layers, convolutional layers, attention layers, recurrent layers, pooling layers, and so forth), in any appropriate number (e.g., 5 layers, or 10 layers, or 50 layers), and connected in any appropriate configuration (e.g., as a directed graph of layers).

[0119] For example, the speech generation model can include an autoregressive model, a diffusion model, or a generative adversarial network (GAN). As a particular example, the speech generation model can have a similar architecture as the speech processing model 140 described above.

[0120] The language model neural network 130 can have any appropriate neural network architecture that allows the model to map an input sequence of tokens from a vocabulary to an output sequence of tokens from the vocabulary. The output sequence of tokens can be decoded using one or more decoder neural networks to generate an output.

[0121] The language model neural network 130 can have any appropriate Transformer-based architecture, e.g., encoder-only Transformer architectures, encoder-decoder Transformer architectures, decoder-only Transformer architectures, other attention-based architectures, and so on. Examples of such Transformer-based neural network architectures include those described in Rohan Anil, et al. Palm 2 technical report. arXiv preprint arXiv:2305.10403 (2023), Devlin, Jacob, et al. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv: 1810.04805 (2018), and Gemini Team, et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805 (2023).

[0122] In some examples, the language model neural network 130 can include a multimodal language model neural network. For example, the input sequence can include tokens representing data of one or more modalities, and the output sequence can include tokens representing data of one or more modalities. As a particular example, the input sequence can include tokens representing text and audio, and the output sequence can include tokens representing text.

[0123] In general a Transformer-based architecture can be one which is characterized by having a succession of self-attention neural network layers. A self-attention neural network layer has an attention layer input for each element of the input and is configured to apply an attention Attorney Docket No.: 45288-0457WO1 mechanism over attention layer inputs to generate an attention layer output for each element of the input. There are many different attention mechanisms that may be used.

[0124] In some examples, the audio prompts are fixed. For example, the audio prompts can be generated or obtained prior to using the system 100 to generate the output audio signal 150 at inference. In some examples, the corresponding prompt style descriptions are also fixed. In these examples, because the audio prompts and corresponding prompt style descriptions are not updated during inference, the language model neural network 130 can precompute the prefill, e.g., keys and values for each self-attention network layer. Thus the system 100 can reduce the amount of computing time and resources used at inference.

[0125] The vocabulary of tokens can include any of a variety of tokens that represent text symbols or other symbols. For example, the vocabulary of tokens can include one or more of characters, sub-words, words, punctuation marks, numbers, or other symbols that appear in a corpus of natural language text and / or computer code.

[0126] Additionally, or alternatively, the vocabulary of tokens can include tokens that can represent data other than text. For example, the vocabulary of tokens can include image tokens that represent a discrete set of image patch embeddings of an image that can be generated by an image encoder neural network based on processing the image patches of the image. As another example, the vocabulary of tokens can include audio tokens that represent code vectors in a codebook of a quantizer, e.g., a residual vector quantizer.

[0127] For example, the input 112 can include the style description 102, the corresponding prompt style descriptions 120, and an instruction to identify the audio prompts that have compatible corresponding prompt style descriptions with the style description 102. The input sequence can include text tokens representing the input 112. The output sequence can include tokens representing text that identifies one or more audio prompts. As another example, the input 112 can include the style description 102, the audio prompts 110, and an instruction to identify the audio prompts that have compatible styles with the style description 102. The input sequence can include text tokens and audio tokens, e.g., generated for the audio prompts using the codebook of a quantizer, representing the input 112. The output sequence can include tokens representing text that identifies one or more audio prompts.

[0128] As an example, the language model neural network 130 can generate text sequences, i.e., each output sequence generated by the language model neural network 130 is a sequence of text Attorney Docket No.: 45288-0457WO1 tokens from a vocabulary of text tokens that includes, e.g., one or more of characters, sub-words, words, punctuation marks, numbers, or other symbols that appear in natural language text.

[0129] In some examples, the language model neural network 130 can have been trained to generate images or videos that each have multiple frames (where each frame is an image) by generating images as sequences of pixels. For example, an output sequence generated by the language model neural network 130 can include a sequence of color values for pixels in an image arranged according to a specified order. As another example, an output sequence generated by the language model neural network 130 can include a sequence of tokens that represent image patch embeddings of an image which can then be processed by a decoder neural network to generate the image (pixel values).

[0130] In particular, the language model neural network 130 can be an auto-regressive neural network that auto-regressively generates the output sequence of tokens by generating each particular token in the output sequence conditioned on a current input sequence that includes (i) the input sequence followed by (ii) any tokens that precede the particular token in the output sequence.

[0131] More specifically, to generate a particular token, the language model neural network 130 can process the current input sequence to generate a score distribution, e.g., a probability distribution, that assigns a respective score, e.g., a respective probability, to each token in the vocabulary of tokens. The language model neural network 130 can then select, as the particular token, a token from the vocabulary using the score distribution. For example, the language model neural network 130 can greedily select the highest-scoring token or can sample, e.g., using top-k sampling, nucleus sampling or another sampling technique, a token from the distribution.

[0132] As a particular example, the language model neural network 130 can be an autoregressive Transformer-based neural network that includes a plurality of layers that each apply a self-attention operation. The language model neural network 130 can have any of a variety of Transformer-based neural network architectures. Examples of such architectures include those described in J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al. Training compute-optimal large language models, arXiv preprint arXiv:2203.15556, 2022; J.W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, H. F. Song, J. Aslanides, S. Henderson, R. Ring, S. Young, E. Rutherford, T. Hennigan, J. Menick, A. Cassirer, R. Powell, G. van den Driessche, L. A. Hendricks, M. Rauh, Attorney Docket No.: 45288-0457WO1

[0133] P. Huang, A. Glaese, J. Welbl, S. Dathathri, S. Huang, J. Uesato, J. Mellor, I. Higgins, A. Creswell, N. McAleese, A.Wu, E. Eisen, S. M. Jayakumar, E. Buchatskaya, D. Budden, E. Sutherland, K. Simonyan, M. Paganini, L. Sifre, L. Martens, X. L. Li, A. Kuncoro, A. Nematzadeh, E. Gribovskaya, D. Donato, A. Lazaridou, A. Mensch, J. Lespiau, M. Tsimpoukelli, N. Grigorev, D. Fritz, T. Sottiaux, M. Pajarskas, T. Pohlen, Z. Gong, D. Toyama, C. de Masson d’Autume, Y. Li, T. Terzi, V. Mikulik, I. Babuschkin, A. Clark, D. de Las Casas, A. Guy, C. Jones, J. Bradbury, M. Johnson, B. A. Hechtman, L. Weidinger, I. Gabriel, W. S. Isaac, E. Lockhart, S. Osindero, L. Rimell, C. Dyer, O. Vinyals, K. Ayoub, J. Stanway, L. Bennett, D. Hassabis, K. Kavukcuoglu, and G. Irving. Scaling language models: Methods, analysis & insights from training gopher. CoRR, abs / 2112.11446, 2021; Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv: 1910.10683, 2019; Daniel Adiwardana, Minh-Thang Luong, David R. So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, and Quoc V. Le. Towards a human-like open-domain chatbot. CoRR, abs / 2001.09977, 2020; and Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020.

[0134] Prior to using the language model neural network 130 to generate network outputs 132, the language model neural network 130 is pre-trained, e.g., by the system 100 or by one or more other systems. As a particular example, the system 100 or the other system(s) can pre-train the language model neural network 130 on a token prediction task, e.g., a task that requires predicting, given a current sequence of tokens, the next token that follows the current sequence in the training data. For example, the token prediction task can require, for each training input sequence included in a training data set, predicting by the language model neural network a second portion of a sequence that followed a first portion of the sequence.

[0135] In some examples, the token prediction task can be a language modeling task, e.g., a task that requires predicting, given a current sequence of text tokens, the next token that follows the current sequence in the training data. Equivalently, the language modeling task can require, for each given unlabeled text sequence in a training data set, predicting a text sequence that followed the given unlabeled text sequence in a corresponding document. Attorney Docket No.: 45288-0457WO1

[0136] As a particular example, the generative neural network can be trained based on optimizing on a maximum -likelihood objective on a next token prediction task that requires, for each training input sequence included in a training data set, predicting by the generative neural network a second portion of a sequence that followed a first portion of the sequence. For example, the language model neural network 130 can be pre-trained on a maximum-likelihood objective on a large dataset of text, e.g., text that is publicly available from the Internet or another text corpus.

[0137] As another example, the training input sequences included in the training data set can be generated from a large dataset of text in one or more natural languages, a large dataset of audio samples, e.g., audio recordings or waveforms that represent the audio recordings, a large dataset of images where each image includes an array of pixels, a large dataset of videos where each video includes a temporal sequence of frames, or a large multi-modal dataset that includes a combination of two or more of these datasets.

[0138] FIG. 2 shows an example process 200 for generating synthetic audio prompts. For convenience, the process 200 will be described as being performed by a system of one or more computers located in one or more locations. For example, an audio generation system, e.g., the audio generation system 100 of FIGS. 1A-1B, appropriately programmed in accordance with this specification, can perform the process 200.

[0139] The system can generate synthetic audio prompts for inclusion in the audio prompts 110 described above with reference to FIGS. 1 A-1B using a voice conversion model 210. The voice conversion model 210 can be configured to perform speech-to-speech voice conversion. In speech-to-speech voice conversion, the voice conversion model 210 processes data representing an input audio signal and a target speaker prompt, and generates an output audio signal that preserves the same spoken content, prosody, and timing as the input audio signal, spoken in the voice of the target speaker.

[0140] The audio prompts 110 can include one or more initial audio prompts such as the initial audio prompt 202. In some examples, the one or more initial audio prompts can include real- world, recorded audio signals, or synthetic audio signals.

[0141] Example initial audio prompts for speakers A and B are shown below in Table 1. Attorney Docket No.: 45288-0457WO1

[0142] Table 1. Example audio prompts

[0143] The system can generate a synthetic audio prompt 212 for the initial audio prompt 202 using the voice conversion model 210. For example, the system can provide data representing the initial audio prompt 202 and a target speaker prompt 204 as input to the voice conversion model 210 to generate the synthetic audio prompt 212. The synthetic audio prompt 212 includes an audio signal that represents speech of the initial audio prompt, i.e., the same spoken content that is represented by the initial audio prompt, spoken by the target speaker of the target speaker prompt 204.

[0144] The data representing the initial audio prompt can be an audio signal, or features of the audio signal such as pitch features or energy features.

[0145] The target speaker prompt 204 represents speech of a target speaker. In some examples, the system can derive the target speaker prompt 204 from the initial audio prompts. For example, the system can select an audio prompt from the audio prompts 110 that represents speech by a different corresponding speaker than the corresponding speaker of the initial audio prompt 202. In the example of Table 1, if the initial audio prompt 202 has the identifier 6, the speaker of the initial audio prompt is the speaker A. The system can derive the target speaker prompt 204 from the initial audio prompt with the identifier 12, where the speaker is the speaker B.

[0146] In some examples, the target speaker prompt 204 can include a target speaker prompt audio signal. In some other examples, the target speaker prompt 204 can include a target speaker prompt embedding of a target speaker prompt audio signal. For example, the system can generate the target speaker prompt embedding from the target speaker prompt audio signal using an encoder such as the encoder neural network of a neural audio codec. Attorney Docket No.: 45288-0457WO1

[0147] The voice conversion model 210 can have any appropriate architecture for performing a voice conversion task to convert speech spoken by a first speaker into speech spoken by a second speaker. More generally, the voice conversion model 210 can be implemented as any appropriate type of machine learning model that can perform a voice conversion task. The voice conversion model 210 can include any appropriate types of neural network layers (e.g., fully connected layers, message passing layers, convolutional layers, attention layers, recurrent layers, pooling layers, and so forth), in any appropriate number (e.g., 5 layers, or 10 layers, or 50 layers), and connected in any appropriate configuration (e.g., as a directed graph of layers).

[0148] As a particular example, the voice conversion model 210 can be configured to generate an audio signal representing speech spoken by the second speaker given an input audio signal representing speech spoken by the first speaker, and an input speaker prompt for the second speaker. In these examples, the voice conversion model 210 can have been trained on training examples that include a first audio signal representing speech by a first speaker, a second audio signal representing speech by a second speaker, and a speaker prompt for the second speaker.

[0149] In this example, the input audio signal can include an initial audio prompt that is an audio signal.

[0150] The input speaker prompt for the second speaker can characterize speech of the second speaker. For example, the input speaker prompt for the second speaker can include the target speaker prompt 204. As described above, the target speaker prompt 204 can include a target speaker prompt audio signal or a target speaker prompt embedding of a target speaker prompt audio signal.

[0151] To generate the audio signal, the voice conversion model 210 can process a semantic representation of the initial audio prompt 202 and a speaker prompt embedding for the target speaker prompt 204 using an audio generation model.

[0152] The semantic representation specifies a respective semantic token at each of multiple first time steps spanning the initial audio prompt 202. Each semantic token is selected from a vocabulary of semantic tokens and represents semantic content of the initial audio prompt 202 at the corresponding first time step. Examples of semantic content represented by the semantic tokens can include linguistic content, phonetics, language syntax, and prosodic features for speech. In some examples, the semantic tokens represent linguistic content, such as phonetics Attorney Docket No.: 45288-0457WO1 and semantics, and do not represent paralinguistic information, such as speaker identity and acoustic information.

[0153] In some examples, the voice conversion model 210 can generate the semantic representation of the initial audio prompt 202. For example, the voice conversion model 210 can use a semantic tokenizer to generate the semantic representation. For example, the voice conversion model 210 can provide the initial audio prompt 202 as input to the semantic tokenizer to generate the semantic representation. The semantic tokenizer can include an audio representation neural network that has been trained to generate representations of input audio. For example, the audio representation neural network can be a self-attention based model, e.g., a Transformer-based model or a Conformer-based model, e.g., a W2v-BERT neural network (described in Chung, Yu-An, et al. “W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training.” 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)). As an example, the audio representation neural network can be trained on a masked language modeling loss or a combination of a masked language modeling loss and a contrastive loss.

[0154] The semantic tokenizer can generate the semantic tokens based on outputs of one or more layers, e.g., of one of the intermediate layers, of an audio representation neural network. For example, the semantic tokenizer can generate the semantic representation by processing the initial audio prompt 202 using the audio representation neural network. The outputs of the one or more layers of the audio representation neural network can include an embedding of the initial audio prompt 202 for each of multiple time steps of the initial audio prompt 202. The semantic tokenizer can generate the semantic representation by assigning each embedding for the source audio signal to the closest semantic token, e.g., having a smallest embedding distance in the embedding space, of a set of semantic tokens. The set of semantic tokens can include the centroids of K clusters of embeddings for an intermediate layer of the audio representation neural network for a set of training audio samples.

[0155] In some examples where the target speaker prompt includes a target speaker prompt audio signal, the voice conversion model 210 can generate the target speaker prompt embedding from the target speaker prompt audio signal using an encoder. As an example, the encoder can include an encoder neural network of a neural audio codec. As a particular example, the encoder can be a SoundStream encoder of the SoundStream neural audio codec described in Zeghidour, Attorney Docket No.: 45288-0457WO1

[0156] Neil, et al. "Soundstream: An end-to-end neural audio codec." IEEE / ACM Transactions on Audio, Speech, and Language Processing 30 (2021): 495-507.

[0157] The audio generation model is configured to generate an output audio signal given at least the semantic representation and the target speaker prompt embedding. For example, the audio generation model can be any appropriate neural network that is configured to generate an output audio signal that preserves the same spoken content, prosody, and timing of the initial audio prompt 202, spoken in the voice characterized by the target speaker prompt embedding.

[0158] In some examples, the audio generation model can be configured to generate the output audio signal by processing a masked representation of the output audio signal using a neural network to generate a sequence of output tokens representing the output audio signal. For example, the masked representation can include a sequence of input tokens that includes conditioning tokens derived from the semantic representation and the target speaker prompt embedding, and masked tokens that represent acoustic tokens of the output audio signal.

[0159] The neural network iteratively updates the sequence of input tokens to unmask the sequence of input tokens. By iteratively updating the sequence of input tokens, the neural network can generate the sequence of output tokens that is an unmasked representation of the output audio signal. The neural network can have been trained using a machine learning training technique, e.g., a gradient descent with backpropagation training technique that uses a conventional optimizer, e.g., stochastic gradient descent, RMSprop, or Adam optimizer, on an audio representation learning task based on optimizing an appropriate objective function for the task. For example, the audio representation learning task can be a masked audio modeling task. For each training example, the masked audio modeling task is a task that requires predicting, given a sequence of input tokens that include masked tokens, a sequence of output tokens that include unmasked tokens in place of the masked tokens and that represent an audio signal.

[0160] As another example, the audio generation model can be configured to generate the output audio signal by processing an encoded representation derived from the semantic representation and the target speaker prompt embedding using a token decoder neural network to generate a sequence of output tokens representing the output audio signal. In this example, the audio generation model can have a similar architecture as the speech processing model described above with reference to FIG. IB. For example, the audio generation model can generate an encoded Attorney Docket No.: 45288-0457WO1 representation from the respective representations for the semantic representation and the target speaker prompt embedding.

[0161] As another example, the voice conversion model 210 can have a similar architecture as the speech processing model described above with reference to FIG. IB. For example, the voice conversion model 210 can be configured to generate an audio signal representing speech spoken by the second speaker given features representing an input audio signal representing speech spoken by the first speaker, and an input target speaker prompt for the second speaker. In some examples, the inputs can also include a transcript of the speech to be represented by the audio signal. For example, to generate the synthetic audio prompt 212, the system can provide features representing the initial audio prompt 202, the target speaker prompt 204, and an initial transcript of speech represented by the initial audio prompt 202 as input to the voice conversion model 210.

[0162] The features can include pitch features or energy features of the initial audio prompt 202. In examples where the initial audio prompts include audio signals, the system can generate the features from the initial audio prompts.

[0163] The input target speaker prompt for the second speaker can include a target speaker prompt embedding for the target speaker prompt 204. In examples where the target speaker prompt includes a target speaker prompt audio signal, the system can generate the speaker prompt embedding from the target speaker prompt audio signal using an encoder. As an example, the encoder can include an encoder neural network of a neural audio codec.

[0164] The transcript of the speech to be represented by the audio signal can include an initial transcript of speech represented by the initial audio prompt 202. For example, the system can generate the initial transcript of speech by providing the initial audio prompt 202 as input to a speech recognition model.

[0165] More generally, the speech recognition model can be implemented as any appropriate type of machine learning model that can perform a speech recognition task. The speech recognition model can include any appropriate types of neural network layers (e.g., fully connected layers, message passing layers, convolutional layers, attention layers, recurrent layers, pooling layers, and so forth), in any appropriate number (e.g., 5 layers, or 10 layers, or 50 layers), and connected in any appropriate configuration (e.g., as a directed graph of layers). Attorney Docket No.: 45288-0457WO1

[0166] The voice conversion model 210 can be configured to generate the synthetic audio prompt 212 by processing an encoded representation derived from features for the initial audio prompt 202, the target speaker prompt embedding for the target speaker prompt 204, and in some examples, the transcript, using a token decoder neural network to generate a sequence of output tokens representing the synthetic audio prompt 212.

[0167] The system can include the synthetic audio prompt 212 in the audio prompts 110. For example, the system can maintain the synthetic audio prompt 212 in the audio prompts 110. In some examples where the system maintains a corresponding prompt style description for each audio prompt, the system also assigns, as the corresponding prompt style description for the synthetic audio prompt 212, the corresponding style description of the initial audio prompt 202.

[0168] The system can perform the process 200 for each of the initial audio prompts, and for each of one or more target speaker prompts. Thus, the system can generate a set of audio prompts 110 that have a diverse range of styles for each of the target speakers. The system can then use the set of audio prompts 110 to generate output audio signals as described with reference to FIG. IB for each of the target speakers.

[0169] Table 2 shows the example audio prompts after performing the process 200 for each of the initial audio prompts and for each of one or more target speaker prompts derived from the initial audio prompts of Table 1. Attorney Docket No.: 45288-0457WO1

[0170] Table 2. Example audio prompts

[0171] In some examples, the system generates one or more synthetic audio prompts from input audio prompts or input target speaker prompts, e.g., received from a user. For example, the system can receive an input audio prompt with a corresponding style description of “News, neutral.” The system can generate synthetic audio prompts for speaker A and speaker B given a target speaker prompt for speaker A and a target speaker prompt for speaker B using the voice conversion model 210.

[0172] By generating the set of audio prompts 110 that have a diverse range of styles for each of the target speakers prior to generating output audio signals as described with reference to FIG. IB, the system can reduce the amount of computing resources used to generate audio signals representing speech of a given style and spoken by a particular speaker when the particular speaker is represented in the audio prompts 110. As an example, the audio prompts 110 can include audio prompts spoken by a speaker A and in a range of different styles. To generate an output audio signal 150 representing speech spoken by the speaker A according to the style description 102 as described in FIG. IB requires fewer computing resources at inference time when compared to generating the converted audio signal 312 as described with reference to FIG. 3, e.g., by obtaining a target speaker prompt representing speech by the speaker A and performing voice conversion using the target speaker prompt. In examples where the audio prompts 110 include audio prompts spoken by the speaker A, the system 100 does not need to perform voice conversion at inference to generate speech spoken by the speaker A.

[0173] As a particular example, the system can perform text to speech for dialogue agents and speech-based dialogue interfaces. The dialogue agent can be part of a dialogue system in which the user interacts with the dialogue agent using natural language text, audio, or other types of data, and the dialogue agent generates responses for the user. The system can generate audio responses for users that lead to more natural and engaging interactions with users. For example, the speakers of the audio prompts 110 can be speakers associated with the dialogue agent. The Attorney Docket No.: 45288-0457WO1 system can generate synthetic audio prompts for inclusion in the audio prompts 110 that have a diverse range of styles for the speakers associated with the dialogue agent. During a conversation with a user, the system can generate speech representing the text of a response from the dialogue agent, in the voice of the dialogue agent, and in a style that is relevant to the context of the conversation. In these examples, the system can receive the style description from the dialogue agent. Because the audio prompts 110 are generated prior to the conversation with the user, the system can generate and present the speech to the user quickly, resulting in an efficient and conversation-like user experience.

[0174] FIG. 3 is a block diagram of the example audio generation system 100 described with reference to FIGS. 1A-1B. In particular, the audio generation system 100 generates a prediction of a converted audio signal 312 given the style description 102, the transcript 104, and a target speaker prompt 306. The converted audio signal 312 represents speech specified by the transcript 104, spoken in a style that is described by the style description 102, and spoken by the speaker of the target speaker prompt 306.

[0175] In some examples, the target speaker prompt 306 includes a target speaker prompt audio signal representing speech by the target speaker. In some examples, the target speaker prompt 306 includes a target speaker prompt embedding of a target speaker audio signal representing speech by the target speaker.

[0176] The system 100 generates the output audio signal 150 as described above with reference to FIGS. 1A-1B. The output audio signal 150 represents speech specified by the transcript 104, in the style described by the style description 102, and spoken by the speaker of the selected audio prompt.

[0177] To generate the converted audio signal 312, the system 100 provides at least data representing the output audio signal 150 and the target speaker prompt 306 as input to a voice conversion model 310. An example voice conversion model 310 is described above with reference to FIG. 2.

[0178] In examples where the voice conversion model 310 is configured to generate an audio signal given an input audio signal representing speech spoken by a first speaker, an input speaker prompt for a second speaker, the data representing the output audio signal 150 includes the output audio signal. The input speaker prompt for the second speaker includes the target speaker prompt 306. Attorney Docket No.: 45288-0457WO1

[0179] In examples where the voice conversion model 310 is configured to generate an audio signal given features representing an input audio signal representing speech spoken by the first speaker, and an input speaker prompt for the second speaker, the features representing the output audio signal 150 can include energy features of the output audio signal 1 0, pitch features of the output audio signal 150, or both. The input speaker prompt for the second speaker can include the target speaker prompt 306. In some of these examples, the system can also provide the transcript 104 as input to the voice conversion model 310 to generate the converted audio signal 312.

[0180] In some examples, the system 100 provides the converted audio signal 312 for presentation to the user. For example, the system 100 can provide data representing the audio signal 312 to the user device and cause playback of the audio signal 312. Alternatively or in addition, the system 100 may provide data representing the audio signal 312 for storage.

[0181] The system 100 can thus generate a converted audio signal representing speech specified by the transcript 104, spoken in a style that is described by the style description 102, and spoken by a particular speaker represented by the target speaker prompt 306. The particular speaker can be any particular speaker. Thus the system 100 can provide for flexibility and control in the generation of audio signals representing speech.

[0182] In some implementations, the system 100 can generate a training example for training a speech generation model. For example, the training example can include the transcript 104, the style description 102, the target speaker prompt 306, and the converted audio signal 312. The system 100 can include the training example in a set of training data.

[0183] A training system of the system 100 or another training system can train a speech generation model on the set of training data. The speech generation model can be configured to process one or more inputs in accordance with current values of parameters of the speech generation model to generate an output audio signal. For example, the speech generation model can be configured to receive a transcript, a style description, and a target speaker prompt, to generate an output audio signal. The speech generation model can have any appropriate architecture for performing a speech generation task. As an example, the speech generation model can have a similar architecture as the speech processing model described above with reference to FIG. IB. Attorney Docket No.: 45288-0457WO1

[0184] FIG. 4 is a flow diagram of an example process 400 for generating audio. For convenience, the process 400 will be described as being performed by a system of one or more computers located in one or more locations. For example, an audio generation system, e.g., the audio generation system 100 depicted in FIGS. 1A-1B, appropriately programmed in accordance with this specification, can perform the process 400.

[0185] The system maintains (e.g., obtains, stores, keeps, accesses, etc.) multiple audio prompts (step 402). Each audio prompt represents speech spoken by a corresponding speaker in a corresponding style.

[0186] In some examples, the system maintains (e.g., obtains, stores, keeps, accesses, etc.), for each audio prompt, a corresponding prompt style description. The corresponding prompt style description for each audio prompt is a description of the corresponding style of the audio prompt.

[0187] In some implementations, the audio prompts can include initial audio prompts and synthetic audio prompts. The system can generate one or more of the synthetic audio prompts. For example, the system can generate each synthetic audio prompt by providing at least data representing an initial audio prompt and a target speaker prompt representing speech of a target speaker as input to a voice conversion model. The synthetic audio prompt represents speech of the initial audio prompt, spoken by the target speaker. Generating synthetic audio prompts is described in further detail above with reference to FIG. 2.

[0188] The system obtains a transcript (step 404). The transcript includes text to be spoken in speech represented by an output audio signal.

[0189] The system obtains a style description (step 406). The style description includes text specifying a target style for the speech represented by the output audio signal.

[0190] The system provides an input including the style description and data characterizing the multiple audio prompts as input to a language model neural network to generate a network output (step 408). The network output identifies a subset of the multiple audio prompts that are compatible with the target style.

[0191] In some examples, data characterizing the multiple audio prompts includes the corresponding prompt style descriptions for the multiple audio prompts. In some examples, data characterizing the multiple audio prompts includes the audio prompts. In these examples, the language model neural network can include a multimodal language model neural network. Attorney Docket No.: 45288-0457WO1

[0192] In some examples, the input also includes a respective identifier for each of the multiple audio prompts. In these examples, the network output includes the respective identifier for each audio prompt in the subset.

[0193] The system provides i) the transcript and ii) one or more of the identified audio prompts as input to a speech processing model to generate the output audio signal (step 410). The output audio signal represents speech specified by the transcript, spoken in the target style and by the speaker of one of the identified audio prompts.

[0194] In some implementations, the system obtains a target speaker prompt. The system can provide data representing the output audio signal and the target speaker prompt as input to a voice conversion model to generate a converted audio signal. The converted audio signal represents speech specified by the transcript, spoken in the target style, and by the speaker of the target speaker prompt.

[0195] Example voice conversion models are described above with reference to FIG. 2. In some examples, the data representing the output audio signal includes the output audio signal, and the input speaker prompt for the second speaker includes the target speaker prompt. In some examples, the data representing the output audio signal includes energy features or pitch features of the output audio signal, and the input speaker prompt for the second speaker includes the target speaker prompt. In some of these examples, the system can also provide the transcript as input to the voice conversion model.

[0196] This specification uses the term “configured” in connection with, or in relation to, systems and environments, as well as computer program components. For a system of one or more computers or environment to be configured to perform particular operations or actions means that the system has installed on it software, firmware, hardware, or a combination of them that in operation cause the system to perform the operations or actions. For instance, configuring a system might involve installing a software library with specific algorithms, updating firmware with new instructions for handling data, or adding a hardware component for enhanced processing capabilities. For one or more computer programs to be configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by data processing apparatus, cause the apparatus to perform the operations or actions. Attorney Docket No.: 45288-0457WO1

[0197] Embodiments of the subject matter and the functional operations described in this specification can be implemented in various forms, including digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non transitory storage medium for execution by, or to control the operation of, data processing apparatus. The computer storage medium can be a storage device such as a machine-readable storage device, a hard drive or solid-state drive (SSD), a storage medium, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode, e.g., carry, information for transmission to suitable receiver apparatus, e.g., a receiving device or system, for execution by a data processing apparatus. Furthermore, implementations may leverage emerging technologies like quantum computing or neuromorphic computing for specific applications, and may be deployed in distributed or cloud-based environments where components reside on different machines or within a cloud infrastructure.

[0198] The term “data processing apparatus” refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor or processing unit, a computer, multiple processors or computers, e.g., working together, graphics processing units (GPUs), or tensor processing units (TPUs). The apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. Embodiments may particularly benefit from utilizing the parallel processing capabilities of GPUs, in a General-Purpose computing on Graphics Processing Units (GPGPU) context, where code specifically designed for GPU execution, often called kernels or shaders, is employed. Similarly, TPUs excel at running optimized tensor operations crucial for many Attorney Docket No.: 45288-0457WO1 machine learning algorithms. By leveraging these accelerators and their specialized programming models, the system can achieve significant speedups and efficiency gains for tasks involving artificial intelligence and machine learning, particularly in areas such as computer vision, natural language processing, and robotics.

[0199] A computer program, which may also be referred to or described as a program, software, a software application, an app, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it can be deployed in any form, including as a stand alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communication network. The specific implementation of the computer programs may involve a combination of traditional programming languages and specialized languages or libraries designed for GPGPU programming or TPU utilization, depending on the chosen hardware platform and desired performance characteristics.

[0200] In this specification, the term “database” is used broadly to refer to any collection of data: the data does not need to be structured in any particular way, or structured at all, and it can be stored on storage devices in one or more locations. Thus, for example, the index database can include multiple collections of data, each of which may be organized and accessed differently.

[0201] Similarly, in this specification the term “engine” is used broadly to refer to a softwarebased system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components, installed on one or more computers in one or more locations, e.g., located at a single site or distributed across multiple locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and running on the same computer or computers. Examples of engine functions within the context of Al and machine learning could include data pre-processing and cleaning, feature engineering and Attorney Docket No.: 45288-0457WO1 extraction, model training and optimization, inference and prediction generation, and postprocessing of results. The specific design and implementation of engines will depend on the overall architecture and the distribution of computational tasks across various hardware components, including CPUs, GPUs, TPUs, and other specialized processors.

[0202] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. Additionally, graphics processing units (GPUs) and tensor processing units (TPUs) can be utilized to enable concurrent execution of aspects of these processes and logic flows, significantly accelerating performance. This approach offers significant advantages for computationally intensive tasks often found in Al and machine learning applications, such as matrix multiplications, convolutions, and other operations that exhibit a high degree of parallelism. By leveraging the parallel processing capabilities of GPUs and TPUs, significant speedups and efficiency gains compared to relying solely on CPUs can be achieved. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA or an ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers, for even greater performance or energy efficiency in specific use cases.

[0203] Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit (CPU). Additionally, graphics processing units (GPUs), tensor processing units (TPUs), and other machine learning accelerators can be employed to enhance performance, particularly for tasks involving artificial intelligence and machine learning. These accelerators often work in conjunction with CPUs, handling specialized computations while the CPU manages overall system operations and other tasks. Generally, a central processing unit will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The specific configuration of processing units and memory will depend on factors like the complexity of the Al model, the volume of data being processed, and the desired performance and latency requirements. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer will also include, or be operatively coupled to receive data from or Attorney Docket No.: 45288-0457WO1 transfer data to, or both, one or more mass storage devices for storing data, e.g., hard drives, SSDs, flash memory for persistent data storage, magnetic, magneto optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few. Embodiments can be implemented on a wide range of computing platforms, from small embedded devices with limited resources to large-scale data center systems with high- performance computing capabilities.

[0204] Computer readable media suitable for storing computer program instructions and data include all forms of non volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, read-only memory (ROM), solid-state drives (SSDs), and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks or hard disk drives (HDDs); magneto optical disks; and optical discs such as CD ROM and DVD-ROM disks and Blu-ray discs. The specific type of computer-readable media used will depend on factors such as the size of the data, access speed requirements, cost considerations, and the desired level of portability or permanence.

[0205] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor or an organic light-emitting diode (OLED) display, for displaying information to the user and a key vectorboard and a pointing device, e.g., a mouse or a trackball, keyboard, touchscreens, voice commands, gesture recognition, or other input modalities depending on the specific device and application, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device or application that is used by the user; for example, by sending web pages to a web browser on a user’s device in response to requests received from the web browser. Also, a computer can interact with a user by sending text messages or other forms of message to a Attorney Docket No.: 45288-0457WO1 personal device, e.g., a smartphone that is running a messaging application, and receiving responsive messages from the user in return. The selection of input and output modalities will depend on the specific application and the desired form of user interaction.

[0206] Data processing apparatus for implementing machine learning models can also include, for example, special-purpose hardware accelerator units for processing common and computeintensive parts of machine learning training or production, i.e., inference, workloads.

[0207] Machine learning models can be implemented and deployed using a machine learning framework, e.g., a TensorFlow framework, or a Jax framework. These frameworks offer comprehensive tools and libraries that facilitate the development, training, and deployment of machine learning models.

[0208] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back end component, e.g., as a data server, a back-end server, or cloud-based infrastructure, or that includes a middleware component, e.g., an application server, a middleware server, or application programming interface (API), to facilitate communication and data exchange, or that includes a front end component, e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back end, middleware, or front end components. For instance, the described functionality could be implemented solely on a client device (e g., for on-device machine learning) or deployed as a combination of front-end and back-end components for more complex applications. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet. The specific system architecture and choice of components will depend on factors such as the scale of the application, the need for real-time processing, data security requirements, and the desired user experience.

[0209] The computing system can include clients and servers. A client and server are generally remote from each other, e.g., geographically separated, and typically interact through a communication network. The specific type of network, such as a local area network (LAN), a wide area network (WAN), or the Internet, will depend on the reach and scale of the application. The relationship of client and server arises by virtue of computer programs running on the Attorney Docket No.: 45288-0457WO1 respective computers and having a client-server relationship to each other, e.g., designed to communicate with each other using appropriate protocols. These protocols may include HTTP, TCP / IP, or other specialized protocols depending on the nature of the data being exchanged and the security requirements of the system. In some embodiments, a server transmits data, e.g., an HTML page, to a user device such as a computer, smartphone, or tablet, e.g., for purposes of displaying data to and receiving user input from a user interacting with the device, which acts as a client. The client device can then process the received information and display results to the user, and potentially send data or feedback back to the server. Data generated at the user device, e.g., a result of the user interaction, can be received at the server from the device, e.g., for further processing or storage. This allows for dynamic interactions between the user and the system, enabling a wide range of applications and functionalities.

[0210] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.

[0211] Similarly, while operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products. Attorney Docket No.: 45288-0457WO1

[0212] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.

[0213] This specification also includes the following clauses:

[0214] Clause 1. A computer-implemented method comprising: maintaining a plurality of audio prompts, wherein each audio prompt represents speech spoken by a corresponding speaker in a corresponding style; obtaining a transcript that comprises text to be spoken in speech represented by an output audio signal; obtaining a style description that comprises text specifying a target style for the speech represented by the output audio signal; providing an input comprising the style description and data characterizing the plurality of audio prompts as input to a language model neural network to generate a network output identifying a subset of the plurality of audio prompts that are compatible with the target style; and providing i) the transcript and ii) one or more of the identified audio prompts as input to a speech processing model to generate the output audio signal.

[0215] Clause 2. The method of clause 1, further comprising maintaining, for each audio prompt, a corresponding prompt style description, wherein the corresponding prompt style description is a description of the corresponding style of the audio prompt.

[0216] Clause 3. The method of clause 2, wherein the data characterizing the plurality of audio prompts comprises the corresponding prompt style descriptions for the plurality of audio prompts.

[0217] Clause 4. The method of any preceding clause, wherein the data characterizing the plurality of audio prompts comprises the plurality of audio prompts, and wherein the language model neural network comprises a multimodal language model neural network.

[0218] Clause 5. The method of any preceding clause, wherein the input further comprises a respective identifier for each of the plurality of audio prompts. Attorney Docket No.: 45288-0457WO1

[0219] Clause 6. The method of clause 5, wherein the network output comprises the respective identifier for each of the subset of the plurality of audio prompts that are compatible with the target style.

[0220] Clause 7. The method of any preceding clause, wherein the plurality of audio prompts comprise one or more initial audio prompts and one or more synthetic audio prompts, and wherein the one or more synthetic audio prompts are generated by: for each of one or more of the initial audio prompts: providing at least data representing the initial audio prompt and a target speaker prompt representing speech of a target speaker as input to a voice conversion model to generate a synthetic audio prompt, wherein the synthetic audio prompt represents speech of the initial audio prompt, spoken by the target speaker.

[0221] Clause 8. The method of clause 7, further comprising: maintaining, for each audio prompt, a corresponding prompt style description, wherein the corresponding prompt style description is a description of the corresponding style of the audio prompt; and for each of the one or more initial audio prompts: assigning, as the corresponding prompt style description for the synthetic audio prompt, the corresponding style description of the initial audio prompt.

[0222] Clause 9. The method of any one of clauses 7-8, wherein the target speaker prompt is derived from an initial audio prompt that represents speech spoken by a different corresponding speaker than the corresponding speaker of the initial audio prompt.

[0223] Clause 10. The method of any one of clauses 7-8, wherein data representing the initial audio prompt comprises any one or more of: an audio signal, energy features of the audio signal, or pitch features of the audio signal.

[0224] Clause 11. The method of any one of clauses 7-8, wherein providing at least data representing the initial audio prompt and a target speaker prompt as input to a voice conversion model to generate a synthetic audio prompt comprises providing the data representing the initial audio prompt, the target speaker prompt, and an initial transcript of speech represented by the initial audio prompt as input to the voice conversion model to generate the synthetic audio prompt. Attorney Docket No.: 45288-0457WO1

[0225] Clause 12. The method of any preceding clause, further comprising: providing at least data representing the output audio signal and a target speaker prompt as input to a voice conversion model to generate a converted audio signal.

[0226] Clause 13. The method of clause 12, wherein data representing the output audio signal comprises any one or more of: the output audio signal, energy features of the output audio signal, or pitch features of the output audio signal.

[0227] Clause 14. The method of any one of clauses 12-13, wherein providing at least data representing the output audio signal and a target speaker prompt as input to a voice conversion model to generate a converted audio signal comprises providing data representing the output audio signal, the target speaker prompt, and the transcript as input to the voice conversion model to generate the converted audio signal.

[0228] Clause 15. The method of any one of clauses 12-14, further comprising: generating a training example comprising i) the transcript, ii) the style description, iii) the target speaker prompt, and iv) the converted audio signal; and including the training example in a first set of training data.

[0229] Clause 16. The method of clause 15, further comprising training a speech generation model on the first set of training data.

[0230] Clause 17. The method of any preceding clause, further comprising: generating a training example comprising i) the transcript, ii) the style description, and iii) the output audio signal; and including the training example in a second set of training data.

[0231] Clause 18. The method of clause 17, further comprising training a speech generation model on the second set of training data.

[0232] Clause 19. A system comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform the respective operations of any one of clauses 1-18. Attorney Docket No.: 45288-0457WO1

[0233] Clause 20. One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform the respective operations of any one of clauses 1-18.

[0234] What is claimed is:

Claims

Attorney Docket No.: 45288-0457WO1CLAIMS1. A computer-implemented method comprising: maintaining a plurality of audio prompts, wherein each audio prompt represents speech spoken by a corresponding speaker in a corresponding style; obtaining a transcript that comprises text to be spoken in speech represented by an output audio signal; obtaining a style description that comprises text specifying a target style for the speech represented by the output audio signal; providing an input comprising the style description and data characterizing the plurality of audio prompts as input to a language model neural network to generate a network output identifying a subset of the plurality of audio prompts that are compatible with the target style; and providing i) the transcript and ii) one or more of the identified audio prompts as input to a speech processing model to generate the output audio signal.

2. The method of claim 1, further comprising maintaining, for each audio prompt, a corresponding prompt style description, wherein the corresponding prompt style description is a description of the corresponding style of the audio prompt.

3. The method of claim 2, wherein the data characterizing the plurality of audio prompts comprises the corresponding prompt style descriptions for the plurality of audio prompts.

4. The method of any preceding claim, wherein the data characterizing the plurality of audio prompts comprises the plurality of audio prompts, and wherein the language model neural network comprises a multimodal language model neural network.

5. The method of any preceding claim, wherein the input further comprises a respective identifier for each of the plurality of audio prompts.Attorney Docket No.: 45288-0457WO16. The method of claim 5, wherein the network output comprises the respective identifier for each of the subset of the plurality of audio prompts that are compatible with the target style.

7. The method of any preceding claim, wherein the plurality of audio prompts comprise one or more initial audio prompts and one or more synthetic audio prompts, and wherein the one or more synthetic audio prompts are generated by: for each of one or more of the initial audio prompts: providing at least data representing the initial audio prompt and a target speaker prompt representing speech of a target speaker as input to a voice conversion model to generate a synthetic audio prompt, wherein the synthetic audio prompt represents speech of the initial audio prompt, spoken by the target speaker.

8. The method of claim 7, further comprising: maintaining, for each audio prompt, a corresponding prompt style description, wherein the corresponding prompt style description is a description of the corresponding style of the audio prompt; and for each of the one or more initial audio prompts: assigning, as the corresponding prompt style description for the synthetic audio prompt, the corresponding style description of the initial audio prompt.

9. The method of any one of claims 7-8, wherein the target speaker prompt is derived from a further initial audio prompt that represents speech spoken by a different corresponding speaker than the corresponding speaker of the initial audio prompt.

10. The method of any one of claims 7-9, wherein data representing the initial audio prompt comprises any one or more of: an audio signal, energy features of the audio signal, or pitch features of the audio signal.

11. The method of any one of claims 7-10, wherein providing at least data representing the initial audio prompt and a target speaker prompt as input to a voice conversion model to generate a synthetic audio prompt comprises providing the data representing the initial audio prompt, theAttorney Docket No.: 45288-0457WO1 target speaker prompt, and an initial transcript of speech represented by the initial audio prompt as input to the voice conversion model to generate the synthetic audio prompt.

12. The method of any preceding claim, further comprising: providing at least data representing the output audio signal and a target speaker prompt as input to a voice conversion model to generate a converted audio signal.

13. The method of claim 12, wherein data representing the output audio signal comprises any one or more of: the output audio signal, energy features of the output audio signal, or pitch features of the output audio signal.

14. The method of any one of claims 12-13, wherein providing at least data representing the output audio signal and a target speaker prompt as input to a voice conversion model to generate a converted audio signal comprises providing data representing the output audio signal, the target speaker prompt, and the transcript as input to the voice conversion model to generate the converted audio signal.

15. The method of any one of claims 12-14, further comprising: generating a first training example comprising i) the transcript, ii) the style description, iii) the target speaker prompt, and iv) the converted audio signal; and including the first training example in a first set of training data.

16. The method of claim 15, further comprising training a speech generation model on the first set of training data.

17. The method of any preceding claim, further comprising: generating a second training example comprising i) the transcript, ii) the style description, and iii) the output audio signal; and including the second training example in a second set of training data.Attorney Docket No.: 45288-0457WO118. The method of claim 17, further comprising training a speech generation model on the second set of training data.

19. A system comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform the respective operations of any one of claims 1-18.

20. One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform the respective operations of any one of claims 1-18.

Citation Information

Cited By

  • Creating context-specific, versatile expert ai personas

    US20260134011A1