Systems and methods for cross-speaker style transfer in text-to-speech and for training data generation.

By aligning source speaker data with speech posterior maps and extracting additional prosodic features to generate spectrogram data, the efficiency and cost issues of training multi-target speaker-style text-to-speech models are solved, enabling rapid training and efficient generation of target speaker speech data.

CN114203147BActive Publication Date: 2026-03-06MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010885556.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-28
Publication Date
2026-03-06
Estimated Expiration
2040-08-28

AI Technical Summary

Technical Problem

Existing technologies struggle to quickly and efficiently generate speech data with diverse target speaker styles when training text-to-speech models, especially due to the time-consuming and costly nature of training data collection and the potential inability of the target speaker to adapt to the desired speech style.

Method used

Additional prosodic features are extracted by aligning the waveforms of the source speaker data with the speech posterior graph data, and spectrogram data is generated based on the timbre of the target speaker's voice, which is then used to train a text-to-speech model across speaker styles.

Benefits of technology

It enables more efficient generation of spectrogram data, reduces computation time and data collection costs, and can quickly train text-to-speech models with the timbre of the target speaker and the prosodic patterns of the source speaker.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114203147B_ABST
    Figure CN114203147B_ABST
Patent Text Reader

Abstract

Each system is configured to generate spectrogram data characterized by the timbre of the target speaker's voice and the prosodic pattern of the source speaker by: converting the waveform of the source speaker data into speech posterior graph (PPG) data; extracting additional prosodic features from the source speaker data; and generating a spectrogram based on the PPG data and the extracted prosodic features. Each system is configured to utilize / train machine learning models for generating the spectrogram data and for training neural, text-to-speech models using the generated spectrogram data.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] Text-to-speech (TTS) models are configured to convert arbitrary text into speech data that sounds like human speech. Sometimes referred to as voice-font TTS models, they typically consist of a front-end module, an acoustic model, and a speech encoder. The front-end module is configured to perform text normalization (e.g., converting unit symbols into readable words) and typically converts the text into a corresponding sequence of phonemes. The acoustic model is configured to convert the input text (or converted phonemes) into a spectral sequence, while the speech encoder is configured to convert the spectral sequence into speech waveform data. Furthermore, the acoustic model determines how the text will be pronounced (e.g., with what rhythm, timbre, etc.).

[0002] Prosody generally refers to patterns of rhythm and sound, or patterns of stress and / or intonation in a language. For example, in linguistics, prosody is responsible for the properties of syllables and larger speech units (i.e., larger than individual speech segments). Prosody is typically characterized by variations in loudness, pauses, and rhythm (e.g., speaking rate). Speakers can also express prosody by altering pitch (i.e., the quality of the sound, determined by the rate of vibration that produces the sound, or in other words, the degree to which the pitch is high or low). In some instances, pitch refers to the fundamental frequency associated with a particular segment of speech. Prosody is also expressed by altering the energy of speech. Energy generally refers to the energy of the speech signal (i.e., the power fluctuations of the speech signal). In some instances, energy is based on the volume or amplitude of the speech signal.

[0003] In music, timbre (i.e., pitch quality) generally refers to the characteristic or quality of a musical sound or tone, and the characteristics associated with timbre are different from the pitch and intensity of a musical sound. Timbre is the quality that allows the human ear to distinguish a violin from a flute (or even the more subtle viola). In the same way, the human ear can distinguish different sounds with different timbres.

[0004] The source acoustic model is configured as a multi-speaker model trained on multi-speaker data. In some cases, the source acoustic model is further refined or adapted using target speaker data. Typically, the acoustic model is speaker-dependent, meaning that the acoustic model is trained directly on speaker data from a specific target speaker, or the source acoustic model is refined using speaker data from a specific target speaker.

[0005] When well-trained, this model can convert any text into speech that closely mimics how the target speaker speaks, i.e., with the same timbre and similar rhythm. Training data for TTS models typically includes audio data obtained by recording the specific target speaker while they are speaking, along with a set of text corresponding to that audio data (i.e., a textual representation of what the target speaker said to produce that audio data).

[0006] In some instances, the text used to train the TTS model is generated by a speech recognition model and / or a natural language understanding model, which is specifically configured to recognize and interpret speech and provide text representations of the words identified in the audio data. In other instances, the speaker is given a pre-recorded text to read aloud, and this pre-recorded text and the corresponding audio data are used to train the TTS model.

[0007] Note that the target speaker can produce speech in various ways and styles. For example, an individual may speak very quickly when excited or stutter when nervous. Additionally, an individual may speak differently when conversing with a friend compared to when reciting to an audience.

[0008] If a user wants a trained model to speak in a specific style or with a particular emotional impact, such as being happy or sad, speaking like a news anchor, a speaker, or a storyteller, then it is necessary to train the model using training data with that corresponding target style. For example, first, recordings of the target speaker using the target style must be collected, and then the user can use the training data with that style to construct the corresponding voice font.

[0009] Initially, thousands of hours are required to build the source acoustic model. Then, a large amount of training data is needed to correctly train the TTS model for a specific style. In some instances, training / refinement of the source acoustic model for a specific style may require hundreds, sometimes thousands, of sentences of speech training data. Therefore, to correctly train (or similar) TTS models for multiple different styles, a proportional amount of training data must be collected for each of the different target speaker styles. This is an extremely time-consuming and costly process for recording and analyzing data for each desired style. Furthermore, in some instances, the target speaker is unable or ill-suited to producing speech in the desired target style, further complicating the training of the acoustic model. This poses a significant obstacle to quickly and efficiently training TTS models with voice fonts of (or similar) different target speaking styles.

[0010] In light of the above, there is a continued need for improved systems and methods for generating training data for TTS models and training models to produce speech data of multiple speaking styles for one or more target speakers.

[0011] The subject matter claimed herein is not limited to the embodiments that address any shortcomings or operate only in environments such as those described above. Rather, this background is provided merely to illustrate an exemplary technical field in which some of the embodiments described herein may be practiced.

[0012] Brief Overview

[0013] The disclosed embodiments relate to various embodiments for cross-speaker style transfer in text-to-speech and for generating training data. In some instances, the disclosed embodiments include generating and utilizing spectrogram data in which the target speaker employs a specific prosodic style. In some instances, the spectrogram data is used to train a machine learning model for text-to-speech (TTS) conversion.

[0014] Some embodiments include methods and systems for receiving electronic content, which includes source speaker data from a source speaker. In these embodiments, a computing system converts the waveform of the source speaker data into PPG data by aligning the waveform of the source speaker data with phonetic posterior gram (PPG) data, wherein the PPG data defines one or more features corresponding to the prosodic pattern of the source speaker data.

[0015] In addition to one or more features defined by the PPG data, one or more additional prosodic features are extracted from the source speaker data. The computational system then generates a spectrogram based on (i) the PPG data, (ii) the one or more additional prosodic features extracted, and (iii) the target speaker's vocal timbre. Using this technique, the generated spectrogram is characterized by the prosodic pattern of the source speaker and the vocal timbre of the target speaker.

[0016] In some instances, the disclosed embodiments relate to embodiments for training a voice translation machine learning model to generate spectrogram data for cross-speaker style delivery. Additionally, some embodiments relate to systems and methods for training a neural TTS model on training data generated from spectrogram data.

[0017] This summary is provided to introduce, in a simplified form, a selection of concepts further described in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to help determine the scope of the claimed subject matter.

[0018] Additional features and advantages will be set forth in the following description, and in part will be obvious from the description, or may be learned by practice of the teachings herein. The features and advantages of the invention may be realized and obtained by means of the tools and combinations particularly pointed out in the appended claims. The features of the invention will become more apparent from the following description and the appended claims, or may be learned by practice of the invention as set forth below. Brief description of the attached diagram

[0019] To describe how the above-described and other advantages and features can be obtained, a more specific description of the subject matter briefly described above will be presented with reference to various specific embodiments, which are illustrated in the accompanying drawings. It should be understood that these drawings only describe typical embodiments and are therefore not intended to limit the scope of the invention. The embodiments will be described and explained with additional specificity and detail using the drawings, in which:

[0020] Figure 1 An example is illustrated of a computing environment in which a computing system is incorporated and / or used to perform the disclosed aspects of the various embodiments. The illustrated computing system is configured for voice conversion and includes hardware storage devices and multiple machine learning engines. The computing system communicates with remote / third-party systems.

[0021] Figure 2 An embodiment is illustrated with a flowchart of multiple actions associated with a method for generating machine learning training data, which includes spectrogram data of a target speaker.

[0022] Figure 3 An embodiment is illustrated with a diagram showing multiple actions associated with various methods for aligning waveforms of source speaker data with corresponding speech posterior graph (PPG) data.

[0023] Figure 4 An embodiment is illustrated with diagrams showing multiple actions associated with various methods for extracting additional prosodic features from source speaker data.

[0024] Figure 5 An embodiment is illustrated with flowcharts of multiple actions associated with various methods for training a voice conversion machine learning model (including a PPG spectrogram component of the voice conversion machine learning model).

[0025] Figure 6 An embodiment is illustrated with flowcharts of various actions associated with methods for training a neural TTS model on spectrogram data generated for a target speaker employing a specific prosodic pattern.

[0026] Figure 7An example is illustrated with a flowchart of multiple actions for generating speech data from text using a trained TTS model.

[0027] Figure 8 An embodiment of a process flowchart illustrating a high-level view of generating training data and training a neural TTS model is shown.

[0028] Figure 9 An embodiment of an example process flowchart including training a voice conversion model within a speech recognition module is illustrated. This voice conversion model includes an MFCC-PPG component and a PPG-Mel component.

[0029] Figure 10 An example configuration of a neural TTS model according to the various embodiments disclosed herein is illustrated.

[0030] Figure 11 An embodiment of example waveform-to-PPG component (e.g., MFCC-PPG) is illustrated, wherein the computing system generates PPG data.

[0031] Figure 12 An example of a sample PPG-spectrum (PPG-Mel) component for a sound conversion model is illustrated. Detailed Implementation

[0032] Some embodiments of the disclosed examples include generating spectrogram data having a specific vocal timbre of a first speaker (e.g., a target speaker) and a specific prosodic pattern transmitted from a second speaker (e.g., a source speaker).

[0033] For example, in some embodiments, the computing system receives electronic content comprising source speaker data obtained from a source speaker. The waveform of the source speaker data is converted into speech data (e.g., a speech posterior graph, PPG). The PPG data is aligned with the waveform speaker data and defines one or more features corresponding to the prosodic pattern of the source speaker. In addition to the one or more features defined by the PPG data, the computing system also extracts one or more additional prosodic features from the source speaker data. Then, based on (i) the PPG data, (ii) the additional extracted prosodic features, and (iii) the timbre of the target speaker, the computing system generates a spectrogram having the timbre of the target speaker and the prosodic pattern of the source speaker.

[0034] There are numerous technical benefits associated with the disclosed embodiments. For example, spectrogram data can be generated at a more efficient rate because it can be generated based on datasets obtained from both the target speaker and source speakers. Furthermore, spectrogram data can be generated based on any of the prosodic patterns obtained from multiple source speakers. In this way, spectrogram data can incorporate the timbre of the target speaker's voice as well as any prosodic pattern defined by the source speaker datasets. In some embodiments, the timbre is configured as a speaker vector that takes into account the timbre of a particular speaker (e.g., the target speaker). This is a highly versatile approach and reduces both the computational time required for data generation and the time and cost of initial speaker data collection.

[0035] The technical benefits of the disclosed embodiments also include the training of neural TTS models and the generation of speech output from text-based input using neural TTS models. For example, because the disclosed methods are used to generate spectrogram data, it is preferable that large datasets for correctly training the TTS model can be obtained more quickly than conventional methods.

[0036] Additional benefits and functionality of the disclosed embodiments will be described below, including training of the voice conversion model and methods for aligning PPG data (and other prosodic feature data) with source speaker data at a frame-based granularity.

[0037] Now let's turn our attention to... Figure 1 , Figure 1 Examples of components of a computing system 110 that may include and / or be used to implement aspects of the disclosed invention are illustrated. As shown, the computing system includes multiple machine learning (ML) engines, models, and data types associated with the inputs and outputs of the machine learning engines and models.

[0038] First, shift your attention Figure 1 , Figure 1 A computing system 100, as part of a computing environment 100, is illustrated. The computing environment 100 also includes remote / third-party systems 120 communicating with the computing system 110 (via network 130). The computing system 110 is configured to train multiple machine learning models for speech recognition, natural language understanding, text-to-speech, and more specifically, cross-speaker style delivery applications. The computing system 110 is also configured to generate training data configured to train the machine learning models to generate speech data for a target speaker, characterized by the target speaker's timbre and the prosodic patterns of a specific source speaker. Additionally or alternatively, the computing system is configured to run the trained machine learning models for text-to-speech generation.

[0039] The computing system 110 includes, for example, one or more processors 112 (such as one or more hardware processors) and storage 140 (i.e., hardware storage devices) storing computer-executable instructions 118, wherein the storage 140 is capable of accommodating any number of data types and any number of computer-executable instructions 118, the computing system 110 being configured to implement one or more aspects of the disclosed embodiments by means of the computer-executable instructions 118 when the computer-executable instructions 118 are executed by the one or more processors 112. The computing system 110 is also shown to include user interfaces and input / output (I / O) devices 116.

[0040] Storage 140 is shown as a single storage unit. However, it will be appreciated that in some embodiments, storage 140 is a distributed storage system distributed across several separate and sometimes remote / third-party systems 120. In some embodiments, system 110 may also include a distributed system in which one or more system 110 components are maintained / operated by different discrete systems that are geographically isolated from each other and each performs different tasks. In some instances, multiple distributed systems perform similar and / or shared tasks to achieve the disclosed functionality, such as in a distributed cloud environment.

[0041] In some embodiments, storage 140 is configured to store one or more of the following: target speaker data 141, source speaker data 142, PPG data 143, spectrogram data 144, prosodic feature data 145, neural TTS model 146, voice conversion model 147, (various) executable instructions 118, or prosodic patterns 148.

[0042] In some instances, storage 140 includes computer-executable instructions 118 for instantiating or executing one or more of the models and / or engines shown in computing system 110. In some instances, the one or more models are configured as machine learning models or machine learning-enabled models. In some instances, the one or more models are configured as deep learning models and / or algorithms. In some instances, the one or more models are configured as engines or processing systems (e.g., computing systems integrated within computing system 110), wherein each engine (i.e., model) includes one or more processors (e.g., hardware processor 112) and corresponding computer-executable instructions 118.

[0043] In some embodiments, target speaker data 141 includes electronic content / data obtained from a target speaker, while source speaker data 142 includes electronic content / data from a source speaker. In some instances, target speaker data 141 and / or source speaker data 142 include audio data, text data, and / or visual data. Additionally or alternatively, in some embodiments, target speaker data 141 and / or source speaker data 142 include metadata (i.e., attributes, information, speaker identifiers, etc.) corresponding to a specific speaker from whom data was collected. In some embodiments, this metadata includes attributes associated with the speaker's identity, characteristics of the speaker and / or the speaker's voice, and / or information about where, when, and / or how the speaker data was obtained.

[0044] In some embodiments, the target speaker data 141 and / or the source speaker data 142 are raw data (e.g., direct recordings). Additionally or alternatively, in some embodiments, the target speaker data 141 and / or the source speaker data 142 include processed data (e.g., waveform format of the speaker data and / or PPG data corresponding to the target and / or source speaker (e.g., PPG data 143)).

[0045] In some embodiments, PPG data 143 includes speech information about speech data from a specific speaker (e.g., a source speaker and / or a target speaker). In some instances, the speech information is obtained at a determined granularity, such as frame-based granularity. In other words, a speech posterior map is generated for each frame so that the source speaker's speech diachronic information (i.e., the source prosodic pattern) is accurately maintained during voice transformation and pattern transfer.

[0046] In some embodiments, the frame length of each voice message includes a complete voice phrase, a complete voice word, a specific voice phoneme, and / or a predetermined time duration. In some examples, the frame includes a time duration selected between 1 millisecond and 10 seconds, or preferably between 1 millisecond and 1 second, or even more preferably between 1 millisecond and 50 milliseconds, or even more preferably about 12.5 milliseconds.

[0047] In some embodiments, PPG data 143 is generated by a voice conversion model or a component of a voice conversion model (e.g., an MFCC-PPG model) in which speech information (e.g., waveforms of the source speaker data) is extracted from source speaker data. In some embodiments, PPG data 143 is input to a voice conversion model configured to generate spectrogram data (e.g., spectrogram data 144), more specifically to a PPG-Mel model.

[0048] The generated spectrogram data will have the same content as the source data, while maintaining the integrity of the timing alignment between PPG data 143 and spectrogram data 144. Thus, in some instances, PPG data 143 includes one or more prosodic features (i.e., prosodic attributes), wherein the one or more prosodic attributes include diachronic information (e.g., speech duration, timing information, and / or speech rate).

[0049] In some embodiments, prosodic attributes extracted from PPG data are included in prosodic feature data 145. Additionally or alternatively, prosodic feature data 145 includes additional prosodic features or prosodic attributes. For example, in some instances, additional prosodic features include attributes corresponding to pitch and / or energy profiles of the speech waveform data.

[0050] In some embodiments, spectrogram data 144 includes multiple spectrograms. Typically, a spectrogram is a visual representation of the spectrum of a signal's frequencies as they change over time (e.g., the spectrum of frequencies constituting speaker data). In some instances, spectrograms are sometimes referred to as spectrographs, voiceprints, or voice reports. In some embodiments, the spectrograms included in spectrogram data 144 are characterized by the timbre and prosodic pattern of the target speaker's voice. Additionally or alternatively, the spectrograms included in spectrogram data 144 are characterized by the timbre of the target speaker's voice and the prosodic pattern of the source speaker.

[0051] In some embodiments, the spectrogram is converted to a Mel scale. The Mel scale is a non-linear scale of pitch determined by listeners at equal distances from each other and more closely mimics human response to sound / human sound recognition than a linear scale of frequency. In such embodiments, the spectrogram data includes Mel frequency cepstral (MFC) data (i.e., a representation of the short-term power spectrum of sound based on the linear cosine transform of the logarithmic power spectrum on a non-linear Mel scale of frequency). Therefore, Mel frequency cepstral coefficients (MFCC) are coefficients that include the MFC. For example, frequency bands are evenly spaced on the Mel scale for the MFC.

[0052] In some embodiments, hardware storage device 140 stores a neural TTS model 146, which is configured to be trained or trained to convert input text into speech data. For example, a portion of an email containing one or more sentences (e.g., a specific number of machine-recognizable words) is applied to the neural TTS model, wherein the model is capable of recognizing words or portions of words (e.g., phonemes) and is trained to produce sounds corresponding to those phonemes or words.

[0053] In some embodiments, the neural TTS model 146 is adapted for a specific target speaker. For example, target speaker data (e.g., target speaker data 141) includes audio data, including spoken words and / or phrases obtained and / or recorded from the target speaker. An example of the neural TTS model 1000 is referenced below. Figure 10 To describe in more detail.

[0054] In some instances, target speaker data 141 is formatted as training data, wherein neural TTS model 146 is trained (or pre-trained) on the target speaker training data to enable neural TTS model 146 to generate speech data adopting the vocal timbre and prosodic style of the target speaker based on the input text. In some embodiments, neural TTS model 146 is speaker-independent, meaning that the model generates arbitrary speech data based on one or a combination of target speaker datasets (e.g., target speaker data 141). In some embodiments, neural TTS model 146 is a multi-speaker neural network, meaning that the model is configured to generate speech data corresponding to multiple discrete speakers / speaker profiles. In some embodiments, neural TTS model 146 is speaker-dependent, meaning that the model is configured to generate speech primarily for a specific target speaker.

[0055] In some embodiments, the neural TTS model 146 is further trained and / or adapted such that the model is trained on training data including and / or based on a combination of target speaker data 141 and source speaker data 142, so that the neural TTS model 146 is configured to produce speech data that adopts the vocal timbre of the target speaker and the prosodic patterns of the source speaker data.

[0056] In some embodiments, a database is provided that stores multiple voice timbre profiles (e.g., voice timbre 149) corresponding to multiple target speakers and multiple prosodic patterns (e.g., prosodic pattern 148) corresponding to multiple source speakers. In some instances, a user can select a specific voice timbre profile from the multiple voice timbre profiles and a prosodic pattern from the multiple prosodic patterns, wherein a neural TTS model 146 is configured to convert input text into speech data based on the specific voice timbre and the specific prosodic pattern. In such embodiments, it should be understood that any number of combinations of voice timbre 149 and prosodic pattern 148 exist.

[0057] In some embodiments, the newly generated prosodic pattern is based on a combination of previously stored prosodic patterns and / or a combination of source speaker datasets. In some embodiments, the newly generated vocal timbre is based on a combination of previously stored vocal timbres and / or a combination of target speaker datasets.

[0058] In some embodiments, a prosodic style refers to a set or subset of prosodic attributes. In some instances, prosodic attributes correspond to a specific speaker (e.g., a target speaker or a source speaker). In some instances, a specific prosodic style is assigned an identifier, such as a name identifier. For example, a prosodic style is associated with a name identifier that identifies the speaker from whom the prosodic style was generated / obtained. In some examples, prosodic styles include descriptive identifiers such as a storytelling style (e.g., the way of speaking typically used when reading a novel aloud or associating a story as part of a speech or dialogue), a news anchor style (e.g., the way of speaking typically used by news anchors when delivering news in an objective, unemotional, and direct style), a speech style (e.g., a formal speaking style typically used when a person is giving a speech), and a conversational style (e.g., a colloquial speaking style typically used when a person is speaking to a friend or relative). Additional styles include, but are not limited to, serious, casual, and customer service styles. It will be understood that any other type of speaking style besides those listed may also be used to train an acoustic model using corresponding training data for said styles(s).

[0059] In some embodiments, prosodic patterns are attributed to emotions typically expressed by humans, such as happiness, sadness, excitement, tension, or other emotions. Often, a particular speaker is experiencing a specific emotion, and therefore the way that speaker speaks is influenced by that specific emotion in a way that indicates to the listener that the speaker is experiencing such an emotion. As an example, a speaker who is feeling angry might speak in a highly energetic manner, at a loud volume, and / or with a truncated tone. In some embodiments, a speaker may wish to convey a specific emotion to the audience, in which case the speaker will consciously choose to speak in a certain way. For example, a speaker might wish to instill a sense of awe in the audience and will speak in a quiet, respectful tone with a slower, more fluent voice. It should be understood that in some embodiments, prosodic patterns are not further categorized or defined by descriptive identifiers.

[0060] In some embodiments, hardware storage device 140 stores a voice conversion model 147 configured to convert speech data from a first speaker (e.g., a source speaker) into speech that sounds like that of a second speaker (e.g., a target speaker). In some embodiments, the converted speech data is adapted to the timbre of the target speaker while maintaining the prosodic pattern of the source speaker. In other words, the converted speech mimics the voice of the target speaker (i.e., timbre) but retains one or more prosodic attributes of the source speaker (e.g., speech duration, pitch, energy, etc.).

[0061] Additional storage units for storing the Machine Learning (ML) Engine 150 Figure 1The system is demonstrated to store multiple machine learning models and / or engines. For example, the computing system 110 includes one or more of the following: a data retrieval engine 151, a transformation engine 152, a feature extraction engine 153, a training engine 154, an alignment engine 155, an implementation engine 156, a refinement engine 157, or a decoding engine 158, which are individually and / or collectively configured to implement the different functionalities described herein.

[0062] For example, in some instances, data retrieval engine 151 is configured to locate and access data sources, databases, and / or storage devices from which it may extract datasets or subsets of data to be used as training data, including one or more data types. In some instances, data retrieval engine 151 receives data from databases and / or hardware storage devices, wherein data retrieval engine 151 is configured to reformat or otherwise augment the received data for use as training data. Additionally or alternatively, data retrieval engine 151 communicates with remote / third-party systems (e.g., remote / third-party system 120) that include remote / third-party datasets and / or data sources. In some instances, these data sources include audio / video services that record speech, text, images, and / or video to be used in cross-speaker style delivery applications.

[0063] In some embodiments, the data retrieval engine 151 accesses electronic content including target speaker data 141, source speaker data 142, PPG data 143, spectrogram data 144, prosodic feature data 145, prosodic pattern 148, and / or voice timbre 149.

[0064] In some embodiments, the data retrieval engine 151 is an intelligent engine capable of learning the optimal dataset extraction process and providing sufficient data in a timely manner, as well as retrieving data most suitable for the machine learning model / engine to be trained for the desired application. For example, the data retrieval engine 151 can learn which databases and / or datasets will generate training data to train a model (e.g., for a specific query or task) to improve the accuracy, efficiency, and effectiveness of that model in the desired natural language understanding application.

[0065] In some instances, data retrieval engine 151 locates, selects, and / or stores raw, unstructured source data (e.g., speaker data), wherein data retrieval engine 151 communicates with one or more other ML engines and / or models (e.g., transformation engine 152, feature extraction engine 153, training engine 154, etc.) included in computing system 110. In such instances, other engines communicating with data retrieval engine 151 are able to receive data that has been retrieved (i.e., extracted, pulled, etc.) from one or more data sources, so that the received data can be further augmented and / or applied to downstream processes.

[0066] For example, in some embodiments, data retrieval engine 151 communicates with conversion engine 152. Conversion engine 152 is configured to convert between data types and transform raw data into training data that can be used to train any machine learning model described herein. The conversion model beneficially transforms data to improve the efficiency and accuracy of model training. In some embodiments, conversion engine 152 is configured to receive speaker data (e.g., source speaker data 142) and transform the raw speaker data into waveform data. Additionally, conversion engine 152 is configured to transform the waveform of the source speaker data into PPG data. Additionally or alternatively, in some embodiments, conversion engine 152 is configured to facilitate the conversion of speech data from a first speaker to a second speaker (e.g., speech conversion via a speech conversion model).

[0067] In some embodiments, computing system 110 stores and / or accesses feature extraction engine 153. Feature extraction engine 153 is configured to extract features and / or attributes from target speaker data 141 and source speaker data 142. These extracted attributes include attributes corresponding to speech information, prosodic information, and / or timbre information. In some embodiments, feature extraction engine 153 extracts one or more additional prosodic features from the source speaker data, including pitch profiles and / or energy profiles of the source speaker data. In such embodiments, the extracted attributes are included in a training dataset configured to train a machine learning model.

[0068] In some embodiments, the feature extraction engine 153 is configured to receive electronic content comprising a plurality of prosodic features and / or attributes, wherein the feature extraction engine 153 is configured to detect discrete attributes and distinguish specific attributes from one another. For example, in some instances, the feature extraction engine 153 is capable of distinguishing between a pitch attribute corresponding to a pitch profile of the source speaker data and an energy attribute corresponding to an energy profile of the source speaker data.

[0069] In some embodiments, the data retrieval engine 151, the transformation engine 152, and / or the feature extraction engine 153 communicate with the training engine 154. The training engine 154 is configured to receive one or more training datasets from the data retrieval engine 151, the transformation engine 152, and / or the feature extraction engine 153. After receiving training data relevant to a specific application or task, the training engine 154 trains one or more models on that training data for a specific natural language understanding application, speech recognition application, speech generation application, and / or cross-speaker style delivery application. In some embodiments, the training engine 154 is configured to train the model via unsupervised or supervised training.

[0070] In some embodiments, based on attributes extracted by feature extraction engine 153, training engine 154 can adapt the training process and method such that the training process produces a trained model configured to generate specialized training data that reflects specific features and attributes that contribute to the desired prosodic pattern. For example, including pitch attributes helps determine the fundamental frequency from which to generate spectrogram data, while including energy attributes helps determine at what volume (or volume variation) to generate the spectrogram data. Each attribute contributes differently to the overall prosodic pattern.

[0071] For example, in some embodiments, training engine 154 is configured to train a model (e.g., neural TTS model 146, see also) using training data (e.g., spectrogram data 144). Figure 10 Model 1000 is configured such that the machine learning model is configured to generate speech from arbitrary text as described in the various embodiments herein. In some examples, training engine 154 is configured to train voice conversion model 147, or components of voice conversion model, on speaker data (e.g., target speaker data 141, source speaker data 142, or multi-speaker data).

[0072] In some embodiments, the conversion engine 152 and / or training engine 154 communicate with the alignment engine 155. The alignment engine 155 is configured to align the waveform of the source speaker data 142 with the PPG data 143 at a specific granularity (e.g., frame-based granularity). The alignment engine 155 is also configured to align one or more additional prosodic features (e.g., pitch, energy, speech rate, speech duration) extracted from the source speaker data with the PPG data 143 at the same granularity used to align the PPG data 143 with the source speaker data 142. Aligning the data in this manner advantageously maintains the integrity of the source speaker's prosodic pattern during pattern transfer.

[0073] In some embodiments, the computing system 110 includes a refinement engine 157. In some instances, the refinement engine 157 communicates with a training engine. The refinement engine 157 is configured to refine the voice conversion model or its components (e.g., PPG-spectrum components) by adapting model components (or sub-components) of the voice conversion model to the target speaker using target speaker data 141.

[0074] In some embodiments, the computing system 110 includes a decoding engine 158 (or an encoder-decoder engine) configured to encode and decode data. Generally, a decoder takes feature maps, vectors, and / or tensors from an encoder and generates a neural network that best matches the expected input. In some embodiments, the encoder / decoder engine 158 is configured to encode text input to the neural TTS model 146 and decode the encoding to convert the input text into a Mel spectrum. (See also...) Figure 10 In some embodiments, the encoding / decoding engine 158 is configured to encode the PPG data 143 as part of the spectrogram generation process. (See also...) Figure 12 ).

[0075] In some embodiments, the computing system 110 includes a separate encoding engine (not shown) configured to learn and / or run a shared encoder among one or more models. In some embodiments, the encoder is a neural network that takes input and outputs feature maps, vectors, and / or tensors. In some embodiments, the shared encoder is part of an encoder-decoder network.

[0076] In some embodiments, the decoding engine 158 communicates with the refinement engine 157, which is configured to refine the encoder / decoder network of the neural TTS model 146 by employing a feedback loop between the encoder and decoder. The neural TTS model 146 is then trained and refined by iteratively minimizing the reconstruction loss incurred in converting the input text into speech data and the speech data back into text data. In some embodiments, the refinement engine 157 is also configured to refine and / or optimize any or a combination of machine learning engines / models included in the computing system 110 to improve the efficiency, effectiveness, and accuracy of the engine / model.

[0077] In some embodiments, the computing system 110 includes an implementation engine 156 that communicates with any (or all) of the models and / or ML engines 150 included in the computing system 110, such that the implementation engine 156 is configured to implement, initiate, or run one or more functions of the plurality of ML engines 150. In one example, the implementation engine 156 is configured to run a data retrieval engine 151 such that the data retrieval engine 151 retrieves data at appropriate times to generate training data for training engine 154.

[0078] In some embodiments, implementation engine 156 facilitates communication processes and timing between one or more of the ML engines 150. In some embodiments, implementation engine 156 is configured to implement a speech conversion model to generate spectrogram data. Additionally or alternatively, implementation engine 156 is configured to perform natural language understanding tasks by performing text-to-speech data conversion (e.g., via a neural TTS model).

[0079] In some embodiments, the computing system communicates with a remote / third-party system 120 including one or more processors 122 and one or more computer-executable instructions 124. In some instances, the remote / third-party system 120 may further include a database containing data that can be used as training data (e.g., external speaker data). Additionally or alternatively, the remote / third-party system 120 includes a machine learning system external to the computing system 110. In some embodiments, the remote / third-party system 120 is a software program or application.

[0080] Now let's turn our attention to... Figure 2 , Figure 2 Examples include those that can be computed by computing systems (such as those mentioned above). Figure 1 The flowchart 200 describes the various actions associated with the exemplary method implemented by the computing system 110. (As described in...) Figure 2 As shown, flowchart 200 includes multiple actions (actions 210, 220, 230, 240, and 250) associated with methods for generating training data and training machine learning models for natural language understanding tasks (e.g., converting text into speech data). Example references to computing systems (e.g., [example references would be inserted here]) are provided for the components claimed in each action. Figure 1 The characteristics of the computing system 110 are described.

[0081] like Figure 2As shown, flowchart 200 and the corresponding method include an action (action 210) in which a computing system (e.g., computing system 110) receives electronic content including source speaker data (e.g., source speaker data 142) from a source speaker. Upon receiving the source speaker data, the computing system converts the waveform of the source speaker data into PPG data (e.g., PPG data 143) by aligning the waveform of the source speaker data with speech posterior graph (PPG) data, wherein the PPG data defines one or more features corresponding to the prosodic pattern of the source speaker data (action 220).

[0082] Flowchart 200 also includes an action (action 230) to extract one or more additional prosodic features (e.g., prosodic feature data 145) from the source speaker data. Subsequently, the computing system generates a spectrogram (e.g., spectrogram data 144) based on the PPG data, the extracted one or more additional prosodic features, and the target speaker's timbre, wherein the spectrogram is characterized by the source speaker's prosodic pattern (e.g., prosodic pattern 148) and the target speaker's timbre (e.g., timbre 149) (action 240). Spectrograms, such as audio or video spectrograms, are known to those skilled in the art and include digital representations of sound attributes, such as the frequency spectrum as the frequency of a particular sound or other signal changes over time. In the current embodiment, the spectrogram is characterized by a specific prosodic pattern (e.g., the source speaker's prosodic pattern) and timbre (e.g., the target speaker's timbre).

[0083] In some embodiments, the computing system 250 uses the generated spectrogram to train a neural text-to-speech (TTS) model (e.g., neural TTS model 146), wherein the neural TTS model is configured to generate speech data from arbitrary text, the speech data being characterized by the prosodic pattern of the source speaker and the timbre of the target speaker (action 250).

[0084] An example of a TTS model that can be trained is the Neural TTS Model 100, such as Figure 10 As shown, it includes a text encoder 1020 and a decoder 1040, wherein the model uses attention 1030 to guide and inform the encoding-decoding at various layers of the model (e.g., phoneme and / or frame layers, and context layers). The neural TTS model 1000 is capable of generating outputs in a Mel spectrum (e.g., spectrogram data or speech waveform data) such that the generated output is based on the speech data of the input text 1010. The Mel spectrum 1050 is characterized by the timbre of the voice of a first speaker (e.g., the target speaker) and the prosodic patterns of a second speaker (e.g., the source speaker).

[0085] about Figure 2The actions described herein will be understood to be capable of being performed in a different order than that explicitly shown in flowchart 200. For example, actions 210 and 220 may be performed in parallel with action 230, and in some alternative embodiments, actions 210 and 220 may be performed serially with actions 230, 240 and 250.

[0086] It will also be understood that the actions to perform natural language understanding tasks can be performed by the same computer device that performs the aforementioned actions (e.g., actions 210-250), or alternatively by one or more different computer devices in the same distributed system.

[0087] Now let's turn our attention to... Figure 3 , Figure 3 Examples include those that can be computed by computing systems (such as those mentioned above). Figure 1 Illustration 300 shows variations of actions associated with the exemplary methods implemented by the described computing system 110. For example... Figure 3 As shown, Figure 300 includes multiple actions (actions 320, 330, and 340) associated with various methods for performing the action (action 310) of converting waveform data into PPG data. Examples of the components claimed in each action are referenced to a computing system (e.g., Figure 1 The various features of the computing system 110 are described. It should be understood that in some embodiments, action 310 represents... Figure 2 Action 220.

[0088] For example, Figure 300 includes an action of converting the waveform of the source speaker data into PPG data (e.g., PPG data 143) by aligning the waveform of the source speaker data with speech posterior graph (PPG) data, wherein the PPG data defines one or more features corresponding to the prosodic pattern of the source speaker data (action 310).

[0089] In some embodiments, the computing system aligns the waveform of the source speaker data with a granularity that is narrower than that of phoneme-based granularity (Action 320). In some embodiments, the computing system aligns the waveform of the source speaker data with PPG data with a frame-based granularity (Action 330). In some embodiments, the computing system aligns the waveform data of the source speaker data with PPG data with a frame-based granularity based on a specific frame rate, such as a frame rate of 12.5 milliseconds, or a frame rate with a shorter or longer duration, for example (Action 340).

[0090] Now let's turn our attention to... Figure 4 , Figure 4 Examples include those that can be computed by computing systems (such as those mentioned above). Figure 1Illustration 400 shows variations of actions associated with the exemplary methods implemented by the described computing system 110. For example... Figure 4 As shown, Figure 400 includes multiple actions (actions 420, 430, 440, 450, and 460) associated with various methods for performing actions (actions 410) to extract additional prosodic features. In some instances, action 410 represents... Figure 2 Action 230.

[0091] For example, Figure 400 includes the action (action 410) of extracting one or more additional prosodic features (e.g., prosodic feature data 145) from source speaker data (e.g., source speaker data 142) in addition to one or more features defined by PPG data (e.g., PPG data 143). In some embodiments, the computing system extracts additional prosodic features including pitch (action 420). Additionally or alternatively, the computing system extracts additional prosodic features including speech duration (action 430). Additionally or alternatively, the computing system extracts additional prosodic features including energy (action 440). Additionally or alternatively, the computing system extracts additional prosodic features including speech rate (action 450). Furthermore, in some embodiments, the computing system extracts one or more additional prosodic features at a frame-based granularity (action 460).

[0092] Now let's turn our attention to... Figure 5 , Figure 5 Examples include those that can be computed by computing systems (such as those mentioned above). Figure 1 The flowchart 500 describes the various actions associated with the exemplary methods implemented by the computing system 110. For example... Figure 5 As shown, flowchart 500 includes multiple actions (actions 510, 520, 530, 540, and 550) associated with various methods for training machine learning models for natural language understanding tasks, such as training and generating spectrograms using a PPG-spectrum graph component for voice conversion machine learning models.

[0093] like Figure 5 As shown, flowchart 500 and the corresponding method include actions of a computing system (e.g., computing system 110) training a speech posterior map (PPG) to spectrogram component of a speech conversion machine learning model, wherein the PPG to spectrogram component of the speech conversion machine learning model is initially trained on multi-speaker data during training and configured to convert PPG data into spectrogram data (action 510).

[0094] After training the PPG-to-spectrum component, the computational system refines the PPG-to-spectrum component (Action 520) by adapting the PPG data from the target speaker's data, which has a specific vocal timbre and a specific prosodic pattern, into spectrogram data with the specific prosodic pattern of the target speaker, using target speaker data from the target speaker.

[0095] Flowchart 500 further includes: an action of receiving electronic content comprising new PPG data converted from the waveform of the source speaker data, wherein the new PPG data is aligned with the waveform of the source speaker data (action 530); and an action of receiving one or more prosodic features extracted from the waveform of the source speaker data (action 540). In some embodiments, actions 510, 520, 530 and / or 540 are performed sequentially. In some embodiments, as shown, actions 510 and 520 are performed sequentially, and actions 530 and 540 are performed independently of each other and independently of actions 510 and 520.

[0096] After performing actions 510-540, the computing system applies the source speaker data to the voice conversion machine learning model, wherein the refined PPG-to-spectrum component of the voice conversion machine learning model is configured to generate a spectrogram that adopts the specific vocal timbre of the target speaker but has a new prosodic pattern of the source speaker rather than the specific prosodic pattern of the target speaker (action 550). In some embodiments, action 530 represents Figure 2 Actions 210 and 220. In some embodiments, action 540 represents... Figure 2 Action 230. In some embodiments, Figure 2 The action 240 for generating spectrogram data is performed by one or more actions included in method 500.

[0097] Now let's turn our attention to... Figure 6 , Figure 6 Examples include those that can be computed by computing systems (such as those mentioned above). Figure 1 , 8 -12 describes a flowchart 600 of the various actions associated with the exemplary methods implemented by the computing system 110. For example... Figure 6 As shown, flowchart 600 includes multiple actions (actions 610, 620, 630, 640, and 650) associated with various methods for training machine learning models for natural language understanding tasks (e.g., generating training data configured to train neural TTS models).

[0098] As an example, method 600 includes the action (action 610) of receiving electronic content including source speaker data (e.g., source speaker data 142) from a source speaker. The computing system then converts the waveform of the source speaker data into speech posterior map (PPG) data (e.g., PPG data 143). Method 600 further includes the action (action 630) of extracting one or more prosodic features (e.g., prosodic feature data 145) from the source speaker data.

[0099] In (for example, using) Figure 9 After the MFCC-PPG independent of the speaker model converts the waveform into PPG data and extracts additional prosodic features, the computational system applies at least the PPG data and one or more of the extracted prosodic features to the pre-trained PPG-to-spectrum component of the voice conversion module (e.g., MFCC-PPG independent of the speaker model). Figure 9 The PPG-Mel model, which is pre-trained, is configured to generate spectrograms (e.g., spectrogram data) that adopt the specific vocal timbre of the target speaker (e.g., vocal timbre 149) but have a new prosodic pattern of the source speaker (e.g., prosodic pattern 148) instead of the specific prosodic pattern of the target speaker (action 640).

[0100] The computing system also generates training data configured to train a neural TTS model (e.g., neural TTS model 146), which includes multiple spectrograms (actions 650) characterized by the specific vocal timbre of the target speaker and the new prosodic pattern of the source speaker.

[0101] Optionally, in some embodiments, the computing system trains a neural TTS model on the generated training data such that the neural TTS model is configured to generate speech data from arbitrary text by performing cross-speaker style transfer, wherein the speech data is characterized by the prosodic style of the source speaker and the vocal timbre of the target speaker (action 650).

[0102] Now let's turn our attention to... Figure 7 , Figure 7 Examples include those that can be computed by computing systems (such as those mentioned above). Figure 1 and Figure 10 The flowchart 700 describes the various actions associated with the exemplary methods implemented by the computing system 110. For example... Figure 7 As shown, flowchart 700 includes multiple actions (actions 710, 720, and 730) associated with various methods for generating speech output from a TTS model based on input text.

[0103] For example, flowchart 700 includes an action (action 710) of receiving electronic content comprising arbitrary text (e.g., text 1010). The computing system then applies the arbitrary text as input to a trained neural TTS model (e.g., TTS model 1000). Using the trained neural TTS model, the computing system generates an output comprising speech data based on the arbitrary text (e.g., Mel spectrogram data 1040), wherein the speech data is characterized by the prosodic pattern of the source speaker and the timbre of the target speaker (action 730). It should be understood that in some embodiments, the trained neural TTS model is trained on spectrogram data generated by the methods disclosed herein (e.g., method 200 and / or method 600).

[0104] Now let's turn our attention to... Figure 8 . Figure 8 An embodiment of a process flowchart illustrating a high-level view of generating training data and training a neural TTS model is illustrated. For example, the process for generating speech data characterized by the timbre of the target speaker and the prosodic patterns of the source speaker is at least partially based on a two-step process.

[0105] First, source speaker data 810 (e.g., source speaker data 142, such as audio / text) corresponding to a specific source prosodic pattern and a specific source timbre is obtained. This data is applied to a voice conversion module 820 (e.g., voice conversion model 147), which is configured to convert the source speaker speech data into target speaker speech data 830 by converting the source speaker's timbre into the target speaker's timbre while preserving the source speaker's prosodic pattern. In step two, this data (target speaker data 830) is used to train a neural TTS model (e.g., TTS model 146) (see Neural TTS Training 840), wherein the neural TTS model is capable of generating speech data 850 from text input. This speech data is TTS data employing the target speaker's timbre and patterns inherited from the source speaker.

[0106] Now let's turn our attention to... Figure 9 , Figure 9 Examples include those included in the training speech recognition module (see...). Figure 10This is an embodiment of an example process flowchart 900 for a voice conversion model 930 within a speech recognition system. The speech conversion model includes an MFCC-PPG component 934 and a PPG-Mel component 938. For example, source speaker audio (e.g., source speaker data 142) is obtained from a source speaker, which includes corresponding text 920 corresponding to the source speaker's audio 910. The source speaker's audio 910 is received by a speech recognition (SR) front-end 932, which is configured to perform signal processing on the input speech, including but not limited to signal denoising and feature extraction, such as MFCC extraction. The speech is also converted into a waveform format or other signal-based audio representation. In some embodiments, the waveform is converted into a Mel scale.

[0107] The speech conversion model 930 also includes an MFCC to PPG model configured to convert speech data into PPG data 936 (e.g., PPG data 143). In some embodiments, the MFCC to PPG model 934 is speaker-independent, wherein this component 934 is pre-trained using multi-speaker data. Advantageously, the model does not need to be further refined or adapted for the source speaker's audio.

[0108] refer to Figure 11 , Figure 11 An embodiment of example waveform-to-PPG component (e.g., MFCC-PPG) is illustrated, wherein the computing system generates PPG data.

[0109] In some embodiments, the MFCC-PPG model 1130 is part of a speech recognition (SR) model (e.g., SR front-end 1120, SR acoustic model (AM) 1122, and SR language model (LM) 1124). During the training of the sub-model or component, the complete SR AM is trained. Once the SR AM is trained, only the MFCC-PPG model 1130 is used during the spectrogram generation process and TTS model training. For example, waveform 1110 obtained from source speaker data is received by the SR front-end, which is configured to perform signal processing on the input speech included in the source speaker audio. After processing by the SR front-end 1120, this data is input to the MFCC-PPG module 1130. The MFCC-PPG module 1130 includes several components and / or layers, such as a front-end layer 1132, multiple LC-BLSTM (latency-controlled bidirectional long short-term memory) layers, and a first projection 1136. PPG data 1140 is then extracted from the first projection (the output of the LC-BLSTM layer). PPG data includes frame-based granular speech information and prosodic information (e.g., speech duration / speech rate).

[0110] Once PPG data 936 is generated, the PPG-Mel model 938 receives it. The PPG-Mel model 938, or more generally, the PPG-spectrum model, is configured to generate spectrogram data based on the received PPG data 936. The PPG-to-Mel model is initially a source PPG-to-Mel model, where the source PPG-to-Mel model 938 is trained on multi-speaker data. After initial training, the PPG-to-Mel model 938 is then refined and / or adapted to be speaker-dependent for a specific (or additional) target speaker. This is accomplished by training the PPG-to-Mel model 938 on target speaker data (e.g., target speaker data 141). In this way, the PPG-Mel model is able to generate spectrograms or Mel spectrograms of the target speaker's timbre with enhanced quality attributable to the speaker-dependent adaptation.

[0111] In some embodiments, the source PPG to spectrogram is always speaker-dependent; for example, a multi-speaker source model is configured to generate spectrograms for many speakers (e.g., speakers appearing in the training data), or has been refined with target speaker data (in this case, it is configured to generate spectrograms primarily configured for the target speaker). In some alternative embodiments, it is possible to train a speaker-independent multi-speaker source PPG to spectrogram model, wherein the generated spectrogram is generated for averaged sounds.

[0112] Therefore, the refined / adapted PPG-to-Mel model is now used to convert the PPG data 936 obtained from the source speaker's audio 910, and a spectrogram is generated using the target speaker's Mel spectrum (a spectrogram where frequencies are converted according to a Mel scale, known to those skilled in the art) but with the source speaker's prosodic pattern. In some embodiments, other spectra besides the Mel spectrum may also be used. The target speaker's Mel spectrum 954 (with a prosodic pattern inherited from the source speaker) along with the corresponding text 920 is configured as training data 950, which can be used to train a neural TTS model (e.g., neural TTS model 1000) to generate speech data with the same characteristics as the newly generated spectrogram (e.g., the target speaker's timbre and the source speaker's prosodic pattern). In some embodiments, the spectrogram data is converted to a Mel scale so that it becomes a Mel spectrogram (e.g., the target speaker's Mel spectrogram 940).

[0113] Now for reference Figure 12 , Figure 12An embodiment of an example PPG-spectrum (e.g., PPG-Mel) component of a sound conversion model is illustrated. For example, a PPG-to-Mel module 1200 (also referred to as a PPG-spectrum model) is shown having an encoder-decoder network (e.g., a PPG encoder 1212 configured to encode PPG data 1210, an lf0 encoder 1222 configured to encode lf0 or pitch data 1220, an energy encoder 1232 configured to encode energy data 1230, and a decoder 1260 configured to decode the encoded data output by the multiple encoders) and an attention layer 1250. The PPG-to-Mel module 1200 is configured to receive multiple data types, including PPG 1210 (e.g., PPG data 143, PPG 1140) from the source speaker, lf0 / uv data 1220 (e.g., pitch data / attributes), energy data 1230, and a speaker ID 1240 corresponding to the target speaker. Using speaker ID 1240, the computing system can use speaker lookup table (LUT) 1242 to identify a specific target speaker. Speaker lookup table (LUT) 1242 is configured to store multiple speaker IDs corresponding to multiple target speakers and associated target speaker data (including target speaker mell spectrum data).

[0114] The PPG-to-Mel module 1200 is thus configured to receive as input a PPG 1210 extracted from the source speaker data and one or more prosodic features including pitch data 1220 and / or energy data 1230 extracted from the source speaker data. In some embodiments, the PPG 1210, pitch data 1220, and energy data 1230 are extracted from the source speaker data on a frame-based granularity. In some embodiments, the PPG 1210, pitch data 1220, and energy data 1230 correspond to the target speaker Mel spectrum (e.g., the generated target Mel spectrum or the actual target speaker Mel spectrum), and the generated (or converted) Mel spectrum accurately follows / matches the prosodic features(s) of the source speaker.

[0115] based on Figure 12As shown in the input diagram, the PPG-to-Mel module 1200 is capable of generating spectrogram data (e.g., Mel spectrogram 1270) characterized by the timbre of the voice of the sound speaker based on data obtained from speaker ID 1240 and speaker LUT 1242. Additionally, the spectrogram data is characterized by the prosodic pattern of the source speaker based on data transformed and / or extracted from the source speaker data (e.g., PPG, pitch profile, and / or energy profile). It should be understood that the PPG-to-Mel module 1200 is configured to receive any number of prosodic features extracted from the source speaker audio data, including speech rate and speech duration, as well as other rhythmic and acoustic properties that contribute to the overall prosodic pattern expressed by the source speaker.

[0116] In some embodiments, the PPG-to-Mel module 1200 can distinguish between various prosodic attributes (e.g., pitch relative to energy) and select a specific attribute to improve the efficiency and effectiveness of the module in generating spectrogram data. Furthermore, it should be understood that the training process and the training data generation process are performed differently based on which prosodic features or attributes are detected and selected for the various processes described herein.

[0117] The more prosodic features available during the training and data generation processes, the more accurate and higher quality the generated data will be (e.g., more closely aligned with the prosodic patterns of the source speaker and sound more like the timbre of the target speaker).

[0118] In light of the foregoing, it will be appreciated that the disclosed embodiments offer numerous technical benefits compared to conventional systems and methods for generating machine learning training data configured to train machine learning models for generating spectrogram data in cross-speaker prosodic transmission applications, thereby eliminating the need to record large amounts of data from the target speaker to capture multi-speaker prosodic patterns. Furthermore, it provides a system for generating spectrograms and corresponding text-to-speech data in an efficient and rapid manner. This contrasts with conventional systems that rely solely on target speaker data and struggle to generate large amounts of training data.

[0119] In some instances, the disclosed embodiments offer technical benefits compared to conventional systems and methods for training machine learning models to perform text-to-speech data generation. For example, by training a TTS model on spectrogram data generated via the methods described herein, the TTS model can be rapidly trained to produce speech data employing the vocal timbre of the target speaker and any number of prosodic patterns from the source speaker. Furthermore, it improves the availability and access to previously inaccessible sources of natural language data.

[0120] Various embodiments of the present invention may include or use a dedicated or general-purpose computer (e.g., computing system 110) including computer hardware, as discussed in more detail below. Embodiments within the scope of the present invention also include entities and other computer-readable media for implementing or storing computer-executable instructions and / or data structures. Such computer-readable media may be any available media accessible by a general-purpose or dedicated computer system. Storing computer-executable instructions (e.g., Figure 1 Computer-readable media (e.g., component 118) Figure 1 The storage (140) is a physical storage medium. A computer-readable medium carrying computer-executable instructions is a transmission medium. Thus, by way of example and not limitation, embodiments of the invention may include at least two significantly different computer-readable media: a physical computer-readable storage medium and a transmission computer-readable medium.

[0121] Physical computer-readable storage media are hardware and include RAM, ROM, EEPROM, CD-ROM or other optical disc storage (such as CD, DVD, etc.), magnetic disk storage or other magnetic storage devices, or any other hardware that can be used to store desired program code in the form of computer-executable instructions or data structures and is accessible by a general-purpose or special-purpose computer.

[0122] "Network" (e.g., Figure 1 A network (130) is defined as one or more data links that enable the transmission of electronic data between computer systems and / or modules and / or other electronic devices. When information is transmitted or provided to a computer via a network or another communication connection (hardwired, wireless, or a combination of hardwired and wireless), the computer appropriately regards that connection as a transmission medium. The transmission medium may include networks and / or data links that can be used to carry desired program code in the form of computer-executable instructions or data structures and are accessible to general-purpose or special-purpose computers. Combinations of the above are also included within the scope of computer-readable media.

[0123] Furthermore, upon arrival at various computer system components, program code in the form of computer-executable instructions or data structures can be automatically transferred from the transmission computer-readable medium to a physical computer-readable storage medium (or vice versa). For example, computer-executable instructions or data structures received via a network or data link can be cached in RAM within a network interface module (e.g., a "NIC") and then ultimately transmitted to the computer system RAM and / or a less volatile computer-readable physical storage medium at the computer system. Therefore, a computer-readable physical storage medium can be included in computer system components that also (or even primarily) utilize the transmission medium.

[0124] Computer-executable instructions include, for example, instructions and data that cause a general-purpose computer, a special-purpose computer, or a special-purpose processing device to perform a function or a set of functions. Computer-executable instructions can be, for example, binary code, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in language specific to structural features and / or methodological actions, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the features or actions described above. Rather, the features and actions described above are disclosed as exemplary forms of implementing the claims.

[0125] Those skilled in the art will understand that this invention can be practiced in network computing environments with many types of computer system configurations, including personal computers, desktop computers, laptop computers, message processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics devices, network PCs, minicomputers, mainframes, mobile phones, PDAs, pagers, routers, switches, and so on. This invention can also be implemented in distributed system environments where both local and remote computer systems perform tasks via network links (or via hardwired data links, wireless data links, or a combination of hardwired and wireless data links). In a distributed system environment, program modules can reside on both local and remote memory storage devices.

[0126] Alternatively or additionally, the functionality described herein may be performed at least in part by one or more hardware logic components. For example, but not limited to, illustrative types of hardware logic components that may be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0127] This invention may be embodied in other specific forms without departing from its spirit or characteristics. The described embodiments should be considered in all respects as illustrative rather than restrictive. Thus, the scope of the invention is indicated by the appended claims rather than the foregoing description. All changes falling within the meaning and scope of equivalents of the claims should be covered by the scope of the claims.

Claims

1. A method implemented by a computing system for generating a spectrogram for a target speaker that adopts a prosody pattern of a source speaker and training a neural text-to-speech (TTS) model based on the spectrogram, the method comprising: receiving electronic content that includes source speaker data from the source speaker; converting a waveform of the source speaker data into prosodic posterior graph (PPG) data by aligning the waveform of the source speaker data with the PPG data, wherein the PPG data defines one or more features corresponding to a prosody pattern of the source speaker data; extracting one or more additional prosodic features from the source speaker data in addition to the one or more features defined by the PPG data; generating a spectrogram based on the PPG data, the extracted one or more additional prosodic features, and a vocal timbre of the target speaker, wherein the spectrogram is characterized by the prosody pattern of the source speaker and the vocal timbre of the target speaker; and training the neural TTS model with the generated spectrogram, wherein the neural TTS model is configured to generate speech data from arbitrary text that is characterized by the prosody pattern of the source speaker and the vocal timbre of the target speaker.

2. The method of claim 1, wherein, the waveform of the source speaker data is aligned with the PPG data at a granularity that is narrower than a phoneme-based granularity.

3. The method of claim 1, wherein, the waveform of the source speaker data is aligned with the PPG data at a frame-based granularity.

4. The method of claim 3, wherein, the frame-based granularity is based on a plurality of frames, each frame comprising 12.5 milliseconds.

5. The method of claim 1, wherein, the one or more additional prosodic features extracted from the source speaker data comprise one or more of: pitch or energy.

6. The method of claim 5, wherein, the one or more additional prosodic features extracted from the source speaker data comprise the energy, the energy being measured in a volume of the source speaker data.

7. The method of claim 1, wherein, the one or more additional prosodic features are extracted from the waveform of the source speaker data at a frame-based granularity.

8. The method of claim 1, wherein, the method includes the computing system defining a prosody pattern of the source speaker, the prosody pattern comprising one of: a news anchor pattern, a story telling pattern, a serious pattern, a casual pattern, a customer service pattern, or an emotion-based pattern, the method including the computing system distinguishing the prosody pattern from a plurality of possible prosody patterns, the plurality of possible prosody patterns comprising the news anchor pattern, the story telling pattern, the serious pattern, the casual pattern, the customer service pattern, and the emotion-based pattern.

9. The method of claim 8, wherein, the emotion-based pattern is detected by the computing system as at least one of: a happy emotion, a sad emotion, an angry emotion, an excited emotion, or an embarrassed emotion, the method including the computing system distinguishing the emotion-based pattern from a plurality of possible emotion-based patterns, the plurality of possible emotion-based patterns comprising the happy emotion, the sad emotion, the angry emotion, the excited emotion, or the embarrassed emotion.

10. A method implemented by a computing system for training a voice conversion machine learning model within a sound conversion module to generate a spectrogram for a target speaker with a new prosody pattern of a source speaker and training a neural text-to-speech (TTS) model based on the spectrogram, the method comprising: training a prosody posterior graph (PPG) to spectrogram component of the voice conversion machine learning model, wherein the PPG to spectrogram component of the voice conversion machine learning model is initially trained on multi-speaker data during training and is configured to convert PPG data to spectrogram data; refining the PPG to spectrogram component by adapting the PPG to spectrogram component of the voice conversion machine learning model to convert PPG data to spectrogram data with a particular prosody pattern of a target speaker using target speaker data from the target speaker having a particular voice timbre and the particular prosody pattern of the target speaker; receiving electronic content comprising new PPG data converted from a waveform of source speaker data, wherein the new PPG data is aligned with the waveform of the source speaker data; receiving one or more prosodic features extracted from the waveform of the source speaker data; applying the source speaker data to the voice conversion machine learning model, wherein the refined PPG to spectrogram component of the voice conversion machine learning model is configured to generate a spectrogram that adopts the particular voice timbre of the target speaker but has the new prosody pattern of the source speaker instead of the particular prosody pattern of the target speaker; and training the neural TTS model with the generated spectrogram, wherein the neural TTS model is configured to generate speech data from arbitrary text that is characterized by the new prosody pattern of the source speaker and the particular voice timbre of the target speaker.

11. The method of claim 10, wherein, the new PPG data is aligned with the waveform of the source speaker data at a granularity that is narrower than a phoneme-based granularity.

12. The method of claim 10, wherein, the new PPG data is aligned with the waveform of the source speaker data at a frame-based granularity.

13. The method of claim 12, wherein, the frame-based granularity is based on a plurality of frames, each frame comprising approximately 12.5 milliseconds.

14. The method of claim 10, wherein, the one or more prosodic features extracted from the waveform of the source speaker data comprise at least one of: a pitch contour, an energy contour, a speaking duration, or a speaking rate.

15. The method of claim 14, wherein, the refined PPG to spectrogram component of the voice conversion machine learning model is configured to generate a spectrogram that adopts the particular voice timbre of the target speaker but has the new prosody pattern of the source speaker instead of the particular prosody pattern of the target speaker based on at least one of: the pitch contour, the energy contour, the speaking duration, or the speaking rate extracted from the source speaker data.

16. The method of claim 14, wherein, the pitch contour and / or the energy contour are extracted from the waveform of the source speaker data at a frame-based granularity.

17. A method implemented by a computing system for generating training data for training a neural text-to-speech (TTS) model configured to generate speech data from arbitrary text, the method comprising: receiving electronic content comprising source speaker data from a source speaker; converting a waveform of the source speaker data into phonetic posterior graph (PPG) data, wherein the converting comprises aligning the waveform with the PPG data; extracting one or more prosodic features from the source speaker data; applying at least the PPG data and the one or more extracted prosodic features to a pre-trained PPG-to-spectrogram component of a voice conversion module, the pre-trained PPG-to-spectrogram component configured to generate spectrograms that adopt a particular voice timbre of a target speaker but have a new prosodic style of the source speaker rather than a particular prosodic style of the target speaker; generating training data configured for training a neural TTS model, the training data comprising a plurality of spectrograms characterized by the particular voice timbre of the target speaker and the new prosodic style of the source speaker; and training a neural TTS model with the generated spectrograms, wherein the neural TTS model is configured to generate speech data from arbitrary text that is characterized by the new prosodic style of the source speaker and the particular voice timbre of the target speaker.

18. The method of claim 17, wherein, further comprising: training the neural TTS model on the generated training data such that the neural TTS model is configured to generate speech data from arbitrary text by performing cross-speaker style transfer, wherein the speech data is characterized by the prosodic style of the source speaker and the voice timbre of the target speaker.

19. The method of claim 18, wherein, further comprising: receiving electronic content comprising arbitrary text; applying the arbitrary text as input to the trained neural TTS model; and generating an output comprising speech data based on the arbitrary text, wherein the speech data is characterized by the prosodic style of the source speaker and the voice timbre of the target speaker.

20. A computer system comprising means for performing the method of any one of claims 1-19.

21. A computer-readable storage medium having stored thereon instructions that, when executed, cause a machine to perform the method of any one of claims 1-19.

Citation Information

Patent Citations

  • Many-to-one voice conversion system

    CN110930981A

  • System and method for voice-to-voice conversion

    US10614826B2