Speech synthesis from text in the voice of a target speaker using a neural network
The neural network system efficiently adapts to new speakers and languages by separating speaker modeling and synthesis, using short audio clips for embedding, achieving high-quality speech synthesis with reduced training requirements.
Patent Information
- Application Number
- JP2024008676
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-05-17
- Filing Date
- 2024-01-24
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2039-05-17
AI Technical Summary
Existing speech synthesis systems require extensive training data and time to adapt to new speakers, especially when high-quality data is scarce, and struggle to synthesize speech in different languages without transcripts.
A neural network-based system that separates speaker modeling and speech synthesis, using a speaker verification neural network to generate speaker embeddings from short audio clips, enabling synthesis in the voice of new speakers without additional training, and a spectrogram generation neural network to produce speech in various languages.
Enables rapid adaptation to new speakers and languages with minimal data, producing high-quality speech synthesis with improved accuracy and flexibility, reducing the need for extensive training and fine-tuning.
Smart Images

Figure 0007701490000001 
Figure 0007701490000002 
Figure 0007701490000003
Abstract
Description
Technical Field
[0001] This specification generally relates to speech synthesis from text.
Background Art
[0002] A neural network is a machine learning model that uses multiple layers of operations to predict one or more outputs from one or more inputs. A neural network typically includes one or more hidden layers located between an input layer and an output layer. The output of each hidden layer is used as an input to the next layer, such as the next hidden layer or the output layer.
[0003] Each layer of a neural network specifies one or more transformation operations to be performed on the input to the layer. Some neural network layers perform operations called neurons. Each neuron receives one or more inputs and generates an output that is received by another neural network layer. In many cases, each neuron receives inputs from other neurons, and each neuron provides outputs to one or more other neurons.
[0004] Each layer generates one or more outputs using the current values of the parameter set of that layer. Training a neural network includes continuously performing a forward pass on the input, calculating gradient values, and updating the current values of the parameter set of each layer. Once the neural network is trained, predictions can be made in a production system using the final parameter set.
Summary of the Invention
Problems to be Solved by the Invention
[0005] This specification generally relates to speech synthesis from text.
Means for Solving the Problems
[0006] A neural network-based system for speech synthesis can generate speech audio in the voices of many different speakers, including speakers not used during training. The system can synthesize new speech in the voice of a target speaker using a few seconds of untranscribed reference audio from the target speaker without updating the system's parameters. The system may use a sequence-to-sequence model. This generates an amplitude spectrogram from a sequence of phonemes or graphemes and adjusts the output with a speaker embedding. The embedding can be computed using an independently trained speaker encoder network (also referred to herein as a speaker verification neural network or speaker encoder) that encodes an audio spectrogram of any length into a fixed-dimensional embedding vector. An embedding vector is a set of values that encode or otherwise represent data. For example, an embedding vector can be generated by a hidden or output layer of a neural network. In this case, the embedding vector encodes one or more data values input to the neural network. The speaker encoder can be trained on a speaker verification task using individual datasets of noisy speech from thousands of different speakers. The system can synthesize natural speech from speakers not used during training using only a few seconds of audio from the speaker by leveraging the knowledge about diverse speakers learned by the speaker encoder and performing accurate generalization.
[0007] More specifically, the system may include a separately trained speaker encoder configured for a speaker verification task. The speaker encoder may be trained to discriminate. The speaker encoder is trained using a generalized end-to-end loss on a large dataset of untranscribed audio from thousands of different speakers. The system can separate the networks and train the networks on independent datasets. This may reduce the problem of obtaining high-quality training data for each purpose. That is, by separately training a speaker discrimination embedding network (i.e., a speaker verification neural network) that captures the space of speaker characteristics and a smaller dataset adjusted with the representation learned by the speaker verification neural network, a high-quality text-to-speech model (referred to herein as a spectrogram generation neural network) can be trained, thereby separating speaker modeling and speech synthesis. For example, speech synthesis may have different and more complex data requirements compared to speaker verification that is independent of the text, and may require tens of hours of clean speech along with the associated transcriptions. In contrast, speaker verification can utilize noisy untranscribed speech, including reverberation and background noise, but may require a sufficient number of speakers. Therefore, obtaining a single set of high-quality training data suitable for both purposes may be much more difficult than obtaining two different sets of high-quality training data for each purpose.
[0008] The subject matter of this specification can be implemented to achieve one or more of the following advantages. For example, the system can result in improved adaptive quality and enable the synthesis of completely new speakers different from those used in training by randomly sampling from prior embeddings (points on the unit hypersphere). In another example, the system can synthesize the voice of a target speaker using only a short, limited amount of sample audio, such as 5 seconds of audio. Yet another advantage is that the system may be able to synthesize speech in the voice of a target speaker for which no transcript of the sample audio of the target speaker is available. For example, the system can receive a 5 - second sample of audio from "John Doe" for which no sample audio was previously available and generate speech in the voice of "John Doe" for any text even without a transcript of the sample audio.
[0009] Yet another advantage is that the system may be able to generate speech in a language different from the language for which sample audio of a particular speaker is available. For example, the system can receive a 5 - second sample of audio from "John Doe" in Spanish and may be able to generate speech in the voice of "John Doe" in English even without other samples of audio from "John Doe".
[0010] Unlike conventional systems, by separating the training of speaker modeling and speech synthesis, the described system can effectively adapt the voices of different speakers even when a single set of high - quality voice data containing audio from multiple speakers is not available.
[0011] In conventional systems, it may be necessary to have hours of training and / or fine-tuning before being able to generate voice audio with the voice of a new target speaker, but the described system can generate voice audio with the voice of a new target speaker without requiring additional training or fine-tuning. Thus, the described system can perform more quickly tasks that require generating voice with the voice of a new speaker with minimal latency, such as a voice-to-voice translation where the generated voice audio is the voice of the original speaker.
[0012] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.
[0013] In some aspects, the subject matter described herein can be embodied as a method that can include the actions of obtaining an audio representation of the voice of a target speaker, obtaining input text whose voice is synthesized as the voice of the target speaker, generating a speaker vector by providing the audio representation to a speaker encoder engine trained to distinguish speakers from one another, providing the input text and the speaker vector to a spectrogram generation engine trained using the voice of a reference speaker to generate an audio representation, generating an audio representation of the input text spoken in the voice of the target speaker, and providing the audio representation of the input text spoken in the voice of the target speaker for output.
[0014] The speaker verification neural network is trained to generate speaker embedding vectors of audio representations of the same speaker that are close to each other in the embedding space, and may also be trained to generate speaker embedding vectors of audio representations of different speakers that are far from each other. Alternatively or in addition, the speaker verification neural network may be trained separately from the spectrogram generation neural network. The speaker verification neural network is a long short-term memory (LSTM) neural network.
[0015] Generating a speaker embedding vector may include providing a plurality of overlapping sliding windows of the audio representation to the speaker verification neural network to generate a plurality of individual vector embeddings, and calculating an average of the plurality of individual vector embeddings to generate the speaker embedding vector.
[0016] Providing an audio representation of the input text spoken in the voice of the target speaker for output may include providing the audio representation of the input text spoken in the voice of the target speaker to a vocoder to generate a time-domain representation of the input text spoken in the voice of the target speaker, and providing the time-domain representation for playback to the user. The vocoder may be a vocoder neural network.
[0017] The spectrogram generation neural network may be a sequence-to-sequence attention neural network trained to predict a mel spectrogram from a sequence of phoneme or grapheme inputs. The spectrogram generation neural network may optionally include an encoder neural network, an attention layer, and a decoder neural network. The spectrogram generation neural network can concatenate the speaker embedding vector with the output of the encoder neural network provided as an input to the attention layer.
[0018] The speaker embedding vector may be different from the speaker embedding vector used during the training of the speaker verification neural network or the spectrogram generation neural network. During the training of the spectrogram generation neural network, the parameters of the speaker verification neural network may be fixed.
[0019] According to a further aspect, there is provided a computer-implemented method of training a neural network for use in speech synthesis, including training a speaker verification neural network to distinguish speakers from each other and training a spectrogram generation neural network using the voices of a plurality of reference speakers to generate an audio representation of input text. This aspect may include any of the features of the preceding aspects.
[0020] Other versions include corresponding systems, devices, and computer programs configured to perform the actions of the method, encoded on a computer storage device.
[0021] Details of one or more implementations are set forth in the accompanying drawings and the description below. Other potential features and advantages will be apparent from the description, drawings, and claims.
Brief Description of the Drawings
[0022]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
DETAILED DESCRIPTION OF THE INVENTION
[0023] Similar reference numbers and designations in the various drawings indicate similar elements. FIG. 1 is a block diagram showing an exemplary voice system 100 that can synthesize voice with the voice of a target speaker. The voice synthesis system 100 can be implemented as a computer program on one or more computers located in one or more locations. The voice synthesis system 100 receives input text along with the audio representation of the target speaker, processes the input through a series of neural networks, and generates voice corresponding to the input text in the voice of the target speaker. For example, if the voice synthesis system 100 receives the text of this page as input along with 5 seconds of audio of John Doe saying "Hello, my name is John Doe and I am providing this sample of voice for testing purposes", these inputs can be processed to generate an oral narration of the page in John Doe's voice. In another example, if the voice synthesis system 100 receives the text of this page as input along with 6 seconds of audio of Jane Doe narrating from another book, these inputs can be processed to generate an oral narration of the page in Jane Doe's voice.
[0024] As shown in FIG. 1, the system 100 includes a speaker encoder engine 110 and a spectrogram generation engine 120. For the target speaker, the speaker encoder engine 110 receives the audio representation of the target speaker speaking and outputs a speaker vector, also called a speaker embedding vector or embedding vector. For example, the speaker encoder engine 110 receives an audio recording of John Doe saying "Hello, my name is John Doe", and in response, outputs a vector having a value that identifies John Doe. The speaker vector may also capture the characteristic speaking rate of the speaker.
[0025] The speaker vector can be an embedded vector of a fixed dimension. For example, the speaker vector output by the speaker encoder engine 110 can have a sequence of 256 values. The speaker encoder engine 110 can be a neural network trained to encode an audio spectrogram of any length into an embedded vector of a fixed dimension. For example, the speaker encoder engine 110 can include a long short-term memory (LSTM) neural network trained to encode a mel spectrogram or a log mel spectrogram, a representation of the user's voice, into a vector having a fixed number of elements (e.g., 256 elements). Although the mel spectrogram is referred to throughout this disclosure for consistency and detailed description, it will be understood that other types of spectrograms, or other suitable audio representations, can be used.
[0026] The speaker encoder engine 110 may be trained using labeled training data that includes a pair of audio of a voice and a label identifying the speaker of the audio, whereby the engine 110 learns to classify the audio as corresponding to different speakers from each other. The speaker vector can be the output of the hidden layer of the LSTM neural network, where the more similar the voices are from the speakers, the more similar the speaker vectors obtained from the audio of the speakers are to each other, and the more different the voices are from the speakers, the more different the speaker vectors obtained from the audio of the speakers are from each other.
[0027] The spectrogram generation engine 120 may receive the input text to be synthesized and the speaker vector determined by the speaker encoder engine 110, and accordingly, may generate an audio representation of the voice of the input text in the voice of the target speaker. For example, the spectrogram generation engine 120 may receive the input text of "goodbye" and the speaker vector determined by the speaker encoder engine 110 from the mel-spectrogram representation of John Doe saying "hello, my name is John Doe". Accordingly, a mel-spectrogram representation of the voice of "goodbye" in the voice of John Doe can be generated.
[0028] The spectrogram generation engine 120 is a neural network (also referred to as a sequence-to-sequence synthesizer, a sequence-to-sequence synthesis network, or a spectrogram generation neural network) with an attention network that is trained to predict the mel-spectrogram of the voice of the target speaker from the input text and the speaker vector of the target speaker. The neural network can be trained with training data including a plurality of triplets each including text, an audio representation of the voice of the text by a specific speaker, and a speaker vector regarding the specific speaker. The speaker vector used in the training data may be from the spectrogram generation engine 120 and may not necessarily be from the audio representation of the voice of the text of the triplet. For example, the triplets included in the training data may include the input text of "I like computers", the mel-spectrogram from the audio of John Smith saying "I like computers", and the speaker vector from the mel-spectrogram of John Smith saying "hello, my name is John Smith" output by the speaker encoder engine 110.
[0029] In some implementation forms, the training data of the spectrogram generation engine 120 can be generated using the speaker encoder engine 110 after the speaker encoder engine 110 is trained. For example, in a set consisting of a pair of training data, there may originally be only a pair of input text and the mel spectrogram of the voice of that text. The mel spectrogram in each pair of the paired training data is provided to the trained speaker encoder engine 110, and the trained speaker encoder engine 110 can output a speaker vector corresponding to each mel spectrogram. Next, the system 100 can add each speaker vector to the corresponding pair in the paired training data to generate training data including a triplet consisting of text, the audio representation of the text by a specific speaker, and the speaker vector regarding the specific speaker.
[0030] In some implementation forms, the audio representation generated by the spectrogram generation engine 120 can be provided to a vocoder to generate voice. For example, the mel spectrogram of John Doe saying "goodbye" is in the frequency domain and may be provided to another neural network. The other neural network is trained to receive the frequency domain representation and output a time domain representation, and may output the time domain waveform of "goodbye" in John Doe's voice. Next, the time domain waveform may be provided to a speaker (such as a loudspeaker) that can generate the sound of "goodbye" in John Doe's voice.
[0031] In some implementations, system 100 or another system may be used to perform a process for synthesizing speech in the voice of a target speaker. The process may include obtaining an audio representation of the target speaker's voice, obtaining input text in which the speech is to be synthesized in the voice of the target speaker, generating a speaker vector by providing the audio representation to a speaker encoder engine trained to distinguish speakers from one another, providing the input text and the speaker vector to a spectrogram generation engine trained using the voice of a reference speaker to generate an audio representation, generating an audio representation of the input text spoken in the voice of the target speaker, and providing the audio representation of the input text spoken in the voice of the target speaker for output.
[0032] For example, the process may include the speaker encoder engine 110 obtaining a mel spectrogram from the audio of Jane Doe saying "I like computers", and generating a speaker vector for Jane Doe that is different from the speaker vector generated for the mel spectrogram of John Doe saying "I like computers". The spectrogram generation engine 120 may receive the speaker vector for Jane Doe and obtain input text that may be "Hola como estas", which may mean "Hello, how are you" in Spanish in English, and in response, generate a mel spectrogram that may be converted by a vocoder into speech of "Hola como estas" in the voice of Jane Doe.
[0033] In a more detailed example, system 100 may include three independently trained components. That is, an LSTM speaker encoder for speaker verification that outputs a fixed-dimensional vector from an audio signal of any length, a sequence-to-sequence attention network that predicts a mel spectrogram from a sequence of grapheme or phoneme inputs adjusted by the speaker vector, and an autoregressive neural vocoder network that converts the mel spectrogram into a sequence of time-domain waveform samples. The LSTM speaker encoder may be the speaker encoder engine 110, and the sequence-to-sequence with attention network may be the spectrogram generation engine 120.
[0034] The LSTM speaker encoder is used to adjust the synthesis network with a reference audio signal from a desired target speaker. Appropriate generalization can be achieved using reference audio signals that capture the characteristics of various speakers. With appropriate generalization, these characteristics can be identified using only a short adaptation signal, regardless of the content of the phonetic representation and background noise. These objectives are satisfied using a speaker discrimination model trained on a text-independent speaker verification task. The LSTM speaker encoder can be a speaker discrimination audio embedding network that is not limited to a set of only a few speakers.
[0035] The LSTM speaker encoder maps a sequence of mel spectrogram frames computed from an utterance of any length to a fixed-dimensional embedding vector known as a d-vector or speaker vector. Given an utterance x, the LSTM speaker encoder uses the LSTM network to generate a fixed-dimensional vector embedding e xcan be configured to learn =f(x). A generalized end-to-end loss may be used to train the LSTM network, such that the d-vectors of utterances from the same speaker are close to each other in the embedding space, e.g., the d-vectors of the utterances have a large cosine similarity, while the d-vectors of utterances from different speakers are far from each other. Thus, when an utterance of any length is given, the speaker encoder can be run with overlapping sliding windows, e.g., of length 800 milliseconds, and the average of multiple embeddings of the L2-normalized windows is used as the final embedding of the entire utterance.
[0036] The sequence-to-sequence attention neural network can model multiple specific speakers by concatenating, for each audio example x in the training dataset, a d-dimensional embedding vector associated with the true speaker at each time step with the output of the encoder neural network before the output is provided to the attention neural network. The speaker embeddings provided to the input layer of the attention neural network may be sufficient to converge between different speakers. The synthesizer can be an end-to-end synthesis network that does not rely on intermediate language functions.
[0037] In some implementations, the sequence-to-sequence attention network may be trained with pairs of text transcriptions and target audio. In the input, the text is mapped to a sequence of phonemes. This speeds up convergence and improves the pronunciation of rare words such as names and place names. The network is trained in a transfer learning configuration using a pre-trained speaker encoder (with frozen parameters) to extract speaker embeddings from the target audio. That is, during training, the speaker reference signal is the same as the target voice. No explicit speaker identifier labels are used during training.
[0038] Additionally or alternatively, the decoder of the network may include both the L2 loss in the reconstruction of the spectrogram function and an additional L1 loss. The combined loss may be more robust with noisy training data. Additionally or alternatively, to further clean the synthesized audio, noise reduction by spectral subtraction, e.g., at the 10th percentile, may be performed on the target of the mel spectrogram prediction network.
[0039] System 100 can capture unique characteristics of a speaker not previously seen from a single short audio clip and synthesize new voices with those characteristics. System 100 can achieve the following: (1) highly natural synthesized voices (2) a high degree of similarity to the target speaker. Achieving a highly natural one typically requires a large number of high-quality voice and transcription pairs as training data, while achieving a high degree of similarity usually requires a large amount of training data for each speaker. However, recording a large amount of high-quality data for each individual speaker is very costly or even practically infeasible. System 100 can separate the training of the system from highly natural text to speech and the training of another speaker discrimination embedding network that can well capture speaker characteristics. In some implementations, the speaker discrimination model is trained on a speaker verification task independent of text.
[0040] The neural vocoder inverts the synthesized mel spectrogram produced by the synthesizer into a time-domain waveform. In some implementations, the vocoder can be a sample-by-sample autoregressive Wave Net. The architecture can include multiple dilated convolutional layers. The mel spectrogram predicted by the synthesizer network captures all the relevant details necessary for high-quality synthesis of various voices and can be used to build a multi-speaker vocoder by training on data from multiple speakers without the need for explicit adjustment with speaker vectors. Details of the Wave Net architecture are described in WaveNet: A Generative Model for Raw Audio by van den Oorde et al., CoRR abs / 1609.03499, 2016.
[0041] Figure 2 is a block diagram of an exemplary system 200 during voice synthesis training. The exemplary system 200 includes a speaker encoder 210, a synthesizer 220, and a vocoder 230. The synthesizer 220 includes a text encoder 222, an attention neural network 224, and a decoder 226. During training, a separately trained speaker encoder 210, whose parameters may be frozen, can extract a fixed-length d-vector of the speaker from a variable-length input audio signal. During training, the reference signal or target audio may be ground-truth audio juxtaposed with the text. The d-vector may be concatenated with the output of the text encoder 222 and passed to the attention neural network 224 at each of a plurality of time steps. Except for the speaker encoder 210, the other parts of the system 200 can be driven by a reconstruction loss from the decoder 226. The synthesizer 220 can predict a mel spectrogram from the input text sequence and provide the mel spectrogram to the vocoder 230. The vocoder 230 can convert the mel spectrogram into a time-domain waveform.
[0042] FIG. 3 is a block diagram of an exemplary system 300 during inference for synthesizing speech. System 300 includes a speaker encoder 210, a synthesizer 220, and a vocoder 230. During inference, either of two approaches may be used. In a first approach, a text encoder 222 may be directly adjusted with d-vectors from unseen and / or non-transcribed audio that does not need to match the text whose transcription is to be synthesized. This may enable the network to generate unseen voices from a single audio clip. Since the speaker characteristics used for synthesis are inferred from the audio, it may be adjusted with audio from speakers outside the training set. In a second approach, random sample d-vectors can be obtained and the text encoder 122 may be adjusted with the random sample d-vectors. Since the speaker encoder may be trained based on a large number of speakers, random d-vectors may also generate random speakers.
[0043] FIG. 4 is a flowchart of an exemplary process 400 for generating an audio representation of text spoken in the voice of a target speaker. The exemplary process is described as being executed by a system appropriately programmed according to this specification.
[0044] The system obtains (405) an audio representation of the voice of the target speaker. For example, the audio representation can be in the form of an audio recording file, and the audio can be captured by one or more microphones.
[0045] The system obtains (410) input text in which the speech is synthesized in the voice of the target speaker. For example, the input text can be in the form of a text file. The system generates a speaker embedding vector (415) by providing audio representations to a speaker verification neural network trained to distinguish speakers from each other. For example, the speaker verification neural network can be an LSTM neural network, and the speaker embedding vector can be the output of the hidden layer of the LSTM neural network.
[0046] In some implementations, the system provides a plurality of overlapping sliding windows of the audio representation to the speaker verification neural network to generate a plurality of individual vector embeddings. For example, the audio representation can be divided into windows of about 800 milliseconds in length (e.g., 750 milliseconds or less, 700 milliseconds or less, 650 milliseconds or less), and the overlap can be about 50% (e.g., 60% or more, 65% or more, 70% or more). Next, the system can generate a speaker embedding vector by calculating the average of the individual vector embeddings.
[0047] In some implementations, the speaker verification neural network is trained to generate speaker embedding vectors, such as d-vectors, of audio representations from the same speaker that are close to each other in the embedding space. The speaker verification neural network may be trained to generate speaker embedding vectors of audio representations from different speakers that are far from each other.
[0048] The system generates an audio representation of the input text spoken in the voice of the target speaker (420) by providing the input text and the speaker embedding vector to a spectrogram generation neural network trained using the voice of the reference speaker to generate audio representations.
[0049] In some implementations, during the training of the spectrogram generation neural network, the parameters of the speaker embedding neural network are fixed. In some implementations, the spectrogram generation neural network may be trained separately from the speaker verification neural network.
[0050] In some implementations, the speaker embedding vector is different from the speaker embedding vector used during the training of the speaker verification neural network or the spectrogram generation neural network.
[0051] In some implementations, the spectrogram generation neural network is a sequence-to-sequence attention neural network trained to predict a mel spectrogram from a sequence of phoneme or grapheme inputs. For example, the spectrogram generation neural network architecture can be based on Tacotron 2. Details of the Tacotron 2 neural network architecture are described in "Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions" published in the proceedings by Shen et al. (IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) 2018).
[0052] In some implementations, the spectrogram generation neural network includes an encoder neural network, an attention layer, and a decoder neural network. In some implementations, the spectrogram generation neural network concatenates the speaker embedding vector with the output of the encoder neural network provided as an input to the attention layer.
[0053] In some implementations, the encoder neural network and the sequence-to-sequence attention neural network can be trained on a set of speakers that are unbalanced and have no common elements. The encoder neural network may be trained to distinguish between speakers, thereby enabling the characteristics of the speakers to be more reliably transferred.
[0054] The system provides an audio representation of input text spoken in the voice of a target speaker for output (425). For example, the system can generate a time-domain representation of the input text. In some implementations, the system provides an audio representation of the input text spoken in the voice of the target speaker to a vocoder to generate a time-domain representation of the input text spoken in the voice of the target speaker. The system can provide the time-domain representation for playback to the user.
[0055] In some implementations, the vocoder is a vocoder neural network. For example, the vocoder neural network can be a sample-by-sample autoregressive wave net that can invert a synthesized mel spectrogram generated by a synthesis network into a time-domain waveform. The vocoder neural network can include a plurality of dilated convolutional layers.
[0056] FIG. 5 shows an example of a computing device 500 and a mobile computing device 450 that can be used to implement the techniques described herein. The computing device 500 is intended to represent various forms of digital computers, such as a laptop, desktop, workstation, personal digital assistant, server, blade server, mainframe, and other appropriate computers. The mobile computing device 450 is intended to represent various forms of mobile devices, such as a personal digital assistant, cellular phone, smartphone, and other similar computing devices. The components shown herein, their connections and relationships to each other, and their functions are intended only as examples and are not intended to be limiting.
[0057] Computing device 500 includes a processor 502, a memory 504, a storage device 506, a high-speed interface 508 connected to the memory 504 and a plurality of high-speed expansion ports 510, and a low-speed interface 512 connected to the low-speed expansion port 514 and the storage device 506. Each of the processor 502, the memory 504, the storage device 506, the high-speed interface 508, the high-speed expansion ports 510, and the low-speed interface 512 is interconnected using various buses and may be attached to a common motherboard or, where appropriate, in other manners. The processor 502 is capable of processing instructions for execution within the computing device 500, including instructions stored in the memory 504 or the storage device 506 for displaying graphical information for a graphical user interface (GUI) on an external input / output device such as a display 516 coupled to the high-speed interface 508. In other implementations, multiple processors and / or multiple buses may be used, along with multiple memories and certain types of memory, where appropriate. Further, multiple computing devices may be connected and each device may provide a portion of the necessary operations (e.g., a server bank, a group of blade servers, or a multiprocessor system).
[0058] The memory 504 stores information within the computing device 500. In some implementations, the memory 504 is one or more volatile memory units. In some implementations, the memory 504 is one or more non-volatile memory units. Further, the memory 504 may be another form of computer-readable medium, such as a magnetic disk or an optical disk.
[0059] The memory device 506 can provide large-capacity storage for the computing device 500. In some implementations, the memory device 506 may be a computer-readable medium such as a floppy (registered trademark) disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid-state memory device, or an array of devices (including a storage area network or devices of other configurations). Instructions can be stored in an information carrier. The instructions, when executed by one or more processing devices (e.g., the processor 502), execute one or more methods such as the methods described above. The instructions can also be stored by one or more memory devices such as a computer-readable or machine-readable medium (e.g., the memory 504, the memory device 506, or the memory on the processor 502).
[0060] The high-speed control unit 508 manages bandwidth-intensive operations for the computing device 500, while the low-speed control unit 512 manages relatively low-bandwidth-intensive operations. Such an assignment of functions is merely illustrative. In some implementations, the high-speed control unit 508 is coupled to the memory 504, the display 516 (e.g., through a graphics processor or accelerator), and a high-speed expansion port 510 that can receive various expansion cards (not shown). In that implementation, the low-speed control unit 512 is coupled to the memory device 506 and a low-speed expansion port 514. The low-speed expansion port 514, which may include various communication ports (e.g., USB, Bluetooth (registered trademark), Ethernet (registered trademark), wireless Ethernet), may be coupled to one or more input / output devices such as a keyboard, a pointing device, a scanner, or a networking device such as a switch or router (e.g., through a network adapter).
[0061] As shown in the figure, computing device 500 may be implemented in many different forms. For example, computing device 500 may be implemented as a standard server 520, or multiple times in a group of such servers. Further, computing device 500 may be implemented in a personal computer such as a laptop computer 522. Further, computing device 500 may be implemented as part of a rack server system 524. Alternatively, the components of computing device 500 may be combined with other components in a mobile device (not shown) such as mobile computing device 450. Each of such devices may include one or more of computing device 500 and mobile computing device 450, and the entire system may be composed of a plurality of computing devices that communicate with each other.
[0062] Mobile computing device 450 particularly includes a processor 552, a memory 564, input / output devices such as a display 554, a communication interface 566, and a transceiver 568 as components. A storage device such as a microdrive or other device may be further provided for mobile computing device 450 to provide additional storage. Each of processor 552, memory 564, display 554, communication interface 566, and transceiver 568 is interconnected using various buses, and some of the components may be attached to a common motherboard, or in other appropriate manners in appropriate cases.
[0063] Processor 552 can execute instructions, including those stored in memory 564, within mobile computing device 450. Processor 552 may be implemented as a chipset consisting of a chip that includes a plurality of separate analog and digital processors. Processor 552 enables, for example, the coordination of other components of mobile computing device 450 such as the control of the user interface, applications operated by mobile computing device 450, and wireless communication by mobile computing device 450.
[0064] Processor 552 can communicate with the user through control interface 558 and display interface 556 coupled to display 554. Display 554 may be, for example, a TFT (Thin Film Transistor Liquid Crystal Display) display or an OLED (Organic Light Emitting Diode) display, or other suitable display technology. Display interface 556 may include appropriate circuitry for operating display 554 to present graphical and other information to the user. Control interface 558 may receive commands from the user and convert the commands for passing to processor 552. Additionally, external interface 562 may communicate with processor 552 to enable near-field communication of mobile computing device 450 with other devices. External interface 562 may, for example, enable wired communication in some implementation embodiments or wireless communication in other implementations, and further, multiple interfaces may be used.
[0065] Memory 564 stores information within mobile computing device 450. Memory 564 can be implemented as one or more of one or more computer-readable media, one or more volatile memory units, and one or more non-volatile memory units. Further, an extended memory 574 is provided and may be connected to mobile computing device 450 through an extended interface 572 that may include, for example, a SIMM (Single In-line Memory Module) card interface. The extended memory 574 may provide additional storage space for mobile computing device 450, or may store applications or other information for mobile computing device 450. Specifically, the extended memory 574 may include instructions for performing or complementing the above-described processes, and may further include secure information. Thus, for example, the extended memory 574 may be provided as a security module for mobile computing device 450 and may be programmed with instructions for enabling secure use of mobile computing device 450. Further, secure applications may be provided via the SIMM card along with additional information, such as placing identification information on the SIMM card in a non-hackable manner.
[0066] The memory may include, for example, flash memory and / or NVRAM memory (non-volatile random access memory) as follows. In some implementations, the instructions are stored on an information carrier that, when executed by one or more processing devices (e.g., processor 552), performs one or more methods such as the methods described above. The instructions can also be stored by one or more storage devices such as one or more computer-readable or machine-readable media (e.g., memory 564, extended memory 574, or memory on processor 552). In some implementations, the instructions can be received by a propagated signal, for example, through transceiver 568 or external interface 562.
[0067] The mobile computing device 450 can communicate wirelessly through a communication interface 566 that may include a digital signal processing circuit if necessary. The communication interface 566 can enable communication under various modes or protocols, particularly GSM (Global System for Mobile Communications) (registered trademark) voice calls, SMS (Short Message Service), EMS (Enhanced Message Service), or MMS (Multimedia Message Service), CDMA (Code Division Multiple Access), TDMA (Time Division Multiple Access), PDC (Personal Digital Cellular), WCDMA (Wideband Code Division Multiple Access) (registered trademark), CDMA2000, or GPRS (General Packet Radio Service). Such communication may be performed, for example, through a transceiver 568 using radio frequencies. Further, short-range communication may also be performed by using Bluetooth, WiFi (registered trademark), or other such transceivers (not shown). Additionally, a GPS (Global Positioning System) receiver module 570 can provide additional wireless data related to navigation and location to the mobile computing device 450, and that wireless data may be used by an application operating on the mobile computing device 450 if appropriate.
[0068] Furthermore, the mobile computing device 450 may perform audible communication using a voice codec 560 that can receive voice information from a user and convert it into digital information suitable for use. The voice codec 560 may similarly generate audible sound for the user, for example, through a speaker in the handset of the mobile computing device 450. Such sounds may include sounds from voice calls, may include recorded sounds (such as voice messages, music files, etc.), and may also include sounds generated by an application operating on the mobile computing device 450.
[0069] As shown in the figures, computing device 450 may be implemented in many different forms. For example, computing device 450 may be implemented as a mobile phone 580. Computing device 450 may also be implemented as part of a smartphone 582, a personal digital assistant, or other similar mobile devices.
[0070] The various implementations of the systems and techniques described herein can be realized by digital electronic circuits, integrated circuits, specially designed ASICs, computer hardware, firmware, software, and / or combinations thereof. These various implementations can be executable and / or interpretable in one or more computer programs on a programmable system that includes one or more programmable processors coupled to receive data and instructions from, and to transmit data and instructions to, a memory system, one or more input devices, and one or more output devices.
[0071] These computer programs (also known as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented in high-level procedural and / or object-oriented programming languages and / or in assembly / machine language. The program can be stored in a file that holds other programs or data, such as one or more scripts in a markup language document. The program can be stored in a single file dedicated to the program. Alternatively, the program can be stored in multiple coordinated files, such as files that store one or more modules, subprograms, or portions of code. The computer program can be deployed to be executed on one computer or on computers at one site or distributed across multiple sites and interconnected by a communication network.
[0072] As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., magnetic disks, optical disks, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor that receives the machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0073] To provide for interaction with a user, the systems and techniques described herein may be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide for interaction with a user; for example, feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and input received from the user may be received in any form including acoustic, speech, or tactile input.
[0074] The systems and techniques described herein can be implemented by a computing system that includes backend components (such as a data server), a computing system that includes middleware components (such as an application server), a computing system that includes frontend components (such as a client computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (such as a communication network). Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), and the Internet.
[0075] A computing system can include clients and servers. Clients and servers are generally located remotely from each other and typically interact through a communication network. The relationship between a client and a server arises by computer programs that run on individual computers and have a client-server relationship to each other.
[0076] In addition to the above, controls can be provided to the user to enable the user to first make a selection as to whether, and when, the systems, programs, or functions described herein enable the collection of user information (such as information about the user's social network, social actions or activities, occupation, user preferences, or the user's current location), and whether content or communications are sent from the server to the user. Further, certain data may be processed in one or more ways before being stored or used so that personally identifiable information is removed.
[0077] In some embodiments, for example, the user's identification information may be treated such that it cannot determine personally identifiable information about the user, or the user's geographical location cannot identify the user's specific location (such as city, zip code, state level, etc.) where the location information is obtained. Thus, the user can control what information about the user is collected, how this information is used, and what information is provided to the user.
[0078] Several embodiments have been described. Nevertheless, it will be understood that various modifications may be made without departing from the scope of the present invention. For example, using various forms of the above flow, steps can be rearranged, added, or deleted. Also, although some uses of the system and method have been described, it should be recognized that many other uses are contemplated. Thus, other embodiments are within the scope of the following claims.
[0079] Specific embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve the desired result. As an example, to achieve the desired result, the processes shown in the accompanying figures do not necessarily require the specific order or sequential order shown. In some cases, multitasking and parallel processing may be advantageous.
Claims
1. When executed by the data processing hardware, obtaining an audio recording file including a target speaker's voice; extracting a speaker embedding vector of the target speaker from the audio recording file using a speaker encoder network; obtaining an input text to be synthesized into the target speaker's voice; generating a synthetic speech representation of the input text in the voice of the target speaker using a synthesizer configured to receive as input the speaker embedding vectors and the input text; providing for output the synthetic speech representation of the input text in the voice of the target speaker; causing said data processing hardware to perform operations including The computer-implemented method, wherein the speaker encoder network is trained to extract speaker embedding vectors that are close to each other in an embedding space from audio representations corresponding to utterances of the same speaker.
2. When executed by the data processing hardware, obtaining an audio recording file including a target speaker's voice; extracting a speaker embedding vector of the target speaker from the audio recording file using a speaker encoder network; obtaining an input text to be synthesized into the target speaker's voice; generating a synthetic speech representation of the input text in the voice of the target speaker using a synthesizer configured to receive as input the speaker embedding vectors and the input text; providing for output the synthetic speech representation of the input text in the voice of the target speaker; causing said data processing hardware to perform operations including The computer-implemented method, wherein the speaker encoder network is trained to extract distant speaker embedding vectors from audio representations corresponding to utterances of different speakers.
3. The method of claim 1 or 2, wherein the speaker encoder network comprises a long short-term memory (LSTM) neural network.
4. The method of claim 1 or 2, wherein the speaker encoder network is trained separately from the synthesizer.
5. The method of claim 4 , wherein parameters of the speaker encoder network are fixed during training of the synthesizer.
6. The method of claim 1 or 2, wherein the synthesizer includes a spectrogram generation neural network trained to predict a mel spectrogram from a sequence of phoneme input.
7. The method of claim 6 , wherein the spectrogram generating neural network comprises a sequence-to-sequence attention neural network.
8. The method of claim 6 , wherein the spectrogram generating neural network includes an encoder neural network and a decoder neural network.
9. The method of claim 8 , wherein the spectrogram generating neural network further comprises an attention layer.
10. Data processing hardware; and memory hardware in communication with the data processing hardware and having instructions stored thereon, the instructions, when executed by the data processing hardware, obtaining an audio recording file including a target speaker's voice; extracting a speaker embedding vector of the target speaker from the audio recording file using a speaker encoder network; obtaining an input text to be synthesized into the target speaker's voice; generating a synthetic speech representation of the input text in the voice of the target speaker using a synthesizer configured to receive as input the speaker embedding vectors and the input text; providing for output the synthetic speech representation of the input text in the voice of the target speaker; causing said data processing hardware to perform operations including The speaker encoder network is trained to extract speaker embedding vectors that are close to each other in an embedding space from audio representations corresponding to utterances of the same speaker.
11. Data processing hardware; and memory hardware in communication with the data processing hardware and having instructions stored thereon, the instructions, when executed by the data processing hardware, obtaining an audio recording file including a target speaker's voice; extracting a speaker embedding vector of the target speaker from the audio recording file using a speaker encoder network; obtaining an input text to be synthesized into the target speaker's voice; generating a synthetic speech representation of the input text in the voice of the target speaker using a synthesizer configured to receive as input the speaker embedding vectors and the input text; providing for output the synthetic speech representation of the input text in the voice of the target speaker; causing said data processing hardware to perform operations including The system, wherein the speaker encoder network is trained to extract distant speaker embedding vectors from audio representations corresponding to utterances of different speakers.
12. The system of claim 10 or 11, wherein the speaker encoder network comprises a long short-term memory (LSTM) neural network.
13. The system of claim 10 or 11, wherein the speaker encoder network is trained separately from the synthesizer.
14. The system of claim 13 , wherein parameters of the speaker encoder network are fixed during training of the synthesizer.
15. 12. The system of claim 10 or 11, wherein the synthesizer includes a spectrogram generation neural network trained to predict a mel spectrogram from a sequence of phoneme input.
16. 16. The system of claim 15, wherein the spectrogram generating neural network comprises a sequence-to-sequence attention neural network.
17. 16. The system of claim 15, wherein the spectrogram generating neural network includes an encoder neural network and a decoder neural network.
18. 20. The system of claim 17, wherein the spectrogram generating neural network further comprises an attention layer.
Citation Information
Patent Citations
Acoustic model learning device, voice synthesis device, acoustic model learning method, voice synthesis method, and program
JP2017032839A
Sequence to sequence transformations for speech synthesis via recurrent neural networks
US20180114522A1