Generating Acoustic Sequences via a Neural Network Using Combined Prosodic Information
By receiving linguistic sequences and pronunciation information offsets, and using trained pronunciation information predictors and neural networks to generate acoustic sequences, the limitations of pronunciation control in the existing pronunciation synthesis system are solved, and high-quality pronunciation synthesis and expression control are achieved.
Patent Information
- Application Number
- CN202080056837.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-12
- Filing Date
- 2020-09-07
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2040-09-07
AI Technical Summary
The existing speech synthesis system has limitations in controlling prosody, making it difficult to achieve effective control of speech style, emotional state, speech rate and expression. The existing methods such as semi-supervised and unsupervised methods have problems such as high cost, poor controllability and inconsistent synthesis quality.
By receiving linguistic sequences and pronunciation information offsets, using the trained pronunciation information predictor and neural network, a combined pronunciation information including multiple observations is generated, pronunciation is explicitly modeled and sentence-by-sentence control is performed to generate an acoustic sequence.
It realizes explicit modeling of pronunciation on a continuous scale, improves the quality and expressiveness of synthetic pronunciations, and can perform sentence-by-sentence speaking pace and expressive control over the inference time.
Smart Images

Figure CN114207706B_ABST
Abstract
Description
Background Art
[0001] The present technology relates to controlling prosody. More specifically, these technologies relate to controlling prosody via a neural network. Summary of the Invention
[0002] According to an embodiment described herein, a system may include a processor configured to receive a linguistic sequence and a prosody information offset. The processor may further: generate, via a trained prosody information predictor, prosody information including a combination of multiple observations based on the linguistic sequence. The multiple observations include a linear combination of statistical measures evaluating prosody components within a predetermined time period. The processor may further: generate, via a trained neural network, an acoustic sequence based on the combined prosody information, the prosody information offset, and the linguistic sequence.
[0003] According to another embodiment described herein, a method may include receiving a linguistic sequence and a prosody information offset. The method may further include: generating, via a trained prosody information predictor, combined prosody information including multiple observations based on the linguistic sequence. The multiple observations include a linear combination of statistical measures evaluating prosody components within a predetermined time period. The method may further include: generating, via a trained neural network, an acoustic sequence based on the combined prosody information, the prosody information offset, and the linguistic sequence.
[0004] According to another embodiment described herein, a computer program product for automatically controlling prosody may include a computer-readable storage medium having program code embodied therewith. The computer-readable storage medium is not a transient signal per se. The program code is executable by a processor to cause the processor to receive a linguistic sequence and a prosody information offset. The program code may further cause the processor to generate prosody information including a combination of multiple observations based on the linguistic sequence. The multiple observations include a linear combination of statistical measures evaluating prosody components within a predetermined time period. The program code may further cause the processor to generate an acoustic sequence based on the combined prosody information, the prosody information offset, and the linguistic sequence.
[0005] According to one aspect, there is provided a system including a processor configured to: receive a linguistic sequence and a prosody information offset; generate, via a trained prosody information predictor, prosody information including a combination of multiple observations based on the linguistic sequence, wherein the multiple observations include a linear combination of statistical measures evaluating prosody components within a predetermined time period; and generate, via a trained neural network, an acoustic sequence based on the combined prosody information, the prosody information offset, and the linguistic sequence.
[0006] According to another aspect, there is provided a computer-implemented method, comprising: receiving a linguistic sequence and a prosody information offset; generating, via a trained prosody information predictor, prosody information including a combination of a plurality of observations based on and aligned with the linguistic sequence, wherein the plurality of observations includes a linear combination of statistical measures evaluating prosody components within a predetermined time period; and generating an acoustic sequence based on the combined prosody information, the prosody information offset, and the linguistic sequence via a trained neural network.
[0007] According to another aspect, there is provided a computer program product for automatically controlling prosody, the computer program product comprising a computer-readable storage medium having program code embodied therewith, wherein the computer-readable storage medium itself is not a transient signal, the program code being executable by a processor to cause the processor to: receive a linguistic sequence and a prosody information offset; generate prosody information including a combination of a plurality of observations based on the linguistic sequence, wherein the plurality of observations includes a linear combination of statistical measures evaluating prosody components within a predetermined time period; and generate an acoustic sequence based on the combined prosody information, the prosody information offset, and the linguistic sequence. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Embodiments of the present invention will now be described by way of example only with reference to the accompanying drawings, in which:
[0009] Figure 1 is a block diagram of an example system for training a neural network to automatically control prosody using prosody information;
[0010] Figure 2 is a block diagram of an example system for generating embedded prosody information;
[0011] Figure 3 is a block diagram of an example method for training a neural network to automatically control prosody using prosody information;
[0012] Figure 4 is a block diagram of an example method for generating an acoustic sequence with automatically controlled prosody;
[0013] Figure 5 is a block diagram of an example computing device for automatically controlling prosody using prosody information;
[0014] Figure 6 is a diagrammatic illustration of an example cloud computing environment according to embodiments described herein;
[0015] Figure 7 is a diagrammatic illustration of an example abstract model layer according to embodiments described herein; and
[0016] Figure 8is an example tangible non-transitory computer-readable medium that can automatically control prosody using prosody information. Detailed Description
[0017] A text-to-speech (TTS) system, such as a sequence-to-sequence (seq2seq) neural TTS system, can receive a linguistic sequence as input and output a speech acoustic sequence. For example, the speech acoustic sequence can be represented by frame-by-frame speech parameters or by a speech waveform. Such a system can generate speech with a near-natural speech quality with some variations in prosody. Prosody can include phoneme duration, intonation, and volume. However, such systems implicitly generate speech prosody, and thus prosody control in such systems can be very limited. For example, without being guided, such systems may generate outputs with a random speaking style and prosodic characteristics.
[0018] In addition, in many applications, there may be a request to control prosody, including speaking style, emotional state, speaking rate, and expressiveness at inference time. Semi-supervised methods utilize prosody / speaking style labels, which can be generated partially or fully by human subjects. However, human labeling is expensive, error-prone, and time-consuming. Additionally, there are very few labeled resources for speech synthesis. In example-based prosody control methods, the acoustic / prosodic realization of speech can be transformed from a given spoken example by any speaker using an appropriate latent space representation. However, these methods may be infeasible in most practical TTS applications. In unsupervised methods, a latent space of speech acoustics can be automatically trained. The latent parameters can be disentangled to enable them to operate independently at inference time. However, the automatically trained latent representations may often be uninterpretable and highly data-dependent. Additionally, their controllability and the quality of the synthesized speech may also be inconsistent.
[0019] According to an embodiment of the present disclosure, a system can include a processor for receiving a linguistic sequence and a prosody information offset. The processor can generate prosody information including a combination of multiple observations based on the linguistic sequence via a trained prosody information predictor. The observations can be a linear combination of statistical measurements that evaluate prosodic components over a predetermined time period. The processor can also generate an acoustic sequence based on the combined prosody information, the prosody information offset, and the linguistic sequence via a trained neural network. Thus, embodiments of the present disclosure provide a fully automated method that explicitly models prosody in a system and enables sentence-by-sentence speaking pace and expressiveness control on a continuous scale. The techniques described herein also improve the overall quality and expressiveness of the synthesized speech.
[0020] Now refer to Figure 1, The block diagram illustrates an example system for training a neural network to automatically control prosody using embedded prosody information. System 100 can be used to implement methods 300 and 400, and can be implemented using the computing device 500 of Figure 5 or the computer-readable medium 800 of Figure 8 . As an example, system 100 can be a neural sequence to sort attention. Figure 1 System 100 of
[0021] includes a linguistic encoder 102. For example, the linguistic encoder 102 can include a linear embedding layer, followed by a one-dimensional convolutional layer, and a long short-term memory (LSTM) layer. As used herein, the output of the encoder includes a sequence of embedding vectors, i.e., a sequence of learned continuous vector representations of discrete input vectors. Long short-term memory is an artificial recurrent neural network architecture. LSTM has feedback connections and is designed to process data sequences. System 100 includes a prosody information predictor 104, which is communicatively coupled to the linguistic encoder 102. For example, the prosody information predictor 104 can have an embedded linguistic sequence fed into a stacked LSTM (128×3), followed by a linear fully-connected (FC) layer. System 100 also includes a connector 106 communicatively coupled to the prosody information predictor 104. System 100 also includes a combiner 108 communicatively coupled to the prosody information predictor 104. System 100 includes a prosody information encoder 110, which is communicatively coupled to the prosody information predictor 104 and the connector 106. For example, the prosody information encoder 110 can include an FC layer, followed by a hyperbolic tangent non-linearity. System 100 also includes an acoustic decoder 112 communicatively coupled to the connector 106. For example, the acoustic decoder 112 can include an autoregressive mel-spectrogram predictor. In some instances, the acoustic decoder 112 can include two stacked LSTM layers with an attention mechanism. In various examples, the final layer of the acoustic decoder 112 is a fully-connected layer (FC) that outputs an 80-dimensional sequence of mel-spectrograms and a 1-dimensional sequence of stop bits. System 100 is shown receiving a linguistic sequence 114 and outputting an acoustic sequence 116. The linguistic encoder 102 is shown generating an embedded linguistic sequence 118. The prosody information predictor 104 is shown generating combined prosody information 119. The combiner 108 is shown receiving the combined prosody information 119 and a set of prosody information offsets 120. The prosody information encoder 110 is shown generating embedded prosody information 121. System 100 includes an observed prosody information generator 122, which is shown sending a training target 124 to the prosody information predictor 104 and the prosody information encoder 110. System 100 also includes an observed spectrogram generator 126, which is shown sending a training target 128 to the acoustic decoder 112. Figure 1In an example, system 100 can be trained to receive a linguistic sequence 114 and output an acoustic sequence 116. In particular, the linguistic sequence 114 input to the seq2seq neural TTS system can be augmented with prosodic information. As used herein, prosodic information refers to a set of interpretable temporal observations. For example, the observations can be evaluated globally and / or locally and hierarchically over different time spans. Each observation is a linear combination or a set of linear combinations of statistical measures that evaluate prosodic components over a predetermined time period. In human speech, the same speech information can be conveyed in many ways. The linguistic embedding sequence 118 encapsulates all the speech information used in the system, while the prosodic information observations in the form of training targets 124 extracted from the recordings during training provide additional cues on how to convey that speech information. In various examples, the observations included in the prosodic information can be unraveled and easily interpretable. For example, having different components for pace, pitch, and loudness. In some examples, any number of components can be used for the observations. For example, if the speech corpus has a uniform loudness, the loudness control can be omitted, leaving pace and pitch control as the two components used.
[0022] In various examples, the linguistic sequence 114 can be a sequence of symbols representing speech, represented by one-hot or sparse binary vectors, which describe the input phonemes. As an example, the linguistic sequence 114 can be a sequence of indices corresponding to a discrete alphabet of phonemes. In various examples, the acoustic sequence 116 can be a sequence of acoustic parameters. For example, the acoustic sequence 116 can include a frame-wide spectrogram or a constant-frame spectrogram. In various examples, the spectrogram can be converted to speech using a vocoder. As an example, any suitable vocoder can be used to convert the acoustic sequence 116 to speech. A vocoder is a codec used to analyze and synthesize human voice signals for audio data compression, multiplexing, voice encryption, voice transformation, etc. As an example, the vocoder can be a neural network vocoder.
[0023] Still referring to Figure 1 , during the training and inference phases, the linguistic encoder 102 can receive the linguistic sequence 114 and generate the linguistic embedding sequence 118. The embedding can be a vector representation of a phoneme in a particular speech context. For example, the vector representation can be in the form of 128 numbers. In various examples, the form of the vector representation is learnable during the joint training of the neural network 100. The linguistic embedding sequence 118 can be sent to both the connector 106 and the prosody information predictor 104.
[0024] During the training phase, system 100 can receive training objective 124 and training objective 128 from observed prosody information generator 122 and observed spectrogram generator 126 respectively. For example, the observed prosody information vector can be fed to the system. In various examples, a sequence of prosody information vectors is automatically calculated for a training set of input utterances. The utterances can include both the recording and the transcription of the recording. In some examples, the transcription can be automatically generated. For example, a pitch and energy estimator can be used to calculate pitch and energy trajectories, and automatic speech alignment can be applied to partition the time signal into phonemes, syllables, words, and phrase segments. Then pitch, duration, and energy observations can be derived for various time spans. Then, the observations can be aligned and combined with each other to generate a combined sequence of prosody information vectors. In some examples, the prosody information can be set to zero for the first five epochs of training to facilitate alignment convergence at the initial steps of training. As an example, the prosody information can be set to zero for approximately 1500 mini-batch steps.
[0025] In various examples, after training is completed, the prosody information predictor 104 can be trained separately by minimizing the mean squared error (MSE) loss. For example, the prosody information predictor 104 can be fed with the linguistic embedding sequence 118 and predict the combined prosody information from the linguistic embedding sequence 118. In some examples, the prediction is done using a 3-layer stacked LSTM with 128 units in each layer, followed by a linear layer that produces a prosody information vector with an output size of 2. In some examples, the prosody information predictor 104 can be jointly trained as a sub-network using multi-objective training with the rest of system 100. For example, both sets of training objectives 124 and training objective 128 can be used to jointly train the prosody information predictor 104 and system 100. In various examples, an additional loss can be added to the loss associated with the output acoustic sequence loss to jointly train the prosody information predictor 104. In some examples, the prosody information predictor 104 can be trained separately. For example, the prosody information predictor 104 can be trained separately as a seq2seq acoustic neural network to predict the combined prosody information from the linguistic sequence 114. In some examples, the prosody information observations can also include acoustic observations. For example, the acoustic observations can include observations of other non-linguistic aspects of speech acoustics that may be related to speaking style, such as speech breathiness, hoarseness, voiced effort, etc.
[0026] During the inference phase, the prosody information predictor 104 receives the linguistic embedding sequence 118 and generates the combined prosody information 119. For example, the combined prosody information 119 includes multiple observations. The observation includes a linear combination of statistical measures that evaluate prosodic components over a predetermined time period. In various examples, the observations can be evaluated globally, locally, and hierarchically at different time spans. For example, the global observation can be at the utterance level. The observations evaluated hierarchically locally can be at the level of each paragraph, sentence, phrase, word, syllable, or phoneme segment. As used herein, a segment refers to a time span within such a hierarchical time structure of paragraph / sentence / phrase / word / syllable / phoneme. Then, through concatenation or summation, the observations can be aligned and combined with each other to generate the combined prosody information. Then, the combined prosody information 119 can be embedded via the prosody information encoder 110 to generate the embedded prosody information 121.
[0027] In various examples, the set of observations can at least include in-segment log pitch observations, in-segment sub-segment log duration observations, in-segment log energy observations, or any combination thereof. For example, the log pitch observation can be the span of log pitch evaluated as the 0.95 quantile minus the 0.05 quantile of the utterance log pitch trajectory. As used herein, a sub-segment refers to a segment that is deeper in the hierarchy compared to another segment. For example, the log duration observation can be the logarithm of the average phoneme duration (excluding silences) as a measure of the pace of the utterance. In some examples, the sub-segment log duration observation can measure the duration of words within a phrase. In various examples, each observation in the observations can be a linear combination of statistical measures. Each observation in the observations can include at least some form of statistical measure, such as mean, a set of quantiles, span, standard deviation, variance, or any combination thereof. In various examples, the observations are normalized for each speaker. Regarding Figure 2 Observations are discussed in more detail.
[0028] The prosody information predictor 104 thus generates a set of observations for various prosody parameters that describe the input linguistic sequence. Since these observations are normalized and tractable, one or more prosody information offsets 120 can be applied during inference to adjust the prosody of the final acoustic sequence 116. The prosody information can be deliberately changed by adding a component-wise offset within the range of [-1, 1]. For example, by adjusting the corresponding sub-segment log duration observation value towards -1, an utterance, paragraph, sentence, phrase, or word can be made slower, or by adjusting towards 1, an utterance, paragraph, sentence, phrase, or word can be made faster. Similarly, by modifying the corresponding log pitch observation or log energy observation value towards -1 or 1, the pitch or loudness variation of the entire utterance or any of its paragraphs, sentences, phrases, or words can be adjusted to make the output acoustic sequence 116 more monotonic or more expressive, respectively.
[0029] In various examples, the combined prosody information vector is embedded into a two-dimensional latent space and concatenated with each vector in the output sequence of the linguistic encoder. For example, the prosody information vector can be embedded by a single fully connected unbiased layer with a hyperbolic tangent non-linearity. Thus, the decoder is exposed to the prosody information via the input context vector.
[0030] The combined prosody information observations are thus further fed into the main seq2seq acoustic neural network. The voice decoder 112 can be a neural network that receives the concatenated sequence from the concatenator 106 and generates an acoustic sequence 116.
[0031] As an example, the system 100 can have two-dimensional global (per utterance) observations: the logarithmic pitch span, and the median phoneme logarithmic duration concatenated with the two-dimensional word-level observations: the logarithmic pitch span and the median phoneme logarithmic duration. All observations can be normalized to [-1:1]. Due to the global observations, the system user can control the global speech pace and expressiveness. For example, the user can add a positive global duration modification amount to slow down the speech or make the speech clearer. Additionally, the user can add a positive global pitch span modification amount to increase the speech expressiveness. Using the word-level observations in the combined prosody information, the system 100 can control the desired word emphasis. For example, such word emphasis can be useful in a dialogue application. In some examples, the user can deliberately apply a positive duration modification amount and a positive pitch span modification amount to a subsequence of the observations corresponding to the desired word. In experiments utilizing the proposed prosody information control for several speech corpora, the example system successfully slows down or speeds up in response to per-component prosody information inference time modification as a response to the pace component modification, and increases or decreases expressiveness in response to the pitch component modification.
[0032] It should be understood that Figure 1 the block diagram of Figure 1 is not intended to indicate that the system 100 should include all of the components shown in Figure 1 Rather, the system 100 can include fewer components or additional components (e.g., additional client devices or additional resource servers, etc.) not shown in
[0033] Now referring to Figure 2 , the block diagram illustrates an example system for encoding prosody information. The example system 200 can be used to implement the method of Figure 3 and can be implemented using the computing device 500 of Figure 5 or the computer-readable medium 800 of Figure 8 .
[0034] Figure 2System 200 includes a prosody information encoder 110, which is coupled to an observed prosody information generator 122. System 202 can receive an input utterance 202 and output embedded prosody information 204. For example, the input utterance 202 can be training data for training Figure 1 system 100 using the embedded prosody information 204. In various examples, the input utterance 202 can include recorded paragraphs, sentences, words, etc.
[0035] In Figure 2 an example, the observed prosody information generator 122 receives the input utterance and produces a set of prosody observations. As Figure 2 shown, the prosody observations can include observations at various levels, including sentence prosody observations 206, phrase prosody observations 208, and word prosody observations 210, as well as other possible levels of prosody observations. In various examples, each type of the prosody observations 206, 208, 210 can include at least some of the following types: within-segment logarithmic pitch observations, within-segment sub-segment logarithmic duration observations, and within-segment logarithmic energy observations. For example, other types of prosody observations can be breathiness, noise level, nasality, voice quality, etc. For example, breathiness can be evaluated by the harmonic-to-noise ratio at the voiced speech portion. In some examples, the noise level can be evaluated by the SNR estimate during silence. In various examples, nasality can be evaluated using average formant analysis. In some examples, voice quality can be evaluated using glottal pulse modeling and analysis of glottal closure and opening intervals for the voiced speech portion. For example, the glottal pulse modeling used can be Liljencrants-Fant glottal pulse modeling. Generally, each observation in the observations can be a linear combination of statistical measurements. Each observation can include statistical measurements such as mean, quantile set, standard deviation, variance, or any combination thereof. For example, the quantile set can be in the form: [0.1, 0.5, 0.9]. As described above, the observations can be appropriately normalized for each speaker. For example, the effective span of each observation in the observations can be normalized to [-1, 1]. The effective span can be calculated as: [median - 3*STD, median + 3*STD], where STD is the standard deviation of the set. In some examples, the span can be expressed using quantiles, such as span: 0.95-quantile minus 0.05-quantile.
[0036] In various examples, the aligner and combiner 212 can align and combine the hierarchical observations 206, 208, and 210. For example, the aligner and combiner 212 can align and combine the hierarchical observations 206, 208, and 210 by summing or concatenating to produce combined prosody information synchronized with the input linguistic sequence, which can include a sequence of observation vectors.
[0037] Still referring to Figure 2 , the embedder 214 can embed the combined prosodic information from the aligner and combiner 212 to generate the embedded prosodic information 204. For example, the embedded prosodic information 204 can include a single embedding vector per utterance or a sequence of embedding vectors synchronized with the input linguistic sequence. In various examples, the embedded prosodic information 204 can then be used to train an acoustic decoder, as Figure 1 described in
[0038] It should be understood that Figure 2 the block diagram of Figure 2 is not intended to indicate that the system 200 should include all the components shown Figure 2 , but rather, the system 200 can include fewer components or additional components not shown in
[0039] Figure 3 For example, during inference, instead of the observed prosody generator 122, a prosody information predictor can be fed into the prosody information encoder 110 or the embedder 214. Figure 5 is a process flow diagram of an example method that can train a neural network to automatically control prosody using the embedded prosodic information. The method 300 can be implemented using any suitable computing device, such as Figure 1 the computing device 500 of Figure 2 and described with reference to the systems 100 and 200 of Figure 5 For example, the method 300 can be implemented by the trainer module 536 of the computing device 500 of Figure 8 or the trainer module 818 of the computer-readable medium 800 of
[0040] At block 302, a linguistic sequence and a corresponding acoustic sequence are received. For example, the linguistic sequence can correspond to an input utterance for training.
[0041] At block 304, the observed combined prosodic information is generated based on the linguistic sequence and the corresponding acoustic sequence. For example, the observed combined prosodic information can be a sequence of observed prosodic information for different time spans that is automatically computed and corresponds to the input utterance for training. The observed prosodic information can be temporarily aligned and combined, for example, by using concatenation or summation, to obtain a sequence of the observed combined prosodic information. In various examples, the observed prosodic information can include any combination of observations (including statistical measures associated with the input utterance), such as in-segment log pitch observations, in-segment sub-segment log duration observations, in-segment log energy observations, or any combination thereof.
[0042] At block 306, the observed combined prosodic information is used, along with the linguistic and acoustic sequences, to train a neural network to predict the acoustic sequence. For example, the neural network can include a prosody information encoder, a linguistic encoder, and an acoustic decoder. As an example, the embedded prosodic information and the embedded linguistic sequence are fed into an acoustic decoder that outputs a sequence of mel spectrograms. For example, the mean squared error (MSE) loss of the mel spectrogram can be used to train the neural network.
[0043] At block 308, a prosody information predictor is trained to predict the combined prosodic information observations using the linguistic sequence. In some examples, the prosody information predictor can be trained to predict hierarchical prosodic information observations, which can be further aligned and combined to generate the combined prosodic information. In various examples, the prosody information predictor can be trained to directly predict the combined prosodic information observations. In various examples, the prosody information predictor can be trained separately or jointly with the decoder. As an example, at block 306, the decoder can be trained separately. Then, the prosody information predictor can be trained based on the linguistic sequence and the training objective. In some examples, the prosody information predictor can be trained based on the embedded linguistic sequence from the trained linguistic encoder.
[0044] As an example, the prosody information predictor can be combined with a sequence-to-sequence mel spectrogram feature prediction module. For example, the mel spectrogram feature prediction module can be based on the Tacotron2 architecture released in 2018 and includes a convolutional encoder with a terminal recurrent layer implemented using bidirectional LSTM. The mel spectrogram feature prediction module can encode the linguistic sequence into an embedded linguistic sequence, concatenated with an autoregressive attention decoder that extends the embedded linguistic sequence to a sequence of fixed-frame mel spectrogram feature vectors.
[0045] Specifically, the Tacotron2 decoder takes the previous spectrogram frame s c processed by a pre-net that depends on the input context vector x p generated by the attention module, and predicts one spectrogram frame at a time. The decoder uses a two-layer stacked LSTM network to generate its hidden state vector h c . The hidden state vector h c combined with the input context vector x c is fed into a final linear layer to produce the current mel spectrogram and the end-of-sequence token. Finally, there can also be a post-net convolutional network that refines the entire utterance mel spectrogram to improve the fidelity.
[0046] The Tacotron2 model can directly consume text characters. However, in some examples, the system can be fed a sequence of symbols from an extended speech lexicon to simplify training. For example, the extended speech lexicon can include phone identities, lexical stress, and phrase types, which are rich in different word segmentation and silence symbols. Lexical stress can be a 3-way parameter including primary, secondary, and unstressed. Phrase types can be a 4-way parameter including affirmative, interrogative, exclamatory, and "other" values. In some examples, such a linguistic input sequence can be generated by a TTS front-end module based on external grapheme-to-phoneme rules (e.g., unit selection TTS released in 2006).
[0047] In some instances, better synthetic speech quality can be obtained by incorporating the mean squared error (MSE) applied to the difference between the current and previous mel spectrograms into the final system loss. For example, given the predicted mel spectrogram y at time t before the post-net t , the final predicted mel spectrogram z at time t t , and the mel spectrogram target q at time t t , the spectral loss can be calculated using the following equation:
[0048] Loss spc = 0.5MSE(y t , q t ) + 0.25MSE(z t , q t ) + 0.25MSE(z t - z t-1 , q t - q t-1 , ) Equation 1
[0049] In various examples, opposite to the inference process where the prediction is autoregressive, the training process can follow the teacher-forcing method. For example, the prediction of the current mel spectrogram is performed based on the true previous mel spectrogram and processed by the pre-net. In some examples, double feeding can be applied during training. For example, the pre-net of the decoder can be fed the concatenated true previous mel spectrogram and predicted mel spectrogram. At inference time, when the true frames are not available, the predicted mel spectrogram can simply be copied. Although only increasing the total network size by 0.1%, this modification reduces the total model regression loss by approximately 15%, as tested on two professional recordings of 13-hour and 22-hour American English corpora.
[0050] Figure 3 The process flow diagram of... is not intended to indicate that the operations of method 300 are to be performed in any particular order, or that all operations of method 300 are to be included in each case. Additionally, method 300 can include any suitable number of additional operations.
[0051] Figure 4 is a process flow diagram of an example method that can generate sequences with automatically controlled prosody. Method 400 can be implemented using any suitable computing device, such as Figure 5 computing device 500, and is described with reference to Figure 1 and 2 systems 100 and 200. For example, method 400 can be implemented by Figure 5 and Figure 8 computing device 500 and computer-readable medium 800.
[0052] At block 402, a linguistic sequence and a prosody information offset are received. For example, the linguistic sequence can be a text sequence. The prosody information offset can be a set of external per-component modifications for deliberately shifting the prosodic characteristics of the synthesized speech. For example, the prosody information offset can be used to change speech pacing, pitch variability, volume variability, etc.
[0053] At block 404, combined prosody information is generated based on the linguistic sequence via a trained prosody information predictor. For example, the combined prosody information can include multiple observations. The observations include a linear combination of statistical measurements that evaluate prosodic components over a predetermined time period. For example, the observations can be evaluated at the utterance level. In some examples, the observations are evaluated locally and hierarchically at different time spans. In various examples, the observations can be further aligned and combined in time to obtain combined prosody information observations. Alternatively, the combined prosody information can be directly predicted from the linguistic sequence. In some examples, the prosody information can be generated based on an embedded linguistic sequence. In some examples, the embedded linguistic sequence can be an embedded sequence of discrete variables, i.e., a discrete linguistic sequence that is mapped to a continuous embedding space.
[0054] At block 406, an acoustic sequence is generated based on the combined prosody information, the prosody information offset, and the linguistic sequence via a trained neural network. For example, the trained neural network can include a prosody information encoder, a linguistic encoder, and an acoustic decoder. In some examples, the combined prosody information components are modified based on the prosody information offset. For example, the prosody information offset can be added to the corresponding observations. In some examples, the combined prosody information passes through a prosody information embedder to generate embedded prosody information. For example, the prosody information embedder can align, combine, and embed the observations to generate embedded prosody information. Then, the embedded prosody information can be concatenated with the linguistic sequence or the embedded linguistic sequence and used by the decoder to generate the acoustic sequence.
[0055] Figure 4The process flow diagram of [the method] is not intended to indicate that the operations of method 400 will be performed in any particular order, or that all operations of method 400 will be included in every instance. Additionally, method 400 may include any suitable number of additional operations. For example, method 400 may include generating audio based on an acoustic sequence.
[0056] In some scenarios, the techniques described herein may be implemented in a cloud computing environment. As will be discussed in more detail below, a computing device configured to automatically control prosody using embedded prosody information may be implemented in a cloud computing environment. It is understood in advance that although this disclosure may include a description of cloud computing, the implementation of the teachings recited herein is not limited to a cloud computing environment. Rather, embodiments of the present invention are capable of being implemented in conjunction with any other type of computing environment now known or later developed. Figures 5 - 8 Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with the service provider. The cloud model may include at least five characteristics, at least three service models, and at least four deployment models.
[0057] The characteristics are as follows:
[0058] On-demand self-service: Cloud consumers can unilaterally and automatically provision computing capabilities, such as server time and network storage, as needed, without human interaction with the service provider.
[0059] Broad network access: The capabilities are available over a network and accessed through standard mechanisms that facilitate use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs).
[0060] Resource pooling: The provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, where different physical and virtual resources are dynamically assigned and reassigned as needed. There is a sense of location independence in that consumers generally have no control or knowledge of the exact location of the resources provided, but may be able to specify a location at a higher level of abstraction (e.g., country, state, or data center).
[0061] Rapid elasticity: Capabilities can be rapidly and elastically, in some cases automatically, provisioned to quickly scale down and quickly released to quickly scale up. For consumers, the capabilities available for provisioning generally appear to be unlimited and can be purchased at any time in any quantity.
[0062]
[0063] Measured Services: The cloud system automatically controls and optimizes resource use by leveraging metering capabilities at an abstract level appropriate to the service type (e.g., storage, processing, bandwidth, and active user accounts). Resource use can be monitored, controlled, and reported, providing transparency for both the provider and consumer of the utilized services.
[0064] The service models are as follows:
[0065] Software as a Service (SaaS): The capability provided to the consumer is to use the provider's applications running on the cloud infrastructure. The applications are accessible from different client devices through a thin client interface such as a web browser (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage devices, or even individual application capabilities, with the possible exception of limited user-specific application configuration settings.
[0066] Platform as a Service (PaaS): The capability provided to the consumer is to deploy onto the cloud infrastructure consumer-created or acquired applications that are created using programming languages and tools supported by the provider. The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, or storage devices, but has control over the deployed applications and possibly the application hosting environment configuration.
[0067] Infrastructure as a Service (IaaS): The capability provided to the consumer is to provision processing, storage, networks, and other fundamental computing resources that the consumer can deploy and run arbitrary software, which can include operating systems and applications. The consumer does not manage or control the underlying cloud infrastructure, but has control over the operating systems, storage devices, deployed applications, and possibly limited control over the selection of networking components (e.g., host firewalls).
[0068] The deployment models are as follows:
[0069] Private Cloud: The cloud infrastructure is operated solely for an organization. It can be managed by the organization or a third party and can exist on-premises or off-premises.
[0070] Community Cloud: The cloud infrastructure is shared by several organizations and supports a specific community with shared concerns (e.g., mission, security requirements, policies, and compliance considerations). It can be managed by the organizations or a third party and can exist on-site or off-site.
[0071] Public Cloud: Makes the cloud infrastructure available to the general public or a large industry group and is owned by an organization that sells cloud services.
[0072] Hybrid Cloud: A cloud infrastructure that is a composition of two or more clouds (private, community, or public) that remain unique entities but are bound together by standardized or proprietary technology that enables data and application portability (e.g., cloud bursting for load balancing between clouds).
[0073] The cloud computing environment is service-oriented, focusing on statelessness, low coupling, modularity, and semantic interoperability. The core of cloud computing is the infrastructure that includes a network of interconnected nodes.
[0074] Figure 5 It is a block diagram of an example computing device that can automatically control prosody using embedded prosody information. The computing device 500 can be, for example, a server, a desktop computer, a laptop computer, a tablet computer, or a smart phone. In some examples, the computing device 500 can be a cloud computing node. The computing device 500 can be described in the general context of computer system executable instructions, such as program modules executed by a computer system. Generally, program modules can include routines, programs, objects, components, logic, data structures, etc. that perform particular tasks or implement particular abstract data types. The computing device 500 can be practiced in a distributed cloud computing environment where tasks are performed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules can be located in both local and remote computer system storage media including memory storage devices.
[0075] The computing device 500 can include a processor 502 for executing stored instructions and a memory device 504 for providing temporary memory space for the operation of the instructions during operation. The processor can be a single-core processor, a multi-core processor, a computing cluster, or any other number of configurations. The memory 504 can include random access memory (RAM), read-only memory, flash memory, or any other suitable memory system.
[0076] The processor 502 can be connected through a system interconnect 506 (e.g., PCI, PCIe, etc.) to an input / output (I / O) device interface 508 adapted to connect the computing device 500 to one or more I / O devices 510. The I / O devices 510 can include, for example, a keyboard and a pointing device, where the pointing device can include a touchpad or a touch screen, etc. The I / O devices 510 can be built-in components of the computing device 500 or can be devices externally connected to the computing device 500.
[0077] The processor 502 may also be linked via a system interconnect 506 to a display interface 512 adapted to connect the computing device 500 to a display device 514. The display device 514 may include a display screen that is an in-built component of the computing device 500. The display device 514 may also include a computer monitor, a television set, a projector, etc. that are externally connected to the computing device 500. Additionally, a network interface controller (NIC) 516 may be adapted to connect the computing device 500 to a network 518 via the system interconnect 506. In some embodiments, the NIC 516 may use any suitable interface or protocol to transmit data, such as Small Computer System Interface for Internet, etc. The network 518 may be a cellular network, a radio network, a wide area network (WAN), a local area network (LAN), or the Internet, etc. An external computing device 520 may be connected to the computing device 500 via the network 518. In some examples, the external computing device 520 may be an external network server 520. In some examples, the external computing device 520 may be a cloud computing node.
[0078] The processor 502 may also be linked to a storage device 522 via a system interconnect 506, which may include a hard disk drive, an optical disk drive, a USB flash drive, a drive array, or any combination thereof. In some examples, the storage device may include a receiver module 524, a linguistic encoder module 526, a predictor module 528, a prosody encoder module 530, a linker module 532, an acoustic decoder module 534, and a trainer module 536. The receiver module 524 may receive a linguistic sequence and a prosody information offset. For example, the linguistic sequence may be a text sequence. The linguistic encoder module 526 may generate an embedded linguistic sequence based on the received linguistic sequence. The predictor module 528 may generate prosody information including a combination of multiple observations over respective time periods based on the linguistic sequence or the embedded linguistic sequence. The observations may be aligned with the linguistic sequence and combined by summation or concatenation. The observations include a linear combination of statistical measures evaluating prosody components within a predetermined time period. For example, the observations may be a linear combination or a set of linear combinations of statistical measures evaluating a pacing component, a pitch component, a loudness component, or any combination thereof. In some examples, the observations may include sentence prosody observations, phrase prosody observations, and word prosody observations, or any combination thereof. The prosody encoder module 530 may modify the observations based on the prosody information offset so as to adjust the prosody of an acoustic sequence in a specific predetermined manner. The prosody encoder module 530 may also embed the observations to generate embedded prosody information. The connector module 532 may concatenate the embedded prosody information with the embedded linguistic sequence. The acoustic decoder module 534 may generate an acoustic sequence based on the combined prosody information, the prosody information offset, and the linguistic sequence. For example, the decoder module 534 may generate an acoustic sequence based on the combined prosody information observations and the prosody information offset. The trainer module 536 may train a prosody information predictor based on observed prosody information extracted from unlabeled training data. For example, the trainer module 536 may train the linguistic encoder module 526 and the acoustic decoder module 534 based on observed spectra extracted from recordings during training. In some examples, the trainer module 536 may train the prosody information predictor based on the embedded linguistic sequences generated by a system trained with the observed prosody information.
[0079] It should be understood that Figure 5 the block diagram of Figure 5 is not intended to indicate that the computing device 500 is to include Figure 5Additional components not shown (e.g., additional memory components, embedded controllers, modules, additional network interfaces, etc.). Additionally, any of the functionality of the receiver 524, linguistic encoder module 526, predictor module 528, prosody encoder module 530, linker module 532, acoustic decoder module 534, and trainer module 536 can be implemented in part or in whole in hardware and / or by the processor 502. For example, the functionality can be implemented using application specific integrated circuits, logic implemented in an embedded controller, or logic implemented in the processor 502, etc. In some embodiments, the functionality of the receiver module 524, linguistic encoder module 526, and predictor module 528, prosody encoder module 530, linker module 532, acoustic decoder module 534, and trainer module 536 can be implemented using logic, where, as mentioned herein, logic can include any suitable hardware (e.g., processor, etc.), software (e.g., application, etc.), firmware, or any suitable combination of hardware, software, and firmware.
[0080] Now referring to Figure 6 , an illustrative cloud computing environment 600 is depicted. As shown, the cloud computing environment 600 includes one or more cloud computing nodes 602 with which local computing devices used by cloud consumers can communicate, such as, for example, a personal digital assistant (PDA) or cellular phone 604A, desktop computer 604B, laptop computer 604C, and / or in-vehicle computer system 604N. The nodes 602 can communicate with one another. They can be physically or virtually grouped (not shown) in one or more networks, such as a private cloud, community cloud, public cloud, or hybrid cloud or combinations thereof as described above. This allows the cloud computing environment 600 to provide infrastructure, platforms, and / or software as a service, for which cloud consumers do not need to maintain resources on local computing devices. It should be understood that Figure 6 the types of computing devices 604A-N shown in are only illustrative, and the computing nodes 602 and the cloud computing environment 600 can communicate with any type of computing device via any type of network and / or network addressable connection (e.g., using a web browser).
[0081] Now referring to Figure 7 , a set of functional abstraction layers provided by the cloud computing environment 600 ( Figure 6 ) is shown. It should be understood in advance that Figure 7 the components, layers, and functionality shown in are only illustrative, and embodiments of the present invention are not limited thereto. As described, the following layers and corresponding functionality are provided.
[0082] The hardware and software layer 700 includes hardware and software components. Examples of hardware components include mainframes, in one example System; a server based on the RISC (Reduced Instruction Set Computer) architecture, in one example, IBM System; IBM System; IBM System; storage devices; network and networking components. Examples of software components include web application server software, in one example, IBM Web application server software; and database software, in one instance, IBM Database software. (IBM, zSeries, pSeries, xSeries, BladeCerter, WebSphere, and DB2 are trademarks of International Business Machines Corporation registered in many jurisdictions worldwide).
[0083] The virtualization layer 702 provides an abstraction layer from which the following examples of virtual entities can be provided: virtual servers; virtual storage devices; virtual networks, including virtual private networks; virtual applications and operating systems; and virtual clients. In one example, the management layer 704 can provide the functions described below. Resource provisioning provides the dynamic procurement of computing resources and other resources used to perform tasks within a cloud computing environment. Metering and pricing provide cost tracking for the utilization of resources in a cloud computing environment, as well as billing or invoicing for the consumption of those resources. In one example, these resources can include application software licenses. Security provides authentication for cloud consumers and tasks, as well as protection for data and other resources. The user portal provides access to the cloud computing environment for consumers and system administrators. Service level management provides cloud computing resource allocation and management such that the required service levels are met. Service level agreement (SLA) planning and fulfillment provide the pre-arrangement and procurement of cloud computing resources, where future demands are projected based on the SLA.
[0084] The workload layer 706 provides examples of functionality that can utilize a cloud computing environment. Examples of workloads and functions that can be provided from this layer include: mapping and navigation; software development and lifecycle management; virtual classroom education delivery; data analysis processing; transaction processing; automatic rhythm control.
[0085] The present technology can be a system, method, or computer program product. The computer program product can include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to execute aspects of the present invention.
[0086] A computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer-readable storage medium can be, for example but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer-readable storage medium includes the following: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device (such as a punched card or raised structures in grooves having instructions recorded thereon), and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0087] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network). The network can include a copper transmission cable, an optical transmission fiber, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the corresponding computing / processing device.
[0088] The computer-readable program instructions for performing the operations of the present technology may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, and conventional procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, an electronic circuit including, for example, a programmable logic device, a field-programmable gate array (FPGA), or a programmable logic array (PLA) may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit device in order to perform aspects of the present invention.
[0089] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present technology. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0090] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions executed via the processor of the computer or other programmable data processing apparatus create a means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium having instructions stored therein includes an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0091] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0092] Now referring to Figure 6 , a block diagram of an example tangible non-transitory computer-readable medium 600 that can use embedded prosodic information to automatically control prosody is depicted. The tangible non-transitory computer-readable medium 800 can be accessed by a processor 802 via a computer interconnect 804. Additionally, the tangible non-transitory computer-readable medium 800 can include code for guiding the processor 802 to perform Figure 3 and Figure 4 the operations of methods 300 and 400.
[0093] As Figure 8 shown, the various software components discussed herein can be stored on the tangible, non-transitory computer-readable medium 800. For example, the receiver module 806 includes code for receiving a linguistic sequence and a prosody information offset. The linguistic encoder module 808 includes code for generating an embedded linguistic sequence based on the linguistic sequence. The predictor module 810 also includes code for generating prosody information including a combination of observations over various time periods based on the linguistic sequence. The observations can be aligned with the linguistic sequence and combined by summation or concatenation. The observations include a linear combination of statistical measures evaluating prosodic components within a predetermined time period. The prosody encoder module 812 includes code for encoding the observations to generate embedded prosody information. In some examples, the prosody encoder module 812 includes code for modifying the observations based on the prosody information offset. For example, the prosody encoder module 812 includes code for adding the prosody information offset to the corresponding observations. The concatenator module 814 includes code for concatenating the embedded prosody information with the embedded linguistic sequence. The acoustic decoder module 816 includes code for generating an acoustic sequence based on the embedded prosody information, the prosody information offset, and the linguistic sequence or the embedded linguistic sequence. The trainer module 818 includes code for training a prosody information predictor based on observed prosody information extracted from unlabeled training data. It should be understood that depending on the specific application, Figure 8 any number of additional software components not shown in
[0094] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions. It should be understood that depending on the specific application, Figure 8 any number of additional software components not shown therein may be included within the tangible, non-transitory computer-readable medium 800. For example, the computer-readable medium 800 may also include code for generating audio based on an acoustic sequence.
[0095] The description of the different embodiments of the present technology has been presented for purposes of illustration, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terms used herein have been chosen to best explain the principles of the embodiments, the practical application, or technical improvements made to the technology found in the marketplace, or to enable those of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A system includes a processor, the processor being configured to: Receive a linguistic sequence and a prosody information offset, the prosody information offset including an adjustment for a target observation over a target time span; Encode the linguistic sequence via a linguistic encoder to generate an embedded linguistic sequence, the embedded linguistic sequence including a vector representation of phonemes in a speech context; Generate combined prosody information based on the linguistic sequence via a trained prosody information predictor, the combined prosody information including a plurality of observations, wherein the plurality of observations includes a linear combination of statistical measures evaluating a plurality of prosody components over a plurality of hierarchical time spans, wherein the observations are normalized; Modify the combined prosody information based on the prosody information offset, wherein the prosody information offset adjusts the target observation among the plurality of observations for a specified time span in the hierarchical time spans; Embed the modified combined prosody information into a latent space to generate embedded prosody information; And Generate an acoustic sequence based on the embedded prosody information concatenated with the linguistic embedding via a trained neural network, wherein prosody characteristics of the generated acoustic sequence are adjusted based on the prosody information offset.
2. The system according to claim 1, wherein the processor is configured to: Train the prosody information predictor based on observed prosody information automatically extracted from unlabeled training data via an observed prosody information generator.
3. The system according to claim 1, wherein the processor is configured to: Train the prosody information predictor based on the embedded linguistic sequence, the embedded linguistic sequence being generated by a system trained using the observed prosody information.
4. The system according to claim 1, wherein the processor is configured to: Train the neural network based on observed spectra extracted from recordings during training, the neural network including a sequence-to-sequence neural network, the sequence-to-sequence neural network including a prosody information encoder, a linguistic encoder, and an acoustic decoder.
5. The system according to claim 1, wherein the processor is configured to: Modify the plurality of observations based on the prosody information offset to adjust the prosody characteristics of the acoustic sequence in a specific predetermined manner.
6. The system according to any one of the preceding claims, wherein the prosody components include a pacing component, a pitch component, a loudness component, or any combination thereof.
7. The system according to claim 1, wherein the plurality of observations includes log pitch observations within each segment of the linguistic sequence, sub-segment log duration observations within each segment, and log energy observations within each segment.
8. A computer-implemented method includes: Receive a linguistic sequence and a prosody information offset, the prosody information offset including an adjustment for a target observation over a target time span; Encode the linguistic sequence via a linguistic encoder to generate an embedded linguistic sequence, the embedded linguistic sequence including a vector representation of phonemes in a speech context; Generate combined prosody information via a trained prosody information predictor, based on and aligned with the linguistic sequence, the combined prosody information including a plurality of observations, where the plurality of observations includes a linear combination of statistical measures evaluating a plurality of prosody components over a plurality of hierarchical time spans, and where the observations are normalized; Modify the combined prosody information based on the prosody information offset, where the prosody information offset adjusts the target observation among the plurality of observations for a specified time span in the hierarchical time spans; Embed the modified combined prosody information into a latent space to generate embedded prosody information; and Generate an acoustic sequence via a trained neural network, based on the embedded prosody information concatenated with the linguistic embedding, where the prosody characteristics of the generated acoustic sequence are adjusted based on the prosody information offset.
9. The computer-implemented method according to claim 8, comprising: Combine the plurality of observations by summing or concatenating and encode the plurality of observations to generate the embedded prosody information, and concatenate the embedded prosody information with the embedded linguistic sequence.
10. The computer-implemented method according to claim 8, comprising modifying the plurality of observations based on the prosody information offset.
11. The computer-implemented method according to claim 10, where modifying the plurality of observations includes adding the prosody information offset to the corresponding observation.
12. The computer-implemented method according to claim 8, where the plurality of observations are evaluated at the utterance level.
13. The computer-implemented method according to claim 8, where the plurality of observations are evaluated locally and hierarchically over different time spans.
14. The computer-implemented method according to claim 8, comprising generating audio based on the acoustic sequence.
15. A computer program product for automatically controlling prosody, the computer program product including a computer-readable storage medium having program code embodied therewith, where the computer-readable storage medium itself is not a transient signal, the program code being executable by a processor to cause the processor to: Receive a linguistic sequence and a prosody information offset, the prosody information offset including an adjustment for a target observation at a target time span; Encode the linguistic sequence via a linguistic encoder to generate an embedded linguistic sequence, the embedded linguistic sequence including a vector representation of phonemes in a speech context; Generate combined prosody information based on the linguistic sequence, the combined prosody information including a plurality of observations, where the plurality of observations includes a linear combination of statistical measures evaluating a plurality of prosody components over a plurality of hierarchical time spans, and where the observations are normalized; Modify the combined prosody information based on the prosody information offset, where the prosody information offset adjusts the target observation among the plurality of observations for a specified time span in the hierarchical time spans; Embed the modified combined prosody information into a latent space to generate embedded prosody information; and Generating an acoustic sequence based on prosodic information of the embedding linked to a linguistic embedding, wherein prosodic characteristics of the generated acoustic sequence are adjusted based on an offset of the prosodic information.
16. The computer program product according to claim 15, further comprising program code executable by the processor for: aligning, combining, and embedding the plurality of observations to generate the prosodic information of the embedding, and linking the prosodic information of the embedding with the linguistic sequence of the embedding.
17. The computer program product according to claim 15, further comprising program code executable by the processor for modifying the plurality of observations based on the prosodic information offset.
18. The computer program product according to claim 15, further comprising program code executable by the processor for adding the prosodic information offset to corresponding observations of the prosodic information.
19. The computer program product according to claim 15, further comprising program code executable by the processor for training a prosodic information predictor based on observed prosodic information automatically extracted from unlabeled training data via an observed prosodic information generator.
20. The computer program product according to claim 15, further comprising program code executable by the processor for generating audio based on the acoustic sequence.
21. A computer program comprising program code means which, when the program is run on a computer, is adapted to perform the method according to any one of claims 8 to 14.
Citation Information
Patent Citations
Training method for multiple personalized acoustic models, and voice synthesis method and voice synthesis device
CN105185372A
Method and apparatus for speech synthesis, program, recording medium, method and apparatus for generating constraint information and robot apparatus
US20040019484A1