Cross-sentence Speech Synthesis Method, System and Device Based on Variational Autoencoder
Through the cross-sentence pronunciation synthesis method of the variational autoencoder, the cross-sentence vector representation is generated using phoneme sequences and speaker information, which improves the pronunciation effect of pronunciation synthesis, solves the problems of insufficient pronunciation expressiveness and prior inconsistency in the existing technology, and achieves a more natural and diverse pronunciation synthesis.
Patent Information
- Application Number
- CN202210220764.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-03-03
- Filing Date
- 2022-03-08
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-03-08
AI Technical Summary
In the existing pronunciation synthesis technology, the pronunciation effect of synthesized speech is limited, the expressiveness is insufficient, and the inconsistency is inconsistent between the standard Gaussian prior sampled by the system and the true prior of speech.
Through the cross-sentence pronunciation synthesis method based on the variational autoencoder, the phoneme sequence and speaker information are encoded, cross-sentence representations are combined with the multi-sentence representation and the multi-headed attention layer to generate cross-sentence vector representations, the conditional priors and posterior modules are used to improve rhythm changes, and the phoneme-level cross-sentence representations are generated in combination with the multi-headed attention layer to generate cross-sentence representations, and the cross-sentence information is used as a prior condition of the conditional variational autoencoder.
It improves the naturalness and pronunciation diversity of synthetic pronunciations, solves the problem of inconsistency between standard Gaussian priors and true pronunciation, and generates more natural and expressive pronunciations.
Smart Images

Figure CN114566141B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of speech synthesis, and particularly to a cross-sentence speech synthesis method, system and device based on a variational autoencoder. Background Art
[0002] Speech synthesis technology is the artificial production of human speech, with the goal of converting any input text into clear, natural and expressive speech. The first electronic speech synthesizer was born in 1937, and since then speech synthesis technology has undergone various technical improvements. In the early 1990s, with the proposal of the Pitch Synchronous OverLap and Add (PSOLA) method, the timbre and naturalness of synthetic speech were greatly improved. In recent years, with the rapid development of deep learning, the emergence of end-to-end speech synthesis has simplified the synthesis system while reducing manual intervention and the requirements for linguistics-related background knowledge. With the strong expressive ability of deep learning models, end-to-end speech synthesis systems can generate speech that sounds almost as natural as human speech. However, due to the lack of prosodic information such as pitch, stress and rhythm, the synthetic speech results of basic end-to-end speech synthesis systems for long texts (such as audiobooks or spoken dialogues) lack expressiveness. Therefore, recently, researchers have conducted a lot of research on how to generate more prosodic and emotional speech.
[0003] Some work has used style markers or variational autoencoders (VAEs) to capture prosodic features, achieving fine-grained voice modeling and voice control by extracting phoneme or word-level voice features. However, the speech synthesis system based on variational autoencoders samples from a standard Gaussian prior during the inference process, resulting in unnatural prosodic variations and a lack of effective control over prosodic variations. In addition, researchers have committed to adding cross-sentence information to the input features, applying pre-trained language models, such as the Bidirectional Encoder Representations from Transformers (BERT), to speech synthesis systems, and estimating prosodic features based on text representations pre-trained from discourse or segments. However, existing work only simply utilizes cross-sentence information, and the effect of improving the prosody of synthetic speech is limited.
[0004] With the development of deep learning, non-autoregressive speech synthesis systems have made progress in both efficiency and fidelity. Non-autoregressive speech synthesis systems map the input text sequence to an acoustic feature or waveform sequence without using an autoregressive decomposition of output probabilities. Some non-autoregressive speech synthesis systems, such as FastSpeech and ParaNet, need to be refined from autoregressive models. The latest non-autoregressive speech synthesis systems, such as FastPitch, AlignTTS and FastSpeech2, do not rely on any form of knowledge distillation from pre-trained TTS systems.
[0005] Based on the commonly used non-autoregressive end-to-end TTS system FastSpeech2, FastSpeech2 uses pitch contours and signal amplitudes as labels for supervision during training, and can predict prosodic information including pitch and energy from the encoder output. However, FastSpeech2 does not model cross-sentence information and only extracts pitch and energy information from real speech, failing to fully utilize the rich implicit features in prosody. Therefore, the synthesized speech lacks sufficient expressiveness and prosodic diversity.
[0006] Since prosodic information can be inferred from the linguistic information of the current sentence and surrounding sentences, and this information is usually contained in the vector representations from pre-trained language models (such as the bidirectional encoder BERT), some existing studies have incorporated word- or sub-word-level bidirectional encoder BERT vector representations into autoregressive speech synthesis models. Recent studies have used chunked and pairwise sentence patterns of the bidirectional encoder BERT. There are also some studies that combine the bidirectional encoder BERT with other techniques, including combining the bidirectional encoder BERT with a multi-task learning technique to eliminate the conversion of polyphonic characters to phonemes in Mandarin, and using the bidirectional encoder BERT vector representation as the node input of a relational graph network to extract word-level semantic representations, thereby improving the expressive ability. CU-Tacotron2 uses a pre-trained BERT model to extract sentence embedding vectors of adjacent sentences to improve the prosody generation of each utterance in a paragraph in an end-to-end manner. This method can improve the naturalness and expressiveness of the synthesized speech, but the prosody performance of the synthesized speech is poor and cannot synthesize audio with sufficient expressiveness and prosodic diversity. Summary of the Invention
[0007] In view of the above-mentioned shortcomings of the prior art, the purpose of this application is to provide a cross-sentence speech synthesis method, system and device based on a variational autoencoder, which is used to solve the technical problems in the prior art that the prosody effect of the synthesized speech is limited, the expressiveness is insufficient, and there is an inconsistency between the standard Gaussian prior sampled by the system during inference and the true prior of the speech.
[0008] To achieve the above purpose and other related purposes, this application provides a cross-sentence speech synthesis method based on a variational autoencoder, and the method includes: encoding based on the phoneme sequence and speaker information to obtain the mixed encoding F of the current sentence i ; encoding the context information through cross-sentence representations and multi-head attention layers to obtain the cross-sentence vector representation G i ; connecting the mixed encoding F i and the cross-sentence vector representation G i and outputting the cross-sentence vector representation H containing each phoneme through a linear mapping layer i ; the cross-sentence vector representation H containing each phonemei and obtaining the predicted duration D of each phoneme i After connection, input it into the conditional prior module to obtain the conditional prior z of a specific statement p ; Use the reference Mel spectrogram x i as the input of the conditional posterior module, and combine the conditional prior z p to establish an approximate conditional posterior z i , and add the conditional posterior z i to the cross-sentence vector representation H containing each phoneme i ; According to the predicted duration D of each phoneme i Expand the length of the input reference Mel spectrogram; Convert the context information into a Mel spectrogram sequence through parallel computing.
[0009] In an embodiment of the present application, encoding based on the phoneme sequence and speaker information to obtain the mixed encoding F of the current statement i , including: converting the current statement u i into a phoneme sequence P i =[p1, p2,..., p T , and encoding the phoneme sequence through a transformer to obtain a phoneme encoding; encoding the speaker information into the speaker representation s i ; Adding the speaker representation s i to the phoneme encoding to obtain the mixed encoding F i : F i =[f i (p1), f i (p2),..., f i (p T )]; where T represents the number of phonemes, and f represents the result vector of adding each phoneme encoding and the speaker representation s i .
[0010] In an embodiment of the present application, encoding the context information through cross-sentence representation and multi-head attention layers to obtain the cross-sentence vector representation G i , including: 1) Divide 2L + 1 adjacent statements [u i-L ,..., u i ,..., u i+L into 2L cross-sentence pairs, denoted as C i : C i =[c(u i-L , u i-L+1 ),..., c(u i-1 , u i ),..., c(u i+L-1 , u i+L )]; where c(u k , uk+1 ) = {[CLS], u k , [SBP], u k+1}, where there is a special token [CLS] at the beginning of each sentence pair, and there is another special token [SEP] between the two sentences of each sentence pair to represent the original sentence structure; 2) Send the 2L cross-sentence pairs into the bidirectional encoder BERT to obtain 2L BERT vector representations B i : B i = [b -L , b -L+1 , …, b L-1 ; where the vector b k represents the BERT vector representation of the cross-sentence pair c(u k , u k+1 ); 3) Combine the 2L BERT vector representations B i and the hybrid encoding F i into a cross-sentence vector representation G i to extract the cross-sentence representation of each phoneme: G i = MHA(F i W Q , B i W K , B i W V ); where MHA(·) represents the multi-head attention layer; WQ, W K , W V represent linear projection matrices; F i represents the hybrid encoding sequence of the current sentence and serves as the query matrix in the attention mechanism; and the expression of the cross-sentence vector representation G i is denoted as: G i = [g1, g2, …, g T ; where T represents the length of the multi-head attention layer.
[0011] In an embodiment of the present application, the expression of the cross-sentence vector representation H i containing each phoneme is: H i = [h1, h2, …, h T ; where h t = [g t , f(p t )]W; W represents a linear projection matrix; g t represents the cross-sentence vector representation of the t-th phoneme.
[0012] In an embodiment of the present application, the expression of the conditional prior zp after reparameterization is: where, μ p, σ p represents the approximate prior distribution learned from the conditional prior module ∈ follows a standard Gaussian represent element-wise addition and multiplication operations respectively.
[0013] In an embodiment of the present application, the conditional posterior z i The reparameterized expression is:[[]] where μ and σ represent the approximate posterior distribution estimated by the conditional posterior module z p is sampled from the learned specific statement conditional prior.
[0014] In an embodiment of the present application, the likelihood calculation expression of the reference Mel spectrogram xi is: p θ (x i |H i , D i ) = ∫p θ (x i |z i , H i , D i )p φ (z i |H i , D i )dz; where θ and φ represent the module parameters of the decoder and encoder respectively.
[0015] In an embodiment of the present application, the Mel spectrogram sequence is optimized by minimizing the ELBO loss:
[0016] where φ1 and φ2 represent two parts of the module parameter φ of the encoder; the conditional prior z p is obtained from D i and H i , the conditional posterior z i is obtained from x i and z p , β1 and β2 represent two balance constants, T represents the number of phonemes, follows a standard Gaussian and correspond to the latent representation of the nth phoneme.
[0017] To achieve the above and other related purposes, the present application provides a cross-sentence speech synthesis system based on a variational autoencoder, including: a cross-sentence representation module for encoding based on a phoneme sequence and speaker information to obtain a mixed encoding F of the current sentence i; Encoding context information through cross-sentence representations and multi-head attention layers to obtain cross-sentence vector representations G i ; The mixed encoding F i and the cross-sentence vector representations G i are concatenated and then output through a linear mapping layer to obtain cross-sentence vector representations H containing each phoneme i ; The prosody enhancement module further includes: an encoder for inputting the cross-sentence vector representations H containing each phoneme i and the predicted duration D obtained for each phoneme i after concatenation into a conditional prior module to obtain a conditional prior z for a specific sentence p ; Using the reference Mel spectrogram x i as the input of the conditional posterior module, and combining the conditional prior z p to establish an approximate conditional posterior z i ; A decoder for adding the conditional posterior z i to the cross-sentence vector representations H containing each phoneme i ; According to the predicted duration D of each phoneme i extending the length of the input reference Mel spectrogram; converting context information into a Mel spectrogram sequence through parallel computing
[0018] To achieve the above and other related purposes, the present application provides a computer device, including: a memory and a processor; the memory is used to store computer programs; the processor is used to execute the computer programs stored in the memory so that the device executes the method described above
[0019] In summary, a cross-sentence speech synthesis method, system, and device based on a variational autoencoder provided by the present application have the following beneficial effects: By using a multi-head attention layer to generate phoneme-level cross-sentence representations and using cross-sentence information as the prior condition of a conditional variational autoencoder, the prosody variation of the synthesized speech is improved, and at the same time, the problem of inconsistency between the standard Gaussian prior sampled by the system during inference and the true prior of the speech is solved. The present application organically combines cross-sentence information with a variational autoencoder for enhancing prosody, not only improving the naturalness of the synthesized speech, but also further increasing prosody diversity BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 Shown is a schematic flowchart of a cross-sentence speech synthesis method based on a variational autoencoder in an embodiment of the present application
[0021] Figure 2 Shown is a scenario application diagram of a cross-sentence speech synthesis method based on a variational autoencoder in an embodiment of the present application
[0022] Figure 3 It shows a schematic diagram of the modules of a cross - sentence speech synthesis system based on a variational auto - encoder in an embodiment of the present application.
[0023] Figure 4 It shows a schematic diagram of the structure of a computer device in an embodiment of the present application. Detailed implementation manners
[0024] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0025] It should be noted that in the following description, reference is made to the accompanying drawings, which describe several embodiments of the present application. It should be understood that other embodiments can also be used, and mechanical composition, structure, electrical, and operational changes can be made without departing from the spirit and scope of the present application. The following detailed description should not be considered restrictive, and the scope of the embodiments of the present application is only defined by the claims of the published patent. The terms used here are only for describing specific embodiments and are not intended to limit the present application. Spatially - related terms, such as "upper", "lower", "left", "right", "below", "beneath", "lower part", "above", "upper part", etc., can be used in the text to facilitate the description of the relationship between one element or feature shown in the figure and another element or feature.
[0026] Throughout the specification, unless otherwise clearly defined and limited, terms such as "install", "connect", "join", "fix", "hold" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.
[0027] Furthermore, as used herein, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. The terms "first", "second", "third", "fourth", etc. (if any) in the description and claims of the present invention and in the above drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that the data used in this way may be interchanged where appropriate so that the embodiments described herein can be practiced in an order different from that illustrated or described herein. It should be further understood that the terms "comprising" and "including" indicate the presence of the stated features, operations, elements, components, items, kinds, and / or groups, but do not preclude the presence, occurrence, or addition of one or more other features, operations, elements, components, items, kinds, and / or groups. The terms "or" and "and / or" as used herein are to be construed as inclusive or meaning any one or any combination. Thus, "A, B, or C" or "A, B, and / or C" means "any of the following: A; B; C; A and B; A and C; B and C; A, B, and C". The only exceptions to this definition occur when the combination of elements, functions, or operations is inherently mutually exclusive in some way.
[0028] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the technical solutions in the embodiments of the present invention will be further described in detail below through the following embodiments in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only for explaining the present invention and are not used to limit the invention.
[0029] As Figure 1 shown, it is a schematic flow chart of a cross-sentence speech synthesis method based on a variational autoencoder in an embodiment of the present application. The method includes the following steps:
[0030] Step S1: Encode based on the phoneme sequence and speaker information to obtain the mixed encoding Fi of the current sentence.
[0031] In an embodiment of the present application, the encoding based on the phoneme sequence and speaker information to obtain the mixed encoding F i of the current sentence specifically includes:
[0032] Step S101: Convert the current sentence u i into a phoneme sequence P i = [p1, p2,..., p T , and encode the phoneme sequence through a transducer to obtain a phoneme encoding.
[0033] It should be noted that a sentence generally refers to a connected group of words. A phoneme is the smallest unit of speech, and one pronunciation action constitutes one phoneme. Convert the current sentence u iBy combining the recurrent neural network (RNN) and long short-term memory (LSTM) using grapheme-to-phoneme (G2P), the conversion from grapheme to phoneme can be achieved, thereby converting the current statement u i The corresponding text information into a phoneme sequence P i = [p1, p2,..., p T .
[0034] In some examples, by using a Transformer model to encode the phoneme sequence P i = [p1, p2,..., p T to obtain a phoneme encoding to solve the Seq2Seq (sequence-to-sequence) problem.
[0035] Step S102: Encode the speaker information into the speaker representation s i .
[0036] It should be noted that the speaker information includes the speaker's audio, speech, voiceprint, prosody, speech habits, pronunciation, intonation, speech rate, volume, dialect, etc.; the speaker representation s i is used to represent the features of the speaker information.
[0037] Step S103: Add the speaker representation s i to the phoneme encoding to obtain a mixed encoding F i , and the expression of F i is:
[0038] F i = [f i (p1), f i (p2),..., f i (p T )]; (1)
[0039] where T represents the number of phonemes, and f represents the result vector obtained by adding each phoneme encoding and the speaker representation s i .
[0040] Step S2: Encode the context information through cross-sentence representation and multi-head attention layer to obtain a cross-sentence vector representation G i .
[0041] It should be noted that the content used as the text input includes: speaker information, context information; the context information includes: the current statement u i , the first L statements [u i ,..., u i-L and the last L statements [u i-1 surrounding the current statement u i+1 ,..., ui+L .
[0042] In an embodiment of the present application, the context information is encoded through cross-sentence representation and a multi-head attention layer to obtain a cross-sentence vector representation G i , specifically including:
[0043] 1) Divide 2L + 1 adjacent sentences [u i-L , …, u i , …, u i+L into 2L cross-sentence pairs, denoted as C i :
[0044] C i = [c(u i-L , u i-L+1 ), …, c(u i-1 , u i ), …, c(u i+L-1 , u i+L )]; (2)
[0045] Among them, c(u k , u k+1 ) = {[CLS], u k , [SEP], u k+1}, there is a special marker [CLS] at the beginning of each sentence pair, and there is another special marker [SEP] between the two sentences of each sentence pair to represent the original sentence structure;
[0046] 2) Send the 2L cross-sentence pairs into the bidirectional encoder BERT respectively to obtain 2L BERT vector representations B i :
[0047] B i = [b -L , b -L+1 , …, b L-1 ; (3)
[0048] Among them, the vector b k represents the BERT vector representation of the cross-sentence pair c(u k , u k+1 );
[0049] It should be noted that the essence of the bidirectional encoder BERT is to be used for feature extraction. By mapping the original data to a high-dimensional space, when doing downstream tasks, the output of BERT can be used as the input of the downstream task. For example, in this application, for each cross-sentence pair, the output vector at the [CLS] marker position is projected into a 768-dimensional space to obtain 2L BERT vector representations B i .
[0050] 3) Combine the 2L BERT vector representations B i and the hybrid encoding F i into a cross-sentence vector representation G i to extract the cross-sentence representation of each phoneme:
[0051] G i = MHA(F i W Q , B i W K , B i W V ); (4)
[0052] where MHA(·) represents the multi-head attention layer; W Q , W K , W V represent linear projection matrices; F i represents the hybrid encoding sequence of the current sentence and serves as the query matrix in the attention mechanism;
[0053] In addition, it can be calculated from formula (4) that the expression of the cross-sentence vector representation G i can also be written as:
[0054] G i = [g1, g2,..., g T ; (5)
[0055] where T represents the length of the multi-head attention layer.
[0056] Step S3: Concatenate the hybrid encoding F i and the cross-sentence vector representation G i and output a cross-sentence vector representation H i containing each phoneme through a linear mapping layer.
[0057] In an embodiment of the present application, the expression of the cross-sentence vector representation H i containing each phoneme is:
[0058] H i = [h1, h2,..., h T ; (6)
[0059] where h t = [g t , f(p t )]W; W represents a linear projection matrix; g t represents the cross-sentence vector representation of the t-th phoneme.
[0060] Specifically, the specific implementation principle of steps S1 to S3 can be further understood in combination with Figure 2 the cross-sentence encoding shown.
[0061] Step S4: Input the cross-sentence vector representation H i containing each phoneme and the predicted duration D i obtained for each phoneme after connection into the conditional prior module to obtain the conditional prior z p of a specific sentence.
[0062] It should be noted that a duration predictor is added. By using the cross-sentence vector representation H i containing each phoneme as the input of the duration predictor, the predicted duration D i for each phoneme is obtained as the output.
[0063] In an embodiment of the present application, the conditional prior z p has the following reparameterized expression:
[0064]
[0065] where μ p and σ p represent the approximate prior distribution learned from the conditional prior module ∈ follows the standard Gaussian represent element-wise addition and element-wise multiplication operations respectively.
[0066] Step S5: Use the reference Mel spectrogram x i as the input of the conditional posterior module, and combine it with the conditional prior z p to establish an approximate conditional posterior z i , and add the conditional posterior z i to the cross-sentence vector representation H i containing each phoneme.
[0067] In an embodiment of the present application, the conditional posterior z i has the following reparameterized expression:
[0068]
[0069] where μ and σ represent the approximate posterior distribution estimated by the conditional posterior module z p is sampled from the learned conditional prior of the specific sentence.
[0070] Substituting formula (7) into formula (8) gives:
[0071]
[0072] It should be noted that the conditional posterior z is projected into a high-dimensional space by adding an additional projection layer so as to add the conditional posterior z i to the cross-sentence vector representation H containing each phoneme i i .
[0073] Step S6: Extend the length of the input reference Mel spectrogram x according to the predicted duration D of each phoneme; convert the context information into a Mel spectrogram sequence through parallel computing i i
[0074] It should be noted that the length of the input reference Mel spectrogram x is extended according to the predicted duration D of each phoneme by using a length regulator i . Since the length of the phoneme sequence is usually shorter than that of its Mel spectrogram sequence, that is, each phoneme corresponds to several Mel spectrogram sequences; and the length of the Mel spectrogram sequence aligned with each phoneme is called the phoneme duration. The length regulator tiles the phoneme sequence according to the duration of each phoneme to match the length of the Mel spectrogram sequence. It can not only extend or shorten the duration of the phoneme proportionally to control the sound speed; but also control the pause between words by adjusting the duration of the blank characters in the sentence, thereby adjusting the partial prosody of the sound i
[0075] Specifically, variational autoencoders have been widely used in speech synthesis systems to achieve explicit modeling of prosodic variations. The variational autoencoder projects the input features into a low-dimensional latent space for reconstruction, thereby capturing the data variations in the low-dimensional latent space. The training objective of the variational autoencoder is to maximize the data distribution p θ (x) parameterized by θ, which can be regarded as the marginalization of the latent variable z, as shown in Equation (10):
[0076] p θ (x) = ∫p θ (x|z)p(z)dz; (10)
[0077] For ease of calculation, the evidence lower bound ELBO is used to approximate the marginalization
[0078]
[0079] where q φ (z|x) is the posterior distribution of the latent vector parameterized by φ, β is a hyperparameter, D KL is the Kullback-Leibler divergence. The first term on the right side of Equation (11) measures the expected reconstruction performance of the latent vector and is approximated by Monte Carlo sampling of z according to the posterior distribution. The reparameterization trick is used to make the sampling differentiable; the second term encourages the posterior distribution to be close to the prior distribution sampled during the inference process, and β measures the contribution of this term.
[0080] A large amount of work on speech synthesis uses variational autoencoders to capture and decouple the data variations in various aspects of the latent space, including separating speaker and phoneme information, simulating the speaking style of the speaker, and combining adversarial training to separate prosodic variations and speaker information. Recently, some researchers have adopted fine-grained variational autoencoders to model the prosody in the latent space of each phoneme or word, or applied quantized variational autoencoders to discrete duration modeling.
[0081] The conditional variational autoencoder is a variant of the variational autoencoder. Both the prior distribution and the posterior distribution are conditioned on the additional variable y. The modification of the likelihood calculation of generating data is shown in Equation (12):
[0082] p θ (x|y) = ∫p θ (x|z, y)p φ (z|y)dz; (12)
[0083] Similar to the variational autoencoder, the calculation can be converted into the ELBO form, as shown in Equation (13):
[0084]
[0085] To establish the conditional prior model, a density network is usually used to predict the mean and variance based on the conditional input y.
[0086] In an embodiment of the present application, the likelihood calculation expression of the reference Mel spectrogram x i is:
[0087] p θ (x i |H i , D i ) = ∫p θ (x i |z i , H i , D i )p φ (z i |H i , D i )dz; (14)
[0088] Among them, θ and φ respectively represent the module parameters of the decoder and the encoder.
[0089] In one embodiment of the present application, the Mel spectrum sequence is optimized by minimizing the following ELBO loss:
[0090]
[0091] where φ1 and φ2 respectively represent two parts of the module parameter φ of the encoder; the conditional prior z p is obtained from D i and H i and the conditional posterior z i is obtained from x i and z p β1 and β2 represent two balance constants, T represents the number of phonemes, obeys the standard Gaussian and corresponds to the latent representation of the nth phoneme.
[0092] Specifically, the specific implementation principle of steps S4 to S6 can be further understood in combination with Figure 2 the encoder and decoder for prosody enhancement shown in
[0093] As Figure 3 shown, it shows a module schematic diagram of a cross-sentence speech synthesis system based on a variational autoencoder in an embodiment of the present application. The cross-sentence speech synthesis system 300 based on a variational autoencoder includes:
[0094] A cross-sentence representation module 310, configured to encode based on a phoneme sequence and speaker information to obtain a mixed encoding F i of the current sentence; encode context information through cross-sentence representation and a multi-head attention layer to obtain a cross-sentence vector representation G i ; connect the mixed encoding F i and the cross-sentence vector representation G i and output a cross-sentence vector representation H containing each phoneme through a linear mapping layer i ;
[0095] A prosody enhancement module 320, whose conditional variational autoencoder further includes:
[0096] An encoder 321, configured to connect the cross-sentence vector representation H containing each phoneme i and the predicted duration D obtained for each phoneme i and input them into a conditional prior module to obtain a conditional prior z of a specific sentence p ; use the reference Mel spectrum x i as the input of the conditional posterior module, and combine the conditional prior z pEstablish the approximate conditional posterior z i ;
[0097] A decoder 322 for adding the conditional posterior z i to the cross-sentence vector representation H containing each phoneme i ; Extend the length of the input reference Mel spectrogram according to the predicted duration D of each phoneme i ; Convert the context information into a sequence of Mel spectrograms through parallel computing.
[0098] In an embodiment of the present application, the system 300 can be evaluated through qualitative hearing tests and quantitative measurements. Specifically, the naturalness and intelligibility of speech are measured by using subjective human opinion scores and word error rates. At the same time, the prosodic diversity of the speech generated from the conditional prior can also be evaluated by comparing the standard deviations of the relative fundamental frequency and energy.
[0099] For example, 11 synthetic audio samples are selected for subjective hearing tests. 23 volunteers are recruited to evaluate the naturalness of the speech samples with a subjective opinion score of up to 5 points, and the evaluation results are reported with a 95% confidence interval. In addition to the subjective human opinion scores, the word error rate and the standard deviations of the fundamental frequency and energy are evaluated on 512 test samples. The experimental results on the single-speaker dataset and the multi-speaker dataset are shown in Table 1. Each index of the system 200 proposed in the present application has obvious advantages compared with the baseline, effectively improving the naturalness and prosodic diversity of the generated audio samples.
[0100] Table 1 Qualitative and quantitative test results of samples in single-speaker and multi-speaker datasets
[0101]
[0102] Among them, three indexes are tested for each dataset, namely subjective human opinion scores, word error rates, and prosodic diversity. The prosodic diversity includes the standard deviations of the relative fundamental frequency and energy of phonemes in hertz. "↑" indicates that the higher the value, the better the performance, and "↓" indicates that the lower the value, the better the performance. For example, higher values of subjective human opinion scores and prosodic diversity mean better performance, while lower values of word error rates indicate better performance.
[0103] It should be noted that the cross-sentence representation module 310 takes the speaker information, the BERT vector representations of the current sentence and surrounding sentences as inputs, and uses a multi-head attention layer to generate phoneme-level cross-sentence representations, where the weights of the attention layer come from the encoder outputs of each phoneme and the speaker information. The cross-sentence representation module 310 can generate more natural and expressive audio. At the same time, using the multi-head attention layer can specifically extract the cross-sentence representations of each phoneme.
[0104] It should be noted that the prosody enhancement module 320 is a fine-grained variational autoencoder, which can estimate the posterior of the prosody features of each phoneme based on acoustic features, utterance representations, and speaker information. The prosody enhancement module 320 can solve the problems that the existing FastSpeech2 lacks prosody variations and the standard Gaussian prior distribution sampled by the variational autoencoder-based speech synthesis system is inconsistent with the true prior distribution of speech.
[0105] It should be understood that the division of each module of the above system is only a logical function division. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by a processing element; they can also all be implemented in the form of hardware; or some modules can be implemented in the form of software called by a processing element, and some modules can be implemented in the form of hardware. For example, the decoder 322 can be a separately established processing element, or can be integrated in a certain chip of the above system. In addition, it can also be stored in the memory of the above system in the form of program code, and the function of the above decoder 322 can be called and executed by a certain processing element of the above system. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together or can be independently implemented. The processing element mentioned here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed by the integrated logic circuit in the processor element or the instruction in the form of software.
[0106] For example, the above modules can be one or more integrated circuits configured to implement the above method, such as: one or more application specific integrated circuits (ASICs), or, one or more digital signal processors (DSPs), or, one or more field programmable gate arrays (FPGAs), etc. Again, when a certain module above is implemented in the form of a processing element scheduling program code, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processors that can call program code. Again, these modules can be integrated together and implemented in the form of a system-on-a-chip (SOC).
[0107] Such as Figure 4As shown, it is a schematic structural diagram of a computer device 400 in an embodiment of the present application. The computer device 400 includes: a memory 410 and a processor 420; the memory 410 is used to store computer instructions; the processor 420 runs the computer instructions to implement as Figure 1 the method described.
[0108] In some embodiments, the number of the memory 410 and the processor 420 in the computer device 400 can both be one or more, and Figure 4 one of each is taken as an example herein.
[0109] In an embodiment of the present application, the processor 420 in the computer device 400 will, according to the steps as Figure 1 described, load instructions corresponding to one or more application program processes into the memory 410, and the processor 420 runs the application programs stored in the memory 410, so as to implement as Figure 1 the method described.
[0110] The memory 410 may include a random access memory (Random Access Memory, abbreviated as RAM), and may also include a non-volatile memory, such as at least one disk memory. The memory 410 stores an operating system and operation instructions, executable modules or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions may include various operation instructions for implementing various operations. The operating system may include various system programs for implementing various basic services and processing hardware-based tasks.
[0111] The processor 420 may be a general-purpose processor, including a central processing unit (Central Processing Unit, abbreviated as CPU), a network processor (Network Processor, abbreviated as NP), etc.; it may also be a digital signal processor (Digital Signal Processor, abbreviated as DSP), an application specific integrated circuit (Application Specific Integrated Circuit, abbreviated as ASIC), a field programmable gate array (Field Programmable Gate Array, abbreviated as FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0112] In some specific applications, the various components of the computer device 400 are coupled together through a bus system, and the bus system may include, in addition to a data bus, a power bus, a control bus, a status signal bus, etc. However, for the sake of clarity, inFigure 4 All kinds of buses are collectively referred to as a bus system.
[0113] In summary, the present application provides a cross-sentence speech synthesis method, system, and device based on a variational autoencoder. By organically combining cross-sentence information with a variational autoencoder for enhancing prosody, a cross-sentence speech synthesis system based on a variational autoencoder is proposed. The posterior probability distribution of the potential prosody features of each phoneme is estimated by conditioning on acoustic features, speaker information, and text features obtained from the current and surrounding sentences. The system includes a cross-sentence representation module and a prosody enhancement module. The cross-sentence representation at the phoneme level is generated using a multi-head attention layer, and the output of the cross-sentence representation module is used as the prior condition for a specific sentence in the prosody enhancement module to improve the standard variational autoencoder. The present application not only improves the naturalness of the synthesized speech and the prosodic variations of the synthesized speech but also solves the problem of inconsistency between the standard Gaussian prior sampled by the system during inference and the true prior of the speech.
[0114] The present application effectively overcomes various drawbacks in the prior art and has high industrial utilization value.
[0115] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the relevant technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.
Claims
1. A cross-sentence speech synthesis method based on a variational autoencoder, characterized in that The method includes: Encode based on the phoneme sequence and speaker information to obtain the mixed encoding F of the current statement i ; Encode context information through cross-sentence representations and multi-head attention layers to obtain the cross-sentence vector representation G i ; Connect the hybrid-coded F i and the cross-sentence vector representation G i After connection, output the cross-sentence vector representation H containing each phoneme through a linear mapping layer i ; The cross-sentence vector representation H containing each phoneme i and the predicted duration D obtained for each phoneme i are concatenated and input into the conditional prior module to obtain the conditional prior z for a specific sentence p ; The mel-spectrum x will be referred to i as the input of the conditional posterior module, and combined with the conditional prior z p to establish an approximate conditional posterior z i , and the conditional posterior z i is added to the cross-sentence vector representation H containing each phoneme i ; According to the predicted duration D of each phoneme i Expand the reference Mel spectrogram x of the input i in length; Through parallel computing, the conditional posterior z i of the cross-sentence vector representation H i is converted into a Mel spectrogram sequence; Encoding based on the phoneme sequence and speaker information to obtain the hybrid encoding F of the current utterance i , including: converting the current utterance u i into a phoneme sequence P i = [p1, p2, …, p T , and encoding the phoneme sequence through a transducer to obtain a phoneme encoding; encoding the speaker information into the speaker representation s i ; adding the speaker representation s i to the phoneme encoding to obtain the hybrid encoding F i : F i = [f i (p1), f i (p2), …, f i (p T )]; where T represents the number of phonemes, and f represents the result vector obtained by adding each phoneme encoding and the speaker representation s i . Encoding context information through cross-sentence representation and multi-head attention layers to obtain the cross-sentence vector representation G i , including: 1) Divide 2L + 1 adjacent sentences [u i-L , …, u i , …, u i+L into 2L cross-sentence pairs, denoted as C i : C i = [c(u i-L , u i-L+1 ), …, c(u i-1 , u i ), …, c(u i+L-1 , u i+L )]; where c(u k , u k+1 ) = {[CLS], u k , [SEP], u k+1}, there is a special token [CLS] at the beginning of each sentence pair, and there is another special token [SEP] between the two sentences of each sentence pair to represent the original sentence structure; 2) Send the 2L cross-sentence pairs into the bidirectional encoder BERT respectively to obtain 2L BERT vector representations B i : B i = [b -L , b -L+1 , …, b L-1 ; where the vector b k represents the BERT vector representation of the cross-sentence pair c(u k , u k+1 ); 3) Combine the 2L BERT vector representations B i and the hybrid encoding F i into a cross-sentence vector representation G i through the multi-head attention layer to extract the cross-sentence representation of each phoneme: G i = MHA(F i W Q , B i W K , B i W V ); where MHA(·) represents the multi-head attention layer; W Q , W K , W V represent linear projection matrices; F i represents the hybrid encoding sequence of the current sentence and serves as the query matrix in the attention mechanism; and the expression of the cross-sentence vector representation G i is denoted as: G i = [g1, g2, …, g T ; where T represents the length of the multi-head attention layer.
2. A cross-sentence speech synthesis method based on variational autoencoder according to claim 1, characterized in that The cross-sentence vector representation H containing each phoneme i has the following expression: H i = [h1, h2, …, h T ; where h t = [g t , f(p t )]W; W represents a linear projection matrix; g t represents the cross-sentence vector representation of the t-th phoneme.
3. A cross-sentence speech synthesis method based on variational autoencoder according to claim 1, characterized in that The conditional prior z p The expression after reparameterization is: where, μ p , σ p represent the approximate prior distribution learned from the conditional prior module ∈ follows the standard Gaussian represent the element-wise addition and element-wise multiplication operations, respectively.
4. A cross-sentence speech synthesis method based on variational autoencoder according to claim 3, characterized in that The conditional posterior z i The expression after reparameterization is: where μ and σ represent the approximate posterior distribution estimated by the conditional posterior module z p is sampled from the learned specific statement-conditioned prior 5. The cross-sentence speech synthesis method based on variational autoencoder according to claim 4, characterized in that The likelihood calculation expression of the reference Mel spectrum x i is as follows: p θ (x i ∣H i ,D i ) = ∫ p θ (x i ∣z i ,H i ,D i ) p φ (z i ∣H i ,D i ) dz; Wherein, θ and φ respectively represent the module parameters of the decoder and the encoder.
6. A cross-sentence speech synthesis method based on a variational autoencoder according to claim 5, characterized in that The Mel spectrum sequence is optimized by minimizing the ELBO loss: Among them, φ1 and φ2 respectively represent two parts of the module parameter φ of the encoder; the conditional prior z p is obtained from D i and H i ; the conditional posterior z i is obtained from x i and z p ; β1 and β2 represent two balance constants, T represents the number of phonemes, obeys the standard Gaussian and corresponds to the latent representation of the nth phoneme.
7. A cross-sentence speech synthesis system based on a variational autoencoder, characterized in that, Including: Cross-sentence representation module, which is used to encode based on the phoneme sequence and speaker information to obtain the mixed encoding F of the current sentence i ; Encoding context information through cross-sentence representations and multi-head attention layers to obtain the cross-sentence vector representation G i ; Connect the hybrid-coded F i and the cross-sentence vector representation G i After connection, output, through a linear mapping layer, the cross-sentence vector representation H containing each phoneme i ; The prosody enhancement module further includes: An encoder for generating cross-sentence vector representations H containing each phoneme i and obtaining the predicted duration D of each phoneme i After concatenation, input to the conditional prior module to obtain the conditional prior z of a specific sentence p ; Use the reference Mel spectrogram x i as the input of the conditional posterior module, and combine with the conditional prior z p to establish an approximate conditional posterior z i ; A decoder for adding the conditional posterior z i to the cross-sentence vector representation H containing each phoneme i ; extending the length of the input reference Mel spectrogram according to the predicted duration D of each phoneme i ; converting the cross-sentence vector representation H to which the conditional posterior z is added i to a Mel spectrogram sequence through parallel computation i ; Encoding based on the phoneme sequence and speaker information to obtain the hybrid encoding F of the current utterance i , including: converting the current utterance u i into a phoneme sequence P i = [p1, p2, …, p T , and encoding the phoneme sequence through a transducer to obtain a phoneme encoding; encoding the speaker information into the speaker representation s i ; adding the speaker representation s i to the phoneme encoding to obtain the hybrid encoding F i : F i = [f i (p1), f i (p2), …, f i (p T )]; where T represents the number of phonemes, and f represents the result vector obtained by adding each phoneme encoding and the speaker representation s i . Encoding the context information through cross-sentence representation and multi-head attention layers to obtain the cross-sentence vector representation G i , including: 1) Divide 2L + 1 adjacent sentences [u i-L , …, u i , …, u i+L into 2L cross-sentence pairs, denoted as C i : C i = [c(u i-L , u i-L+1 ), …, c(u i-1 , u i ), …, c(u i+L-1 , u i+L )]; where c(u k , u k+1 ) = {[CLS], u k , [SEP], u k+1}, there is a special token [CLS] at the beginning of each sentence pair, and there is another special token [SEP] between the two sentences of each sentence pair to represent the original sentence structure; 2) Send the 2L cross-sentence pairs into the bidirectional encoder BERT respectively to obtain 2L BERT vector representations B i : B i = [b -L , b -L+1 , …, b L-1 ; where the vector b k represents the BERT vector representation of the cross-sentence pair c(u k , u k+1 ); 3) Combine the 2L BERT vector representations B i and the hybrid encoding F i into a cross-sentence vector representation G i through the multi-head attention layer to extract the cross-sentence representation of each phoneme: G i = MHA(F i W Q , B i W K , B i W V ); where MHA(·) represents the multi-head attention layer; W Q , W K , W V represent linear projection matrices; F i represents the hybrid encoding sequence of the current sentence and serves as the query matrix in the attention mechanism; and the expression of the cross-sentence vector representation G i is denoted as: G i = [g1, g2, …, g T ; where T represents the length of the multi-head attention layer.
8. A computer device, characterized in that, The device includes: a memory and a processor; The memory is used to store a computer program; the processor is used to execute the computer program stored in the memory, so that the device executes the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Phrase-based end-to-end text-to-speech (TTS) synthesis
CN111681641A
Speech synthesis method, device and equipment and computer readable storage medium
CN113838448A