A prosody-controllable speech synthesis method and related device based on VITS

By extracting and controlling multiple prosodic features based on the VITS method, the problem of insufficient prosodic control in existing speech synthesis technology is solved, and the anthropomorphic and naturalness of multi-speaker speech synthesis is improved.

CN120412544BActive Publication Date: 2025-09-12XIANGJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510918273.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-09-12
Estimated Expiration
2045-07-03

AI Technical Summary

Technical Problem

Existing speech synthesis technology has difficulty in effectively controlling various prosodic features, resulting in a large gap between the generated speech and the human voice, making it difficult to achieve an anthropomorphic effect, especially in multi-speaker Chinese speech synthesis.

Method used

A VITS-based method is used to extract prosodic feature information from the input text. Semantic encoding is performed through the BERT pre-trained model, and predictors of duration, energy, pitch period, pause and rhythm are constructed to generate a fused prosodic control embedding. Combined with the speaker embedding vector, a decoder is used to generate a multi-speaker speech spectrum, which is finally converted into a time-domain speech signal to achieve independent control of prosodic features.

Benefits of technology

It improves the anthropomorphic effect of speech synthesis, enhances the naturalness and authenticity of multi-speaker speech synthesis, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120412544B_ABST
    Figure CN120412544B_ABST
Patent Text Reader

Abstract

The present invention provides a prosody-controllable speech synthesis method and related device based on VITS, which relates to the field of speech synthesis technology. The method involves extracting prosody feature information from input text; semantically encoding the input text to generate a text context feature representation; constructing a prosody controller and independently modeling the prosody features to generate prediction results; fusing the text context feature representation with the prediction results to generate a prosody control embedding; generating a multi-speaker speech spectrum based on the VITS model, combining the prosody control embedding with speaker embedding vectors; and converting the multi-speaker speech spectrum into a time-domain speech signal through a decoder. By adjusting prosody feature parameters, the duration, pitch period, energy, pauses, and rhythm of the synthesized speech are independently controlled to improve the subjective quality of the generated speech signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech synthesis technology, and in particular to a prosody-controllable speech synthesis method and related devices based on VITS. Background Art

[0002] Speech synthesis is a challenging task. In practical applications of audio content production, such as audiobooks, audio videos, and news broadcasts, adding sound to corresponding text and videos requires specialized recording equipment and personnel for recording and editing. Reducing costs and time consumption, making the generated speech more human-like, and achieving the rhythmic effect of human speech to enhance user experience are key challenges in speech synthesis. Traditional speech generation technologies struggle to achieve ideal results due to technical limitations. With the rapid development of deep learning in text-to-speech, the quality of generated speech has improved in recent years, but the results remain difficult to translate into practical applications. The performance of multi-speaker Chinese speech synthesis is also unsatisfactory, far behind the quality of human dubbing. Existing models focus on a single prosodic feature. This paper focuses on the refined control of multiple prosodic features in speech synthesis, focusing on how to fuse these extracted prosodic features. Drawing on the inner and outer layer functions of KAN, this paper integrates prosodic features including duration, pitch period, energy, pauses, and rhythm.

[0003] Therefore, how to optimize the speech generation effect has become a technical problem that needs to be solved urgently. Summary of the Invention

[0004] In order to optimize the speech generation effect, the present application provides a prosody-controllable speech synthesis method and related devices based on VITS.

[0005] In the first aspect, the present application provides a prosody-controllable speech synthesis method based on VITS, which adopts the following technical solutions:

[0006] A prosody-controllable speech synthesis method based on VITS, comprising:

[0007] S1: extracting prosodic feature information from an input text, wherein the prosodic feature information includes prosodic features such as duration, pitch period, energy, pause, and rhythm;

[0008] S2: Use the BERT pre-trained model to semantically encode the input text and generate text context feature representation;

[0009] S3: constructing a prosody controller including a duration predictor, an energy predictor, a pitch period predictor, a pause predictor, and a rhythm predictor, and independently modeling the prosody features and generating prediction results through the prosody controller;

[0010] S4: fusing the text context feature representation with the prediction result to generate a fused prosodic control embedding;

[0011] S5: Generate multi-speaker speech spectra based on the VITS model, combining the prosodic control embedding and speaker embedding vectors;

[0012] S6: converting the multi-speaker speech spectrum into a time-domain speech signal through a decoder;

[0013] S7: By adjusting the prosodic feature parameters, independent control of the duration, pitch period, energy, pauses, and rhythm of the synthesized speech is achieved.

[0014] Optionally, in step S1: the duration is extracted by using a forced alignment algorithm to extract the phoneme duration, and a duration predictor is trained using maximum likelihood estimation;

[0015] The pitch period is decomposed into multi-scale components by continuous wavelet transform and quantized as a control condition;

[0016] The energy is encoded by calculating the L2 norm of the Mel spectrum amplitude and quantizing it;

[0017] The pause is generated by analyzing the text punctuation and phrase structure to generate a pause duration sequence;

[0018] The rhythm captures semantic associations through the BERT pre-trained model and adjusts the pronunciation speed in combination with position embedding.

[0019] Optionally, in step S3:

[0020] The rhythm controller is composed of a one-dimensional convolutional network, a bidirectional long short-term memory network and a linear layer. The output embedding of each predictor is superimposed through the feature fusion module, and the Kolmogorov-Arnold representation theorem is applied to realize the unary basis function fusion of multivariate features. The specific formula is:

[0021] ;

[0022] in, is the weight, is the basis function.

[0023] Optionally, in step S5:

[0024] The posterior encoder of the VITS model takes the Mel spectrum as input to generate a latent variable z, and converts the latent variable distribution into a prior distribution through a flow model;

[0025] The speaker embedding vector is concatenated with the text features and the prosody control embedding and input into a decoder to generate a multi-speaker speech spectrum.

[0026] Optionally, in step S7:

[0027] The duration is controlled by adjusting the output multiple of the duration predictor to achieve speech speed changes;

[0028] The control of the fundamental period adjusts the pitch by modifying the amplitude of the wavelet component;

[0029] The control of pauses is done by specifying the pause duration through marker symbols in the input text;

[0030] The control of rhythm dynamically adjusts the pronunciation speed through the semantic weights output by BERT.

[0031] Optionally, a training step is also included:

[0032] Use the AISHELL-3 or BZNSYP Chinese speech dataset to extract the duration, pitch period, energy, and pause duration of real speech as training targets;

[0033] By minimizing the reconstruction error between the predicted value and the true value of each prosodic feature, the text encoder, prosodic controller and VITS model parameters are jointly optimized.

[0034] Optionally, the decoder adopts the HiFiGAN V1 structure and converts the speech spectrum into a time domain waveform through multi-level upsampling.

[0035] In a second aspect, the present application provides a prosody-controllable speech synthesis device based on VITS, comprising:

[0036] An extraction module is used to extract prosodic feature information from an input text, wherein the prosodic feature information includes prosodic features such as duration, pitch period, energy, pause, and rhythm;

[0037] The semantic encoding module is used to semantically encode the input text using the BERT pre-trained model to generate text context feature representation;

[0038] A predictor construction module is used to construct a prosody controller including a duration predictor, an energy predictor, a pitch period predictor, a pause predictor, and a rhythm predictor, and independently model the prosody features and generate prediction results through the prosody controller;

[0039] a fusion module for fusing the text context feature representation with the prediction result to generate a fused prosody control embedding;

[0040] A speech spectrum module, configured to generate a multi-speaker speech spectrum based on the VITS model and in combination with the prosodic control embedding and the speaker embedding vector;

[0041] a conversion module, configured to convert the multi-speaker speech spectrum into a time-domain speech signal through a decoder;

[0042] The control module is used to independently control the duration, pitch period, energy, pause and rhythm of the synthesized speech by adjusting the prosodic feature parameters.

[0043] In a third aspect, the present application provides a computer device, comprising: a memory and a processor, wherein the processor executes the method described above when running computer instructions stored in the memory.

[0044] In a fourth aspect, the present application provides a computer-readable storage medium comprising instructions, which, when executed on a computer, enable the computer to execute the method described above.

[0045] In summary, the present application extracts prosodic feature information from input text; semantically encodes the input text through the BERT pre-trained model to generate a text context feature representation; constructs a prosodic controller including a duration predictor, an energy predictor, a pitch period predictor, a pause predictor, and a rhythm predictor, and independently models the prosodic features through the prosodic controller and generates prediction results; fuses the text context feature representation with the prediction results to generate a fused prosodic control embedding; based on the VITS model, combines the prosodic control embedding and the speaker embedding vector to generate a multi-speaker speech spectrum; converts the multi-speaker speech spectrum into a time-domain speech signal through a decoder; and independently controls the duration, pitch period, energy, pause, and rhythm of the synthesized speech by adjusting the prosodic feature parameters. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 It is a schematic diagram of the computer device structure of the hardware operating environment involved in the embodiment of the present application;

[0047] Figure 2 This is a flowchart of the first embodiment of the prosody-controllable speech synthesis method based on VITS of the present application;

[0048] Figure 3 It is a schematic diagram of the pause design of the prosody-controllable speech synthesis method based on VITS of the present application;

[0049] Figure 4 This is a diagram showing the internal structure of the prosody controller of the VITS-based prosody-controllable speech synthesis method of the present application;

[0050] Figure 5 This is a structural block diagram of the first embodiment of the prosody-controllable speech synthesis device based on VITS of the present application. DETAILED DESCRIPTION

[0051] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is further described in detail below through the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0052] Reference Figure 1 , Figure 1 This is a schematic diagram of the computer device structure of the hardware operating environment involved in the embodiment of the present application.

[0053] like Figure 1 As shown, the computer device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display and an input unit, such as a keyboard. Optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a wireless fidelity (Wi-Fi) interface). The memory 1005 may be a high-speed random access memory (RAM) or a stable non-volatile memory (NVM), such as a disk storage device. The memory 1005 may also be a storage device independent of the processor 1001.

[0054] Those skilled in the art will understand that Figure 1 The structure shown in the figure does not constitute a limitation on the computer device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0055] like Figure 1 As shown, the memory 1005 as a storage medium may include an operating system, a network communication module, a user interface module, and a VITS-based prosody-controllable speech synthesis program.

[0056] exist Figure 1In the computer device shown, the network interface 1004 is mainly used for data communication with the network server; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and the memory 1005 in this application can be set in the computer device, and the computer device calls the VITS-based prosody-controllable speech synthesis program stored in the memory 1005 through the processor 1001, and executes the VITS-based prosody-controllable speech synthesis method provided in the embodiment of this application.

[0057] The present application embodiment provides a prosody controllable speech synthesis method based on VITS, referring to Figure 2 , Figure 2 This is a flowchart of the first embodiment of the prosody-controllable speech synthesis method based on VITS of this application.

[0058] In this embodiment, the prosody-controllable speech synthesis method based on VITS includes the following steps:

[0059] S1: extracting prosodic feature information from an input text, wherein the prosodic feature information includes prosodic features such as duration, pitch period, energy, pause, and rhythm.

[0060] It should be noted that in step S1: the duration is extracted by using a forced alignment algorithm to extract the phoneme duration, and a duration predictor is trained using maximum likelihood estimation;

[0061] The pitch period is decomposed into multi-scale components by continuous wavelet transform and quantized as a control condition;

[0062] The energy is encoded by calculating the L2 norm of the Mel spectrum amplitude and quantizing it;

[0063] The pause is generated by analyzing the text punctuation and phrase structure to generate a pause duration sequence;

[0064] The rhythm captures semantic associations through the BERT pre-trained model and adjusts the pronunciation speed in combination with position embedding.

[0065] In a specific implementation, the duration extraction includes:

[0066] The duration of each input phoneme in the text can be calculated by summing each row in the estimated alignment matrix. The row sum is to count the number of Mel frames occupied by each phoneme and map it to duration. Discrete phonemes are associated with continuous Mel spectra to match the length of text and speech. A duration predictor is designed, usually trained by maximum likelihood estimation. The correspondence between the latent variable z and the distribution is defined as alignment A, and the dimension of A is (Number of phonemes Mel-frame number), calculate the exact log-likelihood of the data:

[0067] ;

[0068] ;

[0069] First Term: Latent Space The logarithmic probability of, the second term: reversible transformation The Jacobian determinant of , ,but:

[0070] ;

[0071] .

[0072] in, is the number of frames in the mel-spectrogram. The above formula implements a reversible mapping from a simple Gaussian distribution to complex speech. Since text and speech are monotonically aligned, we assume that the alignment A is monotonically increasing: this emphasizes the consistency of speech-text temporal order and avoids phoneme reversal. It means that under the given text conditions and alignment, the j-th element of the latent variable z follows a Gaussian distribution with mean given by Prediction, variance is given by Prediction. Find the parameters that maximize the log-likelihood (acoustic model parameters, such as neural network weights) and alignment A, so that the logarithmic probability of speech data x given the text condition is maximized:

[0073] ;

[0074] Assume that the generation of speech frame x obeys a certain parameterized distribution, whose probability density is , the logarithmic transformation converts the continuous multiplication into summation, which is convenient for optimization. The traditional method is not differentiable. This formula realizes gradient back propagation through A.

[0075] The extraction of pitch period includes:

[0076] The continuous wavelet transform is used to decompose the continuous pitch contour and quantize it into 256 possible values ​​per frame on a logarithmic scale as a control condition, which is converted into a control embedding through a control condition encoder. Given a continuous gene contour function F0, it is converted into a gene periodic spectrum. , using the continuous wavelet transform CWT:

[0077] ;

[0078] in, For Mexican hat wavelets, for x The pitch value of the position, and t are the scale and position of the wavelet respectively. The Mexican hat wavelet has the second-order derivative characteristic and can detect instantaneous changes. and pan t Implement multi-resolution analysis, is the energy normalization factor. The traditional STFT cannot adapt to the non-stationarity of the fundamental frequency, and the CWT has adaptive resolution in both the time-frequency domain. The original gene profile F0 can be represented by wavelet Calculated by inverse continuous wavelet transform:

[0079] ;

[0080] Assuming that F0 is decomposed into 10 scales, F0 can be represented by 10 independent components:

[0081] ;

[0082] i =1,…,10 , =5ms, logarithmic scale division: ,coefficient Fitting the sensitivity of the human ear to fundamental frequency changes. Given 10 wavelet components , can be reconstructed by the following formula :

[0083] ;

[0084] like , then its contribution to reconstruction , separating the macro trend and micro fluctuation of the fundamental frequency, is more robust than directly using the original F0.

[0085] Energy extraction includes:

[0086] The energy of each short-time Fourier transform frame is calculated by calculating the L2 norm of the mel_spectrogram amplitude and quantizing it on a logarithmic scale, and encoded using a control conditional encoder. Let the short-time energy of the n-th frame speech signal be represented by En, then its calculation formula is:

[0087] ;

[0088] Among them, M is the frame length, is the sample point in the frame. Short-term energy reflects the sum of the squares of the speech amplitudes, and the square operation enhances the discrimination of high-energy segments.

[0089] Instance parameters:

[0090] Assume frame length M = 400 (25ms@16kHz), sample value =[0.1,-0.2,0.05,...]:

[0091] .

[0092] Stalled extractions include:

[0093] First, the input text is analyzed, and the entire sentence is segmented into phrases using the "," character. In practice, a sequence of elements containing pinyin and tones is used. The example diagram only uses text as an example to briefly convey the idea of ​​pauses. Second, the processed text string is converted into a sequence of IDs corresponding to the symbols in the text. Pause duration is clearly defined: long pauses and short pauses are divided, with different pause lengths corresponding to different markers.

[0094] Rhythm extraction includes:

[0095] Rhythm can be simply defined as the speed of pronunciation of each Chinese character. If the speed is consistent, it will appear smooth. Questioning tone is usually accompanied by changes such as longer speaking speed and higher pitch, so rhythm control is indispensable. The introduction of a BERT-based rhythm pre-training model improves rhythm and rhythm. Similar to the pause design above, assuming that the position is embedded, the pronunciation of the corresponding position can be manipulated to lengthen or shorten. Among them, the schematic diagram of the pause design is as follows Figure 3 As shown, the pause sequence means the corresponding pause duration.

[0096] S2: Use the BERT pre-trained model to semantically encode the input text and generate text context feature representation.

[0097] S3: Constructing a rhythm controller including a duration predictor, an energy predictor, a pitch period predictor, a pause predictor and a rhythm predictor, and independently modeling the rhythm features through the rhythm controller and generating prediction results.

[0098] In a specific implementation, in step S3: the rhythm controller is composed of a one-dimensional convolutional network, a bidirectional long short-term memory network and a linear layer, and the output embedding of each predictor is superimposed through a feature fusion module, and the Kolmogorov-Arnold representation theorem is applied to realize the unary basis function fusion of multivariate features.

[0099] Each multivariate continuous function can be expressed as the superposition of a finite number of univariate continuous functions. The specific formula is:

[0100] ;

[0101] in, Corresponding to the learnable basis function, that is, the inner function of the fusion module, The outer layer function approximates the multivariate function through a unary function. In practical applications, it is converted into a trainable neural network structure (i.e., the KAN fusion module). The inner and outer layer functions are designed as different MLP networks. Decompose the multivariate function into an inner summation, allowing each input dimension x_p to be independently mapped to the feature space. n is the number of features. In practice, the outputs of the inner functions for each prosodic feature are summed and averaged, which serves as the input to the outer function. After passing through the outer function, the output is directly output, eliminating the stacking design in the formula. This improves training speed.

[0102] Therefore, the specific formula in this embodiment is:

[0103] ;

[0104] in, is the weight, is the basis function.

[0105] In the specific implementation, this embodiment designs a rhythm controller, the internal structure of the rhythm controller is as follows Figure 4 As shown in Figure 1, it includes duration prediction, energy prediction, pitch period prediction, pause prediction, and rhythm prediction. The predictor structure consists of a one-dimensional convolutional network and a linear layer with a ReLU activation function and dropout operation.

[0106] The duration predictor (hereinafter referred to as D) outside the prosody predictor uses the duration of the phonemes extracted by the forced alignment MAS as the training target, and inputs the output h_text of the text encoder. The output duration is sent to the prosody predictor and, together with h_text, is normalized by the Length Regulator (length normalization) as the key information for subsequent outputs of other predictors. It is denoted as x .Will x are input into the remaining predictors respectively.

[0107] The pause predictor takes as input the output of the phrase structure encoder, which processes the contextual representation at the word level. It then uses the syntactic representation to predict speaker-dependent pause sequences. The pause predictor consists of two bidirectional long short-term memory (Bi-LSTM) layers and a one-dimensional convolutional network with Reluctant Unit (ReLU) activations, layer normalization, and dropout. A final linear layer projects the hidden representation into a word-level pause sequence. Furthermore, the pause predictor is optimized using a cross-entropy loss between the probability distribution and the target pause label sequence.

[0108] S4: Fusing the text context feature representation with the prediction result to generate a fused prosody control embedding.

[0109] It should be noted that the corresponding embeddings obtained by all predictors are input into the inner function , add up the outputs of all inner functions and average them The outer functions are fused and used as the final output of the rhythm controller.

[0110] in, and :

[0111] The Kolmogorov-Arnold representation theorem states that every continuous function of multiple variables can be represented as a superposition of a finite number of continuous functions of one variable. The key points to understand the Kolmogorov-Arnold representation theorem are:

[0112] 1. Continuous: For any point on the function, no matter from which direction you approach it, the value is equal;

[0113] (1) The multivariate function (the objective function of the fitting) is continuous;

[0114] (2) The decomposed univariate function (hereinafter referred to as basis function or univariate basis function) is also continuous;

[0115] 2. Finite: The number of unary basis functions after decomposition is finite.

[0116] This representation method can be used for feature fusion. The specific steps are as follows:

[0117] 1. Inner function design:

[0118] Use MLP+LayerNorm instead of spline function.

[0119] 2. Inner layer combination:

[0120] The design is an MLP structure containing a linear layer, a LayerNorm layer, and a SiLU activation function. The summation part in the theorem is simplified to averaging after adding features, reducing the computational complexity.

[0121] 3. Outer function :

[0122] The outer layer function is designed with an MLP structure similar to the inner layer function. The MLP implements nonlinear mapping, extending the theorem's purely linear requirements to accommodate speech tasks. Residual connections are added to maintain information integrity. Adding a LayerNorm layer stabilizes feature scaling and addresses issues such as distribution shift and gradient explosion during training caused by MLP spline substitution.

[0123] S5: Based on the VITS model, the prosody control embedding and speaker embedding vector are combined to generate a multi-speaker speech spectrum.

[0124] It should be noted that in step S5: the posterior encoder of the VITS model uses the Mel spectrum as input to generate a latent variable z, and converts the latent variable distribution into a prior distribution through a flow model; the speaker embedding vector is concatenated with the text features and prosody control embedding and input into the decoder to generate a multi-speaker speech spectrum.

[0125] In a specific implementation, the steps of Chinese speech synthesis include: establishing a multi-speaker Chinese speech synthesis model based on VITS, including multiple modules such as an encoder, a discriminator, a decoder, a random duration predictor and a prosody controller.

[0126] The text encoder uses a Transformer structure to convert the input text into an embedding vector, add position encoding to retain the sequential information of the text, extract the contextual information of the text through multiple Transformer layers, and generate a high-dimensional feature representation.

[0127] The rhythmic features are fused with the text features, and the rhythmic features are further extracted and modeled through multiple convolutional layers. The final rhythmic feature prediction value is generated through the prediction layer.

[0128] Speaker embedding vectors are introduced to enable the model to distinguish the timbre and prosody characteristics of different speakers:

[0129] (1) Convert speaker information into an embedding vector;

[0130] (2) Fusion of speaker embedding vectors with text features and prosodic features.

[0131] The latent variable h_text generated by the text encoder and the duration information predicted by the duration predictor D(duration) serve as inputs to the prosody controller. Each predictor is simply denoted as E(energy), P(pitch), R(rhythm), and T(pause). Furthermore, energy and pitch period extracted from the speech dataset serve as target inputs for the E, P, R, and T predictors. The results of each predictor are output as an embedding, which is continuously fused with the LR-derived x as the output of the prosody controller. During inference, appropriate prosodic feature parameters of the generated embeddings can be controlled by adjusting the corresponding parameters.

[0132] The linear projection layer then generates the mean and variance of the prior distribution. A noise is upsampled from the (0, 1) Gaussian and fed into the period predictor along with h_text to generate a period list. A corresponding mapping is constructed based on the duration, and the corresponding Gaussian distribution is replicated to match the audio length.

[0133] During training, the posterior encoder takes a linear spectrum as input and outputs a latent variable z. Audio features undergo a one-dimensional convolution, becoming size-192, consistent with the size of the text embedding. During inference, the flow model generates the latent variable z. The flow layer transforms the Q-distributed z into a P-distributed one, enhancing modeling capabilities. Finally, the decoder of HiFiGAN V1 is trained to slice to ensure consistent sample length. The decoder is a series of upsampling steps, ultimately restoring the audio to a 256-bit size.

[0134] S6: Convert the multi-speaker speech spectrum into a time-domain speech signal through a decoder.

[0135] In a specific implementation, the decoder adopts the HiFiGAN V1 structure and converts the speech spectrum into a time domain waveform through multi-level upsampling.

[0136] S7: By adjusting the prosodic feature parameters, independent control of the duration, pitch period, energy, pauses, and rhythm of the synthesized speech is achieved.

[0137] In step S7, the duration is controlled by adjusting the output multiple of the duration predictor to achieve speech rate change; the pitch period is controlled by modifying the amplitude of the wavelet component to adjust the pitch; the pause is controlled by specifying the pause duration through the marker symbols in the input text; the rhythm is controlled by dynamically adjusting the pronunciation speed through the semantic weight output by BERT.

[0138] In practice, the prosody control step involves adjusting the embeddings output by the corresponding predictor via input parameters to achieve prosody control. The duration of speech directly affects the speed of speech. A shorter duration means a faster rate, while a longer duration results in a slower rate. Speech rate control is achieved by multiplying the output of the duration predictor by a certain value. Faster speech is indicated by a multiplication value less than 1, while slower speech is indicated by a multiplication value greater than 1.

[0139] The fundamental frequency (f0 / F0) is the inverse of the fundamental frequency, which corresponds to the frequency of vocal cord vibration and represents the pitch of the sound. Generally, the faster the vocal cords vibrate, the higher the fundamental frequency.

[0140] Short-time energy reflects the strength of the signal at different times, reflecting the strength of the sound signal.

[0141] Appropriate pauses between adjectives and nouns help listeners identify the modifying relationship. Natural speech contains many subtle pauses, and reproducing these pauses in synthesized speech can make it sound more natural and realistic. At the same time, pauses are also essential for expressing emotions such as surprise and hesitation. By precisely controlling pauses, speech synthesis systems can generate more natural, clear, and expressive speech output. Control over pauses is reflected in different markers in the input text:

[0142] _pause = ["sil", "eos", "sp", "#0", "#1", "#2", "#3"]

[0143] Among them, sil: text start mark, encoded as index 0; eos: silence; sp: comma, short silence, encoded as index 2; "#0" is designed as a short pause mark, and "#1" is a long pause mark, and the pause duration can be entered manually.

[0144] In speech synthesis, rhythm refers to the relative timing and intensity patterns of syllables, words, or sentences in speech, which determines the musicality and fluency of language. It is closely related to pauses, speaking rate, and intensity, and together they constitute the rhythmic structure of language.

[0145] It is understandable that it also includes training steps: using the AISHELL-3 or BZNSYP Chinese speech dataset to extract the duration, pitch period, energy and pause duration of real speech as training targets; by minimizing the reconstruction error between the predicted value and the true value of each prosodic feature, jointly optimize the text encoder, prosodic controller and VITS model parameters.

[0146] In the specific implementation, the steps of rhythm controller training include: in the preprocessing of the data set, first filter out the erhua sounds that are not supported by this project, extract and save information such as energy, fundamental frequency period and spectrum, and use these extracted rhythm feature elements as training targets for the corresponding predictor.

[0147] During the training phase, forced alignment is used to calculate phoneme duration information. The duration information parameters obtained by the duration predictor are input into the prosody controller. The output of the length regulator serves as the input to the remaining rhythm, energy, and pitch period predictors. The extracted intensity, pitch period, and other information are used as the training targets for each predictor. By minimizing the reconstruction error of the prosody features, the prosody controller is trained to accurately model the prosody features. The specific steps are as follows:

[0148] (1) Energy reconstruction error: the error between the predicted energy and the extracted true energy;

[0149] (2) Pitch period reconstruction error: Calculate the error between the predicted pitch period and the extracted true pitch period;

[0150] (3) Rhythm reconstruction error: Calculate the error between the predicted rhythm and the extracted true rhythm;

[0151] (4) Pause reconstruction error: Calculate the error between the predicted pause and the extracted true pause;

[0152] Optimization: Optimize the rhythm controller parameters by minimizing the above reconstruction error.

[0153] It should be noted that the data processing method in this embodiment is: using the AISHELL-3 Hill Shell Chinese Mandarin speech database as training and test data, which can be used as a multi-speaker synthesis system, where the data set contains 174 speakers. Each data set contains full-bandwidth, mono audio, with a speech duration of 85 hours and a sampling frequency of 44.1kHz. The BZNSYP-Biaobei Chinese standard female single-person data set is used to simply replace AISHELL-3 as an alternative data set during development, which is convenient for direct use in model training and reduces the workload of data preprocessing. It contains 10,000 audio data points, with a total duration of approximately 10 hours and a sampling rate of 22050Hz. The sampling rate is unified to 16000Hz during the data preprocessing stage.

[0154] This embodiment extracts prosodic feature information from input text; semantically encodes the input text using a BERT pre-trained model to generate a text context feature representation; constructs a prosodic controller comprising a duration predictor, an energy predictor, a pitch period predictor, a pause predictor, and a rhythm predictor, and independently models the prosodic features using the prosodic controller to generate prediction results; fuses the text context feature representation with the prediction results to generate a fused prosodic control embedding; generates a multi-speaker speech spectrum based on the VITS model, combining the prosodic control embedding with a speaker embedding vector; converts the multi-speaker speech spectrum into a time-domain speech signal using a decoder; and independently controls the duration, pitch period, energy, pauses, and rhythm of the synthesized speech by adjusting prosodic feature parameters.

[0155] In addition, an embodiment of the present application also proposes a computer-readable storage medium, on which a program for prosody-controllable speech synthesis based on VITS is stored. When the program for prosody-controllable speech synthesis based on VITS is executed by a processor, the steps of the method for prosody-controllable speech synthesis based on VITS as described above are implemented.

[0156] Reference Figure 5 , Figure 5This is a structural block diagram of the first embodiment of the prosody-controllable speech synthesis device based on VITS of this application.

[0157] like Figure 5 As shown, the prosody-controllable speech synthesis device based on VITS proposed in the embodiment of the present application includes:

[0158] Extraction module 10, for extracting prosodic feature information from input text, wherein the prosodic feature information includes prosodic features such as duration, pitch period, energy, pause and rhythm;

[0159] Semantic encoding module 20, used to semantically encode the input text using the BERT pre-trained model to generate text context feature representation;

[0160] A predictor construction module 30 is used to construct a prosody controller including a duration predictor, an energy predictor, a pitch period predictor, a pause predictor, and a rhythm predictor, and independently model the prosody features and generate prediction results through the prosody controller;

[0161] a fusion module 40 for fusing the text context feature representation with the prediction result to generate a fused prosody control embedding;

[0162] A speech spectrum module 50 is configured to generate a multi-speaker speech spectrum based on the VITS model and in combination with the prosodic control embedding and the speaker embedding vector;

[0163] A conversion module 60, configured to convert the multi-speaker speech spectrum into a time-domain speech signal through a decoder;

[0164] The control module 70 is used to independently control the duration, pitch period, energy, pause and rhythm of the synthesized speech by adjusting the prosodic feature parameters.

[0165] It should be understood that the above is only an example and does not constitute any limitation to the technical solution of the present application. In specific applications, technicians in this field can make settings as needed, and the present application does not impose any restrictions on this.

[0166] This embodiment extracts prosodic feature information from input text; semantically encodes the input text using a BERT pre-trained model to generate a text context feature representation; constructs a prosodic controller comprising a duration predictor, an energy predictor, a pitch period predictor, a pause predictor, and a rhythm predictor, and independently models the prosodic features using the prosodic controller to generate prediction results; fuses the text context feature representation with the prediction results to generate a fused prosodic control embedding; generates a multi-speaker speech spectrum based on the VITS model, combining the prosodic control embedding with a speaker embedding vector; converts the multi-speaker speech spectrum into a time-domain speech signal using a decoder; and independently controls the duration, pitch period, energy, pauses, and rhythm of the synthesized speech by adjusting prosodic feature parameters.

[0167] It should be noted that the workflow described above is merely illustrative and does not limit the scope of protection of this application. In actual applications, technicians in this field can select part or all of it according to actual needs to achieve the purpose of this embodiment scheme, and no restrictions are imposed here.

[0168] In addition, for technical details not fully described in this embodiment, please refer to the VITS-based prosody-controllable speech synthesis method provided in any embodiment of the present application, and will not be repeated here.

[0169] In addition, it should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or system comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or system. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or system comprising the element.

[0170] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0171] Through the above description of the embodiments, those skilled in the art will clearly understand that the above-mentioned embodiments and methods can be implemented using software plus the necessary general-purpose hardware platform. Of course, hardware can also be used, but in many cases the former is a more preferred embodiment. Based on this understanding, the technical solution of this application, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as read-only memory (ROM) / RAM, a magnetic disk, or an optical disk) and includes several instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of each embodiment of this application. The above are only preferred embodiments of this application and do not limit the scope of the patent application. Any equivalent structure or equivalent process transformation made using the contents of this application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the scope of patent protection of this application.

Claims

1. A prosody-controllable speech synthesis method based on VITS, characterized in that: include: S1: extracting prosodic feature information from an input text, wherein the prosodic feature information includes prosodic features such as duration, pitch period, energy, pause, and rhythm; S2: Use the BERT pre-trained model to semantically encode the input text and generate text context feature representation; S3: constructing a prosody controller including a duration predictor, an energy predictor, a pitch period predictor, a pause predictor, and a rhythm predictor, and independently modeling the prosody features and generating prediction results through the prosody controller; S4: fusing the text context feature representation with the prediction result to generate a fused prosodic control embedding; S5: Generate multi-speaker speech spectra based on the VITS model, combining the prosodic control embedding and speaker embedding vectors; S6: converting the multi-speaker speech spectrum into a time-domain speech signal through a decoder; S7: By adjusting the prosodic feature parameters, the duration, pitch period, energy, pauses and rhythm of the synthesized speech can be independently controlled; Among them, in the step S3: the prosody controller is composed of a one-dimensional convolutional network, a bidirectional long short-term memory network and a linear layer, and the output embedding of each predictor is superimposed through the feature fusion module, and the Kolmogorov-Arnold representation theorem is applied to realize the unary basis function fusion of multivariate features. The specific formula is: in, is the weight, is the basis function.

2. The VITS-based prosody-controllable speech synthesis method according to claim 1, characterized in that: In the step S1: The duration is extracted by a forced alignment algorithm, and a duration predictor is trained using maximum likelihood estimation; The pitch period is decomposed into multi-scale components by continuous wavelet transform and quantized as a control condition; The energy is encoded by calculating the L2 norm of the Mel spectrum amplitude and quantizing it; The pause is generated by analyzing the text punctuation and phrase structure to generate a pause duration sequence; The rhythm captures semantic associations through the BERT pre-trained model and adjusts the pronunciation speed in combination with position embedding.

3. The prosody-controllable speech synthesis method based on VITS according to claim 1, characterized in that: In the step S5: The posterior encoder of the VITS model takes the Mel spectrum as input to generate a latent variable z, and converts the latent variable distribution into a prior distribution through a flow model; The speaker embedding vector is concatenated with the text features and the prosody control embedding and input into a decoder to generate a multi-speaker speech spectrum.

4. The VITS-based prosody-controllable speech synthesis method according to claim 1, characterized in that: In step S7: The duration is controlled by adjusting the output multiple of the duration predictor to achieve speech speed changes; The control of the fundamental period adjusts the pitch by modifying the amplitude of the wavelet component; The control of pauses is done by specifying the pause duration through marker symbols in the input text; The control of rhythm dynamically adjusts the pronunciation speed through the semantic weights output by BERT.

5. The prosody-controllable speech synthesis method based on VITS according to claim 1, characterized in that: Also includes the training steps: Use the AISHELL-3 or BZNSYP Chinese speech dataset to extract the duration, pitch period, energy, and pause duration of real speech as training targets; By minimizing the reconstruction error between the predicted value and the true value of each prosodic feature, the text encoder, prosodic controller and VITS model parameters are jointly optimized.

6. The prosody-controllable speech synthesis method based on VITS according to claim 1, characterized in that: The decoder adopts the HiFiGAN V1 structure and converts the speech spectrum into a time domain waveform through multi-level upsampling.

7. A prosody-controllable speech synthesis device based on VITS, characterized in that: Executing the method according to claim 1, comprising: An extraction module is used to extract prosodic feature information from an input text, wherein the prosodic feature information includes prosodic features such as duration, pitch period, energy, pause, and rhythm; The semantic encoding module is used to semantically encode the input text using the BERT pre-trained model to generate text context feature representation; A predictor construction module is used to construct a prosody controller including a duration predictor, an energy predictor, a pitch period predictor, a pause predictor, and a rhythm predictor, and independently model the prosody features and generate prediction results through the prosody controller; a fusion module for fusing the text context feature representation with the prediction result to generate a fused prosody control embedding; A speech spectrum module, configured to generate a multi-speaker speech spectrum based on the VITS model and in combination with the prosodic control embedding and the speaker embedding vector; a conversion module, configured to convert the multi-speaker speech spectrum into a time-domain speech signal through a decoder; The control module is used to independently control the duration, pitch period, energy, pause and rhythm of the synthesized speech by adjusting the prosodic feature parameters.

8. A computer device, characterized in that: The device comprises: a memory and a processor, and when the processor runs computer instructions stored in the memory, the processor executes the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that The method comprises instructions, which, when executed on a computer, cause the computer to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Speech synthesis model training method, speech synthesis method and related equipment

    CN117953862A

  • Iterative improvement of speech recoginition, voice conversion, and text-to-speech models

    US20240304180A1