Cross-language speech synthesis system and method based on multi-encoder feature decoupling

Through a cross-language speech synthesis system based on multi-encoder feature decoupling, the problems of few language support, serious audio tone change and high training costs in the prior art are solved, and higher speech synthesis accuracy and fluency are achieved, and training costs are reduced.

CN119993119AActive Publication Date: 2025-05-13GUANGDONG LIANTING TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510450232.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-05-13
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

The existing cross-language pronunciation synthesis technology has problems such as few language support, serious audio tone change and high training costs.

Method used

A cross-language speech synthesis system based on multi-encoder feature decoupling is adopted. Through the data collection and processing module, a cross-language speech synthesis model and a model training module, a model composed of text encoder, audio encoder, decoder and discriminator, and a variety of open source monolingual speech data sets are used for training.

Benefits of technology

It improves the accuracy and audio fluency of cross-language pronunciation synthesis, reduces training costs, and expands language support capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993119A_ABST
    Figure CN119993119A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of speech synthesis, in particular to a cross-language speech synthesis system and method based on multi-encoder feature decoupling. According to the technical scheme, an audio encoder with multiple encoders is used for conducting multi-feature decoupling on audio data to obtain multiple audio hidden variables, then the multiple audio hidden variables are spliced, fitting is conducted through text hidden variables, and finally output audio is obtained through decoding of a decoder. The method has the beneficial effects that the text hidden variables can be conveniently fitted, so that the cross-language speech synthesis accuracy is improved; the single-language voice data set of multiple open sources is used, the problems that an existing method depends on multilingual voice audios, the price is high, and more languages are difficult to expand are solved, any language can be added into the model through a unified text and audio processing method by using the single-language open source data set, and the model can be applied to multiple languages. And therefore, the model has higher robustness and universality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech synthesis technology, and in particular to a cross-language speech synthesis system and method based on multi-encoder feature decoupling. Background Art

[0002] Speech synthesis technology, also known as text-to-speech technology, is a technology that aims to make human-understandable speech by inputting text. Speech synthesis technology is an important part of realizing human-computer communication and establishing human-computer interaction systems. It has been widely used in many scenarios such as mobile assistants, video dubbing, and voice navigation. As a branch of speech synthesis technology, cross-language speech synthesis technology is mainly to solve the problem that most current speech synthesis technologies can only achieve single-language speech synthesis, cannot use the voice of a foreign speaker to generate speech in another language, and cannot output corresponding speech when the text contains multiple languages. In recent years, end-to-end speech synthesis technology based on deep neural networks has become a research hotspot. The sounds synthesized by many excellent models have reached the level of being indistinguishable from the real thing. However, the current cross-language speech synthesis models still have problems such as few supported languages ​​and serious pitch shifting of the generated audio.

[0003] In order to make the speech synthesized by the model more natural and fluent, most methods currently use speech in multiple languages ​​as training data, introduce additional feature information into the model, and enhance the model's modeling capabilities for different languages. However, this method has a high training cost and the effect of synthesizing audio is not good enough, so there is still much room for improvement. Summary of the invention

[0004] The purpose of the present invention is to provide a cross-language speech synthesis system and method based on multi-encoder feature decoupling to solve the problems of poor cross-language speech synthesis effect and high training cost in the prior art.

[0005] To achieve the above-mentioned purpose, the present invention adopts the following technical scheme: a cross-language speech synthesis system based on multi-encoder feature decoupling, comprising a data collection and processing module, a cross-language speech synthesis model and a model training module; the data collection and processing module is used to collect and store speech audio data of several target languages ​​and their corresponding texts, and process the texts therein to obtain phoneme sequences, and pre-process the speech audio data therein; the cross-language speech synthesis model is composed of a text encoder, an audio encoder, a decoder and a discriminator, the audio encoder is used to obtain audio latent variables using the speech audio data pre-processed in the data collection and processing module as input, the text encoder is used to convert the input text into text latent variables in combination with the phoneme sequence in the data collection and processing module, the text latent variables are fitted according to the audio latent variables to obtain latent variables and decoded by the decoder to generate output audio, the discriminator is used to discriminate the audio generated by the decoder to verify the audio fluency; the model training module uses the data in the data collection and processing module to train the cross-language speech synthesis model.

[0006] Specifically, the audio encoder includes a pitch encoder, a rhythm encoder and a content encoder, each of which is composed of a first convolutional layer, a group normalization layer, a bidirectional long short-term memory network and a second convolutional layer. The pitch encoder, the rhythm encoder and the content encoder are used to output pitch latent variables, rhythm latent variables and content latent variables respectively. The audio latent variables are obtained by splicing the pitch latent variables, the rhythm latent variables and the content latent variables.

[0007] Specifically, the text encoder consists of an embedding layer, an attention block, and a convolutional layer. The embedding layer is used to convert the input text combined with the phoneme sequence into a feature vector. The feature vector is processed by the attention block and the convolutional layer in turn to obtain the text latent variable.

[0008] Specifically, the decoder adopts the HiFi-GAN V1 structure, which consists of multiple layers of sequentially connected transposed convolutional layers, and a multi-receptive field fusion module is set after each transposed convolutional layer.

[0009] Specifically, the model training module includes an Adam optimizer and a multi-cycle discriminator. The beta values ​​of the Adam optimizer are 0.9 and 0.98, which are used to train the cross-language speech synthesis model. The multi-cycle discriminator is a discriminator of the HiFi-GAN structure, which is used to calculate the loss between the audio output by the cross-language speech synthesis model and the real audio.

[0010] A cross-language speech synthesis method based on multi-encoder feature decoupling comprises the following steps: S1. Collect data. Collect speech and audio data of several target languages ​​and their corresponding text data as the total data set. Preprocess the speech and audio data and text data separately, and divide the total data set into a training set, a test set, and a validation set.

[0011] S2. Build a model, build a cross-language speech synthesis model composed of a text encoder, an audio encoder, a decoder and a discriminator, the audio encoder is used to obtain audio latent variables with the speech audio data preprocessed in the data collection and processing module as input, the text encoder is used to convert the input text into text latent variables based on the phoneme sequence in the data collection and processing module, the text latent variables are fitted according to the audio latent variables to obtain latent variables and decoded by the decoder to generate output audio, the discriminator is used to discriminate the audio generated by the decoder to verify the audio fluency.

[0012] S3, model training, use the training set to train the cross-language speech synthesis model constructed in step S2, use the Adam optimizer for training during the training process, and the beta value of the Adam optimizer is 0.9 and 0.98. After multiple iterations of training convergence, use the text data in the test set as the input of the cross-language speech synthesis model, and use the multi-cycle discriminator of the HiFi-GAN structure to calculate the loss between the predicted audio and the true value. If it meets the requirements, the trained cross-language speech synthesis model is obtained.

[0013] S4, model testing, selecting text data from the validation set and inputting it into the cross-language speech synthesis model trained in step S3, and having people who master the corresponding language score the output audio. If the score meets the requirements, it is judged that the cross-language speech synthesis model meets the usage requirements.

[0014] Specifically, the audio encoder in step S2 includes a pitch encoder, a prosody encoder and a content encoder, each of which is composed of a first convolutional layer, a group normalization layer, a bidirectional long short-term memory network and a second convolutional layer. The pitch encoder, prosody encoder and content encoder are used to output pitch latent variables, prosody latent variables and content latent variables, respectively. The pitch latent variables, prosody latent variables and content latent variables are concatenated to obtain audio latent variables.

[0015] Specifically, the text encoder in step S2 consists of an embedding layer, an attention block and a convolutional layer. The embedding layer is used to convert the input text combined with the phoneme sequence into a feature vector, and the feature vector is processed by the attention block and the convolutional layer in turn to obtain the text latent variable.

[0016] Specifically, when the speech audio data and text data are preprocessed in step S1, the speech audio data is specifically unified: with a frame length of 1024 points, a window length of 1024 points, a frame shift of 256 points, a Mel minimum frequency of 0Hz, and a Mel maximum frequency of 8000Hz as parameters, and a pre-emphasis coefficient of 0.97, the speech audio data is converted into a linear spectrum feature with a unified sampling rate of 16000Hz; the text data of all languages ​​are converted into a phoneme sequence according to a pronunciation dictionary.

[0017] Specifically, in step S2, when the text latent variable output by the text encoder is fitted according to the audio latent variable, the similarity between the text latent variable and the audio latent variable is calculated through the KL loss function, so that the text latent variable is fitted close to the audio latent variable to obtain the latent variable.

[0018] The beneficial effects of the present invention are: 1. By using a multi-encoder audio encoder, the audio data is decoupled into multiple features to obtain multiple audio latent variables. Then, by splicing multiple audio latent variables, it is convenient to fit the text latent variables, thereby improving the accuracy of cross-language speech synthesis; 2. The use of a variety of open source monolingual speech datasets solves the problem that existing methods rely on multilingual speech audio, which is expensive and difficult to expand to more languages. By using monolingual open source datasets, any language can be added to the model through a unified text and audio processing method, making the model more robust and versatile. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Attached Figure 1 A schematic diagram of the connection principle of the cross-language speech synthesis model in the embodiment; Attached Figure 2 It is a schematic diagram of the connection between the prosody encoder / pitch encoder / content encoder in the audio encoder in the embodiment; Attached Figure 3 4 is a connection schematic diagram of an audio encoder in an embodiment. DETAILED DESCRIPTION

[0020] Example 1, reference Figure 1-3A cross-language speech synthesis system based on multi-encoder feature decoupling includes a data collection and processing module, a cross-language speech synthesis model and a model training module; the data collection and processing module is used to collect and store speech audio data of several target languages ​​and their corresponding texts, and process the texts therein to obtain phoneme sequences, and pre-process the speech audio data therein; the cross-language speech synthesis model is composed of a text encoder, an audio encoder, a decoder and a discriminator, the audio encoder is used to obtain audio latent variables using the speech audio data pre-processed in the data collection and processing module as input, the text encoder is used to convert the input text into text latent variables in combination with the phoneme sequence in the data collection and processing module, the text latent variables are fitted according to the audio latent variables to obtain latent variables and decoded by the decoder to generate output audio, the discriminator is used to discriminate the audio generated by the decoder to verify the audio fluency; the model training module uses the data in the data collection and processing module to train the cross-language speech synthesis model.

[0021] Specifically, the audio encoder includes a pitch encoder, a prosody encoder and a content encoder, each of which is composed of a first convolutional layer, a group normalization layer, a bidirectional long short-term memory network and a second convolutional layer. The pitch encoder, the prosody encoder and the content encoder are used to output pitch latent variables, prosody latent variables and content latent variables respectively, and the audio latent variables are obtained by splicing the pitch latent variables, the prosody latent variables and the content latent variables. Among them, the first and second convolutional layers of the pitch encoder have 3 layers, the number of group normalization layers is 16, and the number of bidirectional long short-term memory network layers is 3; the first and second convolutional layers of the prosody latent encoder have 1 layer, the number of group normalization layers is 8, and the number of bidirectional long short-term memory network layers is 1; the first and second convolutional layers of the content encoder have 3 layers, the number of group normalization layers is 32, and the number of bidirectional long short-term memory network layers is 2.

[0022] Specifically, the text encoder consists of an embedding layer, an attention block, and a convolutional layer. The embedding layer is used to convert the input text combined with the phoneme sequence into a feature vector. The feature vector is processed by the attention block and the convolutional layer in turn to obtain the text latent variable.

[0023] Specifically, the decoder adopts the HiFi-GAN V1 structure, which consists of multiple layers of transposed convolutional layers connected in sequence, and a multi-receptive field fusion module is set after each transposed convolutional layer.

[0024] Specifically, the model training module includes an Adam optimizer and a multi-cycle discriminator. The beta values ​​of the Adam optimizer are 0.9 and 0.98, which are used to train the cross-language speech synthesis model. The multi-cycle discriminator is a discriminator of the HiFi-GAN structure, which is used to calculate the loss between the audio output by the cross-language speech synthesis model and the real audio.

[0025] Based on the above cross-language speech synthesis system, this embodiment also provides a cross-language speech synthesis method based on multi-encoder feature decoupling, including the following steps: S1. Collect data, collect speech audio data of several target languages ​​and their corresponding text data as the total data set, pre-process the speech audio data and text data respectively, and divide the total data set into training set, test set and verification set. This embodiment uses Chinese, English, Thai and Polish as target languages, and uses corresponding open source data sets respectively: Chinese data set AIShell3, English data set VCTK, LibriTTS, Thai data set Common Voice, Polish data set, where the Polish data set is obtained by the applicant's own recording. Among them, when pre-processing speech audio data and text data, specifically, the speech audio data is processed in a unified manner: with a frame length of 1024 points, a window length of 1024 points, a frame shift of 256 points, a Mel minimum frequency of 0Hz, and a Mel maximum frequency of 8000Hz as parameters, the speech audio data is converted into a linear spectral feature with a unified sampling rate of 16000Hz with a pre-emphasis coefficient of 0.97; the text data of all languages ​​are converted into phoneme sequences according to the pronunciation dictionary.

[0026] S2. Build a model, build a cross-language speech synthesis model composed of a text encoder, an audio encoder, a decoder and a discriminator, wherein the audio encoder is used to obtain audio latent variables by taking the pre-processed speech audio data in the data collection and processing module as input, the text encoder is used to convert the text latent variables into text latent variables according to the input text combined with the phoneme sequence in the data collection and processing module, the text latent variables are fitted according to the audio latent variables to obtain latent variables and decoded by the decoder to generate output audio, and the discriminator is used to discriminate the audio generated by the decoder to verify the audio fluency. The relevant structure and parameters of the above cross-language speech synthesis model are as described in the above system. It should be noted that the prosody encoder and the content encoder both directly use the pre-processed audio data as input to calculate the corresponding prosody latent variables and content latent variables, and for the pitch encoder, the audio data input is the audio data obtained after pitch contour normalization and alignment resampling based on the pre-processed audio data as input for calculation, so as to obtain the pitch latent variables. In addition, when the text latent variables output by the text encoder are fitted according to the audio latent variables, the similarity between the text latent variables and the audio latent variables is calculated through the KL loss function, so that the text latent variables are fitted close to the audio latent variables to obtain the latent variables.

[0027] S3, model training, use the training set to train the cross-language speech synthesis model constructed in step S2, use the Adam optimizer for training during the training process, the beta value of the Adam optimizer is 0.9 and 0.98, after multiple iterations of training convergence, use the text data in the test set as the input of the cross-language speech synthesis model, use the multi-cycle discriminator of the HiFi-GAN structure to calculate the loss between the predicted audio and the true value, and obtain a trained cross-language speech synthesis model if it meets the requirements. Specifically, in this embodiment, when using the Adam optimizer for training, after 1 million steps of iterative training, a trained cross-language speech synthesis model is obtained. S4, model testing, selecting text data from the validation set and inputting it into the cross-language speech synthesis model trained in step S3, and having people who master the corresponding language score the output audio, and if the score meets the requirements, it is determined that the cross-language speech synthesis model meets the use requirements. Specifically, in this embodiment, the cross-language speech synthesis model obtained in step S3 is tested, and the test process is as follows: 5 to 10 audios of different speakers are randomly selected from the validation set, and then one audio is selected from them and marked as real audio, and the real audio (pre-processed) is input into the audio encoder of the cross-language speech synthesis model (that is, the speaker of the audio is the speaker of the output audio), and then the cross-language speech synthesis model is used to input some pre-set texts (including different languages), and 15 to 20 audios are randomly generated, and the output audios are mixed. Invite 10 speakers who can master the above target languages ​​as testers, and inform them of the 5-point evaluation criteria for speech synthesis audio quality: 4.0~5.0 is very good, clear hearing, small delay, and smooth speech; 3.5~4.0 is slightly worse, clear hearing, small delay, poor communication, and noise; 3.0~3.5 is OK, not very clear, with a certain delay, and communication is possible; 1.5~3.0 is barely, not very clear, with a large delay, and communication requires multiple repetitions; 0~1 is extremely poor, incomprehensible, with large delay, and poor communication. At the beginning of the test, first let the testers listen to the above real audio, and then listen to and score the mixed multiple audios according to the above scoring criteria, average the scores of all testers for the same audio to get the score of the audio, and then average the scores of all audios to get the final score of the cross-language speech synthesis model. At the same time, the effect of the cross-language speech synthesis model obtained in this embodiment is verified. In this embodiment, three other existing cross-language speech synthesis models: vits, GEN, and SANE-TTS are selected to perform scoring in the same way, and the following results are obtained: Model Final score vits 3.55 GEN 3.5 SANE-TTS 3.44 This embodiment 3.89 The tests verified the effectiveness of the system and method of the present application. Compared with the current mainstream cross-language speech synthesis models, the scores of the present invention have obvious advantages.

[0028] Of course, the above are only preferred embodiments of the present invention, and are not intended to limit the scope of use of the present invention. Therefore, any equivalent changes made to the principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A cross-language speech synthesis system based on multi-encoder feature decoupling, characterized by: It includes a data collection and processing module, a cross-language speech synthesis model and a model training module; the data collection and processing module is used to collect and store speech and audio data of several target languages ​​and their corresponding texts, and process the texts therein to obtain phoneme sequences, and pre-process the speech and audio data therein; The cross-language speech synthesis model consists of a text encoder, an audio encoder, a decoder and a discriminator. The audio encoder is used to obtain audio latent variables by taking the speech audio data preprocessed in the data collection and processing module as input. The text encoder is used to convert the input text into text latent variables in combination with the phoneme sequence in the data collection and processing module. The text latent variables are fitted according to the audio latent variables to obtain latent variables and decoded by the decoder to generate output audio. The discriminator is used to discriminate the audio generated by the decoder to verify the audio fluency. The model training module uses the data in the data collection and processing module to train the cross-language speech synthesis model.

2. A cross-language speech synthesis system based on multi-encoder feature decoupling according to claim 1, characterized in that: The audio encoder includes a pitch encoder, a prosody encoder and a content encoder, each of which consists of a first convolutional layer, a group normalization layer, a bidirectional long short-term memory network and a second convolutional layer. The pitch encoder, the prosody encoder and the content encoder are used to output pitch latent variables, prosody latent variables and content latent variables respectively. The audio latent variables are obtained by splicing the pitch latent variables, the prosody latent variables and the content latent variables.

3. The cross-language speech synthesis system based on multi-encoder feature decoupling according to claim 1, characterized in that: The text encoder consists of an embedding layer, an attention block and a convolutional layer. The embedding layer is used to convert the input text combined with the phoneme sequence into a feature vector. The feature vector is processed by the attention block and the convolutional layer in turn to obtain the text latent variable.

4. The cross-language speech synthesis system based on multi-encoder feature decoupling according to claim 1, characterized in that: The decoder adopts the HiFi-GAN V1 structure, which is composed of multiple layers of transposed convolutional layers connected in sequence, and a multi-receptive field fusion module is arranged after each transposed convolutional layer.

5. The cross-language speech synthesis system based on multi-encoder feature decoupling according to claim 1, characterized in that: The model training module includes an Adam optimizer and a multi-cycle discriminator. The beta values ​​of the Adam optimizer are 0.9 and 0.98, which are used to train the cross-language speech synthesis model. The multi-cycle discriminator is a discriminator of a HiFi-GAN structure, which is used to calculate the loss between the audio output by the cross-language speech synthesis model and the real audio.

6. A cross-language speech synthesis method based on multi-encoder feature decoupling, characterized in that: The steps include: S1. Collect data: collect speech and audio data of several target languages ​​and their corresponding text data as a total data set, pre-process the speech and audio data and text data respectively, and divide the total data set into a training set, a test set, and a validation set; S2. Build a model, build a cross-language speech synthesis model composed of a text encoder, an audio encoder, a decoder and a discriminator, wherein the audio encoder is used to obtain audio latent variables by taking the pre-processed speech audio data in the data collection and processing module as input, the text encoder is used to convert the input text into text latent variables in combination with the phoneme sequence in the data collection and processing module, the text latent variables are fitted according to the audio latent variables to obtain latent variables and decoded by the decoder to generate output audio, and the discriminator is used to discriminate the audio generated by the decoder to verify the audio fluency; S3, model training, using the training set to train the cross-language speech synthesis model constructed in step S2, using the Adam optimizer for training during the training process, the beta value of the Adam optimizer is 0.9 and 0.98, after multiple iterations of training convergence, using the text data in the test set as the input of the cross-language speech synthesis model, using the multi-cycle discriminator of the HiFi-GAN structure to calculate the loss between the predicted audio and the true value, and obtaining a trained cross-language speech synthesis model if it meets the requirements; S4, model testing, selecting text data from the validation set and inputting it into the cross-language speech synthesis model trained in step S3, and having people who master the corresponding language score the output audio. If the score meets the requirements, it is judged that the cross-language speech synthesis model meets the usage requirements.

7. The cross-language speech synthesis method based on multi-encoder feature decoupling according to claim 6, characterized in that: The audio encoder in step S2 includes a pitch encoder, a prosody encoder and a content encoder, each of which is composed of a first convolutional layer, a group normalization layer, a bidirectional long short-term memory network and a second convolutional layer. The pitch encoder, the prosody encoder and the content encoder are used to output pitch latent variables, prosody latent variables and content latent variables, respectively. The audio latent variables are obtained by splicing the pitch latent variables, the prosody latent variables and the content latent variables.

8. The cross-language speech synthesis method based on multi-encoder feature decoupling according to claim 6, characterized in that: The text encoder in step S2 is composed of an embedding layer, an attention block and a convolution layer. The embedding layer is used to convert the input text combined with the phoneme sequence into a feature vector, and the feature vector is processed by the attention block and the convolution layer in turn to obtain the text latent variable.

9. The cross-language speech synthesis method based on multi-encoder feature decoupling according to claim 6, characterized in that: When the speech audio data and text data are preprocessed in step S1, specifically, the speech audio data is uniformly processed: with a frame length of 1024 points, a window length of 1024 points, a frame shift of 256 points, a minimum Mel frequency of 0 Hz, and a maximum Mel frequency of 8000 Hz as parameters, the speech audio data is converted into a linear spectrum feature with a uniform sampling rate of 16000 Hz with a pre-emphasis coefficient of 0.97; Convert text data in all languages ​​into phoneme sequences based on pronunciation dictionaries.

10. The cross-language speech synthesis method based on multi-encoder feature decoupling according to claim 6, characterized in that: In step S2, when the text latent variables output by the text encoder are fitted according to the audio latent variables, the similarity between the text latent variables and the audio latent variables is calculated through the KL loss function, so that the text latent variables are fitted close to the audio latent variables to obtain the latent variables.

Citation Information

Patent Citations

  • Voice synthesis method based on MOOC voice data set

    CN113539232A

  • Speech synthesis model method capable of synthesizing multi-emotion audio

    CN116798403A

  • Speech synthesis optimization method and speech synthesis optimization model construction method and device

    CN116994555A

  • System and method for direct speech translation system

    US20200226327A1