Cross-lingual Speech Synthesis System and Method Based on Multi-Encoder Feature Decoupling
By adopting multi-encoder feature decoupling technology in the cross-language speech synthesis system, the problems of less language support and serious audio tone change in the prior art are solved, and more efficient cross-language speech synthesis is achieved, which reduces training costs and improves audio quality and model robustness.
Patent Information
- Application Number
- CN202510450232.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-04-11
AI Technical Summary
The existing cross-language pronunciation synthesis technology has problems such as few language support and serious audio tone change, and the training cost is high, so the synthesized audio effect is poor.
A cross-language speech synthesis system based on multi-encoder feature decoupling is adopted, including a data collection and processing module, a cross-language speech synthesis model and a model training module. The model consists of text encoder, audio encoder, decoder and discriminator. Through the decoupling and splicing of multiple audio features, the fitting ability of text hidden variables is improved, thereby generating more natural and smooth cross-language speech.
It improves the accuracy and audio quality of cross-language pronunciation synthesis, reduces training costs, expands language support capabilities, and makes the model more robust and universal.
Smart Images

Figure CN119993119B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech synthesis, and particularly to a cross-lingual speech synthesis system and method based on multi-encoder feature decoupling. Background Art
[0002] Speech synthesis technology, also known as text-to-speech technology, is a technology that aims to input text and make it emit human-intelligible speech. Speech synthesis technology is an important part of realizing human-machine communication and establishing a human-machine interaction system, and has been widely applied to various scenarios such as mobile phone assistants, video dubbing, and voice navigation. As a branch of speech synthesis technology, cross-lingual speech synthesis technology mainly aims to solve the problems that most current speech synthesis technologies can only achieve single-lingual speech synthesis, cannot use the voices of foreign language speakers to generate speech in another language, and cannot output corresponding speech when the text contains multiple languages. In recent years, end-to-end speech synthesis technology based on deep neural networks has become a research hotspot, and the voices synthesized by many excellent models have reached the level of being indistinguishable from real ones. However, current cross-lingual speech synthesis models still have problems such as supporting few languages and serious pitch variation in the generated audio.
[0003] In order to make the speech synthesized by the model more natural and fluent, most current methods use the speech of multiple languages as training data, introduce additional feature information into the model, and enhance the model's ability to model different languages. However, the cost of training this method is relatively high, and the effect of the synthesized audio is not good enough, and there is still a large room for improvement. Summary of the Invention
[0004] The purpose of the present invention is to provide a cross-lingual speech synthesis system and method based on multi-encoder feature decoupling to solve the problems of poor cross-lingual speech synthesis effect and high training cost in the prior art.
[0005] To achieve the above object, the present invention adopts the following technical solutions: A cross-lingual speech synthesis system based on multi-encoder feature decoupling, comprising a data collection and processing module, a cross-lingual speech synthesis model, and a model training module; the data collection and processing module is used to collect and store speech audio data and their corresponding texts in several target languages, process the texts therein to obtain phoneme sequences, and preprocess the speech audio data therein; the cross-lingual speech synthesis model consists of a text encoder, an audio encoder, a decoder, and a discriminator. The audio encoder is used to obtain an audio latent variable with the preprocessed speech audio data in the data collection and processing module as the input. The text encoder is used to convert the input text into a text latent variable in combination with the phoneme sequence in the data collection and processing module. The text latent variable is fitted according to the audio latent variable to obtain a latent variable, which is decoded by the decoder to generate an output audio. The discriminator is used to discriminate the audio generated by the decoder to verify the audio fluency; the model training module uses the data in the data collection and processing module to train the cross-lingual speech synthesis model.
[0006] Specifically, the audio encoder includes a pitch encoder, a prosody encoder, and a content encoder. The pitch encoder, prosody encoder, and content encoder are all composed of a first convolutional layer, a group normalization layer, a bidirectional long short-term memory network, and a second convolutional layer. The pitch encoder, prosody encoder, and content encoder are respectively used to output a pitch latent variable, a prosody latent variable, and a content latent variable. The pitch latent variable, prosody latent variable, and content latent variable are concatenated to obtain an audio latent variable.
[0007] Specifically, the text encoder is composed of an embedding layer, an attention block, and a convolutional layer. The embedding layer is used to convert the input text into a feature vector in combination with the phoneme sequence. The feature vector is processed by the attention block and the convolutional layer in sequence to obtain a text latent variable.
[0008] Specifically, the decoder adopts the HiFi-GAN V1 structure and is composed of multiple transposed convolutional layers connected in sequence. A multi-receptive field fusion module is set after each transposed convolutional layer.
[0009] Specifically, the model training module includes an Adam optimizer and a multi-cycle discriminator. The beta values of the Adam optimizer are 0.9 and 0.98, which are used to train the cross-lingual speech synthesis model. The multi-cycle discriminator is a discriminator with the HiFi-GAN structure and is used to calculate the loss between the audio output by the cross-lingual speech synthesis model and the real audio.
[0010] A cross-lingual speech synthesis method based on multi-encoder feature decoupling includes the following steps:
[0011] S1. Collect data. Collect the speech audio data of several target languages and their corresponding text data as the total data set. Preprocess the speech audio data and text data respectively, and divide the total data set into a training set, a test set, and a validation set.
[0012] S2. Build a model. Build a cross-lingual speech synthesis model composed of a text encoder, an audio encoder, a decoder, and a discriminator. The audio encoder is used to obtain an audio latent variable with the preprocessed speech audio data in the data collection and processing module as the input. The text encoder is used to convert the input text into a text latent variable according to the input text combined with the phoneme sequence in the data collection and processing module. The text latent variable is fitted according to the audio latent variable to obtain a latent variable and decoded by the decoder to generate an output audio. The discriminator is used to discriminate the audio generated by the decoder to verify the audio fluency.
[0013] S3. Model training. Use the training set to train the cross-lingual speech synthesis model built in step S2. During the training process, use the Adam optimizer for training. The beta values of the Adam optimizer are 0.9 and 0.98. After multiple iterative training converges, use the text data in the test set as the input of the cross-lingual speech synthesis model, and use the multi-period discriminator of the HiFi-GAN structure to calculate the loss between the predicted audio and the ground truth. After meeting the requirements, the trained cross-lingual speech synthesis model is obtained.
[0014] S4. Model testing. Select text data from the validation set and input it into the trained cross-lingual speech synthesis model in step S3. Have the people who master the corresponding language rate the output audio. If the rating meets the requirements, it is determined that the cross-lingual speech synthesis model meets the usage requirements.
[0015] Specifically, the audio encoder in step S2 includes a pitch encoder, a prosody encoder, and a content encoder. The pitch encoder, prosody encoder, and content encoder are all composed of a first convolutional layer, a group normalization layer, a bidirectional long short-term memory network, and a second convolutional layer. The pitch encoder, prosody encoder, and content encoder are respectively used to output a pitch latent variable, a prosody latent variable, and a content latent variable. The pitch latent variable, prosody latent variable, and content latent variable are concatenated to obtain an audio latent variable.
[0016] Specifically, the text encoder in step S2 is composed of an embedding layer, an attention block, and a convolutional layer. The embedding layer is used to convert the input text combined with the phoneme sequence into a feature vector. The feature vector is processed by the attention block and the convolutional layer in sequence to obtain a text latent variable.
[0017] Specifically, when preprocessing the speech audio data and text data in step S1, specifically, the speech audio data is unified: with a frame length of 1024 points, a window length of 1024 points, a frame shift of 256 points, a Mel minimum frequency of 0 Hz, and a Mel maximum frequency of 8000 Hz as parameters, the speech audio data is converted into linear spectral features with a unified sampling rate of 16000 Hz with a pre-emphasis coefficient of 0.97; the text data of all languages is converted into a phoneme sequence according to the pronunciation dictionary.
[0018] Specifically, in step S2, when the text latent variable output by the text encoder is fitted according to the audio latent variable, the similarity between the text latent variable and the audio latent variable is calculated through the KL loss function, and the text latent variable is made to fit and approach the audio latent variable to obtain the latent variable.
[0019] The beneficial effects of the present invention are as follows:
[0020] 1. By using the audio encoder of the multi-encoder, multiple audio latent variables are obtained by multi-feature decoupling of the audio data, and by splicing the multiple audio latent variables, it is convenient for the text latent variable to be fitted, thereby improving the cross-lingual speech synthesis accuracy;
[0021] 2. By using a variety of open-source single-language speech datasets, the problems of the existing methods relying on high-priced multi-language speech audio and being difficult to expand to more languages are solved. By using the open-source datasets of single languages, any language can be added to the model through a unified text and audio processing method, making the model have stronger robustness and versatility. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Attached Figure 1 is the connection schematic diagram of the cross-lingual speech synthesis model in the embodiment;
[0023] Attached Figure 2 is the connection schematic diagram of the prosody encoder / pitch encoder / content encoder in the audio encoder in the embodiment;
[0024] Attached Figure 3 is the connection schematic diagram of the audio encoder in the embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0025] Example 1, referring to Figures 1-3, A cross-lingual speech synthesis system based on multi-encoder feature decoupling, including a data collection and processing module, a cross-lingual speech synthesis model, and a model training module; the data collection and processing module is used to collect and store speech audio data and their corresponding texts in several target languages, process the texts therein to obtain phoneme sequences, and preprocess the speech audio data therein; the cross-lingual speech synthesis model consists of a text encoder, an audio encoder, a decoder, and a discriminator. The audio encoder is used to obtain an audio latent variable with the preprocessed speech audio data in the data collection and processing module as the input. The text encoder is used to convert the input text into a text latent variable in combination with the phoneme sequence in the data collection and processing module. The text latent variable is fitted according to the audio latent variable to obtain a latent variable and decoded by the decoder to generate an output audio. The discriminator is used to discriminate the audio generated by the decoder to verify the audio fluency; the model training module uses the data in the data collection and processing module to train the cross-lingual speech synthesis model.
[0026] Specifically, the audio encoder includes a pitch encoder, a prosody encoder, and a content encoder. The pitch encoder, the prosody encoder, and the content encoder are all composed of a first convolutional layer, a group normalization layer, a bidirectional long short-term memory network, and a second convolutional layer. The pitch encoder, the prosody encoder, and the content encoder are respectively used to output a pitch latent variable, a prosody latent variable, and a content latent variable. After the pitch latent variable, the prosody latent variable, and the content latent variable are concatenated, an audio latent variable is obtained. Among them, the number of layers of the first convolutional layer and the second convolutional layer of the pitch encoder is 3, the number of layers of the group normalization layer is 16, and the number of layers of the bidirectional long short-term memory network is 3; the number of layers of the first convolutional layer and the second convolutional layer of the prosody latent encoder is 1, the number of layers of the group normalization layer is 8, and the number of layers of the bidirectional long short-term memory network is 1; the number of layers of the first convolutional layer and the second convolutional layer of the content encoder is 3, the number of layers of the group normalization layer is 32, and the number of layers of the bidirectional long short-term memory network is 2.
[0027] Specifically, the text encoder consists of an embedding layer, an attention block, and a convolutional layer. The embedding layer is used to convert the input text into a feature vector in combination with the phoneme sequence. The feature vector is processed by the attention block and the convolutional layer in sequence to obtain a text latent variable.
[0028] Specifically, the decoder adopts the HiFi-GAN V1 structure and is composed of multiple transposed convolutional layers connected in sequence. A multi-receptive field fusion module is set after each transposed convolutional layer.
[0029] Specifically, the model training module includes an Adam optimizer and a multi-cycle discriminator. The beta values of the Adam optimizer are 0.9 and 0.98, which are used to train the cross-lingual speech synthesis model. The multi-cycle discriminator is a discriminator with a HiFi-GAN structure, which is used to calculate the loss between the audio output by the cross-lingual speech synthesis model and the real audio.
[0030] Based on the above cross-lingual speech synthesis system, this embodiment also provides a cross-lingual speech synthesis method based on multi-encoder feature decoupling, including the following steps:
[0031] S1. Collect data. Collect the speech audio data of several target languages and their corresponding text data as the total data set. Preprocess the speech audio data and the text data respectively, and divide the total data set into a training set, a test set, and a validation set. In this embodiment, Chinese, English, Thai, and Polish are used as target languages, and the corresponding open-source data sets are used respectively: Chinese data set AIShell3, English data sets VCTK, LibriTTS, Thai data set Common Voice, and Polish data set. The Polish data set is recorded by the applicant himself. When preprocessing the speech audio data and the text data, specifically, the speech audio data is unified: with parameters of frame length 1024 points, window length 1024 points, frame shift 256 points, Mel minimum frequency 0 Hz, and Mel maximum frequency 8000 Hz, the speech audio data is converted into linear spectral features with a unified sampling rate of 16000 Hz with a pre-emphasis coefficient of 0.97; the text data of all languages is converted into phoneme sequences according to the pronunciation dictionary.
[0032] S2. Build a model, that is, build a cross-lingual speech synthesis model composed of a text encoder, an audio encoder, a decoder, and a discriminator. The audio encoder is used to obtain an audio latent variable with the preprocessed speech audio data in the data collection and processing module as the input. The text encoder is used to convert the input text into a text latent variable according to the phoneme sequence in the data collection and processing module. The text latent variable is fitted according to the audio latent variable to obtain a latent variable, which is then decoded by the decoder to generate an output audio. The discriminator is used to discriminate the audio generated by the decoder to verify the audio fluency. The relevant structures and parameters of the above cross-lingual speech synthesis model are as recorded in the foregoing system. It should be noted that both the prosody encoder and the content encoder directly use the preprocessed audio data as the input for calculation to obtain the corresponding prosody latent variable and content latent variable. For the pitch encoder, the input audio data is the audio data obtained by performing pitch contour normalization and alignment resampling on the basis of the preprocessed audio data for calculation, so as to obtain the pitch latent variable. In addition, when the text latent variable output by the text encoder is fitted according to the audio latent variable, the similarity between the text latent variable and the audio latent variable is calculated through the KL loss function, so that the text latent variable is fitted and approximated to the audio latent variable to obtain the latent variable.
[0033] S3. Model training: Use the training set to train the cross-lingual speech synthesis model constructed in step S2. During the training process, the Adam optimizer is used for training. The beta values of the Adam optimizer are 0.9 and 0.98. After multiple iterative trainings converge, the text data in the test set is used as the input of the cross-lingual speech synthesis model, and the multi-period discriminator with the HiFi-GAN structure is used to calculate the loss between the predicted audio and the ground truth. After meeting the requirements, the trained cross-lingual speech synthesis model is obtained. Specifically, in this embodiment, when using the Adam optimizer for training, after 1 million steps of iterative training, the trained cross-lingual speech synthesis model is obtained.
[0034] S4. Model testing: Select text data from the validation set and input it into the cross-lingual speech synthesis model trained in step S3. Have people who master the corresponding language rate the output audio. If the rating meets the requirements, it is determined that the cross-lingual speech synthesis model meets the usage requirements. Specifically, in this embodiment, the cross-lingual speech synthesis model obtained in step S3 is tested as follows: Randomly select 5 - 10 audio clips of different speakers from the validation set, and then select one audio clip and mark it as the real audio. Input this real audio (which has been preprocessed) into the audio encoder of the cross-lingual speech synthesis model (i.e., the speaker of this audio is the speaker of the output audio). Then, use the cross-lingual speech synthesis model to input some pre-set texts (including different languages) and randomly generate 15 - 20 audio clips, and mix the output audio clips. Invite 10 speakers who can master the above target language as testers, and inform them of the 5-point rating standard for the quality of the speech synthesis audio: 4.0 - 5.0 is very good, clear to hear, small delay, and smooth speech; 3.5 - 4.0 is slightly poor, clear to hear, small delay, poor communication, and there is noise; 3.0 - 3.5 is okay, not very clear to hear, there is a certain delay, and can communicate; 1.5 - 3.0 is barely acceptable, not very clear to hear, large delay, and communication requires multiple repetitions; 0 - 1 is extremely poor, unable to understand, large delay, and poor communication. At the beginning of the test, first let the testers listen to the above real audio, and then listen to and rate the mixed multiple audio clips according to the above rating standard. Average the ratings of all testers for the same audio clip to get the score of this audio clip, and then average the scores of all audio clips to get the final score of this cross-lingual speech synthesis model. At the same time, to verify the effect of the cross-lingual speech synthesis model obtained in this embodiment, this embodiment also selects three other existing cross-lingual speech synthesis models: vits, GEN, SANE-TTS and rates them in the same way, and gets the following results:
[0035] model final score vits 3.55 GEN 3.5 SANE-TTS 3.44 this embodiment 3.89
[0036] Through testing, the effectiveness of the system and method of this application is verified. Compared with the current mainstream cross-lingual speech synthesis models, the scores of the present invention have obvious advantages.
[0037] Of course, the above is only the preferred embodiment of the present invention, and it does not limit the scope of use of the present invention. Therefore, all equivalent changes made on the principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A cross-language speech synthesis system based on multi-encoder feature decoupling, characterized by: It includes a data collection and processing module, a cross-language speech synthesis model and a model training module; the data collection and processing module is used to collect and store speech and audio data of several target languages and their corresponding texts, and process the texts therein to obtain phoneme sequences, and pre-process the speech and audio data therein; The cross-language speech synthesis model consists of a text encoder, an audio encoder, a decoder and a discriminator. The audio encoder is used to obtain audio latent variables by taking the speech audio data preprocessed in the data collection and processing module as input. The text encoder is used to convert the input text into text latent variables in combination with the phoneme sequence in the data collection and processing module. The text latent variables are fitted according to the audio latent variables to obtain latent variables and decoded by the decoder to generate output audio. The discriminator is used to discriminate the audio generated by the decoder to verify the audio fluency. The model training module trains the cross-language speech synthesis model using the data in the data collection and processing module; the audio encoder includes a pitch encoder, a prosody encoder and a content encoder, each of which is composed of a first convolutional layer, a group normalization layer, a bidirectional long short-term memory network and a second convolutional layer; the pitch encoder, the prosody encoder and the content encoder are respectively used to output pitch latent variables, prosody latent variables and content latent variables, and the audio latent variables are obtained by splicing the pitch latent variables, the prosody latent variables and the content latent variables; the text encoder is composed of an embedding layer, an attention block and a convolutional layer, the embedding layer is used to convert the input text into a feature vector in combination with a phoneme sequence, and the feature vector is processed by the attention block and the convolutional layer in turn to obtain the text latent variable.
2. A cross-language speech synthesis system based on multi-encoder feature decoupling according to claim 1, characterized in that: The decoder adopts the HiFi-GAN V1 structure, which is composed of multiple layers of transposed convolutional layers connected in sequence, and a multi-receptive field fusion module is arranged after each transposed convolutional layer.
3. The cross-language speech synthesis system based on multi-encoder feature decoupling according to claim 1, characterized in that: The model training module includes an Adam optimizer and a multi-cycle discriminator. The beta values of the Adam optimizer are 0.9 and 0.98, which are used to train the cross-language speech synthesis model. The multi-cycle discriminator is a discriminator of a HiFi-GAN structure, which is used to calculate the loss between the audio output by the cross-language speech synthesis model and the real audio.
4. A cross-language speech synthesis method based on multi-encoder feature decoupling, characterized in that: The steps include: S1. Collect data: collect speech and audio data of several target languages and their corresponding text data as a total data set, pre-process the speech and audio data and text data respectively, and divide the total data set into a training set, a test set, and a validation set; S2. Build a model, build a cross-language speech synthesis model composed of a text encoder, an audio encoder, a decoder and a discriminator, wherein the audio encoder is used to obtain audio latent variables by taking the pre-processed speech audio data in the data collection and processing module as input, the text encoder is used to convert the input text into text latent variables in combination with the phoneme sequence in the data collection and processing module, the text latent variables are fitted according to the audio latent variables to obtain latent variables and decoded by the decoder to generate output audio, and the discriminator is used to discriminate the audio generated by the decoder to verify the audio fluency; The audio encoder comprises a pitch encoder, a prosody encoder and a content encoder, each of which is composed of a first convolutional layer, a group normalization layer, a bidirectional long short-term memory network and a second convolutional layer, and the pitch encoder, the prosody encoder and the content encoder are used to output pitch latent variables, prosody latent variables and content latent variables respectively, and the pitch latent variables, the prosody latent variables and the content latent variables are spliced to obtain audio latent variables; The text encoder is composed of an embedding layer, an attention block and a convolution layer. The embedding layer is used to convert the input text into a feature vector in combination with a phoneme sequence. The feature vector is processed by the attention block and the convolution layer in turn to obtain a text latent variable. S3, model training, using the training set to train the cross-language speech synthesis model constructed in step S2, using the Adam optimizer for training during the training process, the beta value of the Adam optimizer is 0.9 and 0.98, after multiple iterations of training convergence, using the text data in the test set as the input of the cross-language speech synthesis model, using the multi-cycle discriminator of the HiFi-GAN structure to calculate the loss between the predicted audio and the true value, and obtaining a trained cross-language speech synthesis model if it meets the requirements; S4, model testing, selecting text data from the validation set and inputting it into the cross-language speech synthesis model trained in step S3, and having people who master the corresponding language score the output audio. If the score meets the requirements, it is judged that the cross-language speech synthesis model meets the usage requirements.
5. The cross-language speech synthesis method based on multi-encoder feature decoupling according to claim 4, characterized in that: When the speech audio data and text data are preprocessed in step S1, specifically, the speech audio data is uniformly processed: with a frame length of 1024 points, a window length of 1024 points, a frame shift of 256 points, a minimum Mel frequency of 0 Hz, and a maximum Mel frequency of 8000 Hz as parameters, the speech audio data is converted into a linear spectrum feature with a uniform sampling rate of 16000 Hz with a pre-emphasis coefficient of 0.97; Convert text data in all languages into phoneme sequences based on pronunciation dictionaries.
6. The cross-language speech synthesis method based on multi-encoder feature decoupling according to claim 4, characterized in that: In step S2, when the text latent variables output by the text encoder are fitted according to the audio latent variables, the similarity between the text latent variables and the audio latent variables is calculated through the KL loss function, so that the text latent variables are fitted close to the audio latent variables to obtain the latent variables.
Citation Information
Patent Citations
Voice synthesis method based on MOOC voice data set
CN113539232A
Speech synthesis model method capable of synthesizing multi-emotion audio
CN116798403A