Speech synthesis method and device, terminal equipment and storage medium

The number of decoding layers of the student's speech synthesis model is adjusted through distillation technology, which solves the response problem of the existing speech synthesis model in scenarios with high real-time requirements, and achieves efficient speech synthesis and improves user experience.

CN119993120AActive Publication Date: 2025-05-13GUANGDONG KAMFU TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411937534.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-05-13
Estimated Expiration
2044-12-26

AI Technical Summary

Technical Problem

In application scenarios with high real-time requirements, the existing voice synthesis model cannot respond to users' immediate instructions in a timely manner, resulting in poor user experience.

Method used

By obtaining the preset teacher speech synthesis model and student speech synthesis model, the number of decoding layers of the student model is adjusted using distillation technology until the relative entropy between the student model output distribution and the teacher model output distribution reaches the threshold, thereby generating the target speech synthesis model.

Benefits of technology

While ensuring the model expression ability and generation quality, the inference speed of the speech synthesis model is improved, meeting the real-time needs of users in certain scenarios with high real-time requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993120A_ABST
    Figure CN119993120A_ABST
Patent Text Reader

Abstract

The invention discloses a voice synthesis method and device, equipment and a storage medium, and the method comprises the steps: obtaining to-be-processed voice and text data, and inputting the to-be-processed voice and text data into a target voice synthesis model, so as to obtain target synthesis voice data corresponding to acoustic features and corresponding contents; wherein the generation of the target speech synthesis model is that in the model training process, the model parameters of the student speech synthesis model are adjusted by obtaining and according to the relative entropy between the output distributions of the teacher speech synthesis model and the student speech synthesis model, and the target speech synthesis model is generated when the relative entropy reaches a first preset threshold value. The corresponding student speech synthesis model is used as a target speech synthesis model. According to the invention, the expression ability and the generation quality of the target speech synthesis model can be ensured, and the real-time requirement of the user can be met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech synthesis, and in particular to a speech synthesis method, device, equipment and storage medium. Background Art

[0002] With the rapid development of artificial intelligence technology, speech synthesis technology has been widely used in many fields. For existing speech synthesis models, such as the cosyvoice model, the cosyvoice model supports multiple languages, speech sentiment analysis, event recognition, and cross-language synthesis. In order to ensure the expressiveness and generation quality of the model, more decoding layers are required in the model to process and optimize the data. This means that during the inference process of the model, each decoding layer will take up a lot of computing time. For some application scenarios with high real-time requirements, the user's immediate instructions cannot be responded to in time, resulting in unsmooth conversations or feedback between users, which in turn affects the user experience. Therefore, how to meet the user's real-time needs while ensuring the expressiveness and generation quality of the model in certain specific scenarios with high real-time requirements is a technical problem that needs to be solved urgently. Summary of the invention

[0003] The present invention provides a speech synthesis method, apparatus, device and storage medium, which can ensure the expressiveness and generation quality of the model while meeting the real-time requirements of users in certain specific scenarios with high real-time requirements.

[0004] The present invention provides a speech synthesis method, comprising: obtaining speech data to be processed and text data to be processed corresponding to a target language;

[0005] Inputting the speech data to be processed and the text data to be processed into a target speech synthesis model, so that the target speech synthesis model generates target synthesized speech data having the same acoustic features as the speech data to be processed and the same speech content as the text data to be processed according to the speech data to be processed and the text data to be processed;

[0006] The generation of the target speech synthesis model includes:

[0007] Obtain a preset teacher speech synthesis model, a preset student speech synthesis model, a speech data set to be trained corresponding to the target language, and a text data set to be trained; wherein the number of decoding layers in the teacher speech synthesis model is greater than the number of decoding layers in the student speech synthesis model;

[0008] Repeat the following steps until the relative entropy between the output distribution of the student model and the output distribution of the teacher model reaches a first preset threshold, and when the relative entropy reaches the first preset threshold, the corresponding student speech synthesis model is used as the target speech synthesis model;

[0009] Randomly select a speech data set to be trained and a text data set to be trained from the current speech data set to be trained and the text data set to be trained respectively;

[0010] Input the speech data to be trained and the text data to be trained into the teacher speech synthesis model and the student speech synthesis model, respectively, to obtain the teacher model output distribution of the teacher speech synthesis model and the student model output distribution of the student speech synthesis model, respectively;

[0011] Obtaining the relative entropy between the teacher model output distribution and the student model output distribution;

[0012] When it is determined that the relative entropy does not reach the first preset threshold, the model parameters of the student speech synthesis model are adjusted according to the relative entropy, and the current speech data to be trained and the text data to be trained are respectively removed from the speech data set to be trained and the text data set to be trained.

[0013] Furthermore, the target speech synthesis model includes: an input layer, a decoding layer set, a stream matching layer and an output layer;

[0014] The input layer is used to collect data features of the voice data to be processed and the text data to be processed;

[0015] The decoding layer set is used to decode the data features of the collected voice data to be processed and the collected text data to be processed, so as to fuse the data features of the voice data to be processed and the text data to be processed to generate corresponding target voice features;

[0016] The stream matching layer is used to generate a Mel-spectrogram corresponding to the target speech feature according to the target speech feature;

[0017] The output layer is used to convert the Mel-spectrogram into target synthetic speech data having a sound wave signal.

[0018] Furthermore, the input layer includes: a text encoder and a speech segmenter;

[0019] The text encoder is used to extract text semantic features of the text data to be processed;

[0020] The speech segmenter is used to extract acoustic features of speech data to be processed; wherein the acoustic features include: pitch, timbre, volume and speech speed.

[0021] Furthermore, the number of decoding layers of the student speech synthesis model is 6.

[0022] Furthermore, the relative entropy includes: inverse KL divergence.

[0023] Based on the above method embodiment, the present invention provides a corresponding device embodiment;

[0024] The present invention provides a speech synthesis method and device, comprising: a data acquisition module and a speech synthesis module;

[0025] The data acquisition module is used to acquire the speech data to be processed and the text data to be processed corresponding to the target language;

[0026] The speech synthesis module is used to input the speech data to be processed and the text data to be processed into a target speech synthesis model, so that the target speech synthesis model generates target synthesized speech data that is consistent with the acoustic features of the speech data to be processed and has speech content consistent with the content of the text data to be processed according to the speech data to be processed and the text data to be processed;

[0027] The generation of the target speech synthesis model includes:

[0028] Obtain a preset teacher speech synthesis model, a preset student speech synthesis model, a speech data set to be trained corresponding to the target language, and a text data set to be trained; wherein the number of decoding layers in the teacher speech synthesis model is greater than the number of decoding layers in the student speech synthesis model;

[0029] Repeat the following steps until the relative entropy between the output distribution of the student model and the output distribution of the teacher model reaches a first preset threshold, and when the relative entropy reaches the first preset threshold, the corresponding student speech synthesis model is used as the target speech synthesis model;

[0030] Randomly select a speech data set to be trained and a text data set to be trained from the current speech data set to be trained and the text data set to be trained respectively;

[0031] Input the speech data to be trained and the text data to be trained into the teacher speech synthesis model and the student speech synthesis model, respectively, to obtain the teacher model output distribution of the teacher speech synthesis model and the student model output distribution of the student speech synthesis model, respectively;

[0032] Obtaining the relative entropy between the teacher model output distribution and the student model output distribution;

[0033] When it is determined that the relative entropy does not reach the first preset threshold, the model parameters of the student speech synthesis model are adjusted according to the relative entropy, and the current speech data to be trained and the text data to be trained are respectively removed from the speech data set to be trained and the text data set to be trained.

[0034] Furthermore, the target speech synthesis model includes: an input layer, a decoding layer set, a stream matching layer and an output layer;

[0035] The input layer is used to collect data features of the voice data to be processed and the text data to be processed;

[0036] The decoding layer set is used to decode the data features of the collected voice data to be processed and the collected text data to be processed, so as to fuse the data features of the voice data to be processed and the text data to be processed to generate corresponding target voice features;

[0037] The stream matching layer is used to generate a Mel-spectrogram corresponding to the target speech feature according to the target speech feature;

[0038] The output layer is used to convert the Mel-spectrogram into target synthetic speech data having a sound wave signal.

[0039] Furthermore, the input layer includes: a text encoder and a speech segmenter;

[0040] The text encoder is used to extract text semantic features of the text data to be processed;

[0041] The speech segmenter is used to extract acoustic features of speech data to be processed; wherein the acoustic features include: pitch, timbre, volume and speech speed.

[0042] Based on the above method embodiment, the present invention provides a corresponding device embodiment;

[0043] The present invention provides a device, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor implements any one of the speech synthesis methods of the present invention when executing the computer program.

[0044] Based on the above method embodiment, the present invention provides a storage medium embodiment;

[0045] The present invention provides a storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the storage medium is located is controlled to execute any one of the speech synthesis methods described in the present invention.

[0046] The embodiments of the present invention have the following beneficial effects:

[0047] The present invention provides a speech synthesis method, apparatus, device and storage medium; the method, by respectively inputting the speech data to be trained and the text data to be processed corresponding to the target language into the teacher speech synthesis model and the student speech synthesis model, to obtain the corresponding teacher model output distribution and the student model output distribution, and then adjusting the model parameters of the student speech synthesis model according to the relative entropy between the teacher model output distribution and the student model output distribution, when the relative entropy reaches a first preset threshold, the student model output distribution can be made to be similar to the teacher model output distribution, so that the student speech synthesis model has achieved almost the same effect as the teacher speech synthesis model in processing the target language data, and then the student speech synthesis model after adjusting the model parameters can be used as the target speech synthesis model. After the speech data to be processed and the text data to be processed are input into the target speech synthesis model, the target synthesized speech data can be generated by the target speech synthesis model. By implementing the present invention, the target speech synthesis model can achieve the same effect as the existing teacher speech synthesis model with a more complex decoding layer structure when processing speech data corresponding to the target language. In addition, since the target speech synthesis model is only for the target language in a specific scenario, when setting the student speech synthesis model, fewer decoding layers are required, making the data processing of the model more targeted and efficient, thereby improving the reasoning speed of the model. In summary, the target speech synthesis model of the present invention not only retains the processing ability of the teacher speech synthesis model for the target language data, but also speeds up the reasoning speed of the target language data. It can ensure the expressiveness and generation quality of the model in certain specific scenarios with high real-time requirements, while also meeting the real-time needs of users. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 It is a flowchart of a speech synthesis method provided by an embodiment of the present invention;

[0049] Figure 2 is a schematic diagram of the structure of a target speech synthesis model provided by an embodiment of the present invention;

[0050] Figure 3 It is a schematic diagram of a decoding layer simplification method of a student speech synthesis model provided by an embodiment of the present invention;

[0051] Figure 4 It is a schematic diagram of a distillation learning process provided by an embodiment of the present invention;

[0052] Figure 5 It is a structural schematic diagram of a speech synthesis method and device provided by one embodiment of the present invention. DETAILED DESCRIPTION

[0053] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0054] like Figure 1 As shown, an embodiment provides a speech synthesis method, comprising:

[0055] Step S101: Acquire speech data and text data to be processed corresponding to the target language;

[0056] Step S102: inputting the speech data to be processed and the text data to be processed into a target speech synthesis model, so that the target speech synthesis model generates target synthesized speech data having the same acoustic features as the speech data to be processed and the same speech content as the text data to be processed according to the speech data to be processed and the text data to be processed;

[0057] The generation of the target speech synthesis model includes:

[0058] Obtain a preset teacher speech synthesis model, a preset student speech synthesis model, a speech data set to be trained corresponding to the target language, and a text data set to be trained; wherein the number of decoding layers in the teacher speech synthesis model is greater than the number of decoding layers in the student speech synthesis model;

[0059] Repeat the following steps until the relative entropy between the output distribution of the student model and the output distribution of the teacher model reaches a first preset threshold, and when the relative entropy reaches the first preset threshold, the corresponding student speech synthesis model is used as the target speech synthesis model;

[0060] Randomly select a speech data set to be trained and a text data set to be trained from the current speech data set to be trained and the text data set to be trained respectively;

[0061] Input the speech data to be trained and the text data to be trained into the teacher speech synthesis model and the student speech synthesis model, respectively, to obtain the teacher model output distribution of the teacher speech synthesis model and the student model output distribution of the student speech synthesis model, respectively;

[0062] Obtaining the relative entropy between the teacher model output distribution and the student model output distribution;

[0063] When it is determined that the relative entropy does not reach the first preset threshold, the model parameters of the student speech synthesis model are adjusted according to the relative entropy, and the current speech data to be trained and the text data to be trained are respectively removed from the speech data set to be trained and the text data set to be trained.

[0064] For step S101, in a preferred embodiment, it is first necessary to obtain the corresponding speech data to be processed and text data to be processed of the target language; the target language can be any language with a separate language and characters; the speech data to be processed refers to speech data with a specific speaker style; the text data to be processed refers to the text data corresponding to the target language.

[0065] Taking Chinese as the target language as an example, the speech data to be processed refers to speech data with a specific speaker style and pronounced in Chinese; then the text data to be processed refers to Chinese text data.

[0066] For step S102, in a preferred embodiment, the speech data to be processed and the text data to be processed are input into a trained target speech synthesis model, and the target speech synthesis model can generate a piece of target synthesized speech data that is consistent with the acoustic features of the speech data to be processed and the speech content is consistent with the content of the text data to be processed based on the acoustic features of the speech data to be processed and the content of the text data to be processed.

[0067] Taking Chinese as the target language as an example, after the target speech synthesis model obtains the to-be-processed speech data "go eat" of user A and the to-be-processed text data "go travel", the target speech synthesis model can extract the corresponding acoustic features from the to-be-processed speech data "go eat" of user A, including pitch, timbre, volume, and speaking speed, etc., and then combine the extracted acoustic features with the to-be-processed text data "go travel" according to the content of the to-be-processed text data "go travel" to generate a target speech synthesis data that reads "go travel" with the acoustic features of user A.

[0068] For the generation of the target speech synthesis model, first obtain a preset teacher speech synthesis model, which can be a cosyvoice model or any model in the prior art that can achieve the same function; then obtain a preset student speech synthesis model. Compared with the structure of the teacher speech synthesis model, the student speech synthesis model has fewer decoding layers because the student speech synthesis model is only targeted at specific application scenarios or target languages ​​(such as Mandarin speech synthesis tasks). Too many decoding layers will only make the parameters of the student speech synthesis model redundant and increase latency; finally, obtain the speech data set to be trained and the text data set to be trained for model training. Both data sets belong to the same target language type (such as Chinese).

[0069] Then it enters into a cyclic iterative process, which uses the teacher speech synthesis model as a "guide" to continuously train and optimize the student speech synthesis model until its performance in processing the speech data set to be trained and the text data set to be trained is close to that of the teacher speech synthesis model. The specific method is to first randomly select a speech data to be trained and a text data to be trained from the current speech data set to be trained and the text data set to be trained, respectively, and then input these two data into the teacher speech synthesis model and the student speech synthesis model, respectively. This step is equivalent to inputting the same data into the teacher speech synthesis model and the student speech synthesis model, and then use this to obtain the teacher model output distribution of the teacher speech synthesis model and the student model output distribution of the student speech synthesis model, respectively. In order to verify whether the teacher model output distribution and the student model output distribution are approximately equal, the method of the present invention is to calculate the output distribution of the teacher model and the output distribution of the student model. The relative entropy is determined by presetting a first preset threshold, and the calculated relative entropy is compared with the first preset threshold. When it is determined that the relative entropy has not reached the first preset threshold, it means that the output distribution of the teacher model and the output distribution of the student model have not yet reached the required approximately equal condition, and it also means that under certain conditions, the student speech synthesis model has not yet achieved a capability similar to that of the teacher model. At this time, it is necessary to adjust the model parameters of the student speech synthesis model, and then the current speech data to be trained and the text data to be trained that have been used for training are removed from the speech data set to be trained and the text data set to be trained, respectively, to prevent invalid repeated training.

[0070] It should be noted that the first preset threshold can be set according to actual needs, and the specific calculation method of relative entropy also belongs to the prior art, and the present invention will not be specifically limited; the method for adjusting the model parameters of the student speech synthesis model of the present invention may be adjusting the weights and biases in the decoding layer, etc., and the specific adjustment method may also belong to any similar adjustment method in the prior art, and the present invention will not make any specific limitation.

[0071] In a preferred embodiment, as shown in 2, the target speech synthesis model mainly consists of several parts, including: an input layer, a decoding layer set, a stream matching layer and an output layer;

[0072] The input layer refers to the part of the model used to obtain the speech data to be processed and the text data to be processed and collect relevant data features. The decoding layer set includes several decoding layers. The main function of the decoding layer set is to receive the data features from the input layer and start decoding, that is, to decode the relevant speech and text features to effectively fuse them together and generate a target speech feature that integrates speech and text information. The main function of the stream matching layer is to convert the target speech features into a format that is easy to understand and process, that is, the Mel spectrogram. The Mel spectrogram is a spectrum representation method based on the Mel scale, which can better reflect the characteristics of the human auditory system. By generating the Mel spectrogram, the stream matching layer provides a more intuitive and easy-to-operate data form for subsequent speech synthesis. The output layer is mainly responsible for converting the Mel spectrogram into target synthetic speech data with sound wave signals. The method can be to use inverse transformation and signal processing technology to restore the Mel spectrogram to an actual sound wave signal. These sound wave signals are further processed and optimized to finally generate target synthetic speech data with high quality and realistic effects. The specific target synthetic speech data generation method can also be any method with the same function in the prior art, and the present invention is not specifically limited.

[0073] In a preferred embodiment, the input layer of the target speech synthesis model consists of two parts, including: a text encoder and a speech segmenter;

[0074] The text encoder is mainly responsible for parsing and processing text data. Its main function is to extract key text semantic features from the input text data. These features can capture and express the meaning, context, and relationship between words in the text, thereby providing the necessary semantic information for the subsequent speech synthesis process.

[0075] The speech segmenter is mainly responsible for analyzing and processing speech data. Its main function is to extract acoustic features from the speech signal to be processed. These features determine the natural fluency of the target synthesized speech data that is finally generated. Specifically, acoustic features cover multiple aspects, including: pitch, timbre, volume, and speaking speed; pitch refers to the high and low changes in the sound, which plays a decisive role in expressing different emotions, intonations, and changes in tone in language. Timbre refers to the quality or characteristics of the sound, such as sweetness, depth, etc., which mainly affects the user's feeling of hearing the sound. Volume refers to the size or strength of the sound, reflecting the loudness of the sound, and has an important impact on the clarity and audibility of the speech. Speaking speed refers to the speed of speaking, that is, the number of syllables uttered per unit time, which affects the rhythm and fluency of the speech.

[0076] In a preferred embodiment, the number of decoding layers of the student speech synthesis model in the present invention is 6 layers, that is, the decoding layer set in the target speech synthesis model has 6 decoding layers.

[0077] Optionally, the present invention adopts the existing CosyVoice model as the teacher speech synthesis model. The CosyVoice model can generate speech when the real-time factor (RTF) is less than 1, but it still takes 3 to 4 seconds to generate short-term audio (such as 5 seconds). This will have a certain impact on the user experience in some application scenarios with high real-time requirements. In addition, the CosyVoice model supports multiple languages, speech sentiment analysis, event recognition, and cross-language synthesis, which is too redundant in some specific scenarios.

[0078] To address this problem, the present invention adopts a distillation method for speech synthesis models, which aims to simplify the model structure and improve the reasoning speed while retaining the core functions of the original model, so as to meet the higher requirements for real-time performance in specific scenarios. The present invention mainly targets scenarios where only Chinese speech cloning is required. Through distillation technology, the CosyVoice model is optimized, so that it has faster reasoning capabilities while maintaining the original speech synthesis quality.

[0079] Therefore, in the design of the student speech synthesis model structure, the present invention performs refined compression and simplification on the basis of the teacher speech synthesis model structure, performs deep optimization through the decoding layer of the model, reduces the original 14-layer structure of the cosyvoice model to 6 layers, and removes the middle 8 layers. The reason for retaining the front and back 3 decoding layers is that these 6 layers play a key role in the encoding and decoding process. The front 3 decoding layers retained are crucial for receiving and processing signals from the text encoder and speech segmenter, and they ensure the accurate transmission of information. The last 3 decoding layers play a key role in converting feature vectors into standard audio signal formats. Such adjustments not only ensure the effectiveness of the model, but also improve its operating efficiency.

[0080] The reason for reducing the decoding layer of the CosyVoice model is that the CosyVoice model performs well in the field of multilingual speech synthesis and has versatility such as sentiment analysis and voice event recognition. However, when the CosyVoice model focuses on a single Chinese speech synthesis application, the decoding layer accounts for 90% of the model parameters and consumes 80% of the time during inference. Therefore, it is crucial to reduce the decoding layer.

[0081] Optionally, compared with the teacher speech synthesis model, the structure of the simplified student speech synthesis model is as shown in 3. When simplifying the decoding layer structure, it is necessary to retain the front and back decoding layers and simplify the middle part.

[0082] The role of the text encoder in the figure is to extract the semantic features of the text, while the speech segmenter focuses on capturing the timbre features (acoustic features) of the speech. On this basis, this process fuses the extracted text semantic features with the timbre features to form a new token sequence. The sequence is then input into multiple decoding layers for decoding, and then the corresponding speech token sequence is generated through the text encoder. Subsequently, the stream matching model is used to convert the resulting speech token sequence into a mel-spectrogram. Finally, the sound wave synthesis is completed to achieve speech cloning.

[0083] Optionally, during the distillation learning process, except for the decoding layer, the parameters of other modules in the student speech synthesis model remain fixed. This strategy was chosen based on the following two points: first, the number of parameters of other modules is relatively small, so there is no need to compress them; second, the parameters of these modules have reached a high quality standard after training with a large amount of data. If these parameters are allowed to adjust during the distillation process, it may lead to a decrease in model performance. Therefore, fixing the parameters of these modules not only helps to reduce the number of parameters in the training process, thereby saving computing resources, but also effectively shortens the training time.

[0084] By implementing the distillation strategy of the present invention to optimize the student speech synthesis model, the model's generalization ability for speech synthesis of non-target languages ​​is inevitably reduced. However, while ensuring that the target speech synthesis quality remains unchanged, the size of the student speech synthesis model is successfully reduced to 40% of the original volume compared to the size of the teacher speech synthesis model. At the same time, the processing speed of speech synthesis is doubled, significantly improving the computational efficiency of the student speech synthesis model.

[0085] In a preferred embodiment, the present invention calculates the relative entropy between the teacher model output distribution and the student model output distribution, which means calculating the inverse KL divergence between the teacher model output distribution and the student model output distribution. The specific calculation method belongs to the prior art and is not specifically limited by the present invention.

[0086] Alternatively, if Figure 4 As shown, the present invention adopts inverse KL divergence to achieve model distillation by minimizing the KL divergence between the output distribution q of the student speech synthesis model and the output distribution p of the teacher speech synthesis model.

[0087] In the knowledge distillation process, the forward KL divergence (denoted as KL(p||q)) measures the difference between the teacher model output distribution p and the student model output distribution q, and promotes the student model output distribution q to fully cover the important feature areas of the teacher model output distribution p. However, in the actual application scenarios of large-scale language models, due to the huge output space and the complexity of the distribution pattern, the forward KL divergence may cause the student speech synthesis model to pay too much attention to samples that have a very low probability of appearing in actual applications, thereby affecting the overall performance of the final target speech synthesis model.

[0088] In contrast, the inverse KL divergence (denoted as KL(q||p)) of the present invention exhibits a different mechanism of action in knowledge distillation. It forces the output distribution q of the student model to focus on fitting the main distribution mode of the output distribution p of the teacher model. Especially when facing multi-peak distribution, this strategy helps the student speech synthesis model to capture key distribution features more effectively. Therefore, in the distillation learning of large language models, the application of inverse KL divergence helps to improve the learning efficiency and generalization ability of the student speech synthesis model.

[0089] Therefore, for distillation learning of large models, the choice of divergence measurement method is crucial. Inverse KL divergence shows more outstanding performance in processing complex output spaces and multi-modal distribution tasks, while the advantage of forward KL divergence in traditional classification tasks mainly comes from its good adaptability to small-scale output spaces and single distribution patterns. In the distillation learning training process, the present invention uses inverse KL divergence to replace forward KL divergence. This strategy significantly improves training efficiency and shortens the required time.

[0090] The implementation of the above embodiments of the present invention has the following effects:

[0091] 1. The present invention uses distillation training to enable the student speech synthesis model to achieve almost the same effect as the teacher speech synthesis model in processing the target language data. In addition, since the target speech synthesis model is only for the target language in a specific scenario, when setting the student speech synthesis model, fewer decoding layers are required, making the model's data processing more targeted and efficient, thereby improving the model's reasoning speed. Therefore, the target speech synthesis model of the present invention not only retains the teacher speech synthesis model's ability to process the target language data, but also speeds up the reasoning speed of the target language data. In certain specific scenarios with high real-time requirements, it can ensure the model's expressiveness and generation quality while meeting the user's real-time requirements.

[0092] 2. During the distillation training process, the present invention uses the reverse KL divergence (Reverse KL) to replace the forward KL divergence (Forward KL). Reverse KL divergence shows more outstanding performance when dealing with data with complex output space and multimodal distribution. In comparison, the advantage of forward KL divergence in traditional classification problems is mainly reflected in its good adaptability to small-scale output space and unimodal distribution. In the distillation training process of the present invention, reverse KL divergence is adopted as our divergence metric, replacing the traditional forward KL divergence, which significantly improves the training efficiency and shortens the required time investment, thereby optimizing the entire distillation learning process.

[0093] Based on the above method embodiment, the present invention provides a corresponding device embodiment.

[0094] like Figure 5 As shown, an embodiment of the present invention provides a speech synthesis method and device, including: a data acquisition module and a speech synthesis module;

[0095] The data acquisition module is used to acquire the speech data to be processed and the text data to be processed corresponding to the target language;

[0096] The speech synthesis module is used to input the speech data to be processed and the text data to be processed into a target speech synthesis model, so that the target speech synthesis model generates target synthesized speech data that is consistent with the acoustic features of the speech data to be processed and has speech content consistent with the content of the text data to be processed according to the speech data to be processed and the text data to be processed;

[0097] The generation of the target speech synthesis model includes:

[0098] Obtain a preset teacher speech synthesis model, a preset student speech synthesis model, a speech data set to be trained corresponding to the target language, and a text data set to be trained; wherein the number of decoding layers in the teacher speech synthesis model is greater than the number of decoding layers in the student speech synthesis model;

[0099] Repeat the following steps until the relative entropy between the output distribution of the student model and the output distribution of the teacher model reaches a first preset threshold, and when the relative entropy reaches the first preset threshold, the corresponding student speech synthesis model is used as the target speech synthesis model;

[0100] Randomly select a speech data set to be trained and a text data set to be trained from the current speech data set to be trained and the text data set to be trained respectively;

[0101] Input the speech data to be trained and the text data to be trained into the teacher speech synthesis model and the student speech synthesis model, respectively, to obtain the teacher model output distribution of the teacher speech synthesis model and the student model output distribution of the student speech synthesis model, respectively;

[0102] Obtaining the relative entropy between the teacher model output distribution and the student model output distribution;

[0103] When it is determined that the relative entropy does not reach the first preset threshold, the model parameters of the student speech synthesis model are adjusted according to the relative entropy, and the current speech data to be trained and the text data to be trained are respectively removed from the speech data set to be trained and the text data set to be trained.

[0104] In a preferred embodiment, the target speech synthesis model comprises: an input layer, a decoding layer set, a stream matching layer and an output layer;

[0105] The input layer is used to collect data features of the voice data to be processed and the text data to be processed;

[0106] The decoding layer set is used to decode the data features of the collected voice data to be processed and the collected text data to be processed, so as to fuse the data features of the voice data to be processed and the text data to be processed to generate corresponding target voice features;

[0107] The stream matching layer is used to generate a Mel-spectrogram corresponding to the target speech feature according to the target speech feature;

[0108] The output layer is used to convert the Mel-spectrogram into target synthetic speech data having a sound wave signal.

[0109] In a preferred embodiment, the input layer includes: a text encoder and a speech segmenter;

[0110] The text encoder is used to extract text semantic features of the text data to be processed;

[0111] The speech segmenter is used to extract acoustic features of speech data to be processed; wherein the acoustic features include: pitch, timbre, volume and speech speed.

[0112] It should be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the accompanying drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. A person of ordinary skill in the art may understand and implement it without paying any creative effort.

[0113] Those skilled in the art can clearly understand that, for the sake of convenience and brevity, the specific working process of the device described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0114] Based on the above method item embodiments, the present invention provides corresponding device item embodiments.

[0115] Another embodiment of the present invention provides a device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor; when the processor executes the computer program, the speech synthesis method of any embodiment of the present invention is implemented.

[0116] Exemplarily, in this embodiment, the computer program may be divided into one or more modules, and the one or more modules are stored in the memory and executed by the processor to complete the present invention. The one or more module elements may be a series of computer program instruction segments that can complete specific functions, and the instruction segments are used to describe the execution process of the computer program in the device;

[0117] The device may be a computing device such as a desktop computer, a notebook, a PDA, a cloud server, etc. The device may include, but is not limited to, a processor, a memory;

[0118] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the device, and uses various interfaces and lines to connect various parts of the entire device;

[0119] The memory can be used to store the computer program and / or module, and the processor realizes various functions of the device by running or executing the computer program and / or module stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; in addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Med iaCard, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0120] Based on the above method item embodiments, the present invention provides a corresponding storage medium item embodiment.

[0121] Another embodiment of the present invention provides a storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the storage medium is located is controlled to execute the speech synthesis method of any embodiment of the present invention.

[0122] In this embodiment, the storage medium is a computer-readable storage medium, and the computer program includes computer program code, which may be in source code form, object code form, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc.

[0123] By implementing the above-mentioned embodiments of the present invention, it is possible to ensure the expressiveness and generation quality of the target speech synthesis model while also meeting the real-time requirements of users.

[0124] The above is a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several improvements and modifications without departing from the principle of the present invention. These improvements and modifications are also regarded as protection of the present invention.

Claims

1. A speech synthesis method, characterized in that: include: Obtaining the speech data and text data to be processed corresponding to the target language; Inputting the speech data to be processed and the text data to be processed into a target speech synthesis model, so that the target speech synthesis model generates target synthesized speech data having the same acoustic features as the speech data to be processed and the same speech content as the text data to be processed according to the speech data to be processed and the text data to be processed; The generation of the target speech synthesis model includes: Obtain a preset teacher speech synthesis model, a preset student speech synthesis model, a speech data set to be trained corresponding to the target language, and a text data set to be trained; wherein the number of decoding layers in the teacher speech synthesis model is greater than the number of decoding layers in the student speech synthesis model; Repeat the following steps until the relative entropy between the output distribution of the student model and the output distribution of the teacher model reaches a first preset threshold, and when the relative entropy reaches the first preset threshold, the corresponding student speech synthesis model is used as the target speech synthesis model; Randomly select a speech data set to be trained and a text data set to be trained from the current speech data set to be trained and the text data set to be trained respectively; Input the speech data to be trained and the text data to be trained into the teacher speech synthesis model and the student speech synthesis model, respectively, to obtain the teacher model output distribution of the teacher speech synthesis model and the student model output distribution of the student speech synthesis model, respectively; Obtaining the relative entropy between the teacher model output distribution and the student model output distribution; When it is determined that the relative entropy does not reach the first preset threshold, the model parameters of the student speech synthesis model are adjusted according to the relative entropy, and the current speech data to be trained and the text data to be trained are respectively removed from the speech data set to be trained and the text data set to be trained.

2. A speech synthesis method according to claim 1, characterized in that: The target speech synthesis model includes: an input layer, a decoding layer set, a stream matching layer and an output layer; The input layer is used to collect data features of the voice data to be processed and the text data to be processed; The decoding layer set is used to decode the data features of the collected voice data to be processed and the collected text data to be processed, so as to fuse the data features of the voice data to be processed and the text data to be processed to generate corresponding target voice features; The stream matching layer is used to generate a Mel-spectrogram corresponding to the target speech feature according to the target speech feature; The output layer is used to convert the Mel-spectrogram into target synthetic speech data having a sound wave signal.

3. A speech synthesis method according to claim 2, characterized in that: The input layer includes: a text encoder and a speech segmenter; The text encoder is used to extract text semantic features of the text data to be processed; The speech segmenter is used to extract acoustic features of speech data to be processed; wherein the acoustic features include: pitch, timbre, volume and speech speed.

4. A speech synthesis method according to claim 3, characterized in that: The student speech synthesis model has 6 decoding layers.

5. A speech synthesis method according to claim 4, characterized in that: The relative entropy includes: inverse KL divergence.

6. A speech synthesis method and device, characterized in that: include: Data acquisition module and speech synthesis module; The data acquisition module is used to acquire the speech data to be processed and the text data to be processed corresponding to the target language; The speech synthesis module is used to input the speech data to be processed and the text data to be processed into a target speech synthesis model, so that the target speech synthesis model generates target synthesized speech data that is consistent with the acoustic features of the speech data to be processed and has speech content consistent with the content of the text data to be processed according to the speech data to be processed and the text data to be processed; The generation of the target speech synthesis model includes: Obtain a preset teacher speech synthesis model, a preset student speech synthesis model, a speech data set to be trained corresponding to the target language, and a text data set to be trained; wherein the number of decoding layers in the teacher speech synthesis model is greater than the number of decoding layers in the student speech synthesis model; Repeat the following steps until the relative entropy between the output distribution of the student model and the output distribution of the teacher model reaches a first preset threshold, and when the relative entropy reaches the first preset threshold, the corresponding student speech synthesis model is used as the target speech synthesis model; Randomly select a speech data set to be trained and a text data set to be trained from the current speech data set to be trained and the text data set to be trained respectively; Input the speech data to be trained and the text data to be trained into the teacher speech synthesis model and the student speech synthesis model, respectively, to obtain the teacher model output distribution of the teacher speech synthesis model and the student model output distribution of the student speech synthesis model, respectively; Obtaining the relative entropy between the teacher model output distribution and the student model output distribution; When it is determined that the relative entropy does not reach the first preset threshold, the model parameters of the student speech synthesis model are adjusted according to the relative entropy, and the current speech data to be trained and the text data to be trained are respectively removed from the speech data set to be trained and the text data set to be trained.

7. The speech synthesis method and apparatus according to claim 6, wherein: The target speech synthesis model includes: an input layer, a decoding layer set, a stream matching layer and an output layer; The input layer is used to collect data features of the voice data to be processed and the text data to be processed; The decoding layer set is used to decode the data features of the collected voice data to be processed and the collected text data to be processed, so as to fuse the data features of the voice data to be processed and the text data to be processed to generate corresponding target voice features; The stream matching layer is used to generate a Mel-spectrogram corresponding to the target speech feature according to the target speech feature; The output layer is used to convert the Mel-spectrogram into target synthetic speech data having a sound wave signal.

8. The speech synthesis method and device according to claim 7, characterized in that: The input layer includes: a text encoder and a speech segmenter; The text encoder is used to extract text semantic features of the text data to be processed; The speech segmenter is used to extract acoustic features of speech data to be processed; wherein the acoustic features include: pitch, timbre, volume and speech speed.

9. A terminal device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor implements the speech synthesis method according to any one of claims 1 to 5 when executing the computer program.

10. A storage medium, characterized in that: The storage medium includes a stored computer program, wherein when the computer program is executed, the device where the storage medium is located is controlled to execute the speech synthesis method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Speech synthesis method, device and equipment and computer readable medium

    CN117496946A

  • Speech synthesis method and device, electronic equipment and storage medium

    CN119181349A

  • Synthesis-based speech training system

    US5340316A