Method for generating dialect speech synthesis large model

By introducing a large model architecture combining IPA phonetic notation and diffusion hybrid expert mechanism, the problem of insufficient modeling capabilities in multi-dialect speech synthesis is solved, high-quality and diverse dialect speech generation is achieved, and pronunciation accuracy and speech naturalness are improved.

CN120636367APending Publication Date: 2025-09-12GIANT MOBILE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510984089.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing technologies in the field of dialect speech synthesis lack the ability to uniformly model multiple dialects, making it difficult to achieve effective fusion and sharing of speech features among multiple dialects. This results in room for improvement in the sound quality, stability, and style control of the generated speech. In addition, most models do not introduce diffusion models and hybrid expert mechanisms, making it difficult to meet the needs of high-quality and diverse speech generation.

Method used

It adopts a large model architecture that combines IPA phonetic representation with the Diffusion-Mixed Expert Mechanism (DiT-MoE). By collecting training data and conducting multi-stage training, including initial adaptation of Mandarin data and enhancement of dialect data, combined with a semantic information encoder and a generation module, it generates high-quality and diverse dialect speech.

Benefits of technology

It achieves high-quality generation of multi-dialect speech, improves pronunciation accuracy, style diversity and voice naturalness, reduces word error rate, and improves dialect accent characteristics and classification recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120636367A_ABST
    Figure CN120636367A_ABST
Patent Text Reader

Abstract

The invention relates to a method for generating a dialect speech synthesis large model, and the method comprises the following steps: collecting training data which comprise audio and text; constructing training data, using the target dialect pinyin and IPA mapping to replace the target dialect pinyin precise mark with an IPA precise mark syllabically, and forming the training data; generating a model, wherein the model comprises a semantic information encoder, a voice Mel spectrum and a generation module; training the model, and performing first-stage training on the model by using large-scale mandarin data; and introducing dialect data and mandarin data with a proper ratio to carry out second-stage training on the model to form a dialect speech synthesis large model. According to the invention, high-quality and diversified dialect voices can be generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech synthesis technology, and in particular to a method for generating a large model for dialect speech synthesis. Background Art

[0002] Current TTS technology has achieved a high degree of naturalness and voice quality in Mandarin speech synthesis, but it still has significant shortcomings in the field of dialect speech synthesis. Most existing models use Pinyin, slightly modified Pinyin, or other Pinyin systems designed for the target dialect for front-end tagging. Therefore, they only support one or a few dialects and lack the ability to uniformly model multiple dialects. This makes it difficult to effectively integrate and share speech features across multiple dialects, resulting in a relatively high overall cost for implementing multi-dialect speech synthesis capabilities. At the same time, most current dialect TTS models have not yet introduced advanced large-scale model architectures such as diffusion models and mixture of experts. As a result, the generated speech still has considerable room for improvement in terms of sound quality, stability, and style control, making it difficult to meet the needs of high-quality and diverse speech generation.

[0003] Therefore, it is necessary to provide a method for generating a large dialect speech synthesis model to generate high-quality and diverse dialect speech. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for generating a large model for dialect speech synthesis to generate high-quality and diversified dialect speech.

[0005] In order to solve the problems existing in the prior art, the present invention provides a method for generating a large dialect speech synthesis model, comprising the following steps:

[0006] Collect training data, including audio and text;

[0007] Construct training data. If the training data also contains the target dialect pinyin precise transcription, apply the target dialect pinyin and IPA mapping to replace the target dialect pinyin precise transcription syllable by syllable with the IPA precise transcription. If the training data does not contain the target dialect pinyin precise transcription, first use the open source G2P tool to predict the Mandarin pinyin precise transcription. Then, apply the Mandarin and dialect Chinese character pronunciation mapping and the Mandarin and dialect word pronunciation mapping to replace the Mandarin pinyin precise transcription character by character and word by word with the target dialect pinyin precise transcription. Finally, apply the target dialect pinyin and IPA mapping to replace the target dialect pinyin precise transcription syllable by syllable with the IPA precise transcription.

[0008] The generation model includes a semantic information encoder, a speech mel-spectrogram, and a generation module. The semantic information encoder extracts IPA semantic information and inputs it to the generation module. The speech mel-spectrogram is mixed with Gaussian noise according to the ratio specified by the Timestep to generate MixedMel. The Mixed Mel is concatenated with the IPA semantic information and input into the generation module. The generation module receives the original training data, Timestep information, and all input information to generate a prediction value for the current stream.

[0009] Training the model: Use large-scale Mandarin data to conduct the first phase of training to enable the model to have preliminary adaptability to the IPA front-end; introduce dialect data and an appropriate amount of Mandarin data to conduct the second phase of training to form a large dialect speech synthesis model.

[0010] Optionally, in the method for generating a large dialect speech synthesis model, the full spelling of IPA is International Phonetic Alphabet, and its Chinese translation is International Phonetic Alphabet.

[0011] Optionally, in the method for generating a large dialect speech synthesis model, the full spelling of G2P is grapheme tophoneme, and the G2P tool is a tool for converting Chinese text into pinyin.

[0012] Optionally, in the method of generating a large model for dialect speech synthesis, Mandarin and dialect Chinese character pronunciation mapping, Mandarin and dialect word pronunciation mapping, and dialect polyphone disambiguation model are applied to replace the Mandarin pinyin precise mark with the target dialect pinyin precise mark word by word and phrase by phrase.

[0013] Optionally, in the method for generating a large dialect speech synthesis model, a semantic information encoder is obtained by pre-training based on a Transformer architecture.

[0014] Optionally, in the method for generating a large model for dialect speech synthesis, timestep is a core concept that describes the time discretization processing method. It refers to the smallest unit of time simulation and is also used to mark the relative position of elements in sequence data.

[0015] Optionally, in the method for generating a large dialect speech synthesis model, the method for generating a predicted value of the current stream is as follows:

[0016] The generation module is composed of a dialect mixture expert module and a DiT Block containing N diffusion transformer blocks, where N is a natural number.

[0017] The dialect mixture expert module receives all input information and generates dialect expert information based on the gating mechanism and directly passes it to the DiT Block;

[0018] The original training data is also passed to DiTBlock using residual connections;

[0019] To indicate the current time step to the generation module, the DiT Block also receives Timestep information;

[0020] The generation module generates the predicted value of the current flow.

[0021] Optionally, in the method for generating a large dialect speech synthesis model, DiT Block is a DiffusionTransformerBlock, that is, a diffusion model.

[0022] Optionally, in the method for generating a large dialect speech synthesis model, the amount of Mandarin data and dialect data is appropriately proportioned to be the same.

[0023] Compared with the prior art, the present invention has the following advantages:

[0024] (1) This paper proposes a large dialect speech synthesis model that integrates IPA phonetic representation and the Diffusion-Mixed-MoE mechanism (DiT-MoE), aiming to achieve high-quality speech generation for multiple dialects and improve the pronunciation accuracy, style diversity and synthesis naturalness of the model.

[0025] (2) Using the model provided by the present invention for reasoning, corresponding high-quality dialect speech can be obtained, which reduces the word error rate of the generated speech and improves the naturalness of the generated speech, the characteristics of the dialect accent, and the accuracy of dialect classification and recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 A flow chart of a method provided by an embodiment of the present invention;

[0027] Figure 2 A flowchart of constructing training data provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0028] The following is a more detailed description of the specific embodiments of the present invention with reference to schematic diagrams. The advantages and features of the present invention will become more apparent from the following description. It should be noted that the drawings are greatly simplified and not to exact scale, and are only used for the purpose of conveniently and clearly illustrating the embodiments of the present invention.

[0029] In the description of the present application, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise" and the like to indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present application.

[0030] Hereinafter, if the method described herein includes a series of steps, the order in which the steps are presented herein is not necessarily the only order in which the steps may be performed, and some of the steps described may be omitted and / or some other steps not described herein may be added to the method.

[0031] Current TTS technology has achieved a high degree of naturalness and voice quality in Mandarin speech synthesis, but it still has significant shortcomings in the field of dialect speech synthesis. Most existing models use Pinyin, slightly modified Pinyin, or other Pinyin systems designed for the target dialect for front-end tagging. Therefore, they only support one or a few dialects and lack the ability to uniformly model multiple dialects. This makes it difficult to effectively integrate and share speech features across multiple dialects, resulting in a relatively high overall cost for implementing multi-dialect speech synthesis capabilities. At the same time, most current dialect TTS models have not yet introduced advanced large-scale model architectures such as diffusion models and mixture of experts. As a result, the generated speech still has considerable room for improvement in terms of sound quality, stability, and style control, making it difficult to meet the needs of high-quality and diverse speech generation.

[0032] In order to solve the problems existing in the prior art, the present invention provides a method for generating a large model of dialect speech synthesis, such as Figure 1 As shown, the method includes the following steps:

[0033] S1: Collect training data, which includes audio and text;

[0034] S2: Reference Figure 2 , construct training data. If the training data also contains the target dialect pinyin precise mark, apply the target dialect pinyin and IPA mapping to replace the target dialect pinyin precise mark syllable by syllable with the IPA precise mark; the full spelling of IPA is International Phonetic Alphabet, and its Chinese translation is International Phonetic Alphabet.

[0035] If the training data does not contain the pinyin mark of the target dialect, first use the open source G2P tool (grapheme to phoneme, G2P tool is a tool for converting Chinese text into pinyin) to predict the Mandarin pinyin mark, and then apply the Mandarin and dialect Chinese character pronunciation mapping (that is, "a certain pronunciation of a Chinese character in Mandarin corresponds to a certain pronunciation of this Chinese character in the target dialect"), Mandarin and dialect word pronunciation mapping (that is, "a certain pronunciation of a certain word in Mandarin corresponds to a certain pronunciation of this word in the target dialect") to replace the Mandarin pinyin mark with the target dialect pinyin mark word by word, and compare the pronunciation differences of polyphones with those of Mandarin. For a minority of dialects with a relatively large number of pronunciations, the self-developed polyphone disambiguation module can be used to further correct the pinyin of the target dialect. Then, it is determined whether the distribution of different pronunciations of polyphones in the context is very different from that of Mandarin. If not, the pinyin of the target dialect and IPA mapping are applied to replace the pinyin of the target dialect with the IPA precise transcription syllable by syllable. If it is very different, the pinyin of the polyphones is first adjusted through the dialect polyphone disambiguation model, and then the pinyin of the target dialect and IPA mapping are applied to replace the pinyin of the target dialect with the IPA precise transcription syllable by syllable.

[0036] S3: Generative model, which includes a semantic information encoder (text encoder), a speech mel-spectrogram, and a generation module. The semantic information encoder is obtained by pre-training based on the Transformer architecture.

[0037] S31: The semantic information encoder extracts IPA semantic information and inputs it to the generation module. S32: The speech Mel spectrum is mixed with Gaussian noise according to the ratio specified by Timestep to generate Mixed Mel. The Mixed Mel is concatenated with the IPA semantic information and input into the generation module together. S33: The generation module receives the original training data, Timestep information and all input information to generate the predicted value of the current stream. Among them, timestep is the core concept that describes the time discretization processing method. It refers to the smallest unit of time simulation and is also used to mark the relative position of elements in sequence data.

[0038] Specifically, the method for generating the predicted value of the current flow is as follows:

[0039] The generation module is composed of a dialect mixture expert module and a DiT Block containing N diffusion Transformer blocks, where N is a natural number and DiT Block is a DiffusionTransformerBlock, i.e., a diffusion model.

[0040] The dialect mixture expert module receives all input information and generates dialect expert information based on the gating mechanism and directly passes it to the DiT Block;

[0041] The original training data is also passed to DiTBlock using residual connections;

[0042] To indicate the current time step to the generation module, the DiT Block also receives Timestep information;

[0043] The generation module generates the predicted value of the current flow.

[0044] S4: Training model:

[0045] S41: The first phase of model training is performed using large-scale Mandarin data to ensure initial adaptability to the IPA frontend. This step uses flow matching loss to train the model. Because no dialect data is used, the MoE module in the first step is replaced with an identity mapping, and the residual connection is discarded.

[0046] S42: The model undergoes a second phase of training, introducing dialect data and an appropriate proportion of Mandarin data to form a large dialect speech synthesis model. In addition to the stream matching loss, an additional dialect auxiliary classification loss is added to enhance the MoE module's ability to extract dialect expert information. The amount of Mandarin data used is equal to or comparable to the amount of dialect data.

[0047] In summary, the present invention has the following advantages compared with the prior art:

[0048] (1) This paper proposes a large dialect speech synthesis model that integrates IPA phonetic representation and the Diffusion-Mixed-MoE mechanism (DiT-MoE), aiming to achieve high-quality speech generation for multiple dialects and improve the pronunciation accuracy, style diversity and synthesis naturalness of the model.

[0049] (2) Using the model provided by the present invention for reasoning, corresponding high-quality dialect speech can be obtained, which reduces the word error rate of the generated speech and improves the naturalness of the generated speech, the characteristics of the dialect accent, and the accuracy of dialect classification and recognition.

[0050] The above description is merely a preferred embodiment of the present invention and does not limit the present invention in any way. Any person skilled in the art who, without departing from the scope of the present invention, makes any equivalent substitution, modification, or other changes to the technical solution and technical content disclosed in the present invention shall be deemed to be within the scope of the present invention and still fall within the scope of protection of the present invention.

Claims

1. A method for generating a large dialect speech synthesis model, characterized in that: The following steps are involved: Collect training data, including audio and text; Construct training data. If the training data also contains the target dialect pinyin precise transcription, apply the target dialect pinyin and IPA mapping to replace the target dialect pinyin precise transcription syllable by syllable with the IPA precise transcription. If the training data does not contain the target dialect pinyin precise transcription, first use the open source G2P tool to predict the Mandarin pinyin precise transcription. Then, apply the Mandarin and dialect Chinese character pronunciation mapping and the Mandarin and dialect word pronunciation mapping to replace the Mandarin pinyin precise transcription character by character and word by word with the target dialect pinyin precise transcription. Finally, apply the target dialect pinyin and IPA mapping to replace the target dialect pinyin precise transcription syllable by syllable with the IPA precise transcription. The generation model includes a semantic information encoder, a speech mel-spectrogram, and a generation module. The semantic information encoder extracts IPA semantic information and inputs it to the generation module. The speech mel-spectrogram is mixed with Gaussian noise according to the ratio specified by the Timestep to generate MixedMel. The MixedMel is concatenated with the IPA semantic information and input into the generation module. The generation module receives the original training data, Timestep information, and all input information to generate a prediction value for the current stream. Training the model: Use large-scale Mandarin data to conduct the first phase of training to enable the model to have preliminary adaptability to the IPA front-end; introduce dialect data and an appropriate amount of Mandarin data to conduct the second phase of training to form a large dialect speech synthesis model.

2. The method for generating a large dialect speech synthesis model according to claim 1, wherein: The full spelling of IPA is International Phonetic Alphabet, and its Chinese translation is International Phonetic Alphabet.

3. The method for generating a large dialect speech synthesis model according to claim 2, wherein: The full spelling of G2P is grapheme to phoneme, and the G2P tool is a tool that converts Chinese text into pinyin.

4. The method for generating a large dialect speech synthesis model according to claim 3, wherein: Apply the Mandarin and dialect Chinese character pronunciation mapping, Mandarin and dialect word pronunciation mapping, and dialect polyphone disambiguation model to replace the Mandarin pinyin precise symbols with the target dialect pinyin precise symbols word by word.

5. The method for generating a large dialect speech synthesis model according to claim 1, wherein: The semantic information encoder is obtained by pre-training based on the Transformer architecture.

6. The method for generating a large dialect speech synthesis model according to claim 1, wherein: Timestep is a core concept that describes the time discretization processing method. It refers to the smallest unit of time simulation and is also used to mark the relative position of elements in sequence data.

7. The method for generating a large dialect speech synthesis model according to claim 1, wherein: The prediction value of the current stream is generated as follows: The generation module is composed of a dialect mixture expert module and a DiT Block containing N diffusion transformer blocks, where N is a natural number. The dialect mixture expert module receives all input information and generates dialect expert information based on the gating mechanism and directly passes it to the DiT Block; The original training data is also passed to DiTBlock using residual connections; To indicate the current time step to the generation module, the DiT Block also receives Timestep information; The generation module generates the predicted value of the current flow.

8. The method for generating a large dialect speech synthesis model according to claim 7, wherein: DiT Block is Diffusion TransformerBlock, which is a diffusion model.

9. The method for generating a large dialect speech synthesis model according to claim 1, wherein: The amount of Mandarin data and dialect data are appropriately proportioned and equal.