End-to-End Vietnamese Speech Synthesis Method Guided by Dependency Structure Knowledge
Through the end-to-end Vietnamese pronunciation synthesis method guided by dependent structure knowledge, the Vietnamese dependency structure tree and graph are used to guide on the encoding and decoding ends, the problem of the lack of rhythm and naturalness of the audio of Vietnamese pronunciation synthesis is solved, and a higher quality Vietnamese pronunciation synthesis is achieved.
Patent Information
- Application Number
- CN202210803194.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-09
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-07-09
AI Technical Summary
The existing general speech synthesis model has poor application effect in Vietnamese, and the synthetic audio lacks rhythm, fluency and naturalness.
The end-to-end Vietnamese pronunciation synthesis method guided by dependent structure knowledge is adopted to construct Vietnamese dependent structure trees and graphs through data preprocessing, and these structures are integrated on the encoding and decoding ends to improve the rhythmic expression and model stability of synthetic audio.
It significantly improves the rhythmic expressiveness and naturalness of the audio synthesized in Vietnamese and improves the application effect of the model in Vietnamese.
Smart Images

Figure CN115101049B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an end-to-end Vietnamese speech synthesis method guided by dependency structure knowledge, belonging to the technical field of natural language processing. Background Art
[0002] Vietnamese speech synthesis aims to synthesize human-like natural voices for the input Vietnamese text. Vietnamese speech synthesis is of great significance to the cooperation and exchanges between China and Vietnam. The Vietnamese speech synthesis system has relatively broad practical applications, such as intelligent voice assistants, dubbing for movies and games, online education, and smart homes. However, due to the unique language and speech characteristics of Vietnamese, there is currently a lack of a speech synthesis system that meets the actual needs, and the synthesized audio lacks rhythm, fluency, and naturalness.
[0003] At present, good research results have been achieved in the field of speech synthesis for major languages such as Chinese and English. The performance basically meets the requirements of industrial applications. The end-to-end speech synthesis architecture can better implicitly learn the alignment information between text and speech, requires less manual annotation, and the model has good modeling ability on high-dimensional speech features. The end-to-end speech synthesis network architectures are classified as follows: (1) Based on the recurrent network architecture, Wang et al. proposed the Tacotron speech synthesis model, an end-to-end neural network architecture with an encoder-attention mechanism-decoder. The model takes characters as input and outputs a linear spectrogram, and the vocoder uses the Griffin Lim algorithm. Subsequently, Shen et al. proposed the Tacotron2 model architecture, which generates a mel spectrogram and uses Wavenet as the vocoder. Compared with previous speech synthesis methods based on statistical parameters and cascade systems, Tacotron2 greatly improves the quality of the synthesized audio and reduces the model's dependence on a large amount of linguistically annotated data; (2) Based on the Transformer network architecture; Li et al. proposed the TransformerTTS speech synthesis system based on the Transformer architecture, which improves the model's ability to model long-distance dependencies in text and enables parallel training of the model. Ren et al. proposed the non-autoregressive architecture FastSpeech based on the Transformer, which selects a method based on predicted duration alignment, enabling parallel computing during decoding, greatly improving the decoding speed, and solving the robustness problems such as missing words and skipping words in previous speech synthesis models; Ren et al. proposed FastSpeech2, which adds external speech information such as pitch, pitch, and duration on the basis of FastSpeech, making the synthesized speech more fluent and natural. The end-to-end speech synthesis models based on the Transformer have achieved remarkable results in major languages. However, existing general synthesis models rely on large-scale high-quality text-speech alignment training corpora and do not dig deeply enough into the knowledge of their language and speech characteristics, making it difficult to achieve good results in Vietnamese.
[0004] In terms of Vietnamese speech synthesis, Viet et al. adopted methods such as adding prosodic punctuation marks to Vietnamese training data and using different metrics for utterance selection to improve the model's learning of Vietnamese pronunciation knowledge; Nguyen et al. adopted the method of adding a prosodic boundary detection module for Vietnamese to enhance the prosody of the synthesized audio. However, all of the above methods require additional data annotation for Vietnamese, and language experts are needed to annotate the pronunciation pauses of Vietnamese texts. Summary of the Invention
[0005] The present invention provides an end-to-end Vietnamese speech synthesis method guided by dependency structure knowledge, which solves the problem that the speech synthesis method for common major languages has unsatisfactory synthesis effects on Vietnamese and the synthesized audio lacks prosodic expressiveness, and provides a Vietnamese speech synthesis method with better effects.
[0006] The technical solution of the present invention is: an end-to-end Vietnamese speech synthesis method guided by dependency structure knowledge, and the method includes:
[0007] Step1. Preprocessing of Vietnamese speech synthesis data: cleaning the Vietnamese text data, annotating the audio data, constructing a Vietnamese dependency structure tree, constructing a Vietnamese dependency structure graph, etc.;
[0008] Step2. End-to-end Vietnamese speech synthesis guided by dependency structure knowledge: incorporating external guidance knowledge of the Vietnamese dependency structure tree at the encoding end and the decoding end to enhance the prosodic expressiveness of the synthesized audio; at the same time, combining a mask matrix constructed based on the Vietnamese dependency structure graph at the attention mechanism level to guide the monotonic alignment relationship modeling of Vietnamese text-to-speech implicitly and enhance the stability of the synthesis model;
[0009] As a preferred solution of the present invention, the specific steps of Step1 are:
[0010] At the data preprocessing level, cleaning the Vietnamese text data, annotating the audio data, constructing a Vietnamese dependency structure tree, constructing a Vietnamese dependency structure graph, etc., are mainly to enhance the Vietnamese text representation ability and improve the model performance.
[0011] Step1.1. Cleaning of Vietnamese text data: removing the garbled characters in the Vietnamese text, and uniformly encoding and converting the Vietnamese font into Unicode font;
[0012] Step1.2. Annotation of audio data: constructing a Vietnamese audio dataset for the pronunciation of the Vietnamese text, and the audio sampling rate is uniformly 222050hz;
[0013] Step1.3. Construction of Vietnamese dependency structure tree: performing dependency structure parsing on the Vietnamese text of the present invention, and converting the parsing result into a dependency structure tree of the Vietnamese sentence. Parsing the dependency relationship of the Vietnamese sentence to construct a dependency structure tree as shown in Figure 3 with the predicate as the root node, and other words directly or indirectly depending on this word. The finally obtained dependency structure tree is used as the model input as [0,1,1,3,3,5,8,5,5,9,9,11].
[0014] Step1.4. Construct a Vietnamese dependency structure graph: Expand the connection relationships of the Vietnamese dependency structure tree, and change the unidirectional connections in the Vietnamese dependency structure tree to bidirectional connections by adding reverse connections, thereby constructing a Vietnamese dependency structure graph. When there is a dependency relationship between two words, mark it with "0" in the dependency matrix, and mark the rest as "-inf".
[0015] As a preferred solution of the present invention, the specific steps of Step2 are as follows:
[0016] Step2.1. The encoding end of the present invention includes two parts: a Vietnamese dependency structure encoding end and a text encoding end that integrates dependency structure knowledge. Therefore, the hidden state output formula of the encoding end is as follows:
[0017] H e =FFN(cross_attn e (h e ,attn(Q tree ,K tree ,V tree ),attn(Q tree ,K tree ,V tree )))
[0018] Among them, h e represents the hidden state output of the bidirectional long short-term memory layer of the text encoding end; cross_attn e represents the cross-attention mechanism of the text encoding end; Q tree , K tree , V tree represent the information of the Vietnamese dependency structure tree, and attn(Q tree ,K tree ,V tree ) represents the hidden state output of the Vietnamese dependency structure encoding end.
[0019] Step2.2. The dependency structure encoding end includes a dependency structure-aware attention mechanism enhanced by the dependency structure graph, which is used to extract the important features of sparse data constrained by the dependency structure graph and capture the internal correlation of data or features. Through this masked attention method, when the dependency structure tree performs self-attention and cross-attention, each word focuses on the words that depend on it and the words that it depends on, rather than all the words in the sentence.
[0020] Use attn t (Q tree ,K tree ,V tree ) to represent the self-attention mechanism of the dependency layer. The principle of the attention mechanism is as follows:
[0021]
[0022] Among them, Q tree , K tree , V tree all represent the dependency structure tree after passing through the embedding layer, and M tree represents the MASK matrix constructed based on the dependency structure graph.
[0023] The beneficial effects of the present invention are as follows: The present invention performs data cleaning on Vietnamese text data, audio data annotation, constructs a Vietnamese dependency structure tree, and constructs a Vietnamese dependency structure graph; incorporates external guidance knowledge of the Vietnamese dependency structure tree at the encoding end and the decoding end to enhance the prosodic expressiveness of the synthesized audio; at the same time, combines the mask matrix constructed based on the Vietnamese dependency structure graph at the attention mechanism level to guide the monotonic alignment relationship modeling of Vietnamese text-to-speech at the implicit level and improve the stability of the synthesis model. Description of the Drawings
[0024] Figure 1 It is a correlation diagram between the Vietnamese dependency structure tree and the Vietnamese pronunciation pause;
[0025] Figure 2 It is a model diagram of the end-to-end Vietnamese speech synthesis method guided by dependency structure knowledge;
[0026] Figure 3 It is the dependency structure tree constructed in Example 1 of the present invention;
[0027] Figure 4 It is the dependency structure graph constructed from the dependency structure tree in Example 1 of the present invention. Detailed Implementation Modes
[0028] Example 1: An end-to-end Vietnamese speech synthesis method guided by dependency structure knowledge, the method includes:
[0029] Step1. Preprocessing of Vietnamese speech synthesis data: Perform data cleaning on Vietnamese text data, audio data annotation, construct a Vietnamese dependency structure tree, construct a Vietnamese dependency structure graph, etc.;
[0030] Step2. End-to-end Vietnamese speech synthesis guided by dependency structure knowledge: Incorporate external guidance knowledge of the Vietnamese dependency structure tree at the encoding end and the decoding end to enhance the prosodic expressiveness of the synthesized audio; at the same time, combine the mask matrix constructed based on the Vietnamese dependency structure graph at the attention mechanism level to guide the monotonic alignment relationship modeling of Vietnamese text-to-speech at the implicit level and improve the stability of the synthesis model;
[0031] As a preferred solution of the present invention, the specific steps of Step1 are:
[0032] At the data preprocessing level, tasks such as cleaning Vietnamese text data, annotating audio data, constructing a Vietnamese dependency structure tree, and constructing a Vietnamese dependency structure graph are mainly carried out to enhance the representation ability of Vietnamese text and improve the model performance.
[0033] Step1.1. Cleaning Vietnamese text data: Remove the garbled characters in the Vietnamese text, and convert the unified encoding of the Vietnamese font into the Unicode font;
[0034] Step1.2. Annotating audio data: Construct a Vietnamese audio dataset by pronouncing the Vietnamese text, and the audio sampling rate is unified as 222050hz;
[0035] Step1.3. Constructing a Vietnamese dependency structure tree: Conduct dependency structure parsing on the Vietnamese text of the present invention, and convert the parsing result into a dependency structure tree of the Vietnamese sentence. Parse the dependency relationship of the Vietnamese sentence to construct a dependency structure tree as shown in Figure 3 . The predicate is used as the root node, and other words directly or indirectly depend on this word. The finally obtained dependency structure tree is used as the model input as [0,1,1,3,3,5,8,5,5,9,9,11].
[0036] Step1.4. Constructing a Vietnamese dependency structure graph: Expand the connection relationship of the Vietnamese dependency structure tree, and change the unidirectional connection in the Vietnamese dependency structure tree into a bidirectional connection by adding a reverse connection, thereby constructing a Vietnamese dependency structure graph. When there is a dependency relationship between two words, use "0" to mark in the dependency matrix, and the rest are marked as "-inf". The dependency structure graph constructed through the above dependency structure tree is as shown in Figure 4 :
[0037] As a preferred solution of the present invention, the specific steps of Step2 are as follows:
[0038] Step2.1. The encoding end of the present invention includes two parts: a Vietnamese dependency structure encoding end and a text encoding end that integrates dependency structure knowledge. Therefore, the hidden state output formula of the encoding end is as follows:
[0039] H e =FFN(cross_attn e (h e ,attn(Q tree ,K tree ,V tree ),attn(Q tree ,K tree ,V tree )))
[0040] Among them, h e represents the hidden state output of the bidirectional long short-term memory layer at the text encoding end; cross_attn e represents the cross-attention mechanism at the text encoding end; Q tree , K tree , V tree represent the information of the Vietnamese dependency structure tree. attn(Q tree , K tree , V tree ) represents the hidden state output of the Vietnamese dependency structure encoding end.
[0041] Step2.2: The dependency structure encoding end includes a dependency structure-aware attention mechanism enhanced by a dependency structure graph, which is used to extract the important features of sparse data under the constraint of the extracted dependency structure graph and capture the internal correlation of data or features. Through this masked attention method, when the dependency structure tree performs self-attention and cross-attention, each word focuses on the words that depend on it and the words that it depends on, rather than all the words in the sentence.
[0042] Use attn t (Q tree , K tree , V tree ) to represent the self-attention mechanism of the dependency layer. The principle of the attention mechanism is as follows:
[0043]
[0044] Among them, Q tree , K tree , V tree all represent the dependency structure tree after passing through the embedding layer, and M tree represents the MASK matrix constructed based on the dependency structure graph.
[0045] To illustrate the effect of the present invention, the following experiments were carried out in the present invention: The method was verified in the Vietnamese speech synthesis task, and the evaluation indexes were selected as MOS score, MCD value and RMSE value. The experiment used a Vietnamese text-audio parallel data set including 2000 sentences with a cumulative duration of 1 hour. The numbers of the training set and the validation test set were 1800 sentences and 200 sentences respectively. The present invention used the tacotron2 speech synthesis model as the benchmark model. At the same time, the Adam optimizer with a learning rate of 0.0001 was used, the dropout ratio of the encoding and decoding ends was 0.1, the batchsize during the training process was 32, the dimension of the encoding end hidden layer was 512, and the dimension of the decoding end hidden layer was 1024. The sampling rate of the audio in the training data set was uniformly 22050 hz. All experiments were completed on a single NVIDIA Tesla P40.
[0046] Experiment 1: MOS scores of the method of the present invention compared with five other benchmark models. The differences between various model architectures are as follows:
[0047] Tacotron2: The present invention selects Tacotron2 as the benchmark model. This model is an end-to-end general speech synthesis model that uses an encoder-decoder architecture to synthesize audio from text. When synthesizing audio, the acoustic model predicts the Mel spectrogram for the input text, and the vocoder uses WavNet to transform the Mel spectrogram predicted by the acoustic model to obtain the audio file.
[0048] Tacotron2 + defined attention loss function: Improve the loss function of the decoding end attention mechanism based on Tacotron2
[0049] Tacotron2 + pre-trained encoding model: Based on Tacotron2, the method of integrating the pre-trained encoding model refers to the method of Zhu et al. In the task of the present invention, the specific approach is to add the bert pre-trained model of the Vietnamese language after the text input embedding layer at the encoding end.
[0050] Tacotron2 + integrating dependency structure at the encoding end: On the basis of Tacotron2 + defined attention loss function, add the Vietnamese dependency structure encoding end, and use models such as self-attention and linear layer to extract features from the Vietnamese dependency structure tree information. The extracted hidden state is integrated into the encoding end output through a cross-attention mechanism.
[0051] Tacotron2 + integrating dependency structure at the decoding end: On the basis of Tacotron2 + defined attention loss function, add the Vietnamese dependency structure encoding end, and use models such as self-attention and linear layer to extract features from the Vietnamese dependency structure tree information. The extracted hidden state is integrated into the decoding end output through a cross-attention mechanism.
[0052] Tacotron2 + integrating dependency structure at the encoding end + integrating dependency structure at the encoding end: On the basis of Tacotron2 + defined attention loss function, add the Vietnamese dependency structure encoding end, and use models such as self-attention and linear layer to extract features from the dependency structure tree information. The extracted hidden state is integrated into the encoding and decoding end outputs through a cross-attention mechanism. The experimental results of the above 5 models on the dataset are shown in Table 1.
[0053] Table 1: MOS scores of the method of the present invention compared with five other benchmark models
[0054]
[0055] It can be analyzed from the table that the MOS score of the Vietnamese speech audio synthesized by the Vietnamese speech synthesis model guided by integrating dependency structure knowledge reaches 4.03, and a score improvement of 0.20 is obtained compared with the baseline model. This proves that integrating Vietnamese dependency structure knowledge can enhance the speech understanding ability of the encoding end and the speech synthesis ability of the decoding end, making the quality of the synthesized audio higher. At the same time, the synthesis method of the present invention has a better effect than only integrating the pre-trained language model at the encoding end, which proves that there is a certain correlation between Vietnamese dependency structure information and Vietnamese pronunciation pauses. The dependency structure tree constructed by extracting for Vietnamese has a more positive effect on synthesizing higher-quality Vietnamese speech compared with the language pre-trained model. Although the experimental effect of only integrating Vietnamese dependency structure information at the encoding end is better than that of only integrating Vietnamese dependency structure information at the decoding end, integrating Vietnamese dependency structure information at both the encoding and decoding ends achieves the best effect of the model proposed in the present invention, which proves that integrating Vietnamese dependency structure information at the decoding end is helpful for the effect of the synthesized audio and the stability of the synthesis model.
[0056] Experiment 2: Calculate the mel cepstrum distortion score and root mean square error of the synthesized audio and the real audio of the above six models. Mel cepstrum distortion (MCD) measures the difference between two mel cepstrum sequences and can be used to evaluate the quality of a speech synthesis system. Root mean square error (RMSE) is the square root of the ratio of the sum of the squares of the differences between the predicted values and the true values to the number of observations n, and is used to measure the deviation between the observed values and the true values.
[0057] Table 2: Comparison of MCD scores and RMSE calculation results of the method of the present invention and other five models such as the baseline model
[0058]
[0059] The MCD values of the above six models are calculated, and the calculation results are shown in Table 2. The method of the present invention obtains an MCD value of 1.611 and an RMSE value of 1.732, with an MCD value improvement of 1.603 and an RMSE value improvement of 1.920 compared with the baseline model, thus verifying the effectiveness of the method of the present invention.
[0060] Experiment 3: A variety of attempts are made on the integration method of Vietnamese dependency structure information to explore the optimal method of integrating Vietnamese dependency structure information splicing or cross-attention mechanism into the encoding and decoding ends. The specific experimental content is as follows:
[0061] Model A: Integrate the dependency structure by splicing at the encoding end + Integrate the dependency structure by splicing at the decoding end.
[0062] Model B: Integrate the dependency structure by cross-attention at the encoding end + Integrate the dependency structure by splicing at the decoding end.
[0063] Model C: Incorporating dependency structure into the encoding end through splicing + Incorporating dependency structure into the decoding end through cross-attention.
[0064] Model D: Incorporating dependency structure into the encoding end through cross-attention + Incorporating dependency structure into the decoding end through cross-attention.
[0065] Table 3: MOS scores for comparing four methods of incorporating dependency structure
[0066]
[0067] MOS scores were given to the above four models. The experimental results are shown in Table 3. Model D, that is, the model that uses cross-attention to incorporate dependency structure information at both the encoding and decoding ends, achieved the best experimental results. The present invention believes that the possible reason is that the dependency structure information is inherently related to the input text and audio. Using the cross-attention mechanism fusion method can make the input and dependency structure information pay more attention to each other. In contrast, the splicing fusion method simply concatenates two vectors. Compared with this, the hidden state feature distribution finally output by using the cross-attention mechanism fusion method is more concentrated.
[0068] The specific embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the spirit of the present invention.
Claims
1. An end-to-end Vietnamese speech synthesis method guided by dependency structure knowledge, characterized in that: The specific steps of the method are as follows: Step1. Vietnamese speech synthesis data preprocessing: Clean the Vietnamese text data, annotate the audio data, construct a Vietnamese dependency structure tree, and construct a Vietnamese dependency structure graph; Step2. End-to-end Vietnamese speech synthesis guided by dependency structure knowledge: Incorporate external guidance knowledge of the Vietnamese dependency structure tree at the encoding end and the decoding end to enhance the prosodic expressiveness of the synthesized audio; at the same time, combine the mask matrix constructed based on the Vietnamese dependency structure graph at the attention mechanism level to guide the monotonic alignment relationship modeling of Vietnamese text-to-speech implicitly and improve the stability of the synthesis model; The specific steps of Step2 are as follows: Step2.
1. The encoding end includes two parts: the Vietnamese dependency structure encoding end and the text encoding end that integrates dependency structure knowledge. Therefore, the hidden state output formula of the encoding end is as follows: H e = FFN(cross_attn e (h e , attn(Q tree , K tree , V tree ), attn(Q tree , K tree , V tree ))) Among them, h e represents the hidden state output of the bidirectional long short-term memory layer at the text encoding end; cross_attn e represents the cross-attention mechanism at the text encoding end; Q tree , K tree , V tree represent the Vietnamese dependency structure tree information, and attn(Q tree , K tree , V tree ) represents the hidden state output of the Vietnamese dependency structure encoding end; Step2.
2. The dependency structure encoding end contains a dependency structure-aware attention mechanism enhanced by the dependency structure graph, which is used to extract the important features of sparse data constrained by the dependency structure graph and capture the internal correlation of the data or features; through this masked attention method, when the dependency structure tree performs self-attention and cross-attention, each word focuses on the words that depend on it and the words it depends on rather than all the words in the sentence; Use attn t (Q tree ,K tree ,V tree ) represents the self-attention mechanism of the dependency layer, and the principle of the attention mechanism is as follows: Among them, Q tree , K tree , V tree all represent the dependency structure tree after passing through the embedding layer, and M tree represents the MASK matrix constructed based on the dependency structure graph.
2. The end-to-end Vietnamese speech synthesis method guided by the knowledge of dependency structure according to claim 1, characterized in that: The specific steps of Step1 are as follows: Step1.
1. Vietnamese text data cleaning: Remove the garbled characters in the Vietnamese text, and convert the unified encoding of the Vietnamese font into Unicode font; Step1.
2. Audio data annotation: Construct a Vietnamese audio dataset by pronouncing the Vietnamese text, and the audio sampling rate is unified to 222050hz; Step1.
3. Construct a Vietnamese dependency structure tree: Parse the dependency structure of Vietnamese and convert the parsing result into a dependency structure tree of Vietnamese sentences; Step1.
4. Construct a Vietnamese dependency structure graph: Expand the connection relationship of the Vietnamese dependency structure tree, and change the unidirectional connection in the Vietnamese dependency structure tree to a bidirectional connection by adding a reverse connection, so as to construct a Vietnamese dependency structure graph. When two words have a dependency relationship, use "0" to mark them in the dependency matrix, and the rest are marked as "-inf".
Citation Information
Patent Citations
Audio data generating method and system for voice synthesis
CN109036371A
End-to-end voice synthesis method, device, equipment and storage medium
CN109616093A