Multi-granularity prosodic speech synthesis method and device combined with text grammar

By adding a multi-granularity prosody encoder to the FastSpeech2 network and combining grammatical-level, utterance-level, word-level, and phoneme-level encoders, the problem of unreasonable prosody feature granularity processing in the existing technology is solved, achieving more accurate prosody feature prediction and improved synthetic speech quality.

CN119964543BActive Publication Date: 2025-09-26ANHUI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510100805.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-09-26
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

Existing speech synthesis technology does not process the granularity of prosodic features reasonably and ignores the combination of grammatical information and prosodic features, resulting in poor synthesized speech effects.

Method used

Using the FastSpeech2 network as the basis, a multi-granularity prosodic encoder is added, and the grammatical level, discourse level, word level and phoneme level encoders are combined. Speech synthesis is performed by injecting grammar into the multi-granularity prosodic network, integrating text grammatical information and improving the reasonable processing of prosodic features.

Benefits of technology

It achieves more accurate prediction of word-level and phoneme-level prosodic features, improves the naturalness and prosodic diversity of synthesized speech, improves the prediction accuracy of duration, pitch and energy, and solves the problem of unreasonable granularity processing of prosodic features in existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964543B_ABST
    Figure CN119964543B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of speech processing technology, and in particular to a method and device for multi-granularity prosodic speech synthesis combined with text grammar. The method of the present invention comprises: obtaining a text to be processed and an original reference audio of a target speaker; then, on the one hand, preprocessing the text to be processed into a grammar graph and a phoneme sequence, and on the other hand, preprocessing the original reference audio into an original mel-spectrogram, and inputting the trained grammar-injected multi-granularity prosodic network for processing to obtain a synthesized mel-spectrogram; finally, performing voice coding conversion on the synthesized mel-spectrogram to obtain synthesized speech. The present invention uses the FastSpeech2 network as a base network, and adds a multi-granularity prosodic encoder between the phoneme encoder and the variation adapter therein to construct a grammar-injected multi-granularity prosodic network, performs more reasonable multi-granularity processing and combination on the prosodic features, and introduces text grammar information, thereby improving the effect of prosodic speech synthesis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech processing technology, and more specifically to: 1. a multi-granularity prosodic speech synthesis method combined with text grammar; 2. a multi-granularity prosodic speech synthesis device combined with text grammar. Background Art

[0002] Personalized speech synthesis involves learning a speaker's unique speech pattern from reference audio sources to generate speech that reflects the speaker's habitual prosody. In other words, prosodic features are key to personalized speech synthesis.

[0003] Existing researchers have proposed various methods and tools, such as global style tags (GST), variational autoencoders (VAE), and aggregated variational autoencoders (AVAE), to model prosodic features and implement style control in speech synthesis. However, after analysis, the inventors found that existing approaches have the following shortcomings: 1. The granularity of prosodic features is not properly addressed—some focus only on coarse granularity, some only on fine granularity, and some focus on both global and local aspects but poorly integrate them. 2. Research primarily focuses on prosodic features themselves, ignoring the relationship between other factors and prosodic features. Summary of the Invention

[0004] Based on this, it is necessary to provide a multi-granularity prosodic speech synthesis method and device combined with text grammar to address the problems that existing prosodic speech methods handle the granularity of prosodic features unreasonablely and ignore other factors.

[0005] The present invention is achieved by adopting the following technical solutions:

[0006] In a first aspect, the present invention discloses a multi-granularity prosodic speech synthesis method combined with text grammar, comprising:

[0007] Step 1: Get the text to be processed INF and the original reference audio Voice0 of the target speaker;

[0008] Step 2: Perform syntax extraction and phoneme conversion on INF to obtain syntax graph G and phoneme sequence Phoneme;

[0009] Perform audio conversion on Voice0 to obtain the original Mel spectrogram Mel_S0;

[0010] Step 3: Input the trained grammar of Phoneme, G, and Mel_S0 into the multi-granularity prosodic network for processing to obtain the synthesized mel-spectrogram New_Mel;

[0011] Step 4, performing voice coding conversion on New_Mel to obtain the synthesized speech New_Voice;

[0012] The method for constructing a grammar-injected multi-granularity prosodic network includes:

[0013] The FastSpeech2 network is used as the base network, and a multi-granularity prosody encoder is added between the phoneme encoder and the variation adapter, thus obtaining a grammar-injected multi-granularity prosody network.

[0014] The multi-granularity prosody encoder is used to combine G and Mel_S0 to embed the text output by the phoneme encoder into E c Processed into text hidden optimization embedded E h ;E h Multi-granularity prosodic features are fused and used as input to the variational adapter.

[0015] This multi-granularity prosodic speech synthesis method combined with text grammar implements the method or process according to the embodiment of the present disclosure.

[0016] In a second aspect, the present invention discloses a multi-granularity prosodic speech synthesis device combined with text grammar, which uses the multi-granularity prosodic speech synthesis method combined with text grammar disclosed in the first aspect.

[0017] The multi-granularity prosodic speech synthesis device combined with text grammar includes: a data acquisition module, a preprocessing module, a data processing module, and a speech generation module.

[0018] The data acquisition module is used to obtain the text to be processed INF and the original reference audio Voice0 of the target speaker;

[0019] The preprocessing module is used to perform syntax extraction and phoneme conversion on INF to obtain the syntax graph G and the phoneme sequence Phoneme; and perform audio conversion on Voice0 to obtain the original Mel-spectrogram Mel_S0.

[0020] The data processing module is used to inject the trained grammar of Phoneme, G, and Mel_S0 into the multi-granularity prosodic network for processing to obtain the synthetic mel-spectrogram New_Mel.

[0021] The speech generation module is used to perform voice coding conversion on New_Mel to obtain the synthesized speech New_Voice.

[0022] This multi-granularity prosodic speech synthesis device combined with text grammar implements the method or process according to the embodiment of the present disclosure.

[0023] Compared with the prior art, the present invention has the following beneficial effects:

[0024] 1. The present invention uses the FastSpeech2 network as the base network and adds a multi-granularity prosody encoder between the phoneme encoder and the variation adapter to construct a grammar-injected multi-granularity prosody network (GMG-ProsodyNet). It performs more reasonable multi-granularity processing and combination of prosodic features, introduces text grammar information, and improves the effect of prosodic speech synthesis.

[0025] 2. The multi-granularity prosody encoder of the present invention, on the one hand, adopts a grammatical-level encoder to guide the prediction of word-level and phoneme-level prosody features, and can predict word-level and phoneme-level prosody features that are very similar to the reference speech; on the other hand, it adopts a phoneme-level encoder and a word-level encoder to achieve precise control of the expression style during the speech synthesis process; among them, the word-level prosody predictor is used to generate word-level prosody features, so that it retains valuable information related to the tone and speech continuity in the word-level spectrum, so as to improve the network effect.

[0026] 3. The grammatical-level encoder of the present invention uses the GGNN network to extract grammatical-level prosodic features from the grammatical graph, allowing word nodes to interact with each other to capture grammatical features that reflect the dependency relationship between distant words, thereby extracting the prosodic style related to the grammatical information, thereby improving the prediction accuracy of duration, pitch and energy in synthesized speech. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0028] Figure 1 This is a structural diagram of a grammar-injected multi-granularity prosody network provided in Example 1 of the present invention;

[0029] Figure 2 for Figure 1 The structure diagram of the grammar-level encoder;

[0030] Figure 3 for Figure 1 The structure diagram of the Chinese speech-level encoder;

[0031] Figure 4 for Figure 1 The structure diagram of the word-level encoder;

[0032] Figure 5 for Figure 1 The structure diagram of the phoneme-level encoder;

[0033] Figure 6The data flow diagram of the grammar-injected multi-granularity prosody network provided in Example 1 of the present invention during the p-th round of training;

[0034] Figure 7 for Figure 6 Data flow diagram of the medium multi-granularity prosody encoder;

[0035] Figure 8 The data flow diagram of the grammar-injected multi-granularity prosody network provided in Example 1 of the present invention during the qth round of training;

[0036] Figure 9 for Figure 8 Data flow diagram of the medium multi-granularity prosody encoder;

[0037] Figure 10 A data flow diagram of a multi-granularity prosodic speech synthesis method combined with text grammar provided in Example 2 of the present invention;

[0038] Figure 11 for Figure 10 Data flow diagram of the medium multi-granularity prosody encoder;

[0039] Figure 12 The original Mel spectrogram provided in Example 3 of the present invention;

[0040] Figure 13 The synthesized mel-spectrogram obtained by injecting grammar into the multi-granularity prosodic network in Example 3 of the present invention;

[0041] Figure 14 The synthesized mel-spectrogram obtained by using the FG-transformer-TTS network in Example 3 of the present invention;

[0042] Figure 15 The synthesized mel-spectrogram obtained by using the SyntaSpeech network in Example 3 of the present invention;

[0043] Figure 16 This is the synthesized mel-spectrogram obtained using the AdaSpeech network in Example 3 of the present invention. DETAILED DESCRIPTION

[0044] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0045] It should be noted that when a component is referred to as being "mounted on" another component, it may be directly on the other component or there may be a central component. When a component is considered to be "set on" another component, it may be directly set on the other component or there may be a central component. When a component is considered to be "fixed to" another component, it may be directly fixed to the other component or there may be a central component.

[0046] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention pertains. The terms used herein in the specification of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "or / and" as used herein includes any and all combinations of one or more of the associated listed items.

[0047] Example 1

[0048] This embodiment 1 provides a grammar-injected multi-granularity prosody network (GMG-ProsodyNet for short) and discloses its training process.

[0049] In short, the construction method of GMG-ProsodyNet includes:

[0050] The FastSpeech2 network is used as the base network, and a multi-granularity prosody encoder is added between the phoneme encoder and the variation adapter to obtain GMG-ProsodyNet.

[0051] For ease of understanding, Figure 1 The structure of GMG-ProsodyNet is shown, which includes a base network and a multi-granularity prosody encoder. The base network includes a phoneme embedding layer, a phoneme encoder, a variational adapter, a Mel decoder, two parent stacking layers, and a linear layer.

[0052] The FastSpeech2 network is an efficient and high-quality end-to-end TTS model, particularly suitable for scenarios requiring rapid speech synthesis. The structure of the underlying network is well known, so this article will not detail its structure, focusing instead on its functionality.

[0053] The multi-granularity prosody encoder is a new design part of the present invention. Figure 1 ,The multi-granularity prosody encoder includes: 1 grammar-level encoder, 1 discourse-level encoder, 1 phoneme-level encoder, 1 word-level encoder, 4 expanders, 2 sub-overlay layers, and 1 word segmenter.

[0054] For each part of the multi-granularity prosody encoder:

[0055] (i) The syntax-level encoder aims to summarize the node information in the syntax graph constructed by the syntax tree, thereby ensuring that syntactic information is fully integrated into the prosodic modeling process.

[0056] First see Figure 2 , the grammatical level encoder includes: 1 word-level average pooling layer, 2 GGNN network layers, and 1 sub-overlay layer.

[0057] The syntax-level encoder is used to process input data in0-in1 into output data out1.

[0058] Specifically, in the syntax-level encoder:

[0059] The word-level average pooling layer is used to perform word-level average pooling on in1;

[0060] The first GGNN network layer is used to perform preliminary aggregation of syntactic information on the output of the word-level average pooling layer using in0;

[0061] The second GGNN network layer is used to further aggregate the syntactic information of the output of the first GGNN network layer using in0;

[0062] The sub-overlay layer is used to superimpose the input data in1, the output of the first GGNN network layer, and the output of the second GGNN network layer to obtain out1.

[0063] Among them, the GGNN network layer can achieve effective information propagation through iterative updates of node states: in each iteration, the node receives information from its neighbors, thereby enhancing its state information, enabling it to capture the intricate relationships in the graph. This aggregation process helps to extract grammatical features to reflect the dependencies between distant words.

[0064] It should be noted that the length of in1 is at the phoneme level, and the length of out1 is at the word level, which is shorter than the phoneme level. Figure 1 , out1 is input to the first expander to expand its length to the phoneme level.

[0065] (2) The utterance-level encoder aims to extract coarse-grained prosodic features.

[0066] Next see Figure 3 , the utterance-level encoder includes: 2 one-dimensional convolutional layers, 2 ReLU activation functions, 2 normalization layers, 2 Dropout layers, and 1 one-dimensional average pooling layer.

[0067] The utterance-level encoder is used to process the input data in2 into the output data out2.

[0068] Specifically, in the utterance-level encoder:

[0069] The first one-dimensional convolution layer is used to perform one-dimensional convolution on in2;

[0070] The first Relu activation function is used to activate the output of the first one-dimensional convolutional layer;

[0071] The first normalization layer is used to normalize the output of the first Relu activation function;

[0072] The first Dropout layer is used to simplify the output of the first normalization layer;

[0073] The second one-dimensional convolution layer is used to perform one-dimensional convolution on the output of the first Dropout layer;

[0074] The second Relu activation function is used to activate the output of the second one-dimensional convolutional layer;

[0075] The second normalization layer is used to normalize the output of the second Relu activation function;

[0076] The second Dropout layer is used to simplify the output of the second normalization layer;

[0077] The one-dimensional average pooling layer is used to perform one-dimensional average pooling on the output of the second Dropout layer to obtain out2.

[0078] Similarly, the length of out2 is utterance-level, shorter than phoneme-level. Figure 1 , so out2 is input into the second expander to expand its length to the phoneme level.

[0079] (3) Then refer to Figure 1 , the output of the first expander, the output of the second expander, and the in1 input are superimposed on the first sub-overlay layer.

[0080] (4) The word-level encoder aims to extract prosodic features at a finer level than the utterance-level encoder to improve prosodic modeling.

[0081] Next see Figure 4 , the word-level encoder includes: 2 2D convolutional layers, 2 ReLU activation functions, 2 normalization layers, 2 Dropout layers, 1 2D adaptive average pooling layer, and 1 linear layer.

[0082] The word-level encoder is used to process the input data in3 into the output data out3.

[0083] Specifically:

[0084] The first two-dimensional convolution layer is used to perform two-dimensional convolution processing on in3;

[0085] The first Relu activation function is used to activate the output of the first two-dimensional convolutional layer;

[0086] The first normalization layer is used to normalize the output of the first Relu activation function;

[0087] The first Dropout layer is used to simplify the output of the first normalization layer;

[0088] The second two-dimensional convolution layer is used to perform two-dimensional convolution processing on the output of the first Dropout layer;

[0089] The second Relu activation function is used to activate the output of the second two-dimensional convolutional layer;

[0090] The second normalization layer is used to normalize the output of the second Relu activation function;

[0091] The second Dropout layer is used to simplify the output of the second normalization layer;

[0092] The two-dimensional adaptive average pooling layer is used to perform two-dimensional adaptive average pooling on the output of the second Dropout layer;

[0093] The linear layer is used to transform the output of the two-dimensional adaptive average pooling layer to obtain out3.

[0094] Similarly, the length of out3 is word-level, shorter than phoneme-level. Figure 1 , so out3 is input into the third expander to expand its length to the phoneme level.

[0095] (5) The phoneme-level encoder aims to extract prosodic features at a finer level than the word-level encoder and further improve the prosodic modeling.

[0096] Continue reading Figure 5 ,The phoneme-level encoder includes: 2 one-dimensional convolutional layers, 2 Relu activation functions, 2 normalization layers, 2 Dropout layers, and 1 linear layer.

[0097] The phoneme-level encoder is used to process input data in4 into output data out4.

[0098] Specifically:

[0099] The first one-dimensional convolution layer is used to perform one-dimensional convolution processing on the input data in4;

[0100] The first Relu activation function is used to activate the output of the first one-dimensional convolutional layer;

[0101] The first normalization layer is used to normalize the output of the first Relu activation function;

[0102] The first Dropout layer is used to simplify the output of the first normalization layer;

[0103] The second one-dimensional convolution layer is used to perform one-dimensional convolution on the output of the first Dropout layer;

[0104] The second Relu activation function is used to activate the output of the second one-dimensional convolutional layer;

[0105] The second normalization layer is used to normalize the output of the second Relu activation function;

[0106] The second Dropout layer is used to simplify the output of the second normalization layer;

[0107] The linear layer is used to transform the output of the second Dropout layer to obtain the output data out4.

[0108] It should be noted that out4 is already phoneme-level in length and does not need to be expanded further.

[0109] (6) Then refer to Figure 1 The output of the third expander, the output of the first sub-superposition layer, and out4 are input to the second sub-superposition layer and superimposed to obtain the output data OUT of the multi-granularity prosody encoder.

[0110] (7) Then refer to Figure 1 The fourth expander is used to expand the length of the output of the first superposition layer to Mel length; the word segmenter is used to perform word-level segmentation on the output of the fourth expander and input it into the word-level encoder for processing; the output of the first sub-superposition layer is also input into the phoneme-level encoder for processing.

[0111] It should be noted that:

[0112] The process of (VII) does not work in the first half of the entire network training, but works in the second half of the entire network training when the trained network is used.

[0113] The GMG-ProsodyNet based on the above structure needs to be trained to obtain a trained GMG-ProsodyNet in order to achieve good results in application.

[0114] This embodiment 1 simultaneously discloses a method for obtaining a trained GMG-ProsodyNet, namely, a method for training GMG-ProsodyNet, which includes:

[0115] S1, obtain a sample set with true labels;

[0116] The sample set types include: text samples, reference original sample speech, and reference synthesized sample speech. The true labels of the sample set include: 1. The true value of the reference synthesized sample speech; 2. The true value of the duration of the variation adapter; 3. The true value of the pitch of the variation adapter; 4. The true value of the energy of the variation adapter.

[0117] It should be noted that this embodiment 1 adopts a parallel training method, so the reference original sample speech and the reference synthesized sample speech are the same.

[0118] The existing LJSpeech dataset is used in this embodiment 1. The LJSpeech dataset consists of 13,100 short audio clips in which speakers read passages from seven non-fiction books, with a total of approximately 24 hours of audio.

[0119] S2, divide the sample set into training set and validation set - the specific division ratio can be adjusted according to actual situation.

[0120] S3, train GMG-ProsodyNet for N rounds based on the training set; verify the multi-granularity prosody network after each round of training based on the validation set until the GMG-ProsodyNet with the best performance is selected;

[0121] In each round of training, some data are randomly extracted from the training set as a training subset, and a total of N training subsets {U1, ..., U N} to perform N rounds of training.

[0122] It should be noted that the data of the training subset needs to be preprocessed before being input into GMG-ProsodyNet:

[0123] 1. Perform syntax extraction processing (generally implemented using syntax tree building tools such as Stanza) and phoneme conversion processing (generally implemented using morpheme-to-phoneme conversion tools such as G2p) on the text samples to obtain syntax sample graphs and phoneme sample sequences;

[0124] 2. First, perform audio conversion on the reference original sample speech to obtain the original sample Mel spectrogram, and then perform word-level segmentation and phoneme-level segmentation on the original sample Mel spectrogram (based on the actual duration value of the variation adapter) to obtain word-level sample Mel spectrogram and phoneme-level sample Mel spectrogram.

[0125] In summary:

[0126] 1. In the first N / 2 rounds of training, all layers in the base network are working, and all parts of the multi-granularity prosody encoder except the fourth expander and word segmenter are working.

[0127] Then, the loss functions used in the first N / 2 rounds of training include: prediction loss of grammar-injected multi-granularity prosody network, duration loss of variational adapter, pitch loss of variational adapter, and energy loss of variational adapter.

[0128] To illustrate the training process of the first N / 2 rounds of training, take the p-th round of training as an example, p∈[1,N / 2]:

[0129] See Figure 6 , first take the pth training subset U p Processed into the pth phoneme sample sequence Phoneme p , the pth grammar graph G p , the pth original Mel spectrum sample image Mel_S p,0 , the p-th phoneme-level Mel spectrum sample graph Mel_S p,1 , the p-th word-level Mel spectrum sample graph Mel_S p,2 ; then Phoneme p After the phoneme embedding layer, the first parent superposition layer, and the phoneme encoder, the pth text sample hidden embedding E is obtained p,c ;E p,c , G p 、Mel_S p,0 、Mel_S p,1 、Mel_S p,2 Enter the multi-granularity prosody encoder process and obtain the p-th text sample hidden optimized embedding E p,h ; After E p,h First calculate the duration D of the p-th round of training through the variation adapter p , pitch Pi p Energy E p , and then pass through the second parent overlay layer, Mel decoder, and linear layer to obtain the pth synthetic sample Mel spectrogram New_Mel p .

[0130] Among them, see Figure 7 , E p,c , G p The p-th grammatical level sample prosody embedding E is obtained by encoding it through the grammatical level encoder p,s ;E p,s Adjust to the length of E through the first expander p,c Same; Mel_S p,0 The prosodic embedding E of the p-th utterance-level sample is obtained by encoding it through the utterance-level encoder p,u ;E p,uAdjust the length to the same as E through the second expander p,c Same; E p,c The output of the first expander and the output of the second expander are superimposed by the first superimposer to obtain the hidden embedding E' of the pth relay sample p,c ;Mel_S p,1 The p-th phoneme-level sample prosody embedding E is obtained by encoding it through the phoneme-level encoder p,ph ;Mel_S p,2 The p-th word-level sample prosody embedding E is obtained by encoding it through the word-level encoder p,w ;E p,w Adjust the length to E' through the third expander p,c 、E p,ph Same; E' p,c 、E p,ph The output of the third expander is superimposed by the second superimposer to obtain E p,h .

[0131] Then, the loss function Loss used in the p-th round of training P for:

[0132]

[0133] Where, L p,1 represents the prediction loss of GMG-ProsodyNet in the p-th round of training; L p,d represents the duration loss of the deteriorating adapter in the p-th round of training; L p,pi represents the pitch loss of the deviated adapter in the p-th round of training; L p,e represents the energy loss of the degraded adapter in the p-th round of training; ||.||1 represents the L1 loss function; New_Mel p,true represents the true value of the Mel spectrogram corresponding to the reference synthesized sample speech in the p-th round of training; MSE(.,.) represents the mean square error function; D p,true represents the actual value of the duration of the deterioration adapter in the p-th round of training; Pi p,true represents the true value of the pitch of the variant adapter in the p-th round of training; E p,true represents the true value of the energy of the degraded adapter in the p-th round of training.

[0134] Then, based on Loss p Backpropagation is performed to make the p-th update adjustment to the network parameters.

[0135] 2. In the next N / 2 rounds of training: all layers in the base network are working, and all parts in the multi-granularity prosody encoder are working.

[0136] Then, the loss functions used in the next N / 2 rounds of training include: prediction loss of the grammar-injected multi-granularity prosody network, duration loss of the variational adapter, pitch loss of the variational adapter, energy loss of the variational adapter, encoding loss of the phoneme-level encoder, and encoding loss of the word-level encoder.

[0137] To illustrate the training process of the last N / 2 rounds of training, take the qth round of training as an example, q∈[N / 2+1,N]:

[0138] See Figure 8 , first take the qth training subset U q Processed into the qth phoneme sample sequence Phoneme q , the qth grammar graph G q , the qth original Mel spectrum sample image Mel_S q,0 , the qth phoneme-level Mel spectrum sample graph Mel_S q,1 , the qth word-level Mel spectrum sample graph Mel_S q,2 ; then Phoneme q After the phoneme embedding layer, the first parent superposition layer, and the phoneme encoder, the qth text sample hidden embedding E is obtained q,c ;E q,c , G q 、Mel_S q,0 、Mel_S q,1 、Mel_S q,2 Enter the multi-granularity prosody encoder process and get the qth text sample hidden optimized embedding E q,h ; then E q,h First calculate the duration D of the qth round of training through the variation adapter q , pitch Pi q Energy E q , and then pass the second parent overlay layer, Mel decoder, and linear layer to obtain the qth synthetic sample Mel spectrogram New_Mel q .

[0139] Among them, see Figure 9 , E q,c , G q The qth grammatical level sample prosody embedding E is obtained by encoding it through the grammatical level encoder q,s ;E q,s Adjust to the length of E through the first expander q,c Same; Mel_S q,0 The qth utterance-level sample prosody embedding E is obtained by encoding it through the utterance-level encoder q,u ;E q,u Adjust the length to the same as E through the second expander q,c Same; E q,cThe output of the first expander and the output of the second expander are superimposed by the first superimposer to obtain the qth relay sample hidden embedding E' q,c ; E' q,c The qth phoneme-level sample prosody embedding E is obtained by encoding it through the phoneme-level encoder q,ph ;Mel_S q,1 The qth phoneme-level sample prosodic contrast embedding E' is obtained by encoding it through the phoneme-level encoder q,ph ; E' q,c First, the length of the word is adjusted to Mel length through the fourth expander, and then the word segmenter is used for segmentation. Then, the word-level encoder is used to encode the qth word-level sample rhythm embedding E. q,w ;E q,w Adjust the length to E' through the third expander q,c 、E q,ph Same; Mel_S q,2 The qth word-level sample prosody contrast embedding E' is obtained by encoding it through the word-level encoder q,w ; E' q,c 、E q,ph The output of the third expander is superimposed by the second superimposer to obtain E q,h .

[0140] That is to say, in the qth round of training, the word-level encoder and the phoneme-level encoder both have dual-channel data input and output.

[0141] Then, the loss function Loss used in the qth round of training q for:

[0142]

[0143] Where, L q,1 represents the prediction loss of GMG-ProsodyNet in the qth round of training; L q,d represents the duration loss of the qth round of training of the deviating adapter; L q,pi represents the pitch loss of the qth round of training variant adapter; L q,e represents the energy loss of the degraded adapter in the qth round of training; L q,ph represents the encoding loss of the qth round of training phoneme-level encoder; L q,w represents the encoding loss of the word-level encoder in the qth round of training; ||.||1 represents the L1 loss function; New_Mel q,true represents the true value of the Mel spectrogram corresponding to the reference synthesized sample speech in the qth round of training; MSE(.,.) represents the mean square error function; D q,true represents the actual duration of the variation adapter in the qth round of training; Pi q,truerepresents the true value of the pitch of the variant adapter in the qth round of training; E q,true represents the true value of the energy of the degraded adapter in the qth round of training.

[0144] Then, based on Loss q Backpropagation is performed to make the qth update adjustment to the network parameters.

[0145] After the above N rounds of training, the network with the best performance is selected through the validation set, which is the trained GMG-ProsodyNet.

[0146] Example 2

[0147] This embodiment 2 discloses a specific application of the GMG-ProsodyNet trained in embodiment 1—using it for prosodic speech synthesis.

[0148] This embodiment 2 discloses a multi-granularity prosodic speech synthesis method combined with text grammar, including:

[0149] Step 1: Obtain the text to be processed INF and the original reference audio Voice0 of the target speaker.

[0150] Step 2: Perform syntax extraction and phoneme conversion on INF to obtain syntax graph G and phoneme sequence Phoneme;

[0151] Perform audio conversion on Voice0 to obtain the original Mel spectrogram Mel_S0

[0152] The above processing process is similar to the pre-processing process in Reference Example 1 and will not be described in detail.

[0153] It should be noted that since there is no real value of duration at this time, only the calculated duration value calculated by the variation adapter - which cannot be guaranteed to match Mel_S0; therefore, Mel_S0 is no longer directly subjected to word-level segmentation and phoneme-level segmentation processing to obtain word-level mel-spectrogram Mel_S1 and phoneme-level mel-spectrogram Mel_S2, but the network itself is allowed to perform processing based on Phoneme, G, and Mel_S0.

[0154] Step 3: Input the trained grammar of Phoneme, G, and Mel_S0 into the multi-granularity prosodic network for processing to obtain the synthesized mel-spectrogram New_Mel.

[0155] For details, see Figure 10 , the trained grammar in step 3 is injected into the multi-granularity prosodic network:

[0156] The phoneme embedding layer embeds Phoneme;

[0157] The first parent overlay layer adds positional encoding to the output of the phoneme embedding layer;

[0158] The phoneme encoder encodes the output of the first parent stacking layer to obtain the text hidden embedding E c ;

[0159] The multi-granularity prosody encoder is used to combine G and Mel_S0 to embed the text output by the phoneme encoder into E c Processed into text hidden optimization embedded E h ;

[0160] Variation adapter based on E h Calculate the duration, pitch, and energy information;

[0161] The second parent overlay is E h Add position coding and information about duration, pitch, and energy;

[0162] The Mel decoder decodes the output of the second parent stacking layer to obtain a high-dimensional Mel spectrogram;

[0163] The linear layer reduces the dimensionality of the high-dimensional Mel spectrogram to obtain New_Mel.

[0164] Among them, see Figure 11 , for the multi-granularity prosody encoder:

[0165] The syntax-level encoder converts E c , G encoding is integrated into grammatical prosodic embedding E s (That is, at this time in1 is E c ;in0 is G;out1 is E s );

[0166] The first expander will E s The length is extended to E c same;

[0167] The utterance-level encoder encodes Mel_S0 into an utterance-level prosodic embedding E u (That is, at this time in2 is Mel_S0; out2 is E u );

[0168] The second expander will E u The length is extended to E c same;

[0169] The first sub-layer combines the output of the first expander, the output of the second expander, and the E c Superposition is performed to obtain the relay hidden embedding E' c ;

[0170] The phoneme-level encoder converts E' c Encoded into phoneme-level prosodic embedding E ph (That is, in4 is E' c ;out4 is E ph );

[0171] The 4th expander will be E' c The length is extended to Mel length;

[0172] The word segmenter performs word-level segmentation on the output of the fourth expander (based on the duration calculated by the variation adapter);

[0173] The word-level encoder encodes the output of the word segmenter into a word-level prosodic embedding E w (That is, in3 is the output of the word segmenter; out3 is E w );

[0174] The third expander will E w The length is extended to E ph 、E' c same;

[0175] The second sub-layer will output the third expander, E ph 、E' c Superposition to obtain E h .

[0176] Therefore, E h The multi-granularity rhythmic features are integrated and used as the input of the variation adapter (that is, OUT is E h ).

[0177] Step 4: Perform voice coding on New_Mel to obtain the synthesized speech New_Voice.

[0178] Then, the text content of New_Voice is the content of INF, and the prosody is the style of the target speaker.

[0179] This embodiment 2 also simultaneously discloses a multi-granularity prosodic speech synthesis device combined with text grammar, which uses the above-mentioned multi-granularity prosodic speech synthesis method combined with text grammar.

[0180] The multi-granularity prosodic speech synthesis device combined with text grammar includes: a data acquisition module, a preprocessing module, a data processing module, and a speech generation module.

[0181] The data acquisition module is configured to: obtain the text to be processed INF and the original reference audio Voice0 of the target speaker;

[0182] The preprocessing module is configured as follows: performing syntax extraction processing and phoneme conversion processing on INF to obtain a syntax graph G and a phoneme sequence Phoneme; performing audio conversion processing on Voice0 to obtain the original Mel-spectrogram Mel_S0.

[0183] The data processing module is configured as follows: the trained grammar of Phoneme, G, and Mel_S0 is injected into the multi-granularity prosodic network for processing to obtain the synthetic mel-spectrogram New_Mel.

[0184] The voice generation module is configured to perform voice coding on New_Mel to obtain the synthesized voice New_Voice.

[0185] Since this device uses the above-mentioned multi-granularity prosodic speech synthesis method combined with text grammar, it also has the same effect and will not be repeated here.

[0186] Example 3

[0187] 1. In Example 3, the performance of parallel and non-parallel synthesized speech was compared between the GMG-ProsodyNet and three existing networks (FG-transformer-TTS, AdaSpeech, and SyntaSpeech) (MOS and DMOS were used as indicators). The results are shown in Table 1.

[0188] Among them, MOS is the mean opinion score, which is a standard used to evaluate the naturalness of synthesized audio; the larger the MOS value, the closer the synthesized speech is to real human speech in terms of auditory perception, and the better the naturalness and expressiveness.

[0189] DMOS is the differential mean opinion score, which is a standard for evaluating the prosodic similarity between generated audio and reference audio. A larger DMOS value indicates that the learned prosody is more diverse.

[0190] It should be noted that SyntaSpeech does not have the ability to synthesize non-parallel speech, so it does not have non-parallel MOS values ​​or DMOS values.

[0191] Table 1 Network performance comparison

[0192]

[0193] As shown in Table 1, among the four networks, GMG-ProsodyNet has the largest MOS value, which indicates that the naturalness of the synthesized speech has been significantly improved. GMG-ProsodyNet also has the largest DMOS value, which shows that it is very effective in learning and controlling the rich prosodic features in synthesized audio.

[0194] In order to more objectively evaluate the effect of GMG-ProsodyNet in learning rhythmic diversity, this Example 3 compares the original mel spectrogram and the synthetic mel spectrograms generated by four networks, as shown in Figure 3. Figures 12 to 16 As shown: Figure 12 is the original Mel spectrogram; Figures 13 to 16 These are the synthesized mel-spectrograms obtained by GMG-ProsodyNet, FG-transformer-TTS, SyntaSpeech, and AdaSpeech respectively.

[0195] based on Figures 12 to 16 By comparison, it can be seen that the synthesized mel-spectrogram obtained by GMG-ProsodyNet is highly similar to the original mel-spectrogram and has a richer variation trend, which shows that GMG-ProsodyNet can effectively learn rhythmic diversity.

[0196] In addition, Figure 13 、 Figure 14 By comparison, we can see that: Figure 14 The selected part of the box has obvious word duplication, and Figure 13 There are no repeated words in the selected part, which shows that GMG-ProsodyNet also solves the problem of repeated word skipping in synthesized audio.

[0197] The above results verify the performance advantages of GMG-ProsodyNet and also illustrate the effectiveness of the above-mentioned multi-granularity prosodic speech synthesis method combined with text grammar.

[0198] 2. In Example 3, an ablation experiment is performed on the above-mentioned GMG-ProsodyNet to illustrate the effect changes brought about by its network structure.

[0199] Specifically, we gradually removed the grammatical-level encoder and the word-level encoder to examine the network performance (PESQ and duration loss were selected as the indicators). The results are shown in Table 2 below.

[0200] Among them, PESQ is the perceptual evaluation of speech quality. The larger the value, the higher the speech quality.

[0201] Duration loss represents the accuracy of duration prediction. The smaller its value is, the better the network processing effect is.

[0202] Table 2 Ablation experiment results

[0203]

[0204] As can be seen from Table 2, with the gradual removal of the grammatical level encoder and the word level encoder, the speech quality has decreased (the PESQ value has gradually become smaller), which shows the effectiveness of the grammatical level encoder and the word level encoder.

[0205] Moreover, although adding a word-level encoder can improve PESQ, it increases the duration loss; after adding a grammatical encoder, PESQ increases and the duration loss decreases, which shows the importance of incorporating grammatical information into the network.

[0206] Example 4

[0207] This embodiment 4 discloses a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the multi-granularity prosodic speech synthesis method combined with text grammar disclosed in embodiment 1.

[0208] This embodiment 4 also discloses a readable storage medium, which stores computer program instructions. When the computer program instructions are read and executed by a processor, the steps of the multi-granularity prosodic speech synthesis method combined with text grammar disclosed in embodiment 1 are executed.

[0209] This embodiment 4 further discloses a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the multi-granularity prosodic speech synthesis method combined with text grammar disclosed in embodiment 1 are implemented.

[0210] The above-described embodiments merely represent several implementation methods of the present invention. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent. It should be noted that a person skilled in the art would be able to make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the patent for this invention shall be determined by the appended claims.

Claims

1. A multi-granularity prosodic speech synthesis method combined with text grammar, characterized in that: include: Step 1: Get the text to be processed INF and the original reference audio Voice0 of the target speaker; Step 2: Perform syntax extraction and phoneme conversion on INF to obtain syntax graph G and phoneme sequence Phoneme; Perform audio conversion on Voice0 to obtain the original Mel spectrogram Mel_S0; Step 3: Input the trained grammar of Phoneme, G, and Mel_S0 into the multi-granularity prosodic network for processing to obtain the synthesized mel-spectrogram New_Mel; Step 4, performing voice coding conversion on New_Mel to obtain the synthesized speech New_Voice; The method for constructing a grammar-injected multi-granularity prosodic network includes: The FastSpeech2 network is used as the base network, and a multi-granularity prosody encoder is added between the phoneme encoder and the variation adapter, thus obtaining a grammar-injected multi-granularity prosody network. The multi-granularity prosody encoder is used to combine G and Mel_S0 to embed the text output by the phoneme encoder into E c Processed into text hidden optimization embedded E h ;E h Multi-granularity prosodic features are fused and used as input to the variational adapter.

2. The multi-granularity prosodic speech synthesis method combined with text grammar according to claim 1, characterized in that: The multi-granularity prosodic encoder includes: 1 syntax-level encoder, 1 utterance-level encoder, 1 phoneme-level encoder, 1 word-level encoder, 4 expanders, 2 sub-overlay layers, and 1 word segmenter; Inject the trained grammar in step 3 into the multi-granularity prosodic network: The syntax-level encoder converts E c , G encoding is integrated into grammatical prosodic embedding E s ; The first expander will E s The length is extended to E c same; The utterance-level encoder encodes Mel_S0 into an utterance-level prosodic embedding E u ; The second expander will E u The length is extended to E c same; The first sub-layer combines the output of the first expander, the output of the second expander, and the E c Superposition is performed to obtain the relay hidden embedding E' c ; The phoneme-level encoder converts E' c Encoded into phoneme-level prosodic embedding E ph ; The 4th expander will be E' c The length is adjusted to Mel length; The word segmenter performs word-level segmentation on the output of the fourth expander; The word-level encoder encodes the output of the word segmenter into a word-level prosodic embedding E w ; The third expander will E w The length is extended to E ph 、E' c same; The second sub-layer will output the third expander, E ph 、E' c Superposition to obtain E h .

3. The multi-granularity prosodic speech synthesis method combined with text grammar according to claim 2, characterized in that: The grammatical level encoder includes: 1 word-level average pooling layer, 2 GGNN network layers, and 1 sub-overlay layer; In the grammar-level encoder: The word-level average pooling layer is used to perform word-level average pooling on the input data in1; The first GGNN network layer is used to perform preliminary aggregation of syntactic information on the output of the word-level average pooling layer using the input data in0; The second GGNN network layer is used to further aggregate the syntactic information of the output of the first GGNN network layer using in0; The sub-overlay layer is used to superimpose in1, the output of the first GGNN network layer, and the output of the second GGNN network layer to obtain the output data out1; Among them, the trained grammar in step 3 is injected into the multi-granularity prosodic network: in1 is E c ;in0 is G;out1 is E s .

4. The multi-granularity prosodic speech synthesis method combined with text grammar according to claim 2 is characterized in that: The utterance-level encoder consists of two one-dimensional convolutional layers, two ReLU activation functions, two normalization layers, two Dropout layers, and one one-dimensional average pooling layer. In an utterance-level encoder: The first one-dimensional convolution layer is used to perform one-dimensional convolution processing on the input data in2; The first Relu activation function is used to activate the output of the first one-dimensional convolutional layer; The first normalization layer is used to normalize the output of the first Relu activation function; The first Dropout layer is used to simplify the output of the first normalization layer; The second one-dimensional convolution layer is used to perform one-dimensional convolution on the output of the first Dropout layer; The second Relu activation function is used to activate the output of the second one-dimensional convolutional layer; The second normalization layer is used to normalize the output of the second Relu activation function; The second Dropout layer is used to simplify the output of the second normalization layer; The one-dimensional average pooling layer is used to perform one-dimensional average pooling on the output of the second Dropout layer to obtain the output data out2; Among them, the trained grammar in step 3 is injected into the multi-granularity prosodic network: in2 is Mel_S0; out2 is E u .

5. The multi-granularity prosodic speech synthesis method combined with text grammar according to claim 2, characterized in that: The word-level encoder consists of two 2D convolutional layers, two ReLU activation functions, two normalization layers, two Dropout layers, one 2D adaptive average pooling layer, and one linear layer. In an utterance-level encoder: The first two-dimensional convolution layer is used to perform two-dimensional convolution processing on the input data in3; The first Relu activation function is used to activate the output of the first two-dimensional convolutional layer; The first normalization layer is used to normalize the output of the first Relu activation function; The first Dropout layer is used to simplify the output of the first normalization layer; The second two-dimensional convolution layer is used to perform two-dimensional convolution processing on the output of the first Dropout layer; The second Relu activation function is used to activate the output of the second two-dimensional convolutional layer; The second normalization layer is used to normalize the output of the second Relu activation function; The second Dropout layer is used to simplify the output of the second normalization layer; The two-dimensional adaptive average pooling layer is used to perform two-dimensional adaptive average pooling on the output of the second Dropout layer; The linear layer is used to transform the output of the two-dimensional adaptive average pooling layer to obtain the output data out3; Among them, the trained grammar in step 3 is injected into the multi-granularity prosodic network: in3 is the output of the word segmenter; out3 is E w .

6. The multi-granularity prosodic speech synthesis method combined with text grammar according to claim 2, characterized in that: The phoneme-level encoder consists of two one-dimensional convolutional layers, two ReLU activation functions, two normalization layers, two Dropout layers, and one linear layer. In the phoneme-level encoder: The first one-dimensional convolution layer is used to perform one-dimensional convolution processing on the input data in4; The first Relu activation function is used to activate the output of the first one-dimensional convolutional layer; The first normalization layer is used to normalize the output of the first Relu activation function; The first Dropout layer is used to simplify the output of the first normalization layer; The second one-dimensional convolution layer is used to perform one-dimensional convolution on the output of the first Dropout layer; The second Relu activation function is used to activate the output of the second one-dimensional convolutional layer; The second normalization layer is used to normalize the output of the second Relu activation function; The second Dropout layer is used to simplify the output of the second normalization layer; The linear layer is used to transform the output of the second Dropout layer to obtain the output data out4; Among them, the trained grammar in step 3 is injected into the multi-granularity prosodic network: in4 is E' c ;out4 is E ph .

7. The multi-granularity prosodic speech synthesis method combined with text grammar according to claim 1 is characterized in that: The base network includes: 1 phoneme embedding layer, 1 phoneme encoder, 1 variation adapter, 1 Mel decoder, 2 parent stacking layers, and 1 linear layer; In step 3, the phoneme embedding layer embeds Phoneme; the first parent superposition layer adds position encoding to the output of the phoneme embedding layer; the phoneme encoder encodes the output of the first parent superposition layer to obtain E c ; Variation adapter based on E h Calculate the duration, pitch, and energy information; the second parent overlay layer to E h Add position encoding and duration, pitch, and energy information; the Mel decoder decodes the output of the second parent stacking layer to obtain a high-dimensional Mel spectrogram; the linear layer reduces the dimension of the high-dimensional Mel spectrogram to obtain New_Mel.

8. The multi-granularity prosodic speech synthesis method combined with text grammar according to claim 1 is characterized in that: Methods for obtaining a trained grammar-injected multi-granularity prosodic network include: S1, obtain a sample set with true labels; S2, divide the sample set into training set and validation set; S3, training the grammar-injected multi-granularity prosodic network for N rounds based on the training set; verifying the multi-granularity prosodic network after each round of training based on the validation set until the grammar-injected multi-granularity prosodic network with the best performance is selected; In the first N / 2 rounds of training, all layers in the base network work, and all parts of the multi-granularity prosody encoder except the fourth expander and word segmenter work; The loss functions used in the first N / 2 rounds of training include: prediction loss of the grammar-injected multi-granularity prosody network, duration loss of the variational adapter, pitch loss of the variational adapter, and energy loss of the variational adapter; In the next N / 2 rounds of training: all layers in the base network work, and all parts in the multi-granularity prosody encoder work; The loss functions used in the last N / 2 rounds of training include: prediction loss of the grammar-injected multi-granularity prosodic network, duration loss of the variational adapter, pitch loss of the variational adapter, energy loss of the variational adapter, encoding loss of the phoneme-level encoder, and encoding loss of the word-level encoder.

9. The multi-granularity prosodic speech synthesis method combined with text grammar according to claim 8, characterized in that: In each round of training, some data are randomly extracted from the training set as a training subset, and a total of N training subsets {U1, ..., U N } to carry out N rounds of training; Among them, the method of the p-th round training includes: First, the pth training subset U p Processed into the pth phoneme sample sequence Phoneme p , the pth grammar graph G p , the pth original Mel spectrum sample image Mel_S p,0 , the p-th phoneme-level Mel spectrum sample graph Mel_S p,1 , the p-th word-level Mel spectrum sample graph Mel_S p,2 ; then Phoneme p After the phoneme embedding layer, the first parent superposition layer, and the phoneme encoder, the pth text sample hidden embedding E is obtained p,c ;E p,c , G p 、Mel_S p,0 、Mel_S p,1 、Mel_S p,2 Enter the multi-granularity prosody encoder process and obtain the p-th text sample hidden optimized embedding E p,h ; then E p,h First calculate the duration D of the p-th round of training through the variation adapter p , pitch Pi p Energy E p , and then pass through the second parent overlay layer, Mel decoder, and linear layer to obtain the pth synthetic sample Mel spectrogram New_Mel p ;p∈[1,N / 2]; Among them, E p,c , G p The p-th grammatical level sample prosody embedding E is obtained by encoding it through the grammatical level encoder p,s ;E p,s Adjust to the length of E through the first expander p,c Same; Mel_S p,0 The prosodic embedding E of the p-th utterance-level sample is obtained by encoding it through the utterance-level encoder p,u ;E p,u Adjust the length to the same as E through the second expander p,c Same; E p,c The output of the first expander and the output of the second expander are superimposed by the first superimposer to obtain the hidden embedding E' of the pth relay sample p,c ;Mel_S p,1 The p-th phoneme-level sample prosody embedding E is obtained by encoding it through the phoneme-level encoder p,ph ;Mel_S p,2 The p-th word-level sample prosody embedding E is obtained by encoding it through the word-level encoder p,w ;E p,w Adjust the length to E' through the third expander p,c 、E p,ph Same; E' p,c 、E p,ph The output of the third expander is superimposed by the second superimposer to obtain E p,h ; Among them, the loss function Loss used in the p-th round of training P for: Where, L p,1 represents the prediction loss of GMG-ProsodyNet in the p-th round of training; L p,d represents the duration loss of the deteriorating adapter in the p-th round of training; L p,pi represents the pitch loss of the deviated adapter in the p-th round of training; L p,e represents the energy loss of the degraded adapter in the p-th round of training; ||.||1 represents the L1 loss function; New_Mel p,true represents the true value of the Mel spectrogram corresponding to the reference synthesized sample speech in the p-th round of training; MSE(.,.) represents the mean square error function; D p,true represents the actual value of the duration of the deterioration adapter in the p-th round of training; Pi p,true represents the true value of the pitch of the variant adapter in the p-th round of training; E p,true represents the true energy value of the degraded adapter in the p-th round of training; Among them, the method of the qth round of training includes: First, the qth training subset U q Processed into the qth phoneme sample sequence Phoneme q , the qth grammar graph G q , the qth original Mel spectrum sample image Mel_S q,0 , the qth phoneme-level Mel spectrum sample graph Mel_S q,1 , the qth word-level Mel spectrum sample graph Mel_S q,2 ; then Phoneme q After the phoneme embedding layer, the first parent superposition layer, and the phoneme encoder, the qth text sample hidden embedding E is obtained q,c ;E q,c , G q 、Mel_S q,0 、Mel_S q,1 、Mel_S q,2 Enter the multi-granularity prosody encoder process and get the qth text sample hidden optimized embedding E q,h ; then E q,h First calculate the duration D of the qth round of training through the variation adapter q , pitch Pi q Energy E q , and then pass the second parent overlay layer, Mel decoder, and linear layer to obtain the qth synthetic sample Mel spectrogram New_Mel q ;q∈[N / 2+1,N]; Among them, E q,c , G q The qth grammatical level sample prosody embedding E is obtained by encoding it through the grammatical level encoder q,s ;E q,s Adjust to the length of E through the first expander q,c Same; Mel_S q,0 The qth utterance-level sample prosody embedding E is obtained by encoding it through the utterance-level encoder q,u ;E q,u Adjust the length to the same as E through the second expander q,c Same; E q,c The output of the first expander and the output of the second expander are superimposed by the first superimposer to obtain the qth relay sample hidden embedding E' q,c ; E' q,c The qth phoneme-level sample prosody embedding E is obtained by encoding it through the phoneme-level encoder q,ph ;Mel_S q,1 The qth phoneme-level sample prosodic contrast embedding E' is obtained by encoding it through the phoneme-level encoder q,ph ; E' q,c First, the length of the word is adjusted to Mel length through the fourth expander, and then the word segmenter is used for segmentation. Then, the word-level encoder is used to encode the qth word-level sample rhythm embedding E. q,w ;E q,w Adjust the length to E' through the third expander q,c 、E q,ph Same; Mel_S q,2 The qth word-level sample prosody contrast embedding E' is obtained by encoding it through the word-level encoder q,w ; E' q,c 、E q,ph The output of the third expander is superimposed by the second superimposer to obtain E q,h ; Among them, the loss function Loss used in the qth round of training q for: Where, L q,1 represents the prediction loss of GMG-ProsodyNet in the qth round of training; L q,d represents the duration loss of the qth round of training of the deviating adapter; L q,pi represents the pitch loss of the qth round of training variant adapter; L q,e represents the energy loss of the degraded adapter in the qth round of training; L q,ph represents the encoding loss of the qth round of training phoneme-level encoder; L q,w represents the encoding loss of the word-level encoder in the qth round of training; ||.||1 represents the L1 loss function; New_Mel q,true represents the true value of the Mel spectrogram corresponding to the reference synthesized sample speech in the qth round of training; MSE(.,.) represents the mean square error function; D q,true represents the actual duration of the variation adapter in the qth round of training; Pi q,true represents the true value of the pitch of the variant adapter in the qth round of training; E q,true represents the true value of the energy of the degraded adapter in the qth round of training.

10. A multi-granularity prosodic speech synthesis device combined with text grammar, characterized in that: It uses the multi-granularity prosodic speech synthesis method combined with text grammar as described in any one of claims 1-8; The multi-granularity prosodic speech synthesis device combined with text grammar includes: A data acquisition module is used to obtain the text to be processed INF and the original reference audio Voice0 of the target speaker; The preprocessing module is used to perform syntax extraction and phoneme conversion on INF to obtain a syntax graph G and a phoneme sequence Phoneme; perform audio conversion on Voice0 to obtain an original Mel-spectrogram Mel_S0; The data processing module is used to inject the trained grammar of Phoneme, G, and Mel_S0 into the multi-granularity prosodic network for processing to obtain the synthesized mel-spectrogram New_Mel; and a speech generation module, which is used to perform voice coding conversion on New_Mel to obtain the synthesized speech New_Voice.

Citation Information

Patent Citations

  • Low-resource Lao speech synthesis method based on fine-grained rhythm modeling

    CN115910023A

  • Rhythm migration speech synthesis method and system

    CN115910026A