A method, apparatus, electronic device and storage medium for speech synthesis

By constructing a sentence collection and extracting cross-sentence features to train a speech synthesis model, the method enhances the rhythm and intonation of synthesized speech by leveraging contextual chapter structure information.

CN115346510BActive Publication Date: 2025-07-15JINGDONG TECH HLDG CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202110527979.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-14
Publication Date
2025-07-15
Estimated Expiration
2041-05-14

AI Technical Summary

Technical Problem

In the existing pronunciation synthesis technology, the pronunciation effect of pronunciation is poor, mainly because only the current sentence is used for training, and the contextual chapter structure information cannot be effectively utilized.

Method used

By obtaining the context sentences of the current sentence, constructing a sentence set, and using cross-sentence features to train a pronunciation synthesis model, including cross-sentence feature extraction methods such as CSE, PSE, word vector table and knowledge distillation Student model, combined with the self-attention mechanism network structure, the pronunciation effect of the model is improved.

Benefits of technology

It improves the rhythmic effect of pronunciation, and can synthesize more natural and richer pronunciation based on the contextual text structure, enhancing the nature and fluency of pronunciation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115346510B_ABST
    Figure CN115346510B_ABST
Patent Text Reader

Abstract

The present application discloses a speech synthesis method. The speech synthesis method includes: obtaining context sentences of a current sentence, and constructing a sentence set including the current sentence and the context sentences; performing a text feature extraction operation on the sentences in the sentence set to obtain cross-sentence features, and using the cross-sentence features to train a speech synthesis model; and synthesizing speech information of a target text by using the trained speech synthesis model. In the present application, the speech synthesis model is trained by using cross-sentence features. Since the cross-sentence features can describe the discourse structure of the context of the text, the trained speech synthesis model can synthesize speech based on the context discourse structure, thereby improving the prosody effect of speech synthesis. The present application also discloses a speech synthesis device, an electronic device and a storage medium, which have the above beneficial effects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of speech synthesis, and particularly to a speech synthesis method, an apparatus, an electronic device, and a storage medium. Background Art

[0002] Speech synthesis refers to the technology of generating artificial speech through mechanical or electronic means. Through speech synthesis technology, the text information generated by a computer or input externally can be converted into fluent audio content that users can understand.

[0003] Currently, speech synthesis is mainly achieved by training a speech synthesis model. However, in related technologies, only the current sentence of the training data is used to train the speech synthesis model, resulting in poor prosody effects in speech synthesis.

[0004] Therefore, how to improve the prosody effect of speech synthesis is a technical problem that those skilled in the art need to solve currently. Summary of the Invention

[0005] The purpose of this application is to provide a speech synthesis method, an apparatus, an electronic device, and a storage medium, which can improve the prosody effect of speech synthesis.

[0006] To solve the above technical problem, this application provides a speech synthesis method, which includes:

[0007] Obtain the context sentences of the current sentence, and construct a sentence set including the current sentence and the context sentences;

[0008] Perform text feature extraction operations on the sentences in the sentence set to obtain cross-sentence features, and use the cross-sentence features to train a speech synthesis model;

[0009] Use the trained speech synthesis model to synthesize the speech information of the target text.

[0010] Optionally, performing text feature extraction operations on the sentences in the sentence set to obtain cross-sentence features includes:

[0011] Determine a forward sentence set and a backward sentence set; wherein, the forward sentence set includes the current sentence and the context sentences before the current sentence, and the backward sentence set includes the current sentence and the context sentences after the current sentence;

[0012] Calculate the forward cross-sentence features of the forward sentence set and the backward cross-sentence features of the backward sentence set.

[0013] Optionally, calculating the forward cross-sentence features of the forward sentence set and the backward cross-sentence features of the backward sentence set includes:

[0014] Input all the forward sentence sets into the language model to obtain the forward cross-sentence features;

[0015] Input all the backward sentence sets into the language model to obtain the backward cross-sentence features.

[0016] Optionally, training a speech synthesis model using the cross-sentence features includes:

[0017] Concatenate the forward cross-sentence features and the backward cross-sentence features to obtain context cross-sentence features;

[0018] Obtain the phoneme features output by the encoder of the speech synthesis model;

[0019] Concatenate the context cross-sentence features and the phoneme features to obtain a concatenated feature vector, and train the speech synthesis model using the concatenated feature vector.

[0020] Optionally, performing a text feature extraction operation on the sentences in the sentence set to obtain cross-sentence features includes:

[0021] Obtain adjacent sentence pairs in the sentence set;

[0022] Input the adjacent sentence pairs into the language model to obtain the cross-sentence features of the adjacent sentence pairs;

[0023] Correspondingly, training a speech synthesis model using the cross-sentence features includes:

[0024] Perform a weighted sum on the cross-sentence features of all the adjacent sentence pairs to obtain weighted cross-sentence features;

[0025] Control each phoneme feature of the encoder of the speech synthesis model to learn the weighted cross-sentence features through a self-attention mechanism network structure; wherein, the key vector and the value vector of the self-attention mechanism network structure are the weighted cross-sentence features, and the query vector of the self-attention mechanism network structure is the phoneme feature vector of the encoder;

[0026] Train the speech synthesis model using the weighted cross-sentence features.

[0027] Optionally, performing a text feature extraction operation on the sentences in the sentence set to obtain cross-sentence features includes:

[0028] Query the word vectors of each word in the sentence set using a word vector table, and calculate the average value of the word vectors corresponding to the sentence text in the context sentence to obtain a sentence text vector;

[0029] Concatenate the sentence text vectors to obtain the cross-sentence features.

[0030] Optionally, before performing the text feature extraction operation on the sentences in the sentence set to obtain cross-sentence features, it further includes:

[0031] Performing knowledge distillation on the language model to obtain a Student model, so that the Student model learns the cross-sentence features extracted by the language model;

[0032] Correspondingly, performing the text feature extraction operation on the sentences in the sentence set to obtain cross-sentence features includes:

[0033] Performing the text feature extraction operation on the sentences in the sentence set by using the Student model to obtain the cross-sentence features.

[0034] This application also provides a speech synthesis device, which includes:

[0035] A set construction module, configured to obtain the context sentences of the current sentence and construct a sentence set including the current sentence and the context sentences;

[0036] A model training module, configured to perform a text feature extraction operation on the sentences in the sentence set to obtain cross-sentence features, and use the cross-sentence features to train a speech synthesis model;

[0037] A speech synthesis module, configured to synthesize the speech information of the target text by using the trained speech synthesis model.

[0038] This application also provides a storage medium, on which a computer program is stored, and when the computer program is executed, the steps performed by the above speech synthesis method are implemented.

[0039] This application also provides an electronic device, including a memory and a processor, where a computer program is stored in the memory, and when the processor calls the computer program in the memory, the steps performed by the above speech synthesis method are implemented.

[0040] This application provides a speech synthesis method, including: obtaining the context sentences of the current sentence, constructing a sentence set including the current sentence and the context sentences; performing a text feature extraction operation on the sentences in the sentence set to obtain cross-sentence features, and using the cross-sentence features to train a speech synthesis model; synthesizing the speech information of the target text by using the trained speech synthesis model.

[0041] This application obtains the context sentences of the current sentence, and uses the sentence set including the current sentence and the context sentences as samples for training a speech synthesis model. Specifically, this application performs text feature extraction operations on the sentences in the sentence set to obtain cross-sentence features. The cross-sentence features can describe the discourse structure of the text context. Therefore, the speech synthesis model trained using the cross-sentence features can learn the discourse structure features. Since the prosody information when the text is read is related to the context discourse structure, the same sentence may have completely different prosody performances in different context contexts. The speech synthesis model trained by this application using cross-sentence features can synthesize speech based on the context discourse structure, thereby improving the prosody effect of speech synthesis. This application also provides a speech synthesis device, an electronic device, and a storage medium, which have the above beneficial effects and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0043] Figure 1 It is a flowchart of a speech synthesis method provided by an embodiment of the present application;

[0044] Figure 2 It is an architecture diagram of an end-to-end speech synthesis model provided by an embodiment of the present application;

[0045] Figure 3 It is a flowchart of a speech synthesis model training method provided by an embodiment of the present application;

[0046] Figure 4 It is a schematic diagram of the network structure of a CSE cross-sentence feature encoder provided by an embodiment of the present application;

[0047] Figure 5 It is a flowchart of a speech synthesis model training method provided by an embodiment of the present application;

[0048] Figure 6 It is a schematic diagram of the network structure of a PSE cross-sentence feature encoder provided by an embodiment of the present application;

[0049] Figure 7 It is a structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.

[0051] The end-to-end (sequence to sequence) speech synthesis technology is one of the current mainstream speech synthesis technologies. Common end-to-end speech synthesis models include Tacotron1 (an end-to-end text-to-speech deep learning model), Tacotron2, DurIAN (a model that combines traditional parametric speech synthesis technology and end-to-end speech synthesis technology), and various variants of similar seq2seq models. There are also encoder / decoder network structures based on the Transformer network structure, such as FastSpeech (a speech synthesis model), FastSpeech2, etc. However, the above-mentioned end-to-end speech synthesis models do not use discourse structure information for training, but only use the linguistic features of the current sentence for speech synthesis. For example, in the prior art, in order to learn the representation of prosody, Global StyleToken (GST) is usually learned for prosody control. However, the above method uses the audio data of the current sentence to learn GST and does not use any cross-sentence information. Therefore, the prosody effect of the synthesized speech in the related technology is poor. To solve the various defects existing in the above-mentioned related technologies, this application provides a new speech synthesis solution through the following several embodiments, which can improve the prosody effect of speech synthesis.

[0052] Please refer to Figure 1 , Figure 1 which is a flowchart of a speech synthesis method provided by an embodiment of this application.

[0053] The specific steps may include:

[0054] S101: Obtain the context sentence of the current sentence and construct a sentence set including the current sentence and the context sentence;

[0055] Among them, this embodiment can be applied to an end-to-end speech synthesis model. The main speech synthesis framework of this speech synthesis model can be models such as Tacotron, FastSpeech1, FastSpeech2, EAST (a text detection model), DeepVoice, or ClariNet (a neural network speech synthesis model). The current sentence can be any sentence in the training data. In this embodiment, the size of the preset context query window can be set, and the upper and lower context sentences can be queried with the current sentence as the center of the context query window, so as to obtain the context sentences. The upper context sentence is the sentence in the training data before the current sentence, and the lower context sentence is the sentence in the training data after the current sentence. As a feasible implementation manner, if the current sentence is the Nth sentence in the training data and the size of the context query window is 9 in this embodiment, the context sentences are the (N-4)th, (N-3)th, (N-2)th, (N-1)th, (N+1)th, (N+2)th, (N+3)th, and (N+4)th sentences in the training data.

[0056] After obtaining the current sentence and the context sentences, this embodiment can construct a sentence set including the current sentence and the context sentences. After training the speech synthesis model using the sentence set corresponding to the current sentence, this embodiment can also re-determine a new current sentence and construct a sentence set corresponding to the new current sentence, so as to train the speech synthesis model again.

[0057] S102: Perform a text feature extraction operation on the sentences in the sentence set to obtain cross-sentence features, and use the cross-sentence features to train the speech synthesis model;

[0058] Among them, on the basis of obtaining the sentence set, this embodiment can perform a text feature extraction operation on each sentence in the sentence set. Since the sentence set includes the current sentence and its context sentences, the extraction result of performing the text feature extraction on the sentence set is the cross-sentence feature of the current sentence.

[0059] Specifically, the following methods can be used in this embodiment to extract cross-sentence features: Method (1), perform text feature extraction operations on the sentences in the sentence set based on CSE (Chunked Sentence Embedding) to obtain cross-sentence features; Method (2), perform text feature extraction operations on the sentences in the sentence set based on PSE (Paired Sentence Embedding) to obtain cross-sentence features; Method (3), perform text feature extraction operations on the sentences in the sentence set based on a word vector table (such as a character vector list) to obtain cross-sentence features; Method (4), perform text feature extraction operations on the sentences in the sentence set based on the Student model obtained by knowledge distillation to obtain cross-sentence features. Further, this embodiment can use any one or a combination of any several of the above Method (1), Method (2), Method (3), and Method (4) to extract cross-sentence features. If multiple methods are used to extract cross-sentence features, this embodiment can determine the cross-sentence features extracted by each method, and calculate the average value of all the cross-sentence features to obtain the cross-sentence features finally used for training the speech synthesis model.

[0060] After obtaining the cross-sentence features, this application can use the cross-sentence features to train the speech synthesis model, so that the speech synthesis model can learn the context discourse structure information in the training data, thereby improving the prosody effect of speech synthesis.

[0061] S103: Synthesize the speech information of the target text using the trained speech synthesis model.

[0062] Among them, this embodiment can iteratively train the speech synthesis model by updating the current sentence multiple times. After the speech synthesis model is trained, the target text of the speech to be synthesized can be input into the speech synthesis model, so as to synthesize the speech information corresponding to the target text using the trained speech synthesis model.

[0063] This embodiment obtains the context sentences of the current sentence, and uses the sentence set including the current sentence and the context sentences as a sample for training the speech synthesis model. Specifically, this embodiment performs text feature extraction operations on the sentences in the sentence set to obtain cross-sentence features. The cross-sentence features can describe the context discourse structure of the text. Therefore, the speech synthesis model trained using the cross-sentence features can learn the discourse structure features. Since the prosody information when the text is read is related to the context discourse structure, the same sentence may have completely different prosody performances in different context contexts. The speech synthesis model trained by this embodiment using the cross-sentence features can synthesize speech based on the context discourse structure, thereby improving the prosody effect of speech synthesis.

[0064] Although the current end-to-end speech synthesis technology has achieved a relatively natural and rich prosody speech synthesis effect, the related technology does not adopt the discourse structure information but only uses the linguistic features of the current sentence for speech synthesis. Usually, the prosody information is strongly related to the discourse structure of the context. The same text sentence will have completely different prosody performances in different context situations. Therefore, when the end-to-end system that only uses the text features of the current sentence for speech synthesis synthesizes a text, it is very difficult to convert a text into natural and rich prosody speech according to the context information.

[0065] Please refer to Figure 2 , Figure 2 which is an architecture diagram of an end-to-end speech synthesis model provided by an embodiment of the present application. Figure 2 Input1 in it is the phoneme sequence of the current sentence, Input2 is the acoustic feature Mel-Spectrum predict (Mel spectrum prediction value) predicted by the speech synthesis model last time, 1d convolution refers to one-dimensional convolution, Bi-directional LSTM (Long Short-Term Memory) refers to a bidirectional long short-term memory artificial neural network, Attention Mechanism refers to the attention mechanism, FC (fully connected) refers to the fully connected layer, query refers to the query vector, context vector refers to the context vector, the plus sign refers to the parameter adjustment process of the model, Stop token predict refers to the probability prediction module, which is used to calculate the parameters for stopping training the model. In the above speech synthesis model, the phoneme sequence of the current sentence can be subjected to three convolution calculations and input into the bidirectional long short-term memory artificial neural network for processing, and the cross-sentence features are concatenated with the output features of the long short-term memory artificial neural network and input into the Attention Mechanism module. The present application can also use the acoustic feature predicted last time as the input feature, and successively perform calculations through the fully connected layer and LSTM to obtain the query vector, and input the query vector into the Attention Mechanism module. The Attention Mechanism module generates a context vector according to the concatenation result of the cross-sentence features and the output features of the long short-term memory artificial neural network and the query vector, and finally adjusts the parameters of the speech synthesis model according to the context vector and outputs the acoustic feature predicted this time. After the value predicted by Stop token predict is greater than a specific value, the training of the speech synthesis model is stopped.

[0066] The above speech synthesis model combines the text features of the current sentence and the context information corresponding to this sentence to extract cross-sentence features for model training, improving the prosody effect of the model. When training the model, a contextually continuous TTS (Text To Speech) corpus can also be used. For example, the Chinese training data in this embodiment can use contextually continuous male voice novel reading data, and the English training data can use contextually continuous female non-fiction data.

[0067] To obtain the cross-sentence features of the context text, this embodiment can use the BERT model to extract the cross-sentence features of the context sentences. As a feasible implementation, this embodiment can use a pre-trained open-source BERT model to extract cross-sentence features. To verify different usage methods of cross-sentence features, this embodiment can use CSE and PSE methods to extract cross-sentence features. This embodiment uses an end-to-end speech synthesis model similar to Tacotron2 as the basic framework, as Figure 2 shown, and then combines cross-sentence features under this basic framework to improve the prosody effect of the model.

[0068] As a feasible implementation, this application can perform text feature extraction operations on the sentences in the sentence set based on CSE to obtain cross-sentence features. The specific process is as Figure 3 shown, Figure 3 is a flowchart of a speech synthesis model training method provided by an embodiment of this application. This embodiment may include the following steps:

[0069] S301: Determine a forward sentence set and a backward sentence set;

[0070] Among them, the forward sentence set includes the current sentence and the context sentences before the current sentence, and the backward sentence set includes the current sentence and the context sentences after the current sentence;

[0071] S302: Calculate the forward cross-sentence features of the forward sentence set and the backward cross-sentence features of the backward sentence set.

[0072] Specifically, in this embodiment, cross-sentence features can be calculated using a language model, and the process is as follows: Input all the forward sentence sets into the language model to obtain the forward cross-sentence features; Input all the backward sentence sets into the language model to obtain the backward cross-sentence features. Specifically, the above language model can be a BERT model, GPT model, RoBERTa model, EMLo model, or Transformer-XL model, etc. In this embodiment, all the forward sentence sets can be used as input data and input into the language model, so that the language model can extract the semantic information contained in all the texts in the forward sentence sets, that is, obtain the forward cross-sentence features. This embodiment can also use all the backward sentence sets as input data and input into the language model, so that the language model can extract the semantic information contained in all the texts in the backward sentence sets, that is, obtain the backward cross-sentence features.

[0073] S303: Concatenate the forward cross-sentence features and the backward cross-sentence features to obtain context cross-sentence features;

[0074] S304: Obtain the phoneme features output by the encoder of the speech synthesis model;

[0075] Among them, the above phoneme features are the phoneme features obtained after the encoder of the speech synthesis model processes the current sentence.

[0076] S305: Concatenate the context cross-sentence features and the phoneme features to obtain a concatenated feature vector, and use the concatenated feature vector to train the speech synthesis model.

[0077] As a feasible implementation, after obtaining the concatenated feature vector, the concatenated feature vector can also be input into a linear transformation layer to map the concatenated feature vector to the same dimension as the encoder output vector, so as to reduce the number of parameters of the speech synthesis model. This embodiment provides a solution for improving the prosody performance of the model by combining context cross-sentence features. The cross-sentence feature extraction process can use models such as BERT, GPT, RoBERTa, EMLo, or Transformer-XL to extract the cross-sentence features of the current sentence, and the main speech synthesis framework can be a Tacotron, FastSpeech, FastSpeech2, EAST, DeepVoice, or ClariNet model.

[0078] The input of the encoder of the speech synthesis model is a phoneme sequence, and the phoneme sequence is the result of splitting the initials and finals of the pinyin. The encoder usually obtains a vector through a lookup table according to the ID of the phoneme, then performs a one-dimensional convolution operation on all the vectors, and finally passes through a bidirectional LSTM layer.

[0079] The decoder of the speech synthesis model is an autoregressive generative model based on LSTM. The input of the decoder is the acoustic features output at the previous moment. Then, the output of the last layer of LSTM is used as the query vector of the attention module and the concatenated feature vector for attention calculation, and a context vector is obtained. Then, the context vector is concatenated with the vector of the last layer of LSTM, passed through a fully connected layer, and the acoustic features at the current moment are predicted. Further, the acoustic features at the current moment will be used as the input for the next moment of the decoder (i.e., Figure 2 Input2 in

[0080] Please refer to Figure 4 , Figure 4 which is a schematic diagram of a network structure based on a CSE cross-sentence feature encoder provided by an embodiment of the present application, Figure 4 describing the specific process of extracting cross-sentence features based on CSE in the above embodiment. u i represents the text of the i-th sentence of the training data. SEP is the separator used in the BERT model to separate two sentences, and CLS is the classifier used in the BERT model to distinguish whether two sentences are consecutive sentences. p i is the i-th phoneme in the phoneme list corresponding to a piece of text. T is the number of phonemes in a piece of text. CU (Cross Utterance) represents cross-sentence. u0 is the current sentence. G2P refers to the process of converting morphemes in the current sentence into phonemes. Phoneme Encoder is the phoneme encoder. f(p i ) is the phoneme encoding result. cat refers to the concatenation operation. Linear Projection W is the weight linear projection. Decoder is the decoder. u P is the set of forward sentences. u N is the set of backward sentences. e(u P ) is the forward cross-sentence feature. e(u N ) is the backward cross-sentence feature. c is the context cross-sentence feature after concatenating the forward cross-sentence feature and the backward cross-sentence feature. CU Encoder is the cross-sentence encoder.

[0081] In this embodiment, assume that u0 is the current sentence and L is the number of context sentences considered in this embodiment. Then define u p ={CLS, u -L , SEP, u -L+1 , …SEP, u0} to represent the current sentence and the sentences before it. Define u N ={CLS, u0, SEP, u 0+1 , SEP, u 0+2 , …, SEP, u L}{For the current sentence and the sentences following it, in this embodiment, u and u are respectively sent into the BERT model, and then the vectors of the output layer of the corresponding CLS in u and u are extracted as the forward cross-sentence feature and the backward cross-sentence feature of the current sentence, as shown in. In the CSE cross-sentence feature usage method, the cross-sentence feature vectors corresponding to u and u are concatenated, and then concatenated with the phoneme feature vectors output by the encoder of the speech synthesis model. To reduce the number of parameters of the model, the concatenated feature vectors can be mapped to the same dimension as the output of the encoder of the speech synthesis model through a linear transformation layer. p and u N are sent into the BERT model, and then the vectors of the output layer of the corresponding CLS in u and u are extracted as the forward cross-sentence feature and the backward cross-sentence feature of the current sentence, as shown in. p and u N In the CSE cross-sentence feature usage method, the cross-sentence feature vectors corresponding to u and u are concatenated, and then concatenated with the phoneme feature vectors output by the encoder of the speech synthesis model. To reduce the number of parameters of the model, the concatenated feature vectors can be mapped to the same dimension as the output of the encoder of the speech synthesis model through a linear transformation layer. Figure 4 As shown in. In the CSE cross-sentence feature usage method, the cross-sentence feature vectors corresponding to u and u are concatenated, and then concatenated with the phoneme feature vectors output by the encoder of the speech synthesis model. To reduce the number of parameters of the model, the concatenated feature vectors can be mapped to the same dimension as the output of the encoder of the speech synthesis model through a linear transformation layer. p and u N The corresponding cross-sentence feature vectors are concatenated, and then concatenated with the phoneme feature vectors output by the encoder of the speech synthesis model. To reduce the number of parameters of the model, the concatenated feature vectors can be mapped to the same dimension as the output of the encoder of the speech synthesis model through a linear transformation layer.

[0082] As a feasible implementation manner, the present application can perform text feature extraction operations on the sentences in the sentence set based on PSE to obtain cross-sentence features. The specific process is as shown in. Figure 5 As shown in. Figure 5 is a flowchart of a speech synthesis model training method provided by an embodiment of the present application. This embodiment may include the following steps:

[0083] S501: Obtain adjacent sentence pairs in the sentence set;

[0084] Specifically, in this embodiment, two adjacent sentences in the sentence set can be used as adjacent sentence pairs, and the sentence in the upper context and the sentence in the lower context that are farthest from the current sentence can be set as adjacent sentence pairs, so as to obtain the same number of adjacent sentence pairs as the number of sentences in the sentence set.

[0085] S502: Input the adjacent sentence pairs into the language model to obtain the cross-sentence features of the adjacent sentence pairs;

[0086] S503: Perform weighted summation on the cross-sentence features of all the adjacent sentence pairs to obtain weighted cross-sentence features;

[0087] Among them, the weights in the weighted summation process can be obtained in the following manner: By calculating the correlation (such as vector multiplication) between the feature vectors of each phoneme and the cross-sentence features of each sentence pair, the weight information can be obtained.

[0088] S504: Control each phoneme feature of the encoder of the speech synthesis model to learn the weighted cross-sentence features through the self-attention mechanism network structure;

[0089] Among them, the key vector (i.e., the key vector) and the value vector (i.e., the value vector) of the self-attention mechanism network structure are the weighted cross-sentence features, and the query vector (i.e., the query vector) of the self-attention mechanism network structure is the phoneme feature vector of the encoder. The essence of the attention function in the self-attention mechanism network structure can be described as a mapping from a query vector to multiple key vector-value vector pairs. The process of calculating the attention score in the self-attention mechanism network structure can include the following steps: (1) Calculate the similarity between the query vector and each key vector to obtain weights; (2) Normalize the weights obtained in the previous step using the softmax function; (3) Perform a weighted sum of the weights and the corresponding value vectors to obtain the attention score attention. In the field of natural language processing, the key vector and the value vector are usually set to the same vector, that is, the key vector is equal to the value vector.

[0090] S505: Train the speech synthesis model using the weighted cross-sentence features.

[0091] In the above process, the feature vector obtained for each phoneme is used as the query vector, and then the sentence vectors of all the sentence texts obtained by BERT are queried. Then, each sentence vector will obtain a weight, and a weighted sum is performed according to the weights to obtain the cross-sentence features unique to each phoneme (i.e., different). In this embodiment, the self-attention mechanism network structure is used to obtain the cross-sentence features independent of the encoder features of each phoneme, so that each pronunciation unit obtains a fine-grained cross-sentence feature that is helpful for the pronunciation of the current unit, which is used to improve the prosody effect of the model.

[0092] Please refer to Figure 6 , Figure 6 which is a schematic diagram of the network structure of a PSE cross-sentence feature encoder provided by an embodiment of the present application, Figure 6 describing the specific process of extracting cross-sentence features based on PSE in the above embodiment. Figure 6 and Figure 5 The meanings of the same English expressions in Figure 6 are also the same, and will not be elaborated here. In Figure 6 , pair refers to adjacent sentence pairs, and Multi-Head Attention refers to the multi-head self-attention mechanism. As shown in Figure 6As shown. In order to integrate the cross-sentence features obtained from all sentence pairs and apply them to the speech synthesis model, this embodiment adopts a self-attention mechanism network structure to learn an independent cross-sentence feature for the phoneme feature vectors of each speech synthesis encoder model. This cross-sentence feature is the weighted sum of the cross-sentence features of all sentence pairs. In the self-attention mechanism network structure, the sequence of phoneme feature vectors of the speech synthesis encoder serves as the query vector of the self-attention mechanism network structure, and the cross-sentence feature vectors of all sentence pairs obtained by the PSE method serve as the key vector and value vector of the self-attention mechanism network structure. Adopting the self-attention mechanism network structure can enable each phoneme's encoder feature to learn a different cross-sentence feature, allowing phonemes at different positions to learn different cross-sentence features.

[0093] For the two different ways of using cross-sentence features mentioned above, an additional language model is required to extract cross-sentence features. However, due to the large number of parameters of the language model, it will bring an additional huge overhead to the deployment and inference of the model. If the speech synthesis model is deployed on a mobile device, the large number of parameters of the language model is not suitable for deployment on an offline model. Therefore, this application then proposes two simpler ways to extract cross-sentence features, which can be used to make up for the defects of the language model in terms of inference speed and disk space occupied by the model.

[0094] This application also provides a solution for extracting cross-sentence features by performing text feature extraction operations on the sentences in a sentence set based on a character vector table. The specific process is as follows: Use the character vector table to query the character vectors of each character in the sentence set, and calculate the average value of the character vectors corresponding to the sentence text in the context sentence to obtain a sentence text vector; Concatenate the sentence text vectors to obtain the cross-sentence feature.

[0095] Specifically, this embodiment can concatenate the sentence text vectors corresponding to the forward sentence set to obtain a forward cross-sentence feature; It can also concatenate the sentence text vectors corresponding to the backward sentence set to obtain a backward cross-sentence feature. Further, this embodiment can also concatenate the forward cross-sentence feature and the backward cross-sentence feature to obtain a context cross-sentence feature; Obtain the phoneme features output by the encoder of the speech synthesis model; Concatenate the context cross-sentence feature and the phoneme feature to obtain a concatenated feature vector, and use the concatenated feature vector to train the speech synthesis model.

[0096] Specifically, in this embodiment, the sentence text vectors of adjacent sentence texts can be concatenated to obtain the cross-sentence features of adjacent sentence pairs. Further, in this embodiment, the cross-sentence features of all the adjacent sentence pairs can be weighted and summed to obtain the weighted cross-sentence features; the self-attention mechanism network structure is used to control each phoneme feature of the encoder of the speech synthesis model to learn the weighted cross-sentence features; wherein, the key vector and the value vector of the self-attention mechanism network structure are the weighted cross-sentence features, and the query vector of the self-attention mechanism network structure is the phoneme feature vector of the encoder; the speech synthesis model is trained using the weighted cross-sentence features.

[0097] The above embodiment can obtain cross-sentence features without relying on a language model. Specifically, the word vectors of each sentence text are used to obtain the sentence vector of each sentence, and then the sentence vectors of two adjacent sentences are concatenated as the cross-sentence features of the two sentences. In this way, it is necessary to query the vector of each word through a pre-trained word vector table, and then average the word vectors of each word in each sentence to obtain the sentence vector. After obtaining the sentence vector, the cross-sentence features can be used in the PSE and CSE methods mentioned above. If the cross-sentence features are used in the CSE method, the sentence vectors of each context sentence can be concatenated, and then concatenated with the output of the TTS encoder for speech synthesis as shown. If the cross-sentence features are used in the PSE method, first, the sentence vector of each sentence can be used as the Figure 4 key, value sequence in, and then the output of the encoder is used as the query vector to obtain the cross-sentence features based on each phoneme; secondly, the sentence vectors of two adjacent sentences can also be concatenated as the cross-sentence features of the two sentences, and then all the cross-sentence feature vectors are used as the Figure 6 key, value sequence in. Figure 6

[0098] This application also provides a solution for a Student model obtained through knowledge distillation to perform text feature extraction operations on sentences in a sentence set to obtain cross-sentence features. The specific process is as follows: Before performing text feature extraction operations on sentences in the sentence set to obtain cross-sentence features, knowledge distillation is performed on a language model to obtain a Student model, so that the Student model learns the cross-sentence features extracted by the language model; the Student model is used to perform text feature extraction operations on sentences in the sentence set to obtain the cross-sentence features. Among them, knowledge distillation is a way of model compression that can convert a high-precision but bulky Teacher model (such as the above-mentioned language model) into a more compact and deployable Student model, and this Student model has the knowledge learned in the Teacher model. In this embodiment, the Student model is the result obtained after knowledge distillation of the language model. This Student model can extract cross-sentence features in the sentence set, and the Student model has a smaller demand for storage space.

[0099] As a feasible implementation manner, in this embodiment, the Student model can be used to perform text feature extraction operations on sentences in the sentence set to obtain sentence text vectors for each sentence text, and the sentence text vectors are concatenated to obtain the cross-sentence features.

[0100] Specifically, in this embodiment, the sentence text vectors corresponding to the forward sentence set can be concatenated to obtain forward cross-sentence features; the sentence text vectors corresponding to the backward sentence set can also be concatenated to obtain backward cross-sentence features. Further, in this embodiment, the forward cross-sentence features and the backward cross-sentence features can also be concatenated to obtain context cross-sentence features; the phoneme features output by the encoder of the speech synthesis model are obtained; the context cross-sentence features and the phoneme features are concatenated to obtain a concatenated feature vector, and the speech synthesis model is trained using the concatenated feature vector.

[0101] Specifically, in this embodiment, the sentence text vectors of adjacent sentences can be concatenated to obtain cross-sentence features of adjacent sentence pairs. Further, in this embodiment, the cross-sentence features of all adjacent sentence pairs can also be weighted and summed to obtain weighted cross-sentence features; each phoneme feature of the encoder of the speech synthesis model is controlled by a self-attention mechanism network structure to learn the weighted cross-sentence features; among them, the key vector and value vector of the self-attention mechanism network structure are the weighted cross-sentence features, and the query vector of the self-attention mechanism network structure is the phoneme feature vector of the encoder; the speech synthesis model is trained using the weighted cross-sentence features.

[0102] The above method learns the feature vectors obtained by the BERT model through a simple small model, that is, performs knowledge distillation on the cross-sentence features of the BERT model. Since the BERT model is relatively large and has a large number of parameters, it is not suitable for deployment and inference under restricted conditions. For example, if the environment has a poor CPU, the inference speed of BERT may be very slow, or when deploying on a mobile device, the relatively large storage requirements of the BERT model will pose many challenges, and at the same time, the inference speed of the mobile device cannot meet the inference requirements of BERT. Therefore, on this basis, the present invention proposes to learn the cross-sentence features extracted by the BERT model through a simple small model, and then use the cross-sentence features extracted by this small model in combination with the PSE / CSE method for speech synthesis.

[0103] In the above embodiment, the small model is called the Student model. This model takes the cross-sentence features extracted by the above-mentioned BERT as the training target. In this way, the Student model learns the knowledge of the large BERT model, and while reducing the model complexity, it does not seriously affect the extraction of cross-sentence features. The input of the Student model can be words, or it can use phonemes as the input like the TTS model. If the input is words, a relatively large word vector table needs to be maintained; if phonemes are used as the input, the phoneme vector table can be shared with the TTS model. The Student model in this embodiment can be any simple model, such as an RNN model, a CNN model, or a model based on the Transformer network structure. Taking RNN as an example, for the CSE method, the input of the Student model is to input the words or phonemes of a sentence and then output a vector, which will be as similar as possible to the sentence vector obtained by BERT. For PSE, the input of the Student model is the words or phrases of two adjacent sentences, and then output a vector, which will be as similar as possible to the cross-sentence features of two adjacent sentences obtained by BERT. After having the Student model, speech synthesis based on cross-sentence features can be performed in the same way as the previously mentioned PSE or CSE method. At the same time, the Student model can ensure relatively small storage space requirements and relatively fast inference speed requirements, and can achieve speech synthesis based on cross-sentence features under resource-constrained conditions.

[0104] This application also experimentally compared the performance of different models, including a baseline speech synthesis model without cross-sentence features, a speech synthesis model based on CSE cross-sentence features (CSE-based model), and a speech synthesis model based on PSE cross-sentence features (PSE-based model). During the experiment, the present invention used MUSHRAMOS (mean opinion score) and ABN choice to evaluate the performance of the models. A total of 65 people participated in the scoring, including 50 Chinese people and 15 British people.

[0105] The MUSHRA scores of different models are shown in Table 1 below. From the experimental results, it can be seen that the speech synthesis model combined with the cross-sentence features of the PSE method has a significant improvement in Chinese data compared with the baseline model, but the improvement in English data is not very obvious. This is because there are not many prosodic variations in English data, and the prosody of each sentence is relatively consistent. The speech synthesis model combined with the cross-sentence features of the CSE method has a lower score compared with the baseline. This is because in the synthesized speech, the CSE-based model synthesized some incorrect tones, giving a poor impression to the scorers, resulting in a lower comprehensive score.

[0106] Table 1 Comparison table of MUSHRA scores

[0107]

[0108] The ABN prefer test of this application is shown in Table 2. According to Table 2, in the experimental results of Chinese synthetic data, more testers prefer the audio synthesized by the PSE-based model.

[0109] Table 2 Statistical table of ABN prefer test results

[0110]

[0111] From the above experimental results, it can be seen that the end-to-end speech synthesis model based on cross-sentence information proposed in this embodiment can effectively utilize the cross-sentence features of the current text to improve the prosody performance of the model.

[0112] This application embodiment also provides a speech synthesis device, which may include:

[0113] A set construction module, configured to obtain the context sentences of the current sentence and construct a sentence set including the current sentence and the context sentences;

[0114] A model training module for performing text feature extraction operations on sentences in the sentence set to obtain cross-sentence features, and training a speech synthesis model using the cross-sentence features;

[0115] A speech synthesis module for synthesizing speech information of the target text using the trained speech synthesis model.

[0116] In this embodiment, the context sentences of the current sentence are obtained, and the sentence set including the current sentence and the context sentences is used as a sample for training the speech synthesis model. Specifically, this embodiment performs text feature extraction operations on the sentences in the sentence set to obtain cross-sentence features. The cross-sentence features can describe the discourse structure of the text context. Therefore, the speech synthesis model trained using the cross-sentence features can learn the discourse structure features. Since the prosody information when the text is read is related to the context discourse structure, the same sentence may have completely different prosody performances in different context contexts. The speech synthesis model trained using the cross-sentence features in this embodiment can synthesize speech based on the context discourse structure, thereby improving the prosody effect of speech synthesis.

[0117] Furthermore, the model training module includes:

[0118] A set determination unit for determining a forward sentence set and a backward sentence set; wherein, the forward sentence set includes the current sentence and the context sentences before the current sentence, and the backward sentence set includes the current sentence and the context sentences after the current sentence;

[0119] A forward and backward cross-sentence feature calculation unit for calculating the forward cross-sentence feature of the forward sentence set and the backward cross-sentence feature of the backward sentence set.

[0120] Furthermore, the forward and backward cross-sentence feature calculation unit is configured to input all the forward sentence sets into the language model to obtain the forward cross-sentence features; and is also configured to input all the backward sentence sets into the language model to obtain the backward cross-sentence features.

[0121] Furthermore, the process of the model training module training the speech synthesis model using the cross-sentence features includes: concatenating the forward cross-sentence features and the backward cross-sentence features to obtain context cross-sentence features; obtaining the phoneme features output by the encoder of the speech synthesis model; concatenating the context cross-sentence features and the phoneme features to obtain a concatenated feature vector, and training the speech synthesis model using the concatenated feature vector.

[0122] Furthermore, the model training module includes:

[0123] The first feature extraction unit is used to obtain adjacent sentence pairs in the sentence set; and is further used to input the adjacent sentence pairs into a language model to obtain cross-sentence features of the adjacent sentence pairs;

[0124] The training unit is used to perform weighted summation on the cross-sentence features of all the adjacent sentence pairs to obtain weighted cross-sentence features; and is further used to control each phoneme feature of the encoder of the speech synthesis model to learn the weighted cross-sentence features through a self-attention mechanism network structure; wherein, the key vector and value vector of the self-attention mechanism network structure are the weighted cross-sentence features, and the query vector of the self-attention mechanism network structure is the phoneme feature vector of the encoder; and is further used to train the speech synthesis model by using the weighted cross-sentence features.

[0125] Further, the model training module includes:

[0126] The second feature extraction unit is used to query the word vectors of each word in the sentence set by using a word vector table, and calculate the average value of the word vectors corresponding to the sentence text in the context sentence to obtain a sentence text vector; and is further used to splice the sentence text vectors to obtain the cross-sentence features.

[0127] Further, it further includes:

[0128] The knowledge distillation module is used to perform knowledge distillation on the language model to obtain a Student model before performing text feature extraction operations on the sentences in the sentence set to obtain cross-sentence features, so that the Student model learns the cross-sentence features extracted by the language model;

[0129] Correspondingly, the model training module includes:

[0130] The third feature extraction unit is used to perform text feature extraction operations on the sentences in the sentence set by using the Student model to obtain the cross-sentence features.

[0131] Since the embodiments of the apparatus part correspond to the embodiments of the method part, please refer to the description of the embodiments of the method part for the embodiments of the apparatus part, and details are not described herein.

[0132] This application also provides a storage medium, on which a computer program is stored, and when the computer program is executed, the steps provided in the above embodiments can be implemented. The storage medium may include: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.

[0133] This application also provides an electronic device, please refer toFigure 7 , a structural diagram of an electronic device provided by an embodiment of the present application, as Figure 7 shown, may include a processor 710 and a memory 720.

[0134] Among them, the processor 710 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 710 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 710 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 710 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 710 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0135] The memory 720 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 720 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In this embodiment, the memory 720 is at least used to store the following computer program 721. After the computer program is loaded and executed by the processor 710, it can implement the relevant steps in the speech synthesis method disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory 720 may further include an operating system 722 and data 723, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 722 may include Windows, Linux, Android, etc.

[0136] In some embodiments, the electronic device may further include a display screen 730, an input / output interface 740, a communication interface 750, a sensor 760, a power supply 770, and a communication bus 780.

[0137] Of course, Figure 7The structure of the electronic device shown does not constitute a limitation on the electronic device in the embodiments of the present application. In actual applications, the electronic device may include more or fewer components than Figure 7 shown, or combine certain components.

[0138] The various embodiments in the specification are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section. It should be noted that for those of ordinary skill in the art in the technical field of the present application, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

[0139] It should also be noted that in this specification, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including an..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

Claims

1. A voice synthesis method, characterized in that, including: Obtain the context sentences of the current sentence, and construct a sentence set including the current sentence and the context sentences; Perform text feature extraction operations on the sentences in the sentence set to obtain cross-sentence features, and use the cross-sentence features to train a speech synthesis model; Use the trained speech synthesis model to synthesize the speech information of the target text; Among them, performing text feature extraction operations on the sentences in the sentence set to obtain cross-sentence features includes: Determine a forward sentence set and a backward sentence set; wherein, the forward sentence set includes the current sentence and the context sentences before the current sentence, and the backward sentence set includes the current sentence and the context sentences after the current sentence; Calculate the forward cross-sentence features of the forward sentence set and the backward cross-sentence features of the backward sentence set.

2. The voice synthesis method according to claim 1, characterized in that Calculating the forward cross-sentence features of the forward sentence set and the backward cross-sentence features of the backward sentence set includes: Input all the forward sentence sets into a language model to obtain the forward cross-sentence features; Input all the backward sentence sets into the language model to obtain the backward cross-sentence features.

3. The speech synthesis method according to claim 2, wherein Using the cross-sentence features to train a speech synthesis model includes: Concatenate the forward cross-sentence features and the backward cross-sentence features to obtain context cross-sentence features; Obtain the phoneme features output by the encoder of the speech synthesis model; Concatenate the context cross-sentence features and the phoneme features to obtain a concatenated feature vector, and use the concatenated feature vector to train the speech synthesis model.

4. The speech synthesis method according to claim 1, wherein Performing text feature extraction operations on the sentences in the sentence set to obtain cross-sentence features includes: Obtain adjacent sentence pairs in the sentence set; Input the adjacent sentence pairs into a language model to obtain the cross-sentence features of the adjacent sentence pairs; Correspondingly, using the cross-sentence features to train a speech synthesis model includes: Perform weighted summation on the cross-sentence features of all the adjacent sentence pairs to obtain weighted cross-sentence features; Control each phoneme feature of the encoder of the speech synthesis model to learn the weighted cross-sentence features through a self-attention mechanism network structure; wherein, the key vector and the value vector of the self-attention mechanism network structure are the weighted cross-sentence features, and the query vector of the self-attention mechanism network structure is the phoneme feature vector of the encoder; Use the weighted cross-sentence features to train the speech synthesis model.

5. The voice synthesis method according to claim 1, wherein Performing text feature extraction operations on the sentences in the sentence set to obtain cross-sentence features includes: Query the word vectors of each word in the sentence set using a word vector table, and calculate the average value of the word vectors corresponding to the sentence text in the context sentences to obtain a sentence text vector; Concatenate the sentence text vectors to obtain the cross-sentence features.

6. The voice synthesis method according to claim 1, characterized in that Before performing text feature extraction operations on the sentences in the sentence set to obtain cross-sentence features, it further includes: Perform knowledge distillation on the language model to obtain a Student model, so that the Student model learns the cross-sentence features extracted by the language model; Correspondingly, performing text feature extraction operations on the sentences in the sentence set to obtain cross-sentence features includes: Performing a text feature extraction operation on the sentences in the sentence set by using the Student model to obtain the cross-sentence features.

7. A voice synthesis device, characterized in that, Including: A set construction module, configured to obtain context sentences of a current sentence and construct a sentence set including the current sentence and the context sentences. A model training module, configured to perform a text feature extraction operation on the sentences in the sentence set to obtain cross-sentence features, and train a speech synthesis model by using the cross-sentence features. A speech synthesis module, configured to synthesize speech information of a target text by using the trained speech synthesis model. Wherein, the model training module includes: A set determination unit, configured to determine a forward sentence set and a backward sentence set; wherein, the forward sentence set includes the current sentence and context sentences before the current sentence, and the backward sentence set includes the current sentence and context sentences after the current sentence. A forward and backward cross-sentence feature calculation unit, configured to calculate a forward cross-sentence feature of the forward sentence set and a backward cross-sentence feature of the backward sentence set.

8. An electronic device, characterized in that, Including a memory and a processor, where a computer program is stored in the memory, and when the processor calls the computer program in the memory, the steps of the speech synthesis method according to any one of claims 1 to 6 are implemented.

9. A storage medium, characterized in that, Computer-executable instructions are stored in the storage medium, and when the computer-executable instructions are loaded and executed by a processor, the steps of the speech synthesis method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Intention recognition method and device, dialogue robot and computer-readable storage medium

    CN112115702A

  • Context-based dialogue generation method and system

    CN112328756A

  • Neural text-to-speech synthesis utilizing multi-level contextual features

    CN112489618A

  • Parallel neural text-to-speech conversion

    CN112669809A