A speech synthesis method and system
By constructing the target emotional category set and calculating the average value of the emotion coded vector, the problem of inaccurate speech synthesis in the prior art is solved, stable emotional intensity synthesis is achieved, and user experience is improved.
Patent Information
- Application Number
- CN202210238371.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-11
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2042-03-11
AI Technical Summary
The prior art is difficult to synthesize accurate and stable voices corresponding to emotional intensity, resulting in poor user experience.
By obtaining the emotional categories of the target text, constructing the target emotional categories set, and computing the average value of the emotional encoding vector of the target speech sample, generating an average emotional encoding vector, and combining the text sequence encoding to synthesize the target audio.
It realizes the synthesis of accurate and stable voice corresponding to emotional intensity, improving the user experience.
Smart Images

Figure CN114627851B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to speech synthesis technology, and more specifically, to a speech synthesis method and system. Background Art
[0002] Speech synthesis, also known as text-to-speech conversion, primarily converts text into synthesized speech, ensuring that the synthesized speech is as intelligible and natural as possible. Recent advances in speech synthesis technology have resulted in synthesized speech that is increasingly close to the sound quality and naturalness of real human speech. However, human speech is emotionally charged, expressing emotions such as happiness, sadness, confusion, and anger. Therefore, the key to the development of speech synthesis technology is to synthesize emotionally charged speech, ensuring that the synthesized speech is closer to the real human voice.
[0003] In existing technology, a data-driven approach is often used to synthesize emotional speech. This involves modeling each emotion using collected speech data containing various emotion categories, generating acoustic parameter models corresponding to each emotion category, and synthesizing the target emotional speech based on parameter synthesis techniques. However, the same emotion category can be further divided into multiple levels of emotion based on the intensity of the emotion. For example, the emotion category of "sorrow" can be further divided into multiple levels of emotion, such as "dejected," "heartbroken," and "crying with sorrow." Using a data-driven approach makes it difficult to synthesize audio corresponding to accurate and stable emotion intensities, which is detrimental to the user experience. Summary of the Invention
[0004] The exemplary embodiments of the present application provide a speech synthesis method and device to solve the problem in the prior art that it is impossible to synthesize accurate and stable audio corresponding to the emotional intensity during speech synthesis, thereby improving user experience.
[0005] In one aspect, the present application provides a speech synthesis method, comprising:
[0006] According to the emotion category of the target text, a target emotion category set is obtained, wherein the target emotion category set includes a plurality of target speech samples, and the emotion category of the target speech samples is the same as the emotion category of the target text;
[0007] Obtaining an average emotion coding vector based on the emotion coding vectors of each target speech sample, wherein the emotion coding vector of the target speech sample is a vector representation corresponding to the emotion intensity of the target speech sample, and the average emotion coding vector is obtained by adding and averaging all the emotion coding vectors;
[0008] A target audio is obtained according to the text sequence encoding of the target text and the average emotion encoding vector.
[0009] In another aspect, the present application provides a speech synthesis system, comprising:
[0010] A first acquisition module is configured to acquire a target emotion category set according to the emotion category of the target text, wherein the target emotion category set includes a plurality of target speech samples, and the emotion category of the target speech samples is the same as the emotion category of the target text;
[0011] A second acquisition module is configured to acquire an average emotion coding vector based on the emotion coding vectors of the target speech samples, wherein the emotion coding vector of the target speech sample is a vector representation corresponding to the emotion intensity of the target speech sample, and the average emotion coding vector is obtained by adding and averaging all the emotion coding vectors;
[0012] The audio synthesis module is used to obtain the target audio according to the text sequence encoding of the target text and the average emotion encoding vector.
[0013] The present application provides a speech synthesis method and system, which can obtain a target emotion category set according to the emotion category of the target text, wherein the target emotion category set includes a number of target speech samples, and the emotion category of the target speech samples is the same as the emotion category of the target text; and obtain an average emotion coding vector according to the emotion coding vector of each target speech sample, wherein the emotion coding vector of the target speech sample is a vector representation corresponding to the emotion intensity of the target speech sample, and the average emotion coding vector is obtained by adding and averaging all emotion coding vectors; according to the text sequence code of the target text and the average emotion coding vector, accurate and stable audio corresponding to the emotion intensity is synthesized, which is beneficial to user experience BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the implementation methods in the embodiments of the present application or related technologies, the following is a brief introduction to the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0015] Figure 1 A schematic flow chart of a method for speech-to-speech synthesis according to some embodiments is shown;
[0016] Figure 2 A schematic diagram of a process for obtaining text sequence encoding according to some embodiments is shown;
[0017] Figure 3 FIG2 shows a flow chart of determining a target mel spectrum according to some embodiments;
[0018] Figure 4A schematic diagram of a process for adjusting the emotional intensity of target audio according to some embodiments is shown;
[0019] Figure 5 A structural schematic diagram of a speech synthesis system according to some embodiments is shown. DETAILED DESCRIPTION
[0020] In order to make the purpose and implementation of this application clearer, the exemplary implementation of this application will be clearly and completely described below in conjunction with the drawings in the exemplary embodiments of this application. Obviously, the described exemplary embodiments are only part of the embodiments of this application, not all of the embodiments.
[0021] It should be noted that the brief descriptions of terms in this application are only for the purpose of facilitating the understanding of the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their ordinary and usual meanings.
[0022] In the specification and claims of this application and the accompanying drawings, the terms "first," "second," "third," etc. are used to distinguish similar or similar objects or entities, and are not necessarily intended to limit a particular order or sequence, unless otherwise noted. It should be understood that the terms used in this manner are interchangeable under appropriate circumstances.
[0023] First, the professional terms involved in the embodiments of the present invention are explained as follows:
[0024] Mel spectrum: A frequency domain feature extracted from speech audio that can be used to characterize the short-term characteristics of speech signals.
[0025] Encoder: A device that converts readable data into unreadable data through an algorithm is collectively called an encoder.
[0026] Decoder: A device that converts unreadable data into readable data through an algorithm is collectively called a decoder.
[0027] Convolutional neural network: A neural network that relies on convolution calculations. It is one of the representative algorithms of deep learning and can be subdivided into many different types of convolutional neural networks.
[0028] Recurrent Neural Network: A neural network that takes sequence data (such as audio) as input and performs recursive chain-link calculations in the direction of sequence evolution (the direction in audio is time). It can be subdivided into many different types of recurrent neural networks.
[0029] Fully connected network: The most basic neural network calculation method, which connects all inputs and outputs together through multiplication and addition.
[0030] Attention mechanism: A mechanism that makes weighted changes to target data through encoding and decoding, allowing the system to more clearly know where to focus.
[0031] Fundamental Frequency: When a sound-emitting body makes a sound due to vibration, the sound can generally be decomposed into many simple sine waves. That is to say, all natural sounds are basically composed of many sine waves with different frequencies. Among them, the sine wave with the lowest frequency is the fundamental frequency (represented by F0), and other sine waves with higher frequencies are overtones.
[0032] Energy, also known as intensity or volume, represents the size of a sound and can be simulated by the amplitude of the sound signal. The larger the amplitude, the louder the sound waveform.
[0033] Vocoder is a sound signal processing device or software that can encode acoustic features into sound waveforms.
[0034] While existing speech synthesis technologies are capable of synthesizing realistic and natural emotional speech, most of these emotional speech synthesis technologies employ a data-driven approach. Specifically, to synthesize angry speech, they collect data related to the emotion and construct an acoustic model or database representing the emotion label. This model (or database) is then used to synthesize angry speech. However, anger can be further categorized into multiple levels based on intensity, such as "indignant," "furious," and "burning with rage." However, these emotional speech synthesis technologies can only use simple degree-based representations to identify emotional intensity. For example, based on the listener's subjective perception, the emotional intensity of a speech segment is labeled as a few simple levels, such as "mild," "moderate," and "severe." These techniques then model a small amount of speech data with intensity levels corresponding to each emotion category to generate speech of the corresponding level. Therefore, synthesizing speech with varying emotional intensities remains a major technical challenge in the current field of speech synthesis.
[0035] In order to solve the above problems, the present application provides a speech synthesis method, which can synthesize emotional speech and can autonomously regulate the emotional intensity of the synthesized speech, so that the emotional level of the synthesized speech meets the user's expectations and enhances the user experience.
[0036] Figure 1 A flowchart of a speech synthesis method provided in this application is shown in FIG. Figure 1 As shown, the method includes the following steps:
[0037] S101: Acquire a target emotion category set according to the emotion category of the target text, wherein the target emotion category set includes a plurality of target speech samples, and the emotion category of the target speech samples is the same as the emotion category of the target text.
[0038] In some embodiments, the target text can be the entire text content in an e-book, or the entire text content of a chapter, a fragment or a sentence in an e-book, or the text content in other types of texts, such as news text, public account article text, SMS communication record text, chat record text of Internet platform communication APP, etc. The emotion category of the target text is the emotion category of the speech after the target text is converted into speech. The emotion category of the target text can include joy, anger, sorrow, happiness, disgust, surprise, fear, etc.
[0039] In some embodiments, the emotion category of the target text can be manually calibrated, that is, the user directly determines the emotion category of the target text after it is converted into speech. The emotion category of the target text can also be obtained by inputting the target text into an emotion coding network model. The establishment of the emotion coding network model can utilize crawler technology to crawl existing text data in the public network, analyze each text data by reading the emotion words related to the emotion content in the text data, and combine the relationship between the emotion words and the context, and manually label each text data. For example, the analyzed text data can be labeled as emotion categories such as joy, anger, sadness, happiness, disgust, surprise, and fear. The convolutional neural network of the emotion coding network model is established based on the labeled text data, and the backpropagation technology is used to train the convolutional neural network of the emotion coding network model until the convergence condition is reached to complete the training of the convolutional neural network of the emotion coding network model. The trained convolutional neural network of the emotion coding network model can analyze the emotion category of the input target text to output the emotion category of the target text.
[0040] In some embodiments, after determining the emotion category of the target text, a target emotion category set can be obtained from a pre-established speech library. The speech library includes multiple emotion category sets with different emotion category labels, and the target emotion category set is one of the multiple emotion category sets with different emotion category labels. The emotion category labels of the target emotion category set correspond to the emotion category of the target text.
[0041] In some embodiments, a speech library can be established by obtaining a number of speech samples. Specifically, after obtaining a number of speech samples, the emotion category corresponding to each speech sample can be determined, and the speech samples can be classified according to the emotion category corresponding to each speech sample to generate a corresponding emotion category set. The speech library is a collection of emotion category sets. Each emotion category set is a collection of speech samples with the same emotion category; according to the emotion category of the speech samples in each emotion category set, a corresponding emotion category label is generated, and the corresponding emotion category set is labeled according to each emotion category label. For example, there are 6 speech samples, namely speech sample A, speech sample B, speech sample C, speech sample D, speech sample E and speech sample F. If speech samples A and speech sample B correspond to the emotion category of "joy", speech sample C corresponds to the emotion category of "anger", and speech samples D, speech sample E and speech sample F correspond to the emotion category of "surprise", then a total of three emotion category sets can be generated, namely the first emotion category set, the second emotion category set and the third emotion category set. The first emotion category set includes speech samples A and B, the second emotion category set includes speech sample C, and the third emotion category set includes speech samples D, E, and F. Based on the emotion category "joy" corresponding to speech samples A and B, the emotion label "joy" is generated, and the first emotion category set is labeled with the emotion label "joy." Based on the emotion category "anger" corresponding to speech sample C, the emotion label "anger" is generated, and the second emotion category set is labeled with the emotion label "anger." Based on the emotion category "surprise" corresponding to speech samples D, E, and F, the emotion label "surprise" is generated, and the third emotion category set is labeled with the emotion label "surprise." The labeled first, second, and third emotion category sets together constitute the speech database. It should be noted that this embodiment only exemplifies the composition of the speech library. In actual situations, according to the emotion categories corresponding to each speech sample, the speech library can also include more emotion category sets marked with non-pass emotion labels, for example, an emotion category set marked with the "sadness" emotion label, an emotion category set marked with the "joy" emotion label, an emotion category set marked with the "disgust" emotion label, an emotion category set marked with the "fear" emotion label, etc. This application does not limit this.
[0042] In some embodiments, based on the emotion category of the target text, an emotion category set tagged with a target emotion category label may be obtained from the speech database, and the emotion category set tagged with the target emotion category label may be determined as the target emotion category set. For example, if the emotion category of the target text is "joy," an emotion category set tagged with the emotion label "joy" may be obtained from the speech database, and the emotion category set tagged with the emotion label "joy" may be determined as the target emotion category set.
[0043] S102: Obtain an average emotion coding vector based on the emotion coding vector of each target speech sample, wherein the emotion coding vector of the target speech sample is a vector representation corresponding to the emotion intensity of the target speech sample, and the average emotion coding vector is obtained by adding and averaging all the emotion coding vectors.
[0044] In some embodiments, each target speech sample corresponds to a mel spectrum, and the target speech sample can be input into a corresponding neural network model to output the mel spectrum corresponding to the target speech sample.
[0045] The output mel spectra are sequentially input into the preset convolutional neural network, the preset recurrent neural network and the preset fully connected network, and a corresponding number of coding sequences are output. The coding sequences are passed through a multi-head attention mechanism to generate weighted coefficients relative to each preset feature vector, and the preset feature vectors are weighted according to the weighted coefficients to obtain the emotion coding vectors of each target speech sample, wherein the preset feature vectors represent the emotion intensity of the target speech sample.
[0046] In some embodiments, the same emotion category can be further divided into multiple emotion levels based on the intensity of the emotion, and the target emotion category set can include target speech samples of different emotion intensities. For example, the emotion category "joy" can be further divided into multiple emotion levels based on the intensity of the emotion, such as "elated," "radiant," "crying with joy," and "contented." The emotion intensity corresponding to the emotion level "crying with joy" is significantly greater than the emotion intensity corresponding to the emotion level "contented."
[0047] In some embodiments, since the emotion intensity corresponding to each target speech sample is different, the emotion coding vector used to characterize each target speech sample is also different. The emotion coding vectors corresponding to all speech samples can be summed and averaged to obtain an average emotion coding vector, and the average emotion coding vector can be used to characterize the target emotion category set. For example, the speech samples in the target emotion category set can be sorted in ascending order according to the corresponding emotion intensity, and the target speech samples in the target emotion category set can be processed by setting an emotion coding vector processing rule to obtain an average emotion coding vector. The emotion coding vector processing rule is:
[0048]
[0049] Among them, e i Represents the emotion coding vector corresponding to each target speech sample in the target emotion category set, N i represents the number of target speech samples in the target emotion category set, Represents the emotion encoding vector corresponding to the j-th target speech sample in the emotion category set marked with the emotion category "i".
[0050] S103: Obtain target audio according to the text sequence code of the target text and the average emotion coding vector.
[0051] In some embodiments, the text sequence encoding of the target text can be obtained based on the part-of-speech sequence encoding and phoneme sequence encoding of the target text. The part-of-speech sequence encoding is used to represent the vector sequence corresponding to each word in the target text, and the phoneme sequence encoding is used to represent the vector sequence corresponding to each character in the target text. Figure 2 This is a flow chart of obtaining text sequence encoding in the exemplary embodiment provided by this application, which can be achieved through Figure 2 The steps shown obtain the text sequence encoding of the target text.
[0052] S201: Perform word segmentation and character segmentation on the text content in the target text.
[0053] This process can be performed using a word segmentation tool, such as the LAC word segmentation tool. This application does not limit the tools used for word segmentation and character segmentation. After the text content in the target text is processed by word segmentation, multiple words can be obtained. For example, "Today the weather is really good", the word segmentation result is "Today, the weather is really good". After the text content in the target text is processed by character segmentation, multiple characters can be obtained. For example, "Today, the weather is really good", the character segmentation result is "Today, the weather, is really, good".
[0054] S202: Input the target text after word segmentation into a part-of-speech encoding model, such as Google's BERT model. The part-of-speech encoding model can output a vector representation of each word in the target text after word segmentation. Each output vector representation of the word is concatenated in the output order to obtain the part-of-speech sequence encoding of the target text. Input the target text after character segmentation into a phoneme encoding model, such as a VSM model. The phoneme encoding model can output a vector representation of each word in the target text after character segmentation. Each output vector representation of the word is concatenated in the output order to obtain the phoneme sequence encoding of the target text.
[0055] S203: The part-of-speech sequence code and the phoneme sequence code are input into an encoder, and a text sequence code that can be recognized by a computer can be output, wherein the encoder is constructed based on a multi-head self-attention mechanism. After the part-of-speech sequence code and the phoneme sequence code are input into the encoder, the encoder can analyze the input part-of-speech sequence code and the phoneme sequence code to obtain a text sequence after integrating the context relationship, and after converting the text sequence into a text sequence code that can be recognized by a computer, output it from the encoder.
[0056] In some embodiments, the average emotion coding vector may be determined as a target emotion coding vector, and a target mel spectrum may be determined based on the target emotion coding vector and the text sequence encoding to obtain the target audio.
[0057] Optionally, an emotion coding vector with the highest similarity to the average emotion coding vector can be determined as the target emotion coding vector, and the target mel spectrum can be determined based on the target emotion coding vector and the text sequence encoding to obtain the target audio. The above method of determining an emotion coding vector with the highest similarity to the average emotion coding vector as the target emotion coding vector can avoid the low quality of the synthesized speech caused by the weak generalization ability of the encoder.
[0058] Figure 3 This is a flow chart of determining the target Mel spectrum provided by this application as an example. Figure 3 As shown, residual processing can be performed on the text sequence coding and the phoneme sequence coding, and the residual values of the text sequence coding and the phoneme sequence coding are determined as the target sequence coding, and residual processing can be performed on the target emotion coding vector and the target sequence coding, and the residual values of the target emotion coding vector and the target sequence coding are determined as the sequence coding to be predicted, and the sequence coding to be predicted is input into the model for generating Mel spectrum, and the target Mel spectrum is output. The target audio is obtained according to the target Mel spectrum to avoid phoneme loss when the average emotion coding vector is embedded in the text sequence coding output by the encoder, which causes the sound quality of the final synthesized speech (target audio) to be damaged, thereby improving the quality of the synthesized speech.
[0059] In some embodiments, the sequence code to be predicted can be input into the energy prediction model, the fundamental frequency prediction model and the duration prediction model respectively to obtain the corresponding energy prediction code, fundamental frequency prediction code and duration prediction code respectively; wherein, the energy prediction code is a vector representation of the pronunciation intensity (volume) after synthesis corresponding to each phoneme in the predicted target text, the fundamental frequency prediction code is a vector representation of the pronunciation fundamental frequency after synthesis corresponding to each phoneme in the predicted target text, and the duration prediction code is a vector representation of the pronunciation duration after synthesis corresponding to each phoneme in the predicted target text.
[0060] The energy prediction code, the fundamental frequency prediction code, and the duration prediction code are added to the sequence code to be predicted to obtain a target mel spectrum sequence code, and the target mel spectrum sequence code is input into a decoder for decoding to obtain a target mel spectrum, wherein the decoder is constructed based on a multi-head self-attention mechanism.
[0061] It should be noted that adding the target emotion encoding vector before the energy prediction model, the fundamental frequency prediction model, and the duration prediction model can prevent the emotional expression of the synthesized target audio from being affected by energy, fundamental frequency, and duration, so that the synthesized target audio has appropriate prosodic features.
[0062] Optionally, in some embodiments, the Figure 4 The steps shown adjust the emotional intensity of the target audio to achieve autonomous control of the emotional intensity corresponding to the synthesized target audio, and set the emotional adjustment parameters. The emotional intensity corresponding to the target audio can be adjusted by regulating the emotional adjustment parameters. The steps include:
[0063] S301: Obtain a first emotion coding vector and a second emotion coding vector;
[0064] The first emotion coding vector is the emotion coding vector corresponding to the target speech sample with the weakest emotion intensity, and the second emotion coding vector is the emotion coding vector corresponding to the target speech sample with the strongest emotion intensity.
[0065] Optionally, the average emotion coding vector can be determined as the second emotion coding vector, so that when the user adjusts the emotion adjustment parameter α according to the emotion intensity adjustment rule, the emotion intensity corresponding to the synthesized target audio is less than or equal to the emotion intensity corresponding to the audio synthesized according to the average emotion coding vector.
[0066] Optionally, an emotion coding vector closest to the average emotion coding vector can be determined as the second emotion coding vector, so that when the user adjusts the emotion adjustment parameter α according to the emotion intensity adjustment rule, the emotion intensity corresponding to the synthesized target audio is less than or equal to the emotion intensity corresponding to the audio synthesized according to the emotion coding vector closest to the emotion coding vector, so as to avoid the low quality of the synthesized speech due to the weak generalization ability of the encoder.
[0067] S302: Determine an emotion intensity adjustment rule based on the first emotion coding vector, the second emotion coding vector, and a preset emotion adjustment parameter;
[0068] S303: Adjust the average emotion coding vector according to the emotion intensity adjustment rule:
[0069] In some embodiments, the emotion intensity adjustment rule is:
[0070]
[0071] in, is the adjusted average emotion encoding vector, is the first emotion encoding vector, is the second emotion encoding vector, α is the emotion adjustment parameter, α∈[0,1]. When α=0, the target audio with the weakest emotion intensity is synthesized, and when α=1, the target audio with the strongest emotion intensity is synthesized.
[0072] S304: Determine a target mel spectrum according to the text sequence encoding and the adjusted average emotion encoding vector, and convert the target mel spectrum into an audio signal.
[0073] In some embodiments, an emotion coding vector having the highest similarity to the adjusted average emotion coding vector can be determined as a target emotion coding vector, and a target mel spectrum can be determined according to the target emotion coding vector and the text sequence code to obtain a target audio, wherein the target emotion coding vector It can be expressed as:
[0074]
[0075] The above-mentioned method of determining an emotion coding vector having the highest similarity to the average emotion coding vector as the target emotion coding vector can avoid the low quality of the synthesized speech caused by the weak generalization ability of the encoder.
[0076] See also Figure 5 , a speech synthesis system provided in an embodiment of the present application, comprising:
[0077] The first acquisition module 51 is configured to: acquire a target emotion category set according to the emotion category of the target text, wherein the target emotion category set includes a plurality of target speech samples, and the emotion category of the target speech samples is the same as the emotion category of the target text.
[0078] The second acquisition module 52 is used to execute: obtaining an average emotion coding vector based on the emotion coding vector of each target speech sample, wherein the emotion coding vector of the target speech sample is a vector representation corresponding to the emotion intensity of the target speech sample, and the average emotion coding vector is obtained by adding and averaging all the emotion coding vectors.
[0079] The audio synthesis module 53 is used to execute: obtaining the target audio according to the text sequence encoding of the target text and the average emotion encoding vector.
[0080] In some embodiments, the first acquisition module 51 is specifically configured to execute:
[0081] Acquire a number of speech samples; classify each of the speech samples to obtain a corresponding emotion category set, wherein each emotion category set is a collection of the speech samples with the same emotion category; generate corresponding emotion category labels according to the emotion categories of the speech samples in each emotion category set, and mark the corresponding emotion category set with each emotion category label; determine the emotion category set marked with a target emotion category label as a target emotion category set, wherein the target emotion category label is the emotion category label that matches the emotion category of the target text.
[0082] Optionally, the first acquiring module 51 is further configured to execute:
[0083] Obtain a first emotion coding vector and a second emotion coding vector, wherein the first emotion coding vector is the emotion coding vector corresponding to the target speech sample with the weakest emotion intensity, and the second emotion coding vector is the emotion coding vector corresponding to the target speech sample with the strongest emotion intensity; determine an emotion intensity adjustment rule based on the first emotion coding vector, the second emotion coding vector, and a preset emotion adjustment parameter; adjust the average emotion coding vector based on the emotion intensity adjustment rule; determine the target mel spectrum based on the text sequence encoding and the adjusted average emotion coding vector; convert the target mel spectrum into an audio signal to obtain the target audio with the emotion intensity corresponding to the emotion adjustment parameter;
[0084] In some embodiments, before obtaining the average emotion coding vector based on the emotion coding vectors of the target speech samples, the second obtaining module 52 is further configured to perform:
[0085] Obtain the mel spectrum corresponding to each of the target speech samples; input each of the mel spectrums into a preset convolutional neural network, a preset recurrent neural network, and a preset fully connected network in sequence, and output a corresponding number of coding sequences; subject the coding sequences to a multi-head attention mechanism to generate a weighting coefficient relative to each preset feature vector, wherein the preset feature vector represents the emotional intensity of the target speech sample; and perform weighted processing on the preset feature vector according to the weighting coefficient to obtain an emotional coding vector for each of the target speech samples.
[0086] In some embodiments, the audio synthesis module 53 is specifically configured to perform:
[0087] Obtain a text sequence code of the target text; determine a target mel spectrum based on the text sequence code and the average emotion coding vector; and convert the target mel spectrum into an audio signal to obtain a target audio.
[0088] In some embodiments, when the audio synthesis module 53 obtains the text sequence code of the target text, it is specifically configured to perform:
[0089] The target text is subjected to word segmentation and character segmentation processing; the target text after the word segmentation processing is input into a part-of-speech encoding model, and the part-of-speech sequence encoding of the target text is output; the target text after the character segmentation processing is input into a phoneme encoding model, and the phoneme sequence encoding of the target text is output; the part-of-speech sequence encoding and the phoneme sequence encoding are input into an encoder to obtain a text sequence encoding of the target text, wherein the encoder is constructed based on a multi-head self-attention mechanism.
[0090] In some embodiments, when the audio synthesis module 53 determines the target mel spectrum according to the text sequence code and the average emotion code vector, it is specifically configured to perform:
[0091] Obtain a target emotion coding vector, wherein the target emotion coding vector is the emotion coding vector having the highest similarity to the average emotion coding vector; and determine the target mel spectrum according to the target emotion coding vector and the text sequence code.
[0092] In some embodiments, when the audio synthesis module 53 determines the target mel spectrum according to the target emotion coding vector and the text sequence coding, it is specifically configured to perform:
[0093] The residual value between the text sequence code and the phoneme sequence code is determined as the target sequence code; the residual value between the target emotion code vector and the target sequence code is determined as the sequence code to be predicted; and the target mel spectrum is determined according to the sequence code to be predicted.
[0094] In some embodiments, when the audio synthesis module 53 determines the target mel spectrum according to the sequence code to be predicted, it is specifically configured to perform:
[0095] The sequence code to be predicted is input into an energy prediction model to obtain an energy prediction code; the sequence code to be predicted is input into a fundamental frequency prediction model to obtain a fundamental frequency prediction code; the sequence code to be predicted is input into a duration prediction model to obtain a duration prediction code; the energy prediction code, the fundamental frequency prediction code and the duration prediction code are added to the sequence code to be predicted to obtain a target Mel-spectrogram sequence code; the target Mel-spectrogram sequence code is input into a decoder to obtain a target Mel-spectrogram, wherein the decoder is constructed based on a multi-head self-attention mechanism.
[0096] It can be seen from the above technical solutions that the present application provides a speech synthesis method and system, which can obtain a target emotion category set according to the emotion category of the target text, wherein the target emotion category set includes several target speech samples, and the emotion category of the target speech sample is the same as the emotion category of the target text; and according to the emotion coding vector of each target speech sample, an average emotion coding vector is obtained, wherein the emotion coding vector of the target speech sample is a vector representation corresponding to the emotion intensity of the target speech sample, and the average emotion coding vector is obtained by adding and averaging all emotion coding vectors; according to the text sequence encoding of the target text and the average emotion coding vector, accurate and stable audio corresponding to the emotion intensity is synthesized, which is beneficial to user experience
[0097] Optionally, there can be multiple application scenarios for this method, which can be the above-mentioned method for synthesizing speech of specific emotional intensity for a specified text, that is, synthesizing speech of specific emotional intensity for the input text based on the text input by the user and information related to emotional intensity. It can also be a method for generating a customized speech of new emotional intensity for the speech input by the user based on the information related to emotional intensity input by the user. It can also be applied to a human-computer interaction scenario, that is, the user inputs a sentence or speech or text, determines the reply text based on the sentence / speech / text input by the user, and synthesizes the reply speech based on the information related to emotional intensity customized by the user, or in the process of human-computer interaction, the smart device often analyzes and judges by itself and inputs the synthesized speech of the corresponding emotional intensity. At this time, if the user is not satisfied with the speech synthesized and replied by the machine, the user can input the corresponding emotional intensity feature vector to adjust the emotional intensity of the synthesized speech, etc. This application does not limit this.
[0098] In a specific implementation, the present invention further provides a computer storage medium, wherein the computer storage medium may store a program that, when executed, may include some or all of the steps of each embodiment of the speech synthesis method and system provided by the present invention. The storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0099] Those skilled in the art will clearly understand that the techniques in the embodiments of the present invention can be implemented using software plus a necessary general-purpose hardware platform. Based on this understanding, the technical solutions in the embodiments of the present invention, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention or certain portions of the embodiments.
[0100] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.
[0101] For ease of explanation, the above description has been made with reference to specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Based on the above teachings, various modifications and variations are possible. The above embodiments are selected and described to better explain the principles and practical applications, so that those skilled in the art can better utilize the embodiments and various different variations of the embodiments suitable for specific use considerations.
Claims
1. A speech synthesis method, characterized in that: include: According to the emotion category of the target text, a target emotion category set is obtained, wherein the target emotion category set includes a plurality of target speech samples, and the emotion category of the target speech samples is the same as the emotion category of the target text; Obtaining an average emotion coding vector based on the emotion coding vectors of each target speech sample, wherein the emotion coding vector of the target speech sample is a vector representation corresponding to the emotion intensity of the target speech sample, and the average emotion coding vector is obtained by adding and averaging all the emotion coding vectors; Obtaining a text sequence code of the target text; Determining a target mel spectrum according to the text sequence encoding and the average emotion encoding vector; The target mel spectrum is converted into an audio signal to obtain a target audio.
2. The method according to claim 1, characterized in that The step of obtaining a target emotion category set according to the emotion category of the target text further includes: Obtaining several voice samples; Classifying each of the speech samples to obtain a corresponding emotion category set, wherein each emotion category set is a collection of the speech samples with the same emotion category; Generating corresponding emotion category labels according to the emotion categories of the speech samples in each emotion category set, and marking the corresponding emotion category set with each emotion category label; The emotion category set marked with a target emotion category label is determined as a target emotion category set, wherein the target emotion category label is the emotion category label that matches the emotion category of the target text.
3. The method according to claim 1, characterized in that Before obtaining the average emotion coding vector according to the emotion coding vectors of the target speech samples, the method further includes: Obtaining the mel spectrum corresponding to each target speech sample; Input each of the mel spectrograms into a preset convolutional neural network, a preset recurrent neural network, and a preset fully connected network in sequence, and output a corresponding number of coding sequences; The encoded sequence is subjected to a multi-head attention mechanism to generate a weighted coefficient relative to each preset feature vector, wherein the preset feature vector represents the emotional intensity of the target speech sample; The preset feature vectors are weighted according to the weighting coefficients to obtain the emotion coding vectors of the target speech samples.
4. The method according to claim 1, wherein The step of obtaining the text sequence code of the target text further includes: Performing word segmentation and character segmentation processing on the target text; Inputting the target text after the word segmentation processing into a part-of-speech encoding model, and outputting a part-of-speech sequence encoding of the target text; Inputting the target text after the word segmentation processing into a phoneme coding model, and outputting a phoneme sequence code of the target text; The part-of-speech sequence encoding and the phoneme sequence encoding are input into an encoder to obtain a text sequence encoding of the target text, wherein the encoder is constructed based on a multi-head self-attention mechanism.
5. The method according to claim 1, wherein The step of determining a target Mel spectrum according to the text sequence encoding and the average emotion encoding vector further includes: Obtaining a target emotion coding vector, wherein the target emotion coding vector is the emotion coding vector having the highest similarity to the average emotion coding vector; The target mel spectrum is determined according to the target emotion encoding vector and the text sequence encoding.
6. The method according to claim 5, characterized in that The step of determining the target mel spectrum according to the target emotion coding vector and the text sequence coding further includes: Determining a residual value between the text sequence code and the phoneme sequence code as a target sequence code; Determining the residual value between the target emotion coding vector and the target sequence coding as the sequence coding to be predicted; The target mel spectrum is determined according to the encoding of the sequence to be predicted.
7. The method according to claim 6, characterized in that The step of determining the target mel spectrum according to the sequence encoding to be predicted further includes: Inputting the sequence code to be predicted into the energy prediction model to obtain the energy prediction code; Inputting the sequence code to be predicted into a fundamental frequency prediction model to obtain a fundamental frequency prediction code; Inputting the sequence code to be predicted into a duration prediction model to obtain a duration prediction code; Adding the energy prediction code, the fundamental frequency prediction code, and the duration prediction code to the sequence code to be predicted to obtain a target mel spectrum sequence code; The target mel spectrum sequence is encoded and input into a decoder to obtain a target mel spectrum, wherein the decoder is constructed based on a multi-head self-attention mechanism.
8. The method according to claim 1, characterized in that The determining of the target Mel spectrum according to the text sequence encoding and the average emotion encoding vector further includes: Obtaining a first emotion coding vector and a second emotion coding vector, wherein the first emotion coding vector is the emotion coding vector corresponding to the target speech sample with the weakest emotion intensity, and the second emotion coding vector is the emotion coding vector corresponding to the target speech sample with the strongest emotion intensity; Determining an emotion intensity adjustment rule according to the first emotion coding vector, the second emotion coding vector, and a preset emotion adjustment parameter; Adjusting the average emotion coding vector according to the emotion intensity adjustment rule; Determining the target mel spectrum according to the text sequence encoding and the adjusted average emotion encoding vector; The emotion intensity adjustment rule is: is the adjusted average emotion encoding vector, is the first emotion encoding vector, is the second emotion encoding vector, and α is the emotion adjustment parameter.
9. A speech synthesis system, characterized in that: include: A first acquisition module is configured to acquire a target emotion category set according to the emotion category of the target text, wherein the target emotion category set includes a plurality of target speech samples, and the emotion category of the target speech samples is the same as the emotion category of the target text; A second acquisition module is configured to acquire an average emotion coding vector based on the emotion coding vectors of the target speech samples, wherein the emotion coding vector of the target speech sample is a vector representation corresponding to the emotion intensity of the target speech sample, and the average emotion coding vector is obtained by adding and averaging all the emotion coding vectors; The audio synthesis module is used to obtain the text sequence code of the target text, determine the target Mel spectrum according to the text sequence code and the average emotion code vector, and convert the target Mel spectrum into an audio signal to obtain the target audio.
Citation Information
Patent Citations
Voice synthesis method, related equipment and readable storage medium
CN111128118A
Speech synthesis method, electronic equipment and storage device
CN112786004A