Speech Synthesis Method, Apparatus, Device, Storage Medium and Program Product

By obtaining and integrating the phoneme, emotional and timbre characteristics of the target text and performing speech synthesis, the problem of single, natural and authenticity of speech synthesis in the prior art is solved, and a rich and diverse voice effects are achieved.

CN114242033BActive Publication Date: 2025-07-01GUANGZHOU KUGOU COMP TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111601435.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-24
Publication Date
2025-07-01
Estimated Expiration
2041-12-24

AI Technical Summary

Technical Problem

Existing speech synthesis technologies are difficult to generate speech with specific emotions and tones, resulting in a single, natural and authentic style of synthesized speech.

Method used

By obtaining the target phoneme, target emotion and target timbre of the target text, characteristics are fused and predicted to generate speech with specific emotions and tones. The specific steps include: obtaining the target phoneme and emotional characteristics and performing feature fusion; pronunciation prediction based on the fusion characteristics and timbre characteristics; decoding the predicted characteristics to generate the target acoustic characteristics; and finally synthesizing the target speech based on the acoustic characteristics.

Benefits of technology

It realizes the generation of voices with different emotions and different tones, enriches the voice effects and improves the naturalness and authenticity of voice.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114242033B_ABST
    Figure CN114242033B_ABST
Patent Text Reader

Abstract

The present application discloses a speech synthesis method, apparatus, device, storage medium and program product, which relates to the field of artificial intelligence. The method includes: obtaining target phonemes, a target emotion and a target voice color of a target text; performing feature fusion on target phoneme features corresponding to the target phonemes and target emotion features corresponding to the target emotion to obtain phoneme fusion features; performing pronunciation prediction based on the phoneme fusion features and target voice color features corresponding to the target voice color to obtain speech pronunciation features corresponding to the target phonemes; performing feature decoding on the speech pronunciation features to obtain target acoustic features; and synthesizing a target speech based on the target acoustic features, where the target speech corresponds to the target text, and the target speech is an audio with the target emotion and the target voice color. The method provided by the embodiments of the present application can obtain voices with different emotions and different voice colors, enrich the voice effect of the synthesized speech, and help improve the naturalness and authenticity of the synthesized speech.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of artificial intelligence, and particularly to a speech synthesis method, apparatus, device, storage medium, and program product. Background Art

[0002] Speech synthesis refers to the process of converting text into audio. In this process, an acoustic model is usually used for speech synthesis.

[0003] In related technologies, an acoustic model is trained using the phonemes of sample texts and the audio corresponding to the sample texts. Then, the phonemes corresponding to the text to be synthesized are converted into acoustic features corresponding to the audio by using the trained acoustic model to achieve speech synthesis. Among them, a phoneme is the smallest speech unit divided according to the natural attributes of speech. Taking Mandarin Chinese as an example, phonemes can include initials, finals, tones, etc. However, the audio obtained by using the acoustic model has a unified style of expressing text, and the synthesized speech is relatively rigid and has a single style. Summary of the Invention

[0004] Embodiments of the present application provide a speech synthesis method, apparatus, device, storage medium, and program product. The technical solutions are as follows:

[0005] On the one hand, embodiments of the present application provide a speech synthesis method, the method including:

[0006] Obtaining target phonemes, a target emotion, and a target timbre of a target text;

[0007] Performing feature fusion on the target phoneme features corresponding to the target phonemes and the target emotion features corresponding to the target emotion to obtain phoneme fusion features;

[0008] Performing pronunciation prediction based on the phoneme fusion features and the target timbre features corresponding to the target timbre to obtain speech pronunciation features corresponding to the target phonemes;

[0009] Performing feature decoding on the speech pronunciation features to obtain target acoustic features;

[0010] Synthesizing target speech based on the target acoustic features, the target speech corresponding to the target text, and the target speech being an audio with the target emotion and the target timbre.

[0011] On the other hand, embodiments of the present application provide a speech synthesis apparatus, the apparatus including:

[0012] An obtaining module, configured to obtain target phonemes, a target emotion, and a target timbre of a target text;

[0013] A first fusion module for performing feature fusion on the target phoneme feature corresponding to the target phoneme and the target emotion feature corresponding to the target emotion to obtain a phoneme fusion feature;

[0014] A first prediction module for performing pronunciation prediction based on the phoneme fusion feature and the target timbre feature corresponding to the target timbre to obtain a speech pronunciation feature corresponding to the target phoneme;

[0015] A first decoding module for performing feature decoding on the speech pronunciation feature to obtain a target acoustic feature;

[0016] A speech synthesis module for synthesizing a target speech based on the target acoustic feature, where the target speech corresponds to the target text, and the target speech is an audio with the target emotion and the target timbre.

[0017] On the other hand, an embodiment of the present application provides a computer device, which includes a processor and a memory. At least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the speech synthesis method as described in the above aspect.

[0018] On the other hand, a computer-readable storage medium is provided, in which at least one instruction, at least one program, a code set or an instruction set is stored, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the speech synthesis method as described in the above aspect.

[0019] On the other hand, an embodiment of the present application provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the speech synthesis method provided in the above aspect.

[0020] The beneficial effects brought by the technical solution provided by the embodiment of the present application at least include:

[0021] In the embodiments of the present application, when synthesizing the voice corresponding to the text, an emotion feature is obtained, and the emotion feature is fused with the phoneme feature corresponding to the text to obtain a phoneme fusion feature after fusing the emotion. At the same time, a timbre feature is also obtained, and the pronunciation prediction is performed using the phoneme fusion feature after fusing the emotion and the timbre feature to obtain the pronunciation feature corresponding to the phoneme, and the acoustic feature of the synthesized voice is obtained by decoding using the pronunciation feature. Since the phoneme and the emotion feature are fused during the voice synthesis process, the synthesized voice can have a specific emotion, and at the same time, the timbre feature is also used for pronunciation prediction, so that the synthesized voice has a specific timbre. Therefore, voices with different emotions and different timbres can be obtained, enriching the voice effect of the synthesized voice, and helping to improve the naturalness and authenticity of the synthesized voice. Description of the Drawings

[0022] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0023] Figure 1 Shows a schematic diagram of an implementation environment provided by an exemplary embodiment of the present application;

[0024] Figure 2 Shows a flowchart of a voice synthesis method provided by an exemplary embodiment of the present application;

[0025] Figure 3 Shows a flowchart of a voice synthesis method provided by another exemplary embodiment of the present application;

[0026] Figure 4 Shows a schematic diagram of the structure of an acoustic model provided by an exemplary embodiment of the present application;

[0027] Figure 5 Shows a flowchart of an acoustic model training method provided by an exemplary embodiment of the present application;

[0028] Figure 6 Is a structural block diagram of a voice synthesis device provided by an exemplary embodiment of the present application;

[0029] Figure 7 Shows a schematic diagram of the structure of a computer device provided by an exemplary embodiment of the present application. Detailed Embodiments

[0030] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail in conjunction with the drawings.

[0031] Please refer toFigure 1 , which shows a schematic diagram of the implementation environment provided by the exemplary embodiments of the present application. The implementation environment may include: a terminal 101 and a server 102.

[0032] The terminal 101 is an electronic device provided with a speech synthesis function. The terminal 101 may be a smart phone, a tablet computer, a smart TV, a digital player, a laptop computer, a desktop computer, and so on. A client providing a speech synthesis function may run on the terminal 101, and the client may be an instant messaging application, a music playing application, a reading application, etc. The embodiments of the present application do not limit the specific type of the terminal 101.

[0033] The server 102 may be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. In the embodiments of the present application, the server is the background server of the client providing the speech synthesis function in the terminal 101, and can convert text into speech.

[0034] The terminal 101 and the server 102 can be directly or indirectly connected through wired or wireless communication methods, and the present application does not limit this here.

[0035] In a possible implementation manner, as Figure 1 shown, the terminal 101 sends the target text to be converted and the emotional type and timbre type corresponding to the synthesized speech to the server 102. After receiving the target text, emotional type, and timbre type, the server 102 performs speech synthesis based on the features corresponding to the target text, emotional type, and timbre type to obtain the acoustic features of the audio, realizing the conversion of text into speech with a specific emotion and a specific timbre type.

[0036] In another possible implementation manner, the above speech synthesis process may also be executed by the terminal 101. The server 102 trains the acoustic model for speech synthesis, and then sends the trained acoustic model to the terminal 101, so that the terminal 101 realizes speech synthesis locally without relying on the server 102. Alternatively, the acoustic model for speech synthesis may also be trained on the terminal 101 side, and the terminal 101 executes the speech synthesis process. The embodiments of the present application do not limit this.

[0037] For the convenience of description, the following embodiments are described by taking the speech synthesis method as being executed by a computer device as an example.

[0038] The method provided by the embodiments of the present application can be applied to dubbing scenarios, such as article dubbing, novel dubbing, magazine dubbing, etc. By using the method provided by this embodiment, during the dubbing process, voices with specified emotions and specified timbres can be generated according to the text content in the book, enriching the dubbing effect.

[0039] Moreover, it can also be applied to the intelligent education scenario, converting the text content to be learned into voices with specific emotions and specific timbre characteristics, thereby simulating a real-person education scenario, which helps to better understand and learn the text content.

[0040] The above is only a schematic description by taking the application scenario as an example. The method provided by the embodiments of the present application can also be applied to other scenarios that require speech synthesis. The embodiments of the present application do not limit the actual application scenarios.

[0041] Please refer to Figure 2 , which shows a flowchart of a speech synthesis method provided by an exemplary embodiment of the present application. This embodiment is described by taking the method being used in a computer device as an example. The method includes the following steps.

[0042] Step 201, obtain the target phonemes, target emotions, and target timbres of the target text.

[0043] Optionally, the target text refers to the text that needs to be converted into voice. Phonemes are the smallest speech units divided according to the natural attributes of speech. The corresponding phonemes may be different for different languages. For example, the phonemes of Mandarin Chinese corresponding to the text are different from those of dialects, or the phonemes of Chinese corresponding to the text are different from the corresponding English phonemes.

[0044] Taking Mandarin Chinese as an example, phonemes may include initials, finals, tones, etc. For example, when the target text is "The weather is really nice today", the corresponding target phonemes may be "jin1 tian1 de1 tian1 qi4 zhen1 hao3". The target phonemes may be the phonemes of the language to be synthesized for the target text. The target phonemes may be obtained by performing front-end processing on the target text.

[0045] The target emotion and target timbre refer to the performance effects of the voice after speech synthesis. Among them, the target emotion and target timbre can be a single emotion and a single timbre for the target text. For example, the target emotion can be happiness, and the target timbre is the timbre of "Zhang San" speaking.

[0046] Step 202, perform feature fusion on the target phoneme features corresponding to the target phonemes and the target emotion features corresponding to the target emotion to obtain phoneme fusion features.

[0047] Among them, the target phoneme feature is a vectorized representation of the target phoneme, and each phoneme information in the target phoneme is included in the target phoneme feature. The target emotion feature is a vectorized representation of the target emotion and is used to indicate the emotion type corresponding to the target emotion.

[0048] After obtaining the target phoneme and the target emotion, the target phoneme and the target emotion are processed to obtain the corresponding target phoneme feature and target emotion feature, so as to fuse the target phoneme feature and the target emotion feature, integrate the emotion into the phoneme, and obtain the fused phoneme fusion feature, so as to make the pronunciation have the target emotion when predicting the pronunciation based on the phoneme.

[0049] Step 203: Perform pronunciation prediction based on the phoneme fusion feature and the target timbre feature corresponding to the target timbre to obtain the speech pronunciation feature corresponding to the target phoneme.

[0050] Optionally, the speech pronunciation feature refers to the speech pronunciation manner. For example, the pronunciation duration, pitch, energy, etc.

[0051] Since the pronunciation manners corresponding to different timbres are different. For example, the pitches corresponding to different people speaking are different. Therefore, in a possible implementation manner, when the computer device performs pronunciation prediction based on the phoneme fusion feature, the timbre feature corresponding to the target timbre is introduced at the same time, so as to obtain the pronunciation manner with a specific timbre and a specific emotion.

[0052] Step 204: Perform feature decoding on the speech pronunciation feature to obtain the target acoustic feature. The target speech corresponds to the target text, and the target speech is an audio with the target emotion and the target timbre.

[0053] Optionally, after predicting the speech pronunciation feature, feature decoding is required. When the computer device decodes it into the acoustic feature corresponding to the audio, subsequent speech synthesis can be performed based on the target acoustic feature.

[0054] Among them, the acoustic feature is used to represent the spectral feature of the speech, and the target acoustic feature is the spectral feature corresponding to the synthesized target speech, which can be a mel-spectrogram, Mel-scale Frequency Cepstral Coefficients (MFCC), Linear Prediction Cepstral Coefficients (LPCC), Perceptual Linear Predictive (PLP), etc.

[0055] Step 205: Synthesize the target speech based on the target acoustic feature. The target speech corresponds to the target text, and the target speech is an audio with the target emotion and the target timbre.

[0056] The computer device can use a vocoder to convert acoustic features to obtain target speech. The target speech is the pronunciation corresponding to the target text, and the pronunciation has a specific emotion and a specific timbre.

[0057] Among them, the vocoder is used to convert acoustic features into playable speech waveforms, that is, to restore acoustic features to audio. Optionally, the vocoder can be a neural network-based vocoder, such as WaveNet, HIFIGAN or MelGAN, etc. The specific structure of the vocoder is not limited in this embodiment.

[0058] In summary, in the embodiment of the present application, when synthesizing the speech corresponding to the text, the emotion feature is obtained, the emotion feature is fused with the phoneme feature corresponding to the text to obtain the phoneme fusion feature after fusing the emotion, and at the same time, the timbre feature is also obtained. The pronunciation prediction is performed using the phoneme fusion feature after fusing the emotion and the timbre feature to obtain the pronunciation feature corresponding to the phoneme, and the acoustic feature of the synthesized speech is obtained by decoding using the pronunciation feature. Since the phoneme and the emotion feature are fused during the speech synthesis process, the synthesized speech can have a specific emotion, and at the same time, the timbre feature is also used for pronunciation prediction, so that the synthesized speech has a specific timbre, so that voices with different emotions and different timbres can be obtained, enriching the speech effect of the synthesized speech, and helping to improve the naturalness and authenticity of the synthesized speech.

[0059] Optionally, the phoneme fusion feature is obtained by fusing the target phoneme feature and the target emotion feature by an emotion fusion network; the speech pronunciation feature is predicted by a speech prediction network for the phoneme fusion feature and the target timbre feature; the target acoustic feature is obtained by decoding the speech pronunciation feature by a decoding network. The process of speech synthesis based on the emotion fusion network, the speech prediction network and the decoding network will be described below by way of example.

[0060] Please refer to Figure 3 , which shows a flowchart of a speech synthesis method provided by another exemplary embodiment of the present application. This embodiment is described by taking this method being used in a computer device as an example. The method includes the following steps.

[0061] Step 301, obtain the target phoneme, target emotion and target timbre of the target text.

[0062] The implementation manner of this step can refer to the above step 201, and will not be repeated in this embodiment.

[0063] Step 302, perform feature encoding on the target phoneme to obtain a target phoneme sequence.

[0064] In a possible implementation, the target phoneme is input into a phoneme embedding layer for embedding processing to obtain an initial phoneme sequence corresponding to the target phoneme. After obtaining the initial phoneme sequence, the initial phoneme sequence is input into an encoding network for encoding to obtain a corresponding target phoneme sequence, where the target phoneme sequence is the target phoneme sequence obtained after feature encoding of the target phoneme.

[0065] Optionally, the encoding network can be a Convolutional Neural Networks (CNN), a Recurrent Neural Network (RNN), a Transformer model, etc. The specific structure of the encoding network is not limited in this embodiment.

[0066] Step 303: Perform feature encoding on the target emotion to obtain an initial emotion sequence.

[0067] In a possible implementation, the target emotion is input into an emotion embedding layer for embedding processing to obtain an emotion embedding vector, that is, the initial emotion sequence.

[0068] Among them, the implementation timing of this step and step 302 can be sequential execution or synchronous execution. This embodiment only illustrates the implementation manner and does not limit the implementation timing.

[0069] Step 304: Perform expansion processing on the initial emotion sequence to obtain a target emotion sequence, and the sequence length of the target emotion sequence is the same as that of the target phoneme sequence.

[0070] Since it is necessary to fuse the target phoneme features and the target emotion features, it is necessary to make the sequence lengths of the emotion sequence and the phoneme sequence the same, that is, perform sequence expansion on the initial emotion sequence so that the sequence length of the target emotion sequence is the same as that of the target phoneme sequence.

[0071] Schematically, when the sequence length of the target phoneme sequence is 30 and the sequence length of the initial emotion sequence is 1, the initial emotion sequence can be copied to obtain a target emotion sequence with a sequence length of 30.

[0072] Step 305: Input the target phoneme sequence and the target emotion sequence into an emotion fusion network for fusion processing to obtain a phoneme fusion sequence.

[0073] In a possible implementation, the target phoneme sequence and the target emotion sequence can be directly sequence-fused to obtain a fused phoneme fusion sequence. However, the fusion effect of the phoneme fusion sequence obtained by directly performing sequence fusion is poor, and the audio emotion expression is rather rigid after synthesizing the speech. Therefore, in another possible implementation, an emotion fusion network is used to fuse the target phoneme sequence and the target emotion sequence. Optionally, the emotion fusion network includes a Long Short-Term Memory (LSTM) structure and a residual shortcut structure. This step can be replaced by the following steps:

[0074] Step 305a: Sequence-fuse the target phoneme sequence and the target emotion sequence to obtain a first phoneme fusion sequence.

[0075] In a possible implementation, the computer device first sequence-fuses the target phoneme sequence and the target emotion sequence. Among them, the sequence fusion can be sequence addition to the sequences to obtain a first phoneme fusion sequence.

[0076] Step 305b: Input the first phoneme fusion sequence into a Long Short-Term Memory (LSTM) network for sequence processing to obtain a second phoneme fusion sequence. The LSTM network is used to embed emotion information into the phoneme context information.

[0077] Since the LSTM network has a good ability to learn the correlation of information before and after in time series for features, therefore, the LSTM network is introduced to process the first phoneme fusion sequence, so as to fuse emotion features based on the correlation between phoneme frames, that is, embed emotion information into the context information of phonemes, so that the fused feature fusion effect is better, and the synthesized audio emotion expression is more delicate.

[0078] In a possible implementation, the computer device needs to determine the number of layers of the LSTM network. Optionally, the number of network layers of the LSTM network is determined according to at least one of the fusion requirement or the computation requirement. The number of network layers is positively correlated with the fusion ability and negatively correlated with the computation amount.

[0079] When it is necessary to make the fusion effect of the text and emotion information better, the first phoneme fusion sequence can be processed based on more LSTM layers; when it is necessary to accelerate the speech synthesis and reduce the computation amount during the speech synthesis process, the first fusion sequence can be processed based on fewer LSTM layers. Or, the number of LSTM layers can be determined by comprehensively considering the fusion effect and the computation amount, and the computation amount can be reduced while ensuring the fusion effect. For example, the first phoneme fusion sequence is processed based on 2 LSTM layers.

[0080] After the computer device processes the first phoneme fusion sequence using the LSTM network, a second phoneme fusion sequence is obtained.

[0081] Step 305c: Perform sequence fusion on the first phoneme fusion sequence and the second phoneme fusion sequence to obtain a phoneme fusion sequence.

[0082] In this embodiment, in addition to using the LSTM network to fuse text and emotion information, a residual structure is introduced to ensure that each phoneme in the phonemes is fused with the emotion features. In a possible implementation manner, the computer device performs sequence fusion on the first phoneme fusion sequence and the second phoneme fusion sequence after directly performing sequence fusion on the target phoneme sequence and the target emotion sequence to obtain a final phoneme fusion sequence. That is, through secondary fusion, a phoneme fusion sequence is obtained to ensure the emotion fusion effect, so that the emotional expression of the synthesized speech is more delicate.

[0083] Step 306: Perform feature encoding on the target timbre to obtain an initial timbre sequence.

[0084] After obtaining the phoneme fusion sequence, the pronunciation features can be predicted based on the phoneme fusion sequence. Since the tones, pitches, etc. corresponding to different timbres are different, when predicting the pronunciation features, the timbre features are introduced to improve the consistency between the timbre features of the synthesized speech and the target timbre.

[0085] In a possible implementation manner, first perform feature encoding on the target timbre to obtain an initial timbre sequence. That is, perform embedding processing on the target timbre to obtain an initial timbre sequence.

[0086] Step 307: Perform expansion processing on the initial timbre sequence to obtain a target timbre sequence, and the target timbre sequence has the same sequence length as the phoneme fusion sequence.

[0087] Since it is necessary to fuse the phoneme fusion sequence with the target timbre features corresponding to the target timbre, it is necessary to make the sequence length of the timbre sequence the same as that of the phoneme fusion sequence, that is, perform sequence expansion on the initial timbre sequence so that the sequence length of the target timbre sequence is the same as that of the timbre fusion sequence.

[0088] Combined with the above example, when the sequence lengths of the target phoneme sequence and the target emotion sequence are 30, the sequence length of the fused timbre fusion sequence is still 30. Therefore, the initial timbre sequence is copied to obtain a target timbre sequence with a sequence length of 30.

[0089] Step 308: Input the target timbre sequence and the phoneme fusion sequence into a speech prediction network for pronunciation prediction to obtain a speech pronunciation sequence corresponding to the target phoneme, and the speech pronunciation sequence is used to represent at least one of the pronunciation duration, pitch, and energy corresponding to the target speech.

[0090] In a possible implementation, the target timbre sequence and the phoneme fusion sequence are sequence-fused to obtain a fused feature sequence, and the fused feature sequence is input into a speech prediction network for pronunciation prediction to obtain the pronunciation duration, pronunciation pitch, and pronunciation energy magnitude corresponding to the target speech.

[0091] Optionally, the speech prediction network is a Variance Adaptor, which may include a duration predictor, a pitch predictor, and an energy predictor. After the fused feature sequence is input into the speech prediction network, the duration sequence of phonemes can be predicted by the duration predictor, the pitch sequence can be obtained by the pitch predictor, and the energy sequence can be obtained by the energy predictor.

[0092] Step 309: Input the speech pronunciation sequence into a decoding network for sequence decoding to obtain the target acoustic features. The decoding network is a Flow structure.

[0093] After the computer device obtains the speech pronunciation features, it decodes the speech pronunciation features to obtain the final target acoustic features. Among them, the target acoustic features are mel-spectrogram features.

[0094] Optionally, the decoding network can be structures such as CNN, RNN, Transformer, etc. In a possible implementation, a Flow structure is used as the decoding network. Among them, Flow is a reversible structure with strong feature fitting ability.

[0095] Moreover, when performing audio synthesis on long texts, that is, texts with more words, the computational amount of the Flow structure is small. For example, when the text length of the target text is T, the computational complexity of the Flow structure is O(T), while the computational complexity of the Transformer structure is O(T*T).

[0096] Step 310: Synthesize the target speech based on the target acoustic features.

[0097] The implementation manner of this step can refer to the implementation manner of step 205 above, and will not be elaborated in this embodiment.

[0098] In this embodiment, an LSTM structure and a shortcut structure are used to fuse the target phoneme features and the target emotion features, thereby improving the fusion effect of phonemes and emotions, making the synthesized speech more delicate in emotion expression, and improving the speech anthropomorphic authenticity and fluency.

[0099] In this embodiment, during the process of predicting the pronunciation features, the target timbre features are introduced, so that the finally obtained target acoustic features have the timbre characteristics corresponding to the target timbre, thereby improving the speech anthropomorphic authenticity and fluency.

[0100] In this embodiment, when decoding the speech pronunciation features, the Flow structure is used for decoding, which can reduce the computational complexity when synthesizing the target acoustic features corresponding to the long text.

[0101] In a possible implementation manner, in addition to introducing the target timbre features during the speech pronunciation prediction process, to strengthen the pronunciation features of the target timbre, during the decoding process, decoding is performed based on both the speech pronunciation features and the target timbre features to obtain the target acoustic features of the target speech, thereby making the target speech more expressive.

[0102] Optionally, feature decoding of the speech pronunciation features may include the following steps:

[0103] Step 1: Perform feature fusion on the speech pronunciation features and the target timbre features to obtain pronunciation fusion features.

[0104] To make the timbre corresponding to the speech more consistent with the pronunciation features of the target timbre, the speech pronunciation features and the target timbre features are subjected to feature fusion, and then feature decoding is performed based on the fused pronunciation fusion features, that is, sequence decoding is performed using both the speech pronunciation sequence and the target timbre sequence. In a possible implementation manner, the speech pronunciation sequence and the target timbre sequence are subjected to sequence fusion to obtain a pronunciation fusion sequence. Among them, the sequence length of the fused pronunciation fusion sequence is the same as that of the speech pronunciation sequence.

[0105] Step 2: Perform feature decoding on the pronunciation fusion features to obtain the target acoustic features.

[0106] After obtaining the pronunciation fusion sequence, the computer device inputs the pronunciation fusion sequence into the decoding network for feature decoding. That is, the Flow structure is used to perform feature decoding on the pronunciation fusion sequence.

[0107] During the process of decoding using the Flow structure, multiple feature inputs are included. During each feature input process, both the speech pronunciation sequence and the target timbre sequence are fused and input, that is, the pronunciation fusion sequence is input each time, so as to better fit the audio acoustic features of different timbres and different emotions.

[0108] In this embodiment, during the decoding process, the target timbre features are introduced, and the Flow structure is used to decode the timbre features and the pronunciation features, providing the ability to fit the timbre features and the pronunciation features, thereby further strengthening the timbre characteristics corresponding to the target acoustic features obtained by decoding and improving the similarity with the target timbre.

[0109] In a possible implementation manner, the model structure of the acoustic model for speech synthesis may be as Figure 4 shown, and the process of synthesizing the target acoustic features based on this acoustic model may be:

[0110] The target phonemes of the target text are input into the Phoneme Embedding layer 401 for embedding processing to obtain an initial phoneme sequence, and the initial phoneme sequence is input into the Encoder 402 for encoding processing to obtain a target phoneme sequence. Also, the target emotion is input into the Emotion Embedding layer 403 for embedding processing to obtain a target emotion sequence. Then, the computer device fuses the target phoneme sequence and the target emotion sequence, and inputs the fused first phoneme fusion sequence into the Emotion Net 404 to obtain a second phoneme fusion sequence, and fuses the first phoneme fusion sequence and the second phoneme fusion sequence to obtain a phoneme fusion sequence.

[0111] Meanwhile, the computer device inputs the target timbre into the Speaker Embedding layer 405 for embedding processing to obtain a target timbre sequence. The computer device fuses the phoneme fusion sequence and the target timbre sequence, and inputs the fused pronunciation fusion sequence into the Variance Adaptor 406 of the speech prediction network to obtain a speech pronunciation sequence, and fuses the speech pronunciation sequence and the target timbre sequence, and inputs the fused pronunciation fusion sequence into the mel-spectrogram Flow Decoder 407 for decoding processing to obtain the target acoustic features, that is, mel spectrogram features.

[0112] In a possible implementation manner, the acoustic model for speech synthesis is obtained by training with training samples in a training set. Optionally, the emotion fusion network, the speech prediction network, and the decoding network are trained based on sample texts, sample speeches, sample emotions, and sample timbres. The sample speech is an audio with sample emotion and sample timbre characteristics, and the sample speech corresponds to the sample text.

[0113] That is, a set of training samples includes sample texts, sample speeches, sample emotions, and sample timbres. Optionally, the same sample text may correspond to different sample speeches. For example, the same sample text corresponds to sample speeches with different emotions or different timbres. Among them, emotions can include different types such as neutral, happy, angry, sad, fearful, disgusted, and surprised, and the timbre can be the timbre corresponding to different people speaking. Since the same sample text may correspond to different sample speeches, each group of training samples needs to be labeled to distinguish the emotions and timbres corresponding to the sample speeches in the training samples. For example, for different sample speeches, they can be labeled as "<audio>, Zhang San, happy", "<audio>, Li Si, neutral".

[0114] In a possible implementation, the computer device trains an acoustic model based on multiple sets of training samples. Among them, the acoustic model includes an emotion fusion network, a speech prediction network, and a decoding network. The following is a schematic description of the training method of the acoustic model.

[0115] Please refer to Figure 5 , which shows a flowchart of an acoustic model training method provided by an exemplary embodiment of the present application. This embodiment is described by taking the method being used in a computer device as an example. The method includes the following steps.

[0116] Step 501: Determine the sample phonemes corresponding to the sample text. The sample phonemes include the pinyin information corresponding to the sample text and the duration information corresponding to each pinyin.

[0117] After obtaining the sample text, the sample text can be converted into sample phonemes. Optionally, the front-end processing module is used to convert the text into phonemes.

[0118] In a possible implementation, during the training process, in addition to obtaining the phonemes corresponding to the sample text, it is also necessary to obtain the duration information corresponding to each pinyin in the sample phonemes, that is, the timestamp information of the sample phonemes, which is the start position and end position of the initials and finals of each pinyin in the sample speech. Thus, the model is trained according to the duration information of each phoneme in the real audio to improve the accuracy of the model in predicting the pronunciation duration.

[0119] Optionally, the timestamp information of the sample phonemes can be obtained through the output of the alignment model. In a possible implementation, the force alignment alignment tool is used to obtain the timestamp information of the sample phonemes. Schematically, when the sample text is "The weather is really good today" and the sample phonemes are "jin1 tian1 de1 tian1 qi4 zhen1 hao3", the timestamp information of the sample phonemes is (in seconds): "j(0.0,0.2)in1(0.2,0.5)t(0.5,0.6)ian1(0.6,0.8)d(0.8,0.9)e1(0.9,1.2)t(1.2,1.3)ian1(1.3,1.6)q(1.6,1.8)i4(1.8,2.0)zh(2.0,2.2)en1(2.2,2.5)h(2.5,2.6)ao3(2.6,2.8)".

[0120] Step 502: Input the sample phoneme features corresponding to the sample phonemes and the sample emotion features corresponding to the sample emotion into the emotion fusion network for feature fusion to obtain the sample phoneme fusion features.

[0121] Optionally, the sample phoneme features include the timestamp information of the sample phonemes. After obtaining the sample phonemes, perform Embedding and Encoder processing on the sample phonemes to obtain a sample phoneme sequence, and perform Embedding processing on the sample emotion to obtain a sample emotion sequence. Similarly, the sequence lengths of the sample phoneme sequence and the sample emotion sequence need to be the same.

[0122] In a possible implementation, input the sample phoneme sequence and the sample emotion sequence into an emotion fusion network for fusion to obtain a sample phoneme fusion sequence, that is, sample phoneme fusion features.

[0123] Optionally, the emotion fusion network includes an LSTM structure and a shortcut structure. The fusion process of the sample phoneme sequence and the sample emotion sequence can refer to the fusion process of the target phoneme sequence and the target emotion sequence by the emotion fusion network in step 305 above, which will not be elaborated in this embodiment.

[0124] Step 503, input the sample phoneme fusion features and the sample timbre features corresponding to the sample timbre into a speech prediction network for pronunciation prediction to obtain the predicted speech pronunciation features corresponding to the sample phonemes.

[0125] After fusing to obtain the sample phoneme fusion features, use the sample phoneme fusion features and the sample timbre features for pronunciation prediction. Among them, pronunciation prediction is performed based on a speech prediction network, and the speech prediction network is a Variance Adaptor.

[0126] Among them, the sample timbre features are the sample timbre sequences obtained after performing Embedding processing on the sample timbre. Optionally, the sequence length of the sample timbre sequence needs to be the same as the sequence length of the sample phoneme fusion sequence.

[0127] Optionally, the predicted speech pronunciation features include predicted pronunciation duration, predicted pronunciation pitch, and predicted pronunciation energy magnitude.

[0128] Optionally, the process of the speech prediction network performing pronunciation prediction on the sample phoneme fusion sequence and the sample timbre sequence can refer to the process of the speech prediction network performing pronunciation prediction on the phoneme fusion sequence and the target timbre sequence in step 308 above, which will not be elaborated in this embodiment. The computer device performs pronunciation prediction through the speech prediction network to obtain a predicted speech pronunciation sequence.

[0129] Step 504, input the predicted speech pronunciation features into a decoding network for feature decoding to obtain predicted acoustic features.

[0130] Optionally, the computer device inputs the predicted speech pronunciation features, i.e., the predicted speech pronunciation sequence, into the decoding network for feature decoding. The decoding network is of the Flow structure, and feature decoding is performed based on the decoding network to obtain the predicted Mel spectrogram features.

[0131] In another possible implementation, during the feature decoding process, decoding can also be performed based on the timbre features. Therefore, during the training process, the computer device can also input the predicted speech pronunciation features and the sample timbre features into the decoding network for feature decoding to obtain the predicted acoustic features, thereby improving the fitting ability of the trained decoding network to the timbre features.

[0132] When inputting the predicted speech pronunciation features and the sample timbre features into the decoding network for feature decoding, the predicted speech pronunciation sequence and the sample timbre sequence are fused in sequence to obtain the predicted pronunciation fusion sequence, and then the predicted pronunciation fusion sequence is input into the decoding network to obtain the predicted Mel spectrogram features.

[0133] Among them, the process of performing feature decoding on the predicted speech pronunciation sequence and the sample timbre sequence based on the decoding network can refer to the process of using the decoding network to perform feature decoding on the speech pronunciation sequence and the target timbre sequence in the above embodiment, which will not be elaborated in this embodiment.

[0134] Step 505: Train the emotion fusion network, the speech prediction network, and the decoding network based on the predicted acoustic features and the sample acoustic features corresponding to the sample speech.

[0135] After the computer device predicts the predicted acoustic features corresponding to the sample text through the acoustic model, the emotion fusion network, the speech prediction network, and the decoding network are trained using the predicted acoustic features and the sample acoustic features to obtain the trained acoustic model, so that the trained acoustic model can be used to implement speech synthesis. In one possible implementation, the training process may include the following steps:

[0136] Step 505a: Determine the error loss between the predicted Mel spectrogram corresponding to the predicted acoustic features and the sample Mel spectrogram corresponding to the sample acoustic features.

[0137] In one possible implementation, the computer device pre-processes the sample speech to obtain the sample Mel spectrogram corresponding to the sample speech, so that after obtaining the predicted Mel spectrogram of the sample text based on the acoustic model, the acoustic model is trained using the error between the sample Mel spectrogram and the predicted Mel spectrogram.

[0138] Optionally, the computer device can use the L1 loss function to determine the error loss value between the sample Mel spectrogram and the predicted Mel spectrogram, and train the acoustic model based on the error loss value. Alternatively, the L2 loss function, Mean-Square Error (MSE) loss function, etc. can also be used to calculate the error loss value. This embodiment does not limit the calculation method of the error loss value.

[0139] Step 505b: Based on the error loss value, update the network parameters of the emotion fusion network, speech prediction network, and decoding network through backpropagation.

[0140] In a possible implementation, after determining the error loss, the network parameters of each network in the acoustic model can be updated based on the error loss through backpropagation, including the network parameters of the emotion fusion network, speech prediction network, and decoding network, until the network parameters meet the training conditions, that is, until the error loss reaches the convergence condition.

[0141] For example, the Adam optimization algorithm can be used to perform backpropagation on the gradient of the acoustic model, update the network parameters of each network in the acoustic model, and obtain the trained acoustic model.

[0142] After obtaining the trained acoustic model, the computer device can use the trained acoustic model to perform speech conversion on different texts, and can obtain acoustic features with different emotions and different timbres, enriching the speech effects of the synthesized speech.

[0143] Figure 6 It is a structural block diagram of a speech synthesis device provided by an exemplary embodiment of the present application. As Figure 6 shown, the device includes:

[0144] An acquisition module 601, configured to acquire the target phonemes, target emotion, and target timbre of the target text;

[0145] A first fusion module 602, configured to perform feature fusion on the target phoneme features corresponding to the target phonemes and the target emotion features corresponding to the target emotion to obtain phoneme fusion features;

[0146] A first prediction module 603, configured to perform pronunciation prediction based on the phoneme fusion features and the target timbre features corresponding to the target timbre to obtain the speech pronunciation features corresponding to the target phonemes;

[0147] A first decoding module 604, configured to perform feature decoding on the speech pronunciation features to obtain target acoustic features;

[0148] A speech synthesis module 605, configured to synthesize a target speech based on the target acoustic features, where the target speech corresponds to the target text, and the target speech is an audio with the target emotion and the target timbre.

[0149] Optionally, the phoneme fusion feature is obtained by fusing the target phoneme feature and the target emotion feature through an emotion fusion network;

[0150] The speech pronunciation feature is predicted by a speech prediction network for the phoneme fusion feature and the target timbre feature;

[0151] The target acoustic feature is obtained by decoding the speech pronunciation feature through a decoding network.

[0152] Optionally, the first fusion module 602 includes:

[0153] A first encoding unit, configured to perform feature encoding on the target phoneme to obtain a target phoneme sequence;

[0154] A second encoding unit, configured to perform the feature encoding on the target emotion to obtain an initial emotion sequence;

[0155] A first expansion unit, configured to perform expansion processing on the initial emotion sequence to obtain a target emotion sequence, where the sequence length of the target emotion sequence is the same as that of the target phoneme sequence;

[0156] A first fusion unit, configured to input the target phoneme sequence and the target emotion sequence into the emotion fusion network for fusion processing to obtain a phoneme fusion sequence.

[0157] Optionally, the first fusion unit is further configured to:

[0158] Perform sequence fusion on the target phoneme sequence and the target emotion sequence to obtain a first phoneme fusion sequence;

[0159] Input the first phoneme fusion sequence into a long short-term memory (LSTM) network for sequence processing to obtain a second phoneme fusion sequence, where the LSTM network is used to embed emotion information into phoneme context information;

[0160] Perform the sequence fusion on the first phoneme fusion sequence and the second phoneme fusion sequence to obtain the phoneme fusion sequence.

[0161] Optionally, the number of network layers of the LSTM network is determined according to at least one of the fusion requirement or the computational load requirement, and the number of network layers has a positive correlation with the fusion ability and a negative correlation with the computational load.

[0162] Optionally, the first prediction module 603 includes:

[0163] A third encoding unit, configured to perform the feature encoding on the target timbre to obtain an initial timbre sequence;

[0164] A second extension unit, configured to perform the extension processing on the initial timbre sequence to obtain a target timbre sequence, where the sequence length of the target timbre sequence is the same as that of the phoneme fusion sequence;

[0165] A prediction unit, configured to input the target timbre sequence and the phoneme fusion sequence into the speech prediction network to perform the pronunciation prediction, so as to obtain a speech pronunciation sequence corresponding to the target phoneme, where the speech pronunciation sequence is used to characterize at least one of the pronunciation duration, pitch, and energy corresponding to the target speech.

[0166] Optionally, the first decoding module 604 is further configured to:

[0167] Input the speech pronunciation sequence and the target timbre sequence into the decoding network to perform sequence decoding, so as to obtain the target acoustic feature, where the decoding network is a Flow structure.

[0168] Optionally, the first decoding module 604 further includes:

[0169] A second fusion unit, configured to perform feature fusion on the speech pronunciation feature and the target timbre feature to obtain a pronunciation fusion feature;

[0170] A decoding unit, configured to perform feature decoding on the pronunciation fusion feature to obtain the target acoustic feature.

[0171] Optionally, the emotion fusion network, the speech prediction network, and the decoding network are trained based on a sample text, a sample speech, a sample emotion, and a sample timbre, where the sample speech is an audio with the sample emotion and sample timbre features, and the sample speech corresponds to the sample text.

[0172] Optionally, the apparatus further includes:

[0173] A determination module, configured to determine a sample phoneme corresponding to the sample text, where the sample phoneme includes the pinyin information corresponding to the sample text and the duration information corresponding to each pinyin;

[0174] A second fusion module, configured to input the sample phoneme feature corresponding to the sample phoneme and the sample emotion feature corresponding to the sample emotion into the emotion fusion network to perform the feature fusion, so as to obtain a sample phoneme fusion feature;

[0175] A second prediction module, configured to input the sample phoneme fusion feature and the sample timbre feature corresponding to the sample timbre into the speech prediction network for the pronunciation prediction, so as to obtain the predicted speech pronunciation feature corresponding to the sample phoneme;

[0176] A second decoding module, configured to input the predicted speech pronunciation feature into the decoding network for the feature decoding, so as to obtain the predicted acoustic feature;

[0177] A training module, configured to train the emotion fusion network, the speech prediction network, and the decoding network based on the predicted acoustic feature and the sample acoustic feature corresponding to the sample speech.

[0178] Optionally, the acoustic feature is a Mel spectrogram feature.

[0179] The training module includes:

[0180] A loss determination unit, configured to determine the error loss between the predicted Mel spectrogram corresponding to the predicted acoustic feature and the sample Mel spectrogram corresponding to the sample acoustic feature;

[0181] A training unit, configured to update the network parameters of the emotion fusion network, the speech prediction network, and the decoding network based on the error loss through backpropagation.

[0182] In summary, in the embodiment of the present application, when synthesizing the speech corresponding to the text, the emotion feature is obtained, the emotion feature is fused with the phoneme feature corresponding to the text to obtain the phoneme fusion feature after fusing the emotion, and at the same time, the timbre feature is also obtained. The phoneme fusion feature after fusing the emotion and the timbre feature are used for pronunciation prediction to obtain the pronunciation feature corresponding to the phoneme, and the pronunciation feature is used for decoding to obtain the acoustic feature of the synthesized speech. Since the phoneme and the emotion feature are fused during the speech synthesis process, the synthesized speech can have a specific emotion, and at the same time, the timbre feature is also used for pronunciation prediction, so that the synthesized speech has a specific timbre, so that the speech with different emotions and different timbres can be obtained, enriching the speech effect of the synthesized speech, and helping to improve the naturalness and authenticity of the synthesized speech.

[0183] It should be noted that: for the device provided in the above embodiment, only the above-mentioned division of each functional module is used for illustration. In actual application, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiment and the method embodiment belong to the same concept, and the implementation process is detailed in the method embodiment, which will not be repeated here.

[0184] Please refer to Figure 7, which shows a schematic structural diagram of a computer device provided by an exemplary embodiment of the present application. Specifically: The computer device 700 includes a Central Processing Unit (CPU) 701, a system memory 704 including a Random Access Memory 702 and a Read Only Memory 703, and a system bus 705 connecting the system memory 704 and the central processing unit 701. The computer device 700 also includes a Basic Input / Output System (Input / Output, I / O system) 706 for facilitating the transfer of information between various components within the computer, and a mass storage device 707 for storing an operating system 713, application programs 714, and other program modules 715.

[0185] The basic input / output system 706 includes a display 708 for displaying information and input devices 709 such as a mouse, keyboard, etc. for user input of information. Among them, both the display 708 and the input devices 709 are connected to the central processing unit 701 through an input / output controller 710 connected to the system bus 705. The basic input / output system 706 may also include an input / output controller 710 for receiving and processing inputs from a plurality of other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 710 also provides output to a display screen, printer, or other types of output devices.

[0186] The mass storage device 707 is connected to the central processing unit 701 through a mass storage controller (not shown) connected to the system bus 705. The mass storage device 707 and its associated computer-readable medium provide non-volatile storage for the computer device 700. That is to say, the mass storage device 707 may include a computer-readable medium (not shown) such as a hard disk or a drive.

[0187] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes random access memory (RAM), read-only memory (ROM), flash memory or other solid-state storage technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic tape cartridges, tapes, magnetic disk storage or other magnetic storage devices. Of course, those skilled in the art will know that the computer storage media is not limited to the above several types. The above-mentioned system memory 704 and mass storage device 707 may be collectively referred to as memory.

[0188] The memory stores one or more programs, and the one or more programs are configured to be executed by one or more central processing units 701. The one or more programs contain instructions for implementing the above method, and the central processing unit 701 executes the one or more programs to implement the methods provided by the above various method embodiments.

[0189] According to various embodiments of the present application, the computer device 700 may also run by connecting to a remote computer on the network through a network such as the Internet. That is, the computer device 700 may be connected to the network 712 through the network interface unit 711 connected to the system bus 705. Or rather, the network interface unit 711 may also be used to connect to other types of networks or remote computer systems (not shown).

[0190] The memory further includes one or more programs, and the one or more programs are stored in the memory. The one or more programs contain steps for the computer device to execute in the method provided by the embodiments of the present application.

[0191] The embodiments of the present application further provide a computer-readable storage medium, in which at least one instruction, at least one segment of program, code set or instruction set is stored, and the at least one instruction, at least one segment of program, code set or instruction set is loaded and executed by a processor to implement the speech synthesis method described in any of the above embodiments.

[0192] An embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the speech synthesis method provided in the above aspect.

[0193] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the computer-readable storage medium can be the computer-readable storage medium included in the memory in the above embodiments; it can also exist alone and be a computer-readable storage medium not assembled into the terminal. At least one instruction, at least one program, a code set, or an instruction set is stored in the computer-readable storage medium, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the speech synthesis method described in any of the above method embodiments.

[0194] Optionally, the computer-readable storage medium may include: ROM, RAM, solid state drives (SSDs), or optical discs, etc. Among them, RAM may include resistive random access memory (ReRAM) and dynamic random access memory (DRAM). The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.

[0195] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware or by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the storage medium mentioned above can be a read-only memory, a magnetic disk, or an optical disc, etc.

[0196] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A speech synthesis method, characterized in that, The method includes: Obtaining the target phonemes, target emotion, and target timbre of the target text; Performing feature encoding on the target phonemes to obtain a target phoneme sequence; performing feature encoding on the target emotion to obtain an initial emotion sequence; performing expansion processing on the initial emotion sequence to obtain a target emotion sequence, where the target emotion sequence has the same sequence length as the target phoneme sequence; performing fusion processing on the target phoneme sequence and the target emotion sequence to obtain a phoneme fusion sequence; Performing feature encoding on the target timbre to obtain an initial timbre sequence; performing expansion processing on the initial timbre sequence to obtain a target timbre sequence, where the target timbre sequence has the same sequence length as the phoneme fusion sequence; inputting the target timbre sequence and the phoneme fusion sequence into a speech prediction network for pronunciation prediction to obtain a speech pronunciation sequence corresponding to the target phonemes, where the speech pronunciation sequence is used to represent at least one of the pronunciation duration, pitch, and energy corresponding to the target speech; Performing feature decoding on the speech pronunciation sequence to obtain target acoustic features; Synthesizing the target speech based on the target acoustic features, where the target speech corresponds to the target text, and the target speech is an audio with the target emotion and the target timbre.

2. The method according to claim 1, wherein The phoneme fusion sequence is obtained by a sentiment fusion network fusing the target phoneme sequence and the target emotion sequence; The target acoustic features are obtained by a decoding network decoding the speech pronunciation sequence.

3. The method according to claim 2, characterized in that, The method further includes: Inputting the target phoneme sequence and the target emotion sequence into the sentiment fusion network for fusion processing to obtain the phoneme fusion sequence.

4. The method according to claim 3, characterized in that, The step of inputting the target phoneme sequence and the target emotion sequence into the sentiment fusion network for fusion processing to obtain the phoneme fusion sequence includes: Performing sequence fusion on the target phoneme sequence and the target emotion sequence to obtain a first phoneme fusion sequence; Inputting the first phoneme fusion sequence into a long short-term memory (LSTM) network for sequence processing to obtain a second phoneme fusion sequence, where the LSTM network is used to embed emotion information into phoneme context information; Performing the sequence fusion on the first phoneme fusion sequence and the second phoneme fusion sequence to obtain the phoneme fusion sequence.

5. The method according to claim 4, characterized in that, The number of network layers of the LSTM network is determined according to at least one of the fusion requirement or the computational amount requirement. The number of network layers has a positive correlation with the fusion ability and a negative correlation with the computational amount.

6. The method according to claim 3, wherein The step of performing feature decoding on the speech pronunciation sequence to obtain target acoustic features includes: Inputting the speech pronunciation sequence into the decoding network for sequence decoding to obtain the target acoustic features, where the decoding network is a Flow structure.

7. The method according to any one of claims 1 to 6, characterized in that The step of performing feature decoding on the speech pronunciation sequence to obtain target acoustic features includes: Performing feature fusion on the speech pronunciation sequence and the target timbre sequence to obtain a pronunciation fusion feature; Performing feature decoding on the pronunciation fusion feature to obtain the target acoustic features.

8. The method according to any one of claims 2 to 6, characterized in that, The emotion fusion network, the speech prediction network, and the decoding network are trained based on sample texts, sample speeches, sample emotions, and sample timbres. The sample speech is an audio with the sample emotion and sample timbre characteristics, and the sample speech corresponds to the sample text.

9. The method according to claim 8, wherein The method further includes: Determining sample phonemes corresponding to the sample text, where the sample phonemes include pinyin information corresponding to the sample text and duration information corresponding to each pinyin; Inputting the sample phoneme features corresponding to the sample phonemes and the sample emotion features corresponding to the sample emotion into the emotion fusion network for feature fusion to obtain sample phoneme fusion features; Inputting the sample phoneme fusion features and the sample timbre features corresponding to the sample timbre into the speech prediction network for pronunciation prediction to obtain predicted speech pronunciation features corresponding to the sample phonemes; Inputting the predicted speech pronunciation features into the decoding network for feature decoding to obtain predicted acoustic features; Training the emotion fusion network, the speech prediction network, and the decoding network based on the predicted acoustic features and the sample acoustic features corresponding to the sample speech.

10. The method according to claim 9, characterized in that, The acoustic feature is a Mel spectrum feature; the training of the emotion fusion network, the speech prediction network, and the decoding network based on the predicted acoustic features and the sample acoustic features corresponding to the sample speech includes: Determining the error loss between the predicted Mel spectrum corresponding to the predicted acoustic feature and the sample Mel spectrum corresponding to the sample acoustic feature; Updating the network parameters of the emotion fusion network, the speech prediction network, and the decoding network through backpropagation based on the error loss.

11. A voice synthesis device, characterized in that, The device includes: An acquisition module, configured to acquire target phonemes, a target emotion, and a target timbre of a target text; A first fusion module, configured to perform feature encoding on the target phonemes to obtain a target phoneme sequence; perform feature encoding on the target emotion to obtain an initial emotion sequence; perform extension processing on the initial emotion sequence to obtain a target emotion sequence, where the target emotion sequence has the same sequence length as the target phoneme sequence; and perform fusion processing on the target phoneme sequence and the target emotion sequence to obtain a phoneme fusion sequence; A first prediction module, configured to perform feature encoding on the target timbre to obtain an initial timbre sequence; perform extension processing on the initial timbre sequence to obtain a target timbre sequence, where the target timbre sequence has the same sequence length as the phoneme fusion sequence; input the target timbre sequence and the phoneme fusion sequence into the speech prediction network for pronunciation prediction to obtain a speech pronunciation sequence corresponding to the target phoneme, where the speech pronunciation sequence is used to represent at least one of the pronunciation duration, pitch, and energy corresponding to the target speech; A first decoding module, configured to perform feature decoding on the speech pronunciation sequence to obtain target acoustic features; A speech synthesis module, configured to synthesize the target speech based on the target acoustic features, where the target speech corresponds to the target text, and the target speech is an audio with the target emotion and the target timbre.

12. A computer device, characterized in that, The computer device includes a processor and a memory. At least one instruction, at least one program, a code set, or an instruction set is stored in the memory. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the speech synthesis method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that, At least one instruction, at least one program, a code set, or an instruction set is stored in the readable storage medium. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the speech synthesis method according to any one of claims 1 to 10.

14. A computer program product, characterized in that, The computer program product includes computer instructions. The computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to implement the speech synthesis method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Audio synthesis method and device, terminal and storage medium

    CN113314093A

  • Voice synthesis data generating device, voice synthesizing device, voice synthesis data generating program, and voice synthesizing program

    JP2006030609A