An Adaptive Emotion-Driven Tone Cloning Text-to-Speech Method and Device

Through the adaptive emotion-driven tone cloning text to speech method, the problem of single emotional expression and high training data demand in speech cloning technology is solved, the voice signal and text emotion are adapted, the authenticity and personalization of speech synthesis are improved, and the user experience is improved.

CN119580695BActive Publication Date: 2025-07-18GUANGZHOU ZIWEIYUN TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411781860.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-05
Publication Date
2025-07-18
Estimated Expiration
2044-12-05

AI Technical Summary

Technical Problem

The existing voice cloning TTS technology has the problem of single or unnatural emotional expression in personalized speech synthesis, and the high demand for model training data, resulting in limited user experience fluency and natural interaction capabilities, and cannot efficiently meet the market's high requirements for immediacy and flexibility.

Method used

Adaptive emotion-driven tone cloning text to pronunciation method is adopted. By dividing the target text into sentences, semantic features and phoneme sequences are extracted, combined with reference audio features and spectrum, the emotional color of the speech signal is automatically adjusted, and emotional features are extracted and speech synthesis is performed using deep learning networks to achieve the fit between the speech signal and the text emotion.

Benefits of technology

It realizes that the emotional color of the voice signal is more in line with the text, improves the authenticity and expressiveness of speech synthesis, reduces the demand for model training data, improves the convenience and personalization of services, and enhances the vividness and appeal of communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119580695B_ABST
    Figure CN119580695B_ABST
Patent Text Reader

Abstract

The present application relates to an adaptive emotion-driven voice cloning text-to-speech method and device. The target text is segmented into sentences and semantic features are extracted, and each sentence is split into a target phoneme sequence; the reference text is split into a reference phoneme sequence, and the reference phoneme sequences are respectively combined with each target phoneme sequence to obtain a combined phoneme sequence. The reference audio is processed to obtain reference speech features and spectrograms. Based on the semantic features corresponding to the sentences, the combined phoneme sequences, and the reference speech features, corresponding speech features are obtained. Emotional features are extracted from the semantic features of each sentence. Based on the speech features, emotional features, and target phoneme sequences corresponding to each sentence, and the spectrograms, a speech signal corresponding to the target text is obtained. Emotional features are extracted from the semantic features of the target text to make the emotional color of the speech signal fit the text. Based on the mapping relationship between the preset output features and the emotional features, the emotional color of the speech signal is automatically adjusted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and more particularly, to an adaptive emotion-driven voice cloning text-to-speech method and apparatus. Background Art

[0002] Although the existing voice cloning TTS (Text To Speech) technology has made some progress in realizing personalized speech synthesis, it still faces several key challenges and deficiencies.

[0003] On the one hand, there are limitations in emotional expression. Although some existing technologies attempt to incorporate emotional elements to enhance the realism and expressiveness of synthesized speech, these methods generally have problems of single or unnatural emotional expression. The emotional voice of most systems does not match the emotion of the input text content, resulting in monotonous and dull synthesized speech lacking vividness. For those systems that support emotion-driven, it often requires users or developers to manually assign corresponding emotion tags or parameters to each piece of text, which is both cumbersome and inefficient, seriously restricting the fluency of the user experience and the ability of natural interaction.

[0004] On the other hand, there is a high demand for model training data. Traditional voice cloning methods usually require obtaining a large number of audio samples of specific individuals as training data. This means that to provide customized voice services for each user, an independent model needs to be constructed, which not only greatly increases the consumption of computing resources but also extends the time cycle of service deployment. The ability to quickly respond to users' personalized needs is limited, and it is unable to efficiently meet the high requirements of the market for immediacy and flexibility. Summary of the Invention

[0005] To solve the above at least one defect, the present application proposes an adaptive emotion-driven voice cloning text-to-speech method and apparatus for solving the problem that the emotional color of speech does not match the text.

[0006] According to the first aspect of the present application, there is provided an adaptive emotion-driven voice cloning text-to-speech method, including:

[0007] Segment the target text into at least one sentence, extract the semantic features of each sentence, and split each sentence into a target phoneme sequence;

[0008] Split the reference text into a reference phoneme sequence, and splice the reference phoneme sequence with each of the target phoneme sequences to obtain a combined phoneme sequence for each sentence;

[0009] Extract the features and spectrum of the reference audio corresponding to the reference text to obtain the reference speech features and the reference audio spectrum;

[0010] Process and decode based on the semantic features and combined phoneme sequences corresponding to each sentence, as well as the reference speech features, to obtain the speech features of each sentence;

[0011] Extract the emotional features of each sentence from the semantic features of each sentence;

[0012] Based on the speech features, emotional features and target phoneme sequences corresponding to each sentence, as well as the reference audio spectrum, perform feature audio conversion to obtain the speech signal corresponding to the target text.

[0013] Preferably, extracting the emotional features of each sentence from the semantic features of each sentence specifically includes,

[0014] Adjust the feature dimension of the semantic features of each sentence, and extract the corresponding basic features from the semantic features with adjusted feature dimensions;

[0015] Perform residual connection processing on the basic features, add the processed basic features to the basic features before processing, and perform dimensionality reduction processing on the added result to obtain intermediate output features;

[0016] Perform residual connection processing on the intermediate output features, add the processed intermediate output features to the intermediate output features before processing, and perform dimensionality reduction processing on the added result to obtain the output features of each sentence;

[0017] Based on the mapping relationship between the preset output features and emotional features, perform emotional classification on the output features of each sentence to obtain the emotional features of each sentence;

[0018] Preferably, the adjustment of the feature dimension of the semantic features of each sentence and the extraction of the corresponding basic features from the semantic features with adjusted feature dimensions are implemented through an input processing layer, and the input processing layer is composed of a convolutional layer, a batch normalization layer and a ReLU activation function layer connected in sequence.

[0019] Preferably, the residual connection processing of the basic features, adding the processed basic features to the basic features before processing, and performing dimensionality reduction processing on the added result to obtain intermediate output features are implemented through a first residual connection layer and a first feature dimensionality reduction layer connected to the first residual connection layer;

[0020] and / or,

[0021] The residual connection processing of the intermediate output features, adding the processed intermediate output features to the intermediate output features before processing, and performing dimensionality reduction processing on the added result to obtain the output features of each sentence are implemented through a second residual connection layer and a second feature dimensionality reduction layer connected to the second residual connection layer;

[0022] Among them, both the first residual connection layer and the second residual connection layer are composed of a convolutional layer, a batch normalization layer, a ReLU activation function layer, a convolutional layer, a batch normalization layer, and a ReLU activation function layer connected in sequence;

[0023] Both the first feature dimensionality reduction layer and the second feature dimensionality reduction layer are composed of a convolutional layer, a batch normalization layer, and a ReLU activation function layer connected in sequence.

[0024] Preferably, the sentiment classification of the output feature of each sentence to obtain the sentiment feature of each sentence based on the mapping relationship between the preset output feature and the sentiment feature is implemented through a fully connected classification layer.

[0025] Preferably, the splitting of the target text into at least one sentence, the extraction of the semantic feature of each sentence, and the splitting of each sentence into a target phoneme sequence specifically include:

[0026] Splitting the target text into at least one sentence, inputting each sentence into a text processing network to extract deep semantic information, and obtaining the semantic feature of each sentence;

[0027] Splitting each sentence to obtain the target phoneme sequence of each sentence;

[0028] Splitting the reference text into a reference phoneme sequence, and combining the reference phoneme sequence with each of the target phoneme sequences to obtain the combined phoneme sequence of each sentence;

[0029] and / or

[0030] The extraction of the feature and spectrum of the reference audio corresponding to the reference text to obtain the reference speech feature and the reference audio spectrum specifically includes:

[0031] Using an audio processing network to extract the feature of the reference audio corresponding to the reference text to obtain the reference speech feature, and performing spectrum extraction on the reference audio to obtain the reference audio spectrum.

[0032] Preferably, the processing and decoding based on the semantic feature and the combined phoneme sequence corresponding to each sentence, and the reference speech feature to obtain the speech feature of each sentence specifically are:

[0033] Converting the combined phoneme sequence of each sentence into the corresponding phoneme feature in a high-dimensional vector space;

[0034] Performing dimensionality transformation on the semantic feature corresponding to each sentence respectively to obtain the transformed semantic feature with the same dimension as the phoneme feature;

[0035] The transformed semantic features corresponding to each sentence are respectively fused with the corresponding phoneme features to obtain at least one first fusion feature;

[0036] Perform positional encoding on the at least one first fusion feature respectively to obtain at least one second fusion feature;

[0037] Obtain the reference speech feature as the first speech feature, and pre-define a stop feature EOS, and the stop feature EOS is used as a judgment condition for loop termination;

[0038] Perform decoding multiple times in a cyclic iteration manner. The specific process of each decoding loop includes:

[0039] Perform embedding and positional encoding on the first speech feature to generate an encoded speech feature;

[0040] Respectively splice the at least one second fusion feature with the encoded speech feature to generate at least one spliced feature;

[0041] Construct an attention mask according to the length of the spliced feature, and use the attention mask to decode the corresponding spliced feature to generate a decoding output;

[0042] Perform a fully connected process on the decoding output to obtain an intermediate speech feature for this loop;

[0043] Splice the intermediate speech feature with the first speech feature to generate a second speech feature;

[0044] Among them, if the second speech feature does not contain the stop feature EOS, the second speech feature is used as the first speech feature for the next decoding loop; if the second speech feature contains the stop feature EOS, the loop is terminated, and the second speech feature is used as the speech feature corresponding to the sentence.

[0045] Preferably, the feature audio conversion process is performed based on the speech feature, emotion feature, and target phoneme sequence corresponding to each sentence, and the reference audio spectrum to obtain the speech signal corresponding to the target text, specifically including:

[0046] Splice the speech feature and emotion feature corresponding to each sentence to obtain the speech emotion feature of each sentence;

[0047] Perform quantization processing on the speech emotion feature to obtain the quantization feature of each sentence;

[0048] Process the reference audio spectrum to obtain the embedded feature of the reference audio;

[0049] Encode the quantization features, target phoneme sequences corresponding to each sentence, and the embedding features to generate the mean, variance, and mask of the latent variables as the intermediate representation for speech synthesis;

[0050] Decode the intermediate representation to generate the speech signal corresponding to the target text.

[0051] According to the second aspect of the present application, an adaptive emotion-driven voice cloning text-to-speech device is provided, including;

[0052] A target text processing module for splitting the target text into at least one sentence, extracting the semantic features of each sentence, and splitting each sentence into a target phoneme sequence;

[0053] A reference text processing module for splitting the reference text into a reference phoneme sequence, and combining the reference phoneme sequence with the target phoneme sequences of each sentence to obtain the combined phoneme sequence of each sentence;

[0054] A reference audio processing module for performing feature extraction and spectrum extraction on the reference audio corresponding to the reference text to obtain the reference speech features and the reference audio spectrum;

[0055] A speech feature acquisition module for processing and decoding based on the semantic features, combined phoneme sequences corresponding to each sentence, and the reference speech features to obtain the speech features of each sentence;

[0056] An emotion feature acquisition module for extracting the emotion features of each sentence from the semantic features of each sentence;

[0057] A feature audio conversion module for performing feature audio conversion based on the speech features, emotion features, and target phoneme sequences corresponding to each sentence, and the reference audio spectrum to obtain the speech signal corresponding to the target text.

[0058] According to the third aspect of the present application, an electronic device is provided, including:

[0059] A memory for storing one or more computer programs;

[0060] A processor, when the one or more computer programs are executed by the processor, implementing an adaptive emotion-driven voice cloning text-to-speech method according to the first aspect above.

[0061] According to the fourth aspect of the present application, a computer-readable storage medium is provided, and the computer-readable storage medium stores computer instructions for causing a processor to implement an adaptive emotion-driven voice cloning text-to-speech method according to the first aspect above when executed.

[0062] Based on any of the above aspects, an adaptive emotion-driven voice cloning text-to-speech method and apparatus provided by an embodiment of the present application segment a target text into multiple sentences, extract semantic features and target phoneme sequences of each sentence, extract reference phoneme sequences of a reference text, splice the target phoneme sequences and the reference phoneme sequences to obtain combined phoneme sequences, extract speech features and audio spectra of a reference audio corresponding to the reference text, decode to obtain speech features, extract emotion features of semantic features of each sentence, and perform feature audio conversion based on the speech features, emotion features and target phoneme sequences corresponding to each sentence, and the reference audio spectrum to obtain a speech signal corresponding to the target text; by extracting emotion features from the semantic features of the target text, the emotional color of the speech signal is made more consistent with the text, realizing fast and high-quality synthesis of speech with unique timbre features, and improving the convenience and personalization of services; based on a mapping relationship between preset output features and emotion features, the emotional color of the speech signal is adaptively adjusted, and diverse emotions are incorporated into the speech signal, significantly enhancing the realism and expressiveness of the speech and making the communication more vivid and infectious. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0064] Figure 1 FIG. is a schematic application scenario diagram of an adaptive emotion-driven voice cloning text-to-speech method provided in this embodiment.

[0065] Figure 2 FIG. is a flowchart of an adaptive emotion-driven voice cloning text-to-speech method provided in this embodiment.

[0066] Figure 3 FIG. is a schematic diagram of functional modules of an adaptive emotion-driven voice cloning text-to-speech apparatus provided in this embodiment.

[0067] Figure 4 FIG. is a schematic structural diagram of an electronic device provided in this embodiment.

[0068] Icons: 100, server; 200, terminal; 201, preprocessing module; 202, speech feature acquisition module; 203, emotion feature acquisition module; 204, feature audio conversion module; 710, electronic device; 711, memory; 712, processor; 713, communication module; 714, input / output interface; 715, bus. Detailed implementation mode

[0069] The attached drawings of this application are only for illustrative purposes and should not be construed as limitations on this application. To better illustrate the following embodiments, some components in the drawings will be omitted, enlarged, or reduced, which do not represent the dimensions of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0070] In order to enable those skilled in the art of this technology to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the attached drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0071] It should be noted that the terms "first", "second", etc. in the description, claims, and the above-mentioned attached drawings of this application are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.

[0072] At present, the traditional voice generation method for voice cloning with emotional expression is limited in emotional expression and the process is cumbersome, has high requirements for model training data, cannot efficiently meet the high requirements of the market for instantaneity and flexibility, and limits the fluency of the user experience and the ability of natural interaction.

[0073] Exemplarily, it is a schematic diagram of an application scenario of an adaptive emotion-driven voice cloning text-to-speech method provided for the embodiments of this application. As Figure 1 shown, the application scenario at least includes a server 100 and a terminal 200 that can communicate with the server 100. The server 100 has the function of receiving text information from the terminal device 200 and processing the text information through some steps of the adaptive emotion-driven voice cloning text-to-speech method provided for the embodiments of this application to generate corresponding voice signals; the terminal device 200 has the functions of supporting the input of text information, sending the text information to the server 100, and receiving and playing the voice signals from the server 100.

[0074] It is understandable that the server 100 may be an independent electronic device or a cluster composed of multiple electronic devices; the terminal 200 may be a smart phone terminal, a personal computer, a tablet computer, a vehicle-mounted terminal, etc., but is not limited thereto.

[0075] In an implementable manner, the server 100 and the terminal 200 may respectively execute the adaptive emotion-driven voice cloning text-to-speech method provided by the embodiments of the present application. Alternatively, optionally, part of the adaptive emotion-driven voice cloning text-to-speech method provided by the embodiments of the present application is executed in the server 100 and part is executed in the terminal 200.

[0076] On the one hand, an adaptive emotion-driven voice cloning text-to-speech method provided by the embodiments of the present application, as Figure 2 shown, the method at least includes,

[0077] S1. Segment the target text into at least one sentence, extract the semantic features of each sentence, and split each sentence into a target phoneme sequence;

[0078] S2. Split the reference text into a reference phoneme sequence, and splice the reference phoneme sequence with each of the target phoneme sequences to obtain a combined phoneme sequence for each sentence;

[0079] S3. Extract the feature and spectrum of the reference audio corresponding to the reference text to obtain the reference speech feature and the reference audio spectrum;

[0080] S4. Based on the semantic feature and the combined phoneme sequence corresponding to each sentence, and the reference speech feature, perform processing and decoding to obtain the speech feature of each sentence;

[0081] S5. Extract the emotion feature of each sentence from the semantic features of each sentence;

[0082] S6. Based on the speech feature, emotion feature, and target phoneme sequence corresponding to each sentence, and the reference audio spectrum, perform feature audio conversion to obtain the speech signal corresponding to the target text.

[0083] The technical solutions provided by the present application will be described in detail below with specific embodiments.

[0084] In this embodiment, a text processing network may be used to extract the text semantic features when executing step S1, and the text processing network may adopt a BERT network.

[0085] Specifically, first, the target text is segmented into at least one sentence according to punctuation marks, and the at least one sentence is input into a text processing network to extract the deep semantic information of the sentence, obtaining the semantic features of each sentence, ensuring that the subsequent generated speech signal can accurately restore the target text semantics.

[0086] At the same time, the at least one sentence is also split into a target phoneme sequence.

[0087] Exemplarily, when the input target text is in Chinese, first, the target text is segmented into at least one sentence according to punctuation marks, then each segmented sentence is converted into pinyin, and the pinyin is converted into corresponding phoneme numbers, obtaining at least one target phoneme sequence;

[0088] Similarly, in step S2, the same splitting method as in step S1 can be adopted to split the reference text into a reference phoneme sequence.

[0089] Finally, the one reference phoneme sequence is respectively combined with the target phoneme sequences of each sentence to obtain the combined phoneme sequence of each sentence;

[0090] Optionally, in step S3, an audio processing network is used to extract features and spectra of the reference audio, obtaining reference speech features and a reference audio spectrum. The reference speech features may include timbre features, which are used to make the generated speech signal have a more realistic timbre feature in subsequent steps, and the reference audio spectrum is used for subsequent feature audio conversion.

[0091] In this embodiment, the audio processing network uses the HuBERT network, and other audio processing networks with feature extraction and spectrum extraction functions can also be used.

[0092] In this embodiment, the specific process of S4 includes:

[0093] The combined phoneme sequence of each sentence is input into a phoneme feature embedding layer to be converted into corresponding phoneme features in a high-dimensional vector space;

[0094] The semantic features of each sentence are input into a semantic feature dimension transformation layer for dimension transformation to increase the dimension of the semantic features and align their dimensions with the dimensions of the corresponding phoneme features, obtaining transformed semantic features. The corresponding transformed semantic features are added to the phoneme features of each sentence to ensure that the pronunciation of the text is not only accurate but also contains the deep meaning of the target text;

[0095] The fusion feature position encoding layer fuses the transformed semantic features corresponding to each sentence with the corresponding phoneme features to obtain at least one first fusion feature, and at the same time, performs position encoding on the at least one first fusion feature to obtain at least one second fusion feature. Among them, the position encoding process can be expressed as:

[0096]

[0097] where x (p,i) represents the i-th element in the fusion feature, represents the encoded element, and represents the positional encoding, and dmode represents the total dimension of the fusion feature;

[0098] In this embodiment, the core of step S4 is to perform decoding using an autoregressive decoding layer. The autoregressive decoding layer is a decoder composed of a speech feature embedding layer, a positional embedding layer, and multiple Transformer self-attention layers. The specific decoding process includes:

[0099] Obtain the reference speech feature as the first speech feature, and pre-define a stop feature EOS, which is used as the determination condition for loop termination;

[0100] The speech feature embedding layer embeds the first speech feature, and the processed result is positionally encoded through the positional embedding layer to generate an encoded speech feature;

[0101] Concatenate each of the at least one second fusion feature with the encoded speech feature to generate at least one concatenated feature;

[0102] Construct an attention mask according to the length of the concatenated feature, and the attention mask is used to ensure that the decoding process follows the autoregressive principle;

[0103] The multiple Transformer self-attention layers decode the corresponding concatenated feature based on the attention mask to generate a decoding output;

[0104] Perform a fully connected process on the decoding output to obtain the intermediate speech feature for this loop;

[0105] Concatenate the intermediate speech feature with the first speech feature to generate a second speech feature;

[0106] Determine whether the second speech feature contains the pre-defined stop feature EOS,

[0107] If not, the second speech feature is used as the first speech feature for the next decoding loop;

[0108] If so, terminate the loop, and the second speech feature is used as the speech feature of the corresponding sentence.

[0109] The feature interaction and decoding are performed in a cyclic iteration manner. In each decoding cycle, new speech features are generated based on the input first speech features. The step S4 in the embodiments of the present invention only requires a small amount of reference speech as input to achieve the instant timbre cloning function without a complex training process. The cloning process is rapid and efficient, saving the time-consuming step of building a model for each user individually. This means that users can immediately experience the text content read in their unique timbre, greatly improving the convenience and personalization of the service.

[0110] To execute the step S5, a one-dimensional convolutional network composed of multiple one-dimensional convolutional layers connected in series is constructed. In this embodiment, the specific architecture of the one-dimensional convolutional network includes, but is not limited to, an input processing layer, a first residual connection layer, a first feature dimensionality reduction layer, a second residual connection layer, a second feature dimensionality reduction layer, and a fully connected classification layer connected in sequence. By gradually refining information, the recognition and classification of high-level emotional features are gradually transitioned from the extraction of basic features.

[0111] Specifically,

[0112] The input processing layer is used to adjust the feature dimension of the semantic features of each input sentence, extract the corresponding basic features from the semantic features with adjusted feature dimensions, and lay a foundation for subsequent processing;

[0113] The first residual connection layer processes the basic features, adds the processed basic features to the basic features, and the addition result is subjected to dimensionality reduction processing through the first feature dimensionality reduction layer to obtain intermediate output features;

[0114] The second residual connection layer processes the intermediate output features, adds the processed intermediate output features to the intermediate output features, and the addition result is subjected to dimensionality reduction processing through the second feature dimensionality reduction layer to obtain the output features of the sentence.

[0115] Exemplarily, both the first residual connection layer and the second residual connection layer are composed of a convolutional layer, a batch normalization layer, a ReLU activation function layer, a convolutional layer, a batch normalization layer, and a ReLU activation function layer connected in sequence. The first residual connection layer and the second residual connection layer provide a shortcut path connecting the input and the output, strengthening the network's learning of complex emotional features, while maintaining the coherence of the information flow, effectively avoiding the common gradient disappearance problem in deep networks.

[0116] Both the first feature dimensionality reduction layer and the second feature dimensionality reduction layer are composed of a convolutional layer, a batch normalization layer, and a ReLU activation function layer connected in sequence. The convolutional layers of the first feature dimensionality reduction layer and the second feature dimensionality reduction layer will reduce the feature length while increasing the number of feature channels, enabling the network to further focus on more critical emotional patterns, thereby extracting more abstract and expressive emotional features.

[0117] A fully connected classification layer is configured at the end of the one-dimensional convolutional network, which is used to perform sentiment classification on the output features of the sentence according to the mapping relationship between the preset output features and sentiment features, so as to obtain the sentiment features of each sentence.

[0118] Among them, the mapping relationship between the preset output features and sentiment features is obtained by pre-training the one-dimensional convolutional network. In the actual application process, the fully connected classification layer does not participate in the operation, but directly obtains the output features with rich sentiment output by the last feature dimension reduction layer, providing sentiment features for subsequent speech signal synthesis.

[0119] In an alternative implementation, when training the one-dimensional convolutional network, the Softmax function and the cross-entropy loss function Ce can be used, specifically:

[0120]

[0121] where e i is the sentiment feature of the i-th sentence classified by the fully connected classification layer, p i is the predicted value of the sentiment feature, and y i is the true value of the sentiment feature.

[0122] Through the processing of the one-dimensional convolutional network in step S5, the present invention not only efficiently extracts sentiment features from the target text, but also optimizes the performance during actual application while ensuring the training effect, bringing a more delicate and realistic sentiment expression ability to the speech synthesis technology.

[0123] In addition to the architecture methods listed in this embodiment, other architecture methods can also be adopted to achieve the same technical effects as this embodiment.

[0124] Optionally, in order to obtain a sentiment-colored speech signal that fits the text semantics, a decoder needs to be used to comprehensively decode the speech features, sentiment features, text phoneme sequence, and spectrogram of the reference audio into a high-quality speech signal.

[0125] In this embodiment, a feature audio conversion module is adopted to execute the step S6. The feature audio conversion module is developed based on the decoder structure in the VITS variational autoencoder network and is designed specifically for feature audio conversion. Its core function is to comprehensively decode the preprocessed speech features, sentiment features, text phoneme sequence, and spectrogram of the reference audio into a high-quality speech signal, achieving a high degree of unity of content, sentiment, and personalization in speech synthesis by integrating sentiment information and the style of the reference audio.

[0126] The working process of the feature audio conversion module is specifically as follows:

[0127] Combine the voice features and emotional features corresponding to each sentence to obtain the voice-emotion features of each sentence;

[0128] Perform quantization processing on the voice-emotion features to obtain the quantization features of each sentence;

[0129] Use an encoder to process the reference audio spectrum to obtain the embedding features of the reference audio;

[0130] Use an encoder to encode the quantization features and target phoneme sequences corresponding to each sentence, as well as the embedding features, to generate the mean, variance, and mask of the latent variables as the intermediate representation of speech synthesis;

[0131] Finally, use a decoder to decode the intermediate representation generated by the encoder to generate the voice signal corresponding to the target text.

[0132] The feature audio conversion module follows the efficient decoding logic of the VITS network, integrates the influence of emotional features and reference audio, realizes the accurate conversion from highly abstract feature representations to natural speech waveforms, and improves the naturalness and expressiveness of the synthesized speech.

[0133] The present invention integrates advanced natural language processing and emotion analysis technologies, enabling the TTS system based on the method of the present invention to automatically analyze the implicit emotion information in the input text. Through in-depth understanding of the context, the system can adaptively adjust the emotional color of the speech, and naturally integrate diverse emotions such as joy, sadness, and surprise into the output speech without manual intervention, significantly enhancing the realism and expressiveness of the synthesized speech and making the communication more vivid and contagious.

[0134] On the other hand, this embodiment provides an adaptive emotion-driven voice cloning text-to-speech device, as Figure 3 shown, including,

[0135] A target text processing module 201 for splitting the target text into at least one sentence, inputting each sentence into a text processing network to extract deep semantic information, and obtaining the semantic features of each sentence;

[0136] The target text processing module 201 also splits each sentence to obtain the target phoneme sequence of each sentence;

[0137] A reference text processing module 202 for splitting the reference text into a reference phoneme sequence, and combining the reference phoneme sequence with each of the target phoneme sequences to obtain the combined phoneme sequence of each sentence;

[0138] The reference audio processing module 203 is used to extract features from the reference audio corresponding to the reference text by using an audio processing network to obtain reference speech features, and to extract the spectrum of the reference audio to obtain a reference audio spectrum.

[0139] The speech feature acquisition module 204, the specific working process is as follows:

[0140] Convert the combined phoneme sequence of each sentence into corresponding phoneme features in a high-dimensional vector space;

[0141] Respectively perform dimensionality transformation on the semantic features corresponding to each sentence to obtain transformed semantic features with the same dimension as the phoneme features;

[0142] The transformed semantic features corresponding to each sentence are respectively fused with the corresponding phoneme features to obtain at least one first fusion feature;

[0143] Respectively perform position encoding on the at least one first fusion feature to obtain at least one second fusion feature;

[0144] Obtain the reference speech feature as the first speech feature, and pre-define a stop feature EOS, and the stop feature EOS is used as a determination condition for loop termination;

[0145] Perform decoding multiple times in a cyclic iteration manner. The specific process of each decoding loop includes:

[0146] Perform embedding and position encoding on the first speech feature to generate an encoded speech feature;

[0147] Respectively splice the at least one second fusion feature with the encoded speech feature to generate at least one spliced feature;

[0148] Construct an attention mask according to the length of the spliced feature, and use the attention mask to decode the corresponding spliced feature to generate a decoded output;

[0149] Perform a fully connected process on the decoded output to obtain an intermediate speech feature for this loop;

[0150] Splice the intermediate speech feature with the first speech feature to generate a second speech feature;

[0151] Wherein, if the second speech feature does not contain the stop feature EOS, the second speech feature is used as the first speech feature for the next decoding loop; if the second speech feature contains the stop feature EOS, the loop is terminated, and the second speech feature is used as the speech feature corresponding to the sentence.

[0152] An emotion feature acquisition module 205, which is used to extract the emotion feature of each sentence from the semantic features of each sentence, including an output processing layer, a first residual connection layer, a first feature dimensionality reduction layer, a second residual connection layer, a second feature dimensionality reduction layer, and a fully connected classification layer connected in sequence.

[0153] Among them, both the first residual connection layer and the second residual connection layer are composed of a convolutional layer, a batch normalization layer, a ReLU activation function layer, a convolutional layer, a batch normalization layer, and a ReLU activation function layer connected in sequence; both the first feature dimensionality reduction layer and the second feature dimensionality reduction layer are composed of a convolutional layer, a batch normalization layer, and a ReLU activation function layer connected in sequence.

[0154] The specific working process of the emotion feature acquisition module 205 is as follows:

[0155] The input processing layer is used to adjust the feature dimension of the semantic feature of each input sentence, extract the corresponding basic feature from the semantic feature with the adjusted feature dimension, and lay a foundation for subsequent processing.

[0156] The first residual connection layer processes the basic feature, adds the processed basic feature to the basic feature, and the addition result is subjected to dimensionality reduction processing through the first feature dimensionality reduction layer to obtain an intermediate output feature.

[0157] The second residual connection layer processes the intermediate output feature, adds the processed intermediate output feature to the intermediate output feature, and the addition result is subjected to dimensionality reduction processing through the second feature dimensionality reduction layer to obtain the output feature of the sentence.

[0158] The fully connected classification layer performs emotion classification on the output feature of the sentence according to the mapping relationship between the preset output feature and the emotion feature to obtain the emotion feature of each sentence.

[0159] The feature audio conversion module 206 is developed based on the decoder structure in the VITS variational autoencoder network, and is used to perform feature audio conversion based on the speech feature, emotion feature, and target phoneme sequence corresponding to each sentence, and the reference audio spectrum to obtain the speech signal corresponding to the target text. The specific process is as follows:

[0160] The speech feature and emotion feature corresponding to each sentence are concatenated to obtain the speech emotion feature of each sentence.

[0161] The speech emotion feature is quantized to obtain the quantization feature of each sentence.

[0162] The encoder is used to process the reference audio spectrum to obtain the embedding feature of the reference audio.

[0163] Encode the quantization features corresponding to each sentence, the target phoneme sequence, and the embedding features to generate the mean, variance, and mask of the latent variables as the intermediate representation for speech synthesis;

[0164] Finally, use the decoder to decode the intermediate representation generated by the encoder to generate the speech signal corresponding to the target text.

[0165] It can be understood that the above device embodiments and the above method embodiments can correspond to each other, and the descriptions similar to the device embodiments can refer to the method embodiments. An adaptive emotion-driven voice cloning text-to-speech device provided by an embodiment of the present application can execute an adaptive emotion-driven voice cloning text-to-speech method provided by any embodiment of the present application, and has corresponding functional modules and beneficial effects for executing the method. The functional modules of the adaptive emotion-driven voice cloning text-to-speech device can be implemented in the form of hardware, can be implemented by instructions in the form of software, or can also be implemented by a combination of hardware and software modules.

[0166] Optionally, the software module can be located in a random access memory, and storage media such as a read-only memory, a programmable read-only memory, a flash memory, an electrically erasable programmable memory, and a register are all possible. The storage media is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the above method embodiments.

[0167] An embodiment of the present application provides an electronic device 710, the structure of which is as Figure 3 shown. The electronic device 710 can be the server 100 or the terminal 200 shown in this embodiment Figure 1 as shown.

[0168] As Figure 3 shown, the electronic device 710 includes a memory 711, a processor 712, a communication module 713, and an input / output interface 714, etc. Optionally, the memory 711, the processor 712, the communication module 713, and the input / output interface 714 can be connected and communicate through a bus 715.

[0169] The memory 711 is used to store one or more computer programs and transmit the code of the computer programs to the processor 712; when the one or more computer programs are executed by the processor 711, an adaptive emotion-driven voice cloning text-to-speech method in an embodiment of the present application is implemented.

[0170] Optionally, the electronic device 710 may be connected to a network via the communication module 713 to communicate with other devices, such as terminals or servers, via the network to achieve data interaction. The electronic device 710 may be various forms of digital computers, such as, by way of example, desktop computers, servers, workstations, mainframe computers, or other types of computers. The electronic device 710 may also be various forms of mobile terminals, such as, by way of example, smart phones, tablet computers, wearable devices (such as helmets, glasses, watches, etc.) and other similar mobile terminals.

[0171] Optionally, the electronic device 710 may be connected to the required input / output devices, such as keyboards, display devices, etc., via the input / output interface 714. The electronic device 710 itself may have a display device and may also externally connect other display devices via the input / output interface 714. Optionally, a storage device, such as a hard disk, etc., may also be connected via the input / output interface 714 so as to store the data in the electronic device 710 into the storage device, or read the data in the storage device, and may also store the data in the storage device into the memory 711.

[0172] It can be understood that the input / output interface 714 may be a wired interface or a wireless interface. Depending on different actual application scenarios, the devices connected to the input / output interface 714 may be components of the electronic device 710 or external devices connected to the electronic device 710 when needed.

[0173] Optionally, the memory 711 may be a volatile memory and / or a non-volatile memory. The volatile memory may be a random access memory, etc., and the non-volatile memory may be a read-only memory, a programmable read-only memory, an erasable programmable read-only memory, an electrically erasable programmable read-only memory, or a flash memory, etc.

[0174] Optionally, the computer program stored in the processor 711 may be divided into one or more modules. The one or more modules are stored in the memory 711 and are executed by the processor 712 to complete the method provided by the present embodiment itself. The one or more modules may be a series of computer program instruction segments capable of completing specific functions, and the computer program instruction segments are used to describe the execution process of the computer program in the electronic device 710.

[0175] Optionally, the processor 712 may be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 712 include, but are not limited to, a central processing unit, a graphics processing unit, a digital signal processor, various special artificial intelligence computing chips, various processors running machine learning model algorithms, and may also be any suitable controller, microcontroller, processor, etc. The processor 712 executes each method and process of this embodiment. Exemplarily, it is an adaptive emotion-driven voice cloning text-to-speech method according to an embodiment of the present application.

[0176] Optionally, the bus 715 may include a path for transmitting information. The bus 715 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. According to different functions, the bus 715 may be divided into an address bus, a data bus, a control bus, etc.

[0177] In an alternative implementation, an embodiment of the present application also provides a computer storage medium, on which a computer program is stored. When the computer program is executed by a computer, the computer can execute the methods of the above method embodiments. Part or all of the computer program may be loaded and / or installed on the memory 711 of the electronic device 710. When the computer program is executed by the processor 712, one or more steps of an adaptive emotion-driven voice cloning text-to-speech method according to an embodiment of the present application can be executed.

[0178] Optionally, the computer-readable storage medium may be a random access memory, a read-only memory, a programmable read-only memory, an erasable programmable read-only memory, an electrically erasable programmable read-only memory, etc.

[0179] Obviously, the above embodiments of the present invention are merely examples for clearly explaining the technical solutions of the present invention, rather than limitations on the specific implementation manners of the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the claims of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. An adaptive emotion-driven method for voice cloning in text-to-speech, characterized in that, including segmenting the target text into at least one sentence, extracting the semantic features of each sentence, and splitting each sentence into a target phoneme sequence; splitting the reference text into a reference phoneme sequence, and combining the reference phoneme sequence with each of the target phoneme sequences respectively to obtain a combined phoneme sequence for each sentence; performing feature extraction and spectrum extraction on the reference audio corresponding to the reference text to obtain reference speech features and a reference audio spectrum; processing and decoding based on the semantic features and combined phoneme sequences corresponding to each sentence, and the reference speech features to obtain the speech features of each sentence; adjusting the feature dimension of the semantic feature of each sentence, extracting corresponding basic features from the semantic feature with adjusted feature dimension; performing residual connection processing on the basic features, adding the processed basic features to the basic features before processing, and performing dimensionality reduction processing on the added result to obtain intermediate output features; performing residual connection processing on the intermediate output features, adding the processed intermediate output features to the intermediate output features before processing, and performing dimensionality reduction processing on the added result to obtain the output features of each sentence; performing sentiment classification on the output features of each sentence based on the mapping relationship between the preset output features and sentiment features to obtain the sentiment features of each sentence; performing feature audio conversion processing based on the speech features, sentiment features and target phoneme sequences corresponding to each sentence, and the reference audio spectrum to obtain the speech signal corresponding to the target text.

2. The method for converting text to speech by adaptive emotion-driven voice cloning according to claim 1, wherein, The adjustment of the feature dimension of the semantic feature of each sentence and the extraction of corresponding basic features from the semantic feature with adjusted feature dimension are implemented through an input processing layer, and the input processing layer is composed of a convolutional layer, a batch normalization layer and a ReLU activation function layer connected in sequence.

3. An adaptive emotion-driven voice cloning text-to-speech method according to claim 1, characterized in that The residual connection processing of the basic features, adding the processed basic features to the basic features before processing, and performing dimensionality reduction processing on the added result to obtain intermediate output features are implemented through a first residual connection layer and a first feature dimensionality reduction layer connected to the first residual connection layer; and / or The residual connection processing of the intermediate output features, adding the processed intermediate output features to the intermediate output features before processing, and performing dimensionality reduction processing on the added result to obtain the output features of each sentence are implemented through a second residual connection layer and a second feature dimensionality reduction layer connected to the second residual connection layer; wherein, both the first residual connection layer and the second residual connection layer are composed of a convolutional layer, a batch normalization layer, a ReLU activation function layer, a convolutional layer, a batch normalization layer and a ReLU activation function layer connected in sequence; both the first feature dimensionality reduction layer and the second feature dimensionality reduction layer are composed of a convolutional layer, a batch normalization layer and a ReLU activation function layer connected in sequence.

4. An adaptive emotion-driven voice cloning text-to-speech method according to claim 1, characterized in that The sentiment classification of the output features of each sentence based on the mapping relationship between the preset output features and sentiment features to obtain the sentiment features of each sentence is implemented through a fully connected classification layer.

5. An adaptive emotion-driven voice cloning text-to-speech method according to claim 1, characterized in that Slicing the target text into at least one sentence, extracting the semantic features of each sentence, and splitting each sentence into a target phoneme sequence specifically includes: Slicing the target text into at least one sentence, inputting each sentence into a text processing network to extract deep semantic information, and obtaining the semantic features of each sentence; Splitting each sentence to obtain the target phoneme sequence of each sentence; Splitting the reference text into a reference phoneme sequence, and splicing the reference phoneme sequence with each of the target phoneme sequences respectively to obtain the combined phoneme sequence of each sentence; And / or Extracting features and spectrum of the reference audio corresponding to the reference text to obtain reference speech features and reference audio spectrum, specifically including: analyzing the features of the reference audio corresponding to the reference text by using an audio processing network to obtain reference speech features; Performing spectrum extraction on the reference audio corresponding to the reference text to obtain the reference audio spectrum.

6. An adaptive emotion-driven voice cloning text-to-speech method according to claim 1, characterized in that Processing and decoding based on the semantic features and combined phoneme sequences corresponding to each sentence, and the reference speech features, specifically: Converting the combined phoneme sequence corresponding to each sentence into phoneme features in a high-dimensional vector space; Performing dimensionality transformation on the semantic features corresponding to each sentence respectively to obtain transformed semantic features with the same dimension as the phoneme features; Fusing the transformed semantic features corresponding to each sentence with the corresponding phoneme features respectively to obtain at least one first fusion feature; Performing position encoding on the at least one first fusion feature respectively to obtain at least one second fusion feature; Obtaining the reference speech feature as the first speech feature, and predefining the stop feature EOS; Performing decoding in multiple times through a cyclic iteration manner, and the specific process of each decoding cycle includes: Embedding and position encoding the first speech feature to generate an encoded speech feature; Splicing the at least one second fusion feature with the encoded speech feature respectively to generate at least one spliced feature; Constructing an attention mask according to the length of the spliced feature; Using the attention mask to decode the corresponding spliced feature to generate a decoded output; Performing a fully connected process on the decoded output to obtain the intermediate speech feature of this cycle; Splicing the intermediate speech feature with the first speech feature to generate a second speech feature; Wherein, if the stop feature EOS is not contained in the second speech feature, the second speech feature is used as the first speech feature for the next decoding cycle; if the stop feature EOS is contained in the second speech feature, the cycle is stopped, and the second speech feature is used as the speech feature of the corresponding sentence.

7. An adaptive emotion-driven timbre cloning text-to-speech method according to any one of claims 1 to 6, characterized in that Performing feature audio conversion processing based on the speech features, emotional features, and target phoneme sequences corresponding to each sentence, and the reference audio spectrum to obtain the speech signal corresponding to the target text, specifically including: Splicing the speech features and emotional features corresponding to each sentence to obtain the speech emotional feature of each sentence; Performing quantization processing on the speech emotional feature to obtain the quantization feature of each sentence; Process the reference audio spectrum to obtain the embedded features of the reference audio; Encode the quantization features and target phoneme sequences corresponding to each sentence, as well as the embedded features, to generate the mean, variance, and mask of the latent variables as the intermediate representation for speech synthesis; Decode the intermediate representation to generate the speech signal corresponding to the target text.

8. An adaptive emotion-driven voice color cloning text-to-speech device, characterized in that It includes: A target text processing module for splitting the target text into at least one sentence, extracting the semantic features of each sentence, and splitting each sentence into a target phoneme sequence; A reference text processing module for splitting the reference text into a reference phoneme sequence, concatenating the reference phoneme sequence with each target phoneme sequence respectively to obtain the combined phoneme sequence of each sentence; A reference audio processing module for performing feature extraction and spectrum extraction on the reference audio corresponding to the reference text to obtain the reference speech features and the reference audio spectrum; A speech feature acquisition module for processing and decoding based on the semantic features and combined phoneme sequences corresponding to each sentence, as well as the reference speech features to obtain the speech features of each sentence; An emotion feature acquisition module for adjusting the feature dimension of the semantic features of each sentence, extracting the corresponding basic features from the semantic features with adjusted feature dimensions; performing residual connection processing on the basic features, adding the processed basic features to the basic features before processing, and performing dimensionality reduction processing on the added result to obtain the intermediate output features; Performing residual connection processing on the intermediate output features, adding the processed intermediate output features to the intermediate output features before processing, and performing dimensionality reduction processing on the added result to obtain the output features of each sentence; Based on the preset mapping relationship between the output features and emotion features, perform emotion classification on the output features of each sentence to obtain the emotion features of each sentence; A feature audio conversion module for performing feature audio conversion processing based on the speech features, emotion features, and target phoneme sequences corresponding to each sentence, as well as the reference audio spectrum to obtain the speech signal corresponding to the target text.

9. An adaptive emotion-driven voice cloning text-to-speech device according to claim 8, wherein, The processing and decoding based on the semantic features and combined phoneme sequences corresponding to each sentence, as well as the reference speech features to obtain the speech features of each sentence are specifically: Convert the combined phoneme sequence corresponding to each sentence into phoneme features in a high-dimensional vector space; Perform dimensionality transformation on the semantic features corresponding to each sentence respectively to obtain transformed semantic features with the same dimension as the phoneme features; The transformed semantic features corresponding to each sentence are respectively fused with the corresponding phoneme features to obtain at least one first fusion feature; Perform positional encoding on the at least one first fusion feature respectively to obtain at least one second fusion feature; Obtain the reference speech features as the first speech feature and pre-define the stop feature EOS; Perform decoding multiple times in a cyclic iteration manner, and the specific process of each decoding cycle includes: Perform embedding and positional encoding on the first speech feature to generate an encoded speech feature; Respectively splice the at least one second fusion feature with the encoded speech feature to generate at least one spliced feature; Construct an attention mask according to the length of the spliced feature; Use the attention mask to decode the corresponding spliced feature to generate a decoded output; Perform a fully connected process on the decoded output to obtain the intermediate speech feature of this cycle; Splice the intermediate speech feature with the first speech feature to generate a second speech feature; Wherein, if the stop feature EOS is not contained in the second speech feature, the second speech feature is used as the first speech feature for the next decoding cycle; if the stop feature EOS is contained in the second speech feature, the cycle is stopped and the second speech feature is used as the speech feature of the corresponding sentence.

10. An electronic device, comprising: A memory for storing one or more computer programs; A processor, when the one or more computer programs are executed by the processor, implements an adaptive emotion-driven voice cloning text-to-speech method according to claims 1-7 above.

Citation Information

Patent Citations

  • Speech synthesis method and device based on multi-scale emotion, equipment and storage medium

    CN116434730A