End-to-end adaptive controllable emotional speech synthesis method, device and product
By employing an end-to-end adaptive and controllable emotional speech synthesis method, and utilizing multi-scale emotion modeling and conditional adversarial training, the problem of difficulty in characterizing the dynamic evolution of emotions in speech synthesis is solved, thereby improving the naturalness and emotional expressiveness of speech generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANHUI UNIV
- Filing Date
- 2026-03-16
- Publication Date
- 2026-06-09
AI Technical Summary
Existing speech synthesis technology struggles to capture the dynamic evolution of emotional nuances in speech, resulting in overly smooth generated speech that lacks naturalness, emotional expressiveness, and precision in emotional control.
An end-to-end adaptive and controllable emotional speech synthesis method is adopted. Through multi-scale emotion modeling and text-speech modality coordination mechanism, combined with a conditional adversarial training strategy based on Mel spectrum and its time derivative features, the joint modeling and control of emotion category and intensity are achieved.
It achieves more granular emotion control, improves the naturalness and emotional expressiveness of speech generation, and endows synthesized speech with richer emotional color.
Smart Images

Figure CN122177083A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech synthesis technology, and in particular to an end-to-end adaptive and controllable emotional speech synthesis method, apparatus and product. Background Technology
[0002] Emotional speech synthesis systems, as an important research direction to address the limitations of neutral speech synthesis, are dedicated to generating speech with controllable emotional attributes. These systems employ various emotion expression strategies, among which label-based methods and global style transfer methods are two representative paradigms. Label-based methods model the conditional input to reflect the emotional essence; in particular, relative attribute learning, which models differences in emotional intensity by learning a ranking function, has proven effective. Global style transfer methods, on the other hand, extract global static emotion embeddings from reference speech and apply them to the synthesized speech to achieve control over emotional expression.
[0003] In reality, emotional expression in speech is hierarchical, evolving across multiple timescales, from subtle changes at the phoneme level to trends at the sentence level. However, both label-based emotion modeling and global style transfer paradigms largely rely on static, coarse-grained emotion representations, making it difficult to characterize the dynamic evolution of emotion across different timescales, thus limiting the ability to model prosodic-level emotion changes. Secondly, existing fine-grained emotion modeling methods generally face the problem of mismatched conditions between the training and inference phases, making it difficult for the model to stably reconstruct reasonable fine-grained emotion changes during inference. Furthermore, the information asymmetry between speech and text modalities further weakens the generalization ability of emotion representations. In addition, most methods exhibit significant entanglement in modeling emotion categories and intensity, and the emotion control parameters lack clear semantic interpretability, making precise and independent adjustment of emotion intensity difficult to achieve. Summary of the Invention
[0004] To address the technical problems that existing technologies tend to capture coarse or averaged emotional representations, making it difficult to reflect the inherent dynamic characteristics of speech prosody, resulting in overly smooth speech synthesis that affects the naturalness and emotional expressiveness of the generated speech, and insufficient precision in emotional control, this invention provides an end-to-end adaptive and controllable emotional speech synthesis method, device, and product.
[0005] Firstly, the present invention provides an end-to-end adaptive and controllable emotion-based speech synthesis method, which inputs source text into a speech synthesis model. x i and reference audio y j , get and x i Content consistent and with y jEmotionally consistent synthesized speech. Adaptive and controllable emotional speech synthesis methods include: Extract using a text encoder x i Context-dependent text hiding features hid text Then, it is first processed through a coarse-grained emotion encoder and emotion constraint processing. y j Obtaining emotional embedding emo align Then, it is processed through a multi-head attention mechanism. emo align Subsequently, a coarse-grained sentiment representation containing coarse-grained sentiment information and text relevance was obtained. emo c . emo c and hid text The summation yields coarse-grained intermediate features of the sentiment text. hid text_c And it generates phoneme-level sentiment embeddings through a fine-grained sentiment predictor. F p1 Next, a bidirectional potential matching module is used for processing. F p1 Fine-grained sentiment embeddings obtained by bidirectional matching of textual and sentiment latent spaces. emo f .
[0006] Simultaneously obtain speaker identity features spk i and with hid text_c , emo f The summation yields intermediate hidden features that contain text, sentiment, and speaker information. hid text_c_s .right hid text_c_s After distribution alignment processing, a Mel spectrum spectrum is generated using a Mel spectrum decoder. The Mel spectrum spectrum is then converted into a time-domain speech signal to obtain the emotionally synthesized speech.
[0007] Secondly, this invention also proposes an end-to-end adaptive controllable emotional speech synthesis device, comprising: an input device, a synthesizer, and a player; wherein the input device is used to acquire source text. x i and reference audio y j The synthesizer is based on x i and y jEmotionally synthesized speech is synthesized using an end-to-end adaptive controllable emotion-based speech synthesis method as described in the first aspect. A player is used to play the emotionally synthesized speech produced by the synthesizer.
[0008] Thirdly, the present invention also proposes a computer-readable storage medium having a computer program / instructions stored thereon. When executed by a processor, the computer program / instructions implement the steps of the end-to-end adaptive controllable emotion-based speech synthesis method as described in the first aspect.
[0009] Fourthly, the present invention also proposes a computer program product comprising a computer program / instructions. When executed by a processor, the computer program / instructions implement the steps of the end-to-end adaptive controllable emotion-based speech synthesis method as described in the first aspect.
[0010] The beneficial effects of this invention are as follows: This invention controls the category of emotional speech through emotion tags and reference speech, and controls the emotion intensity through a scalar, thereby giving the synthesized speech a richer and more delicate emotional color, achieving finer-grained emotion control precision in speech generation, and improving the naturalness and emotional expressiveness of emotional speech synthesis. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a schematic diagram illustrating the inference process of the adaptive and controllable emotional speech synthesis method. Figure 2 This is a schematic diagram illustrating the principle of the adaptive and controllable emotional speech synthesis method during the training process. Figure 3 This is a scatter plot after the emotional features were processed using the t-SNE algorithm in the experiment. Detailed Implementation
[0013] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0014] It should be noted that when a component is said to be "installed on" another component, it can be directly on the other component or it may be in a component that is centered on it. When a component is said to be "set on" another component, it can be directly set on the other component or it may also be in a component that is centered on it. When a component is said to be "fixed to" another component, it can be directly fixed to the other component or it may also be in a component that is centered on it.
[0015] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the specification of this invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "or / and" as used herein includes any and all combinations of one or more of the associated listed items.
[0016] This embodiment provides an end-to-end adaptive and controllable emotion-based speech synthesis method, which inputs source text into a speech synthesis model. x i and reference audio y j , get and x i Content consistent and with y j Emotionally Consistent Synthetic Speech. This emotional speech synthesis method achieves joint modeling and independent control of emotion category and intensity through multi-scale emotion modeling and text-speech modality coordination mechanism. Simultaneously, a conditional adversarial training strategy based on Mel spectrum and its time derivative features is introduced to enhance the overall performance of the synthesized speech in terms of emotion consistency, dynamic expressiveness, and speech quality. The following section will describe in detail the speech synthesis model and method used.
[0017] Please refer to Figure 1 The speech synthesis model in this embodiment includes: a text encoder, a multi-scale emotion modeling module, a speaker modeling module, a variational adapter, and a Mel-spectrum decoder. The multi-scale emotion modeling module is used to extract emotion from reference speech. y j The system extracts and models sentiment information at different time scales, including: a coarse-grained sentiment modeling module for representing global sentiment trends and a fine-grained sentiment modeling module for characterizing phoneme-level sentiment changes. The coarse-grained sentiment modeling module includes: a coarse-grained sentiment encoder, a sentiment alignment module, and an attention module. The fine-grained sentiment modeling module includes: a fine-grained sentiment encoder, a phoneme-level averaging module, a bidirectional latent matching module, and a fine-grained sentiment predictor.
[0018] In the reasoning process, firstly, the source text... x iThe text is converted into a phoneme sequence and mapped to a phoneme embedding sequence, and then a text encoder is used to extract context-dependent hidden text features. hid text Simultaneously, the reference speech is processed through a coarse-grained emotion encoder and an emotion alignment module. y j Emotional embedding is obtained by imposing emotional constraints. emo align Then, it is processed through the multi-head attention mechanism of the attention module. emo align Subsequently, a coarse-grained sentiment representation containing coarse-grained sentiment information and text relevance was obtained. emo c Then, coarse-grained sentiment representations... emo c With text hiding features hid text The summation yields coarse-grained intermediate features of the sentiment text. hid text_c Coarse-grained sentiment text intermediate features hid text_c Input a fine-grained sentiment predictor to generate phoneme-level sentiment embeddings F p1 Next, a bidirectional potential matching module is used for processing. F p1 Fine-grained sentiment embeddings obtained by bidirectional matching of textual and sentiment latent spaces. emo f Reference audio y j Speaker identity features are obtained from the speaker modeling module. spk i Next, spk i and hid text_c , emo f The summation yields intermediate hidden features that contain text, sentiment, and speaker information. hid text_cf_s Hidden features in the middle hid text_cf_s The input variational adapter predicts and modulates explicit acoustic features of duration, fundamental frequency, and energy to enhance speech prosodic performance. Specifically, the variational adapter contains three predictors: a duration predictor, a pitch predictor, and an energy predictor. The input to each predictor is an intermediate hidden feature. hid text_cf_s The three predictors predict the duration, auxiliary pitch, and energy profile respectively. The predicted duration includes the hidden features. hid text_cf_s From phoneme-level upsampling to frame-level features ftame_hidtext_cf_s The predicted auxiliary pitch and energy profile are then converted into feature embeddings and combined with frame-level features. ftame_hid text_cf_s Adding them together yields the alignment feature. ftame_hid text_cf_s_p_e Alignment features ftame_ hid text_cf_s_p_e A Mel-spectral graph is generated using a Mel-spectral decoder. The Mel-spectral graph is then converted into a time-domain speech signal using a vocoder to obtain the emotionally synthesized speech.
[0019] The purpose of this invention using a multi-scale emotion modeling module is to obtain the emotional representation of reference emotional speech, while simultaneously decoupling the relationship between speaker identity and speech content. This enables control over the emotion category and intensity of any speech during synthesis; that is, given any speaker's reference speech as input, the speech synthesis model can capture the emotion of the reference speech while preserving the original speaker's timbre. Furthermore, a fine-grained emotion encoder achieves a more granular representation of information and generates more nuanced, emotion-controlled speech. This is because for emotional speech, each phoneme carries rich emotional information and acoustic variations.
[0020] Furthermore, in this embodiment, the text encoder used in the above speech synthesis model can consist of six Conformer blocks with a hidden dimension of 384. The coarse-grained sentiment encoder can consist of a stacked convolutional layer module and a recurrent neural network (RNN). Specifically, the input first passes through a stacked structure consisting of six two-dimensional convolutional layers—all convolutional layers use 3×3 kernels and 2×2 strides, and each convolution is followed by batch normalization and a ReLU activation function. The sentiment alignment module can consist of instance normalization computation units. The fine-grained sentiment encoder and the fine-grained sentiment predictor share the same structure, including two convolutional layers (filter size 256, kernel size 3) and one linear layer (compressing the hidden layer to 16 dimensions). The core of the bidirectional latent matching module is to achieve bidirectional matching between the text latent space and the sentiment latent space through an invertible transformation. This can be achieved by a normalized flow composed of stacked affine coupling layers and a stack of WaveNet residual blocks. The normalized flow can be constructed as a volume-conserving transformation with a Jacobian determinant of 1 to simplify the design. The speaker modeling module can encode speaker identities using a learnable speaker lookup table in scenarios where the speaker is known. In scenarios where the speaker is not visible, a general speaker encoder based on the ECAPA architecture can be used to encode the speaker identity from the reference speech. y jThe speaker's identity features are directly extracted to achieve generalized modeling for speakers not yet seen. The three predictors in the variational adapter share the same internal structure, first performing two layers of one-dimensional convolution (1D-Conv) on the input, followed by ReLU activation + LayerNorm + dropout after each convolution. Finally, a linear projection layer maps the convolutional output to a specific sequence of predicted values. The Mel-spectrum decoder consists of six Conformer blocks and one output linear layer. The Conformer blocks have a hidden dimension of 384, and the output linear layer transforms the 384-dimensional hidden layer into an 80-dimensional Mel-spectrum graph.
[0021] Please refer to Figure 2 Another innovation of this invention lies in the training of the speech synthesis model, which needs to be used after training. The improvement in this training method lies in introducing an adversarial training approach between the Mel derivative-enhanced conditional discriminator and the speech synthesis model. The Mel derivative-enhanced conditional discriminator evaluates the predicted Mel spectrogram to distinguish between real and generated samples, and the result is used to influence the parameters of the preceding modules to generate a better Mel spectrogram, thereby constraining the generated speech output by the speech synthesis model. The training process is described in detail below.
[0022] During training, an emotional speech dataset is first acquired. In this embodiment, for the acquired or collected emotional speech dataset, the emotion and speaker of each real emotional speech in the dataset are labeled. For example, the speech “0001_000371_angry.wav” is labeled as “speaker 1” and “emotion: angry”, and “0005_001092_sad.wav” is labeled as “speaker 5” and “emotion: sad”. Then, the text and speech data in the selected corpus are preprocessed as follows: First, the input sentence is converted into a phoneme sequence using a phoneme-to-phoneme (G2P) module, which serves as the input to the speech synthesis model. Second, the Montreal Forced Aligner (MFA) is applied to obtain phoneme boundaries and durations from the speech data. Third, when extracting the Mel spectrogram, a Short Time Fourier Transform (STFT) is used with a step size of 256, a window size of 1024, an FFT size of 1024, and 80 Mel filter banks. The PyWorld vocoder is used to extract auxiliary pitches and energy profiles as additional prosodic features during training and inference. Therefore, the resulting emotional speech dataset includes all phonemes corresponding to each text, the true duration of each phoneme, Mel spectrograms, auxiliary pitches and energy profiles, and the actual emotional speech. In another embodiment, a dataset including all phonemes corresponding to each text, the true duration of each phoneme, Mel spectrograms, auxiliary pitches and energy profiles, and the actual emotional speech can also be directly used as the emotional speech dataset.
[0023] After obtaining the emotional speech dataset, reference speech y j Heyuan Voice y i (Original audio) y i To match the source text x i (Consistent speech signals) are processed by a coarse-grained emotion modeling module to extract their respective global emotion representations, and then processed by an emotion alignment module. y j Global emotional style alignment to y i Within the global emotional style space, thus obtaining emotional embedding. emo align . emo align After processing by the attention module, a coarse-grained sentiment representation is obtained. emo c Source audio y i Frame-level emotional features were extracted using a fine-grained emotion encoder. F f .Will F f Phoneme-level emotional features are obtained by downsampling using a phoneme-level averaging module. F p2 It is worth mentioning that matching is achieved in a unified probability space through a bidirectional latent matching module. F p2 and F p1 Fine-grained emotion embeddings are obtained after latent representation. emo f In this module, a bidirectional latent matching strategy is employed during training. It samples from the prior rather than the posterior and is trained using a bidirectional KL divergence loss to promote consistency between the fine-grained sentiment encoder and the fine-grained sentiment predictor, enhancing robustness under mismatch conditions. This bidirectional KL divergence loss includes complementary front-end losses. L prior Rear section loss L posterior : .
[0024] .
[0025] In the formula, D KL This indicates the calculation of KL divergence loss. q ( z ' | yi ) represents the posterior distribution, derived from the source speech by a fine-grained emotion encoder. y i I learned it from China. q ( z ' | y i The ) represents the intermediate hidden distribution obtained by the fine-grained emotion encoder and the phoneme-level averaging module, i.e., the phoneme-level emotion features. F p2 . q ( z | y i ) represents the posterior distribution, given by q ( z ' | y i The simplified distribution obtained after passing through the bidirectional potential matching module. p ( z ' | hid text_c () represents the intermediate hidden distribution obtained by the fine-grained sentiment predictor, i.e., phoneme-level sentiment embedding. F p1 . p ( z | hid text_c ) represents the prior distribution, by p ( z ' | hid text_c The resulting complex distribution is obtained after passing through the bidirectional latent matching module. The speaker modeling module is used to obtain speaker identity features. spk i . spk i and hid text_c , emo f After addition, the signals are passed sequentially through a variational adapter and a Mel-spectrum decoder to generate the predicted Mel-spectrum. The predicted Mel-spectrum and the source speech are then compared. y i The generated Mel spectrograms are jointly input into the Mel derivative-enhanced conditional discriminator for evaluation. Specifically, Mel spectrograms are generated for both the predicted and real samples, with input consisting of randomly extracted Mel spectrogram segments obtained through a random sliding window. For each selected segment, its Mel derivative-enhanced features are calculated: the Mel spectrogram, its first derivative, and its second derivative. These representations are stacked and concatenated along the channel dimension to form a Mel derivative-enhanced combination, enabling the Mel derivative-enhanced conditional discriminator to simultaneously capture the spectral content and temporal dynamics of the speech signal. To effectively integrate emotional style information, coarse-grained emotional representations are also included. emoc and fine-grained emotional embedding emo f After concatenating with Mel-derivative augmented features, the system is further concatenated with other Mel-derivative augmented features, and finally, the evaluation result (true / false) is output. In this embodiment, the Mel-derivative augmented conditional discriminator can be composed of a series of stacked two-dimensional convolutional layers and fully connected layers. The adversarial loss of this Mel-derivative augmented conditional discriminator... L D for: .
[0026] Adversarial loss of speech synthesis models L G for: .
[0027] In the formula, These represent the acoustic channels corresponding to the Mel spectrum characteristics, first derivative characteristics, and second derivative characteristics, respectively. This indicates that the loss within the same batch is averaged. Y t express t The true Mel-frequency spectrogram at any given time. (Representation) t The Mel spectrum generated by time prediction. This indicates that the Mel derivative-enhanced conditional discriminator is in t The output is the result of inputting a real Mel-frequency spectrogram at any given time. c It conveys emotional information. This indicates that the Mel derivative-enhanced conditional discriminator is in t The output is generated after inputting the Mel-spectrum at each time step. Among these, the adversarial loss... L D In This represents the judgment result of the Mel-derivative-enhanced conditional discriminator on the true sample. That is, the Mel-derivative-enhanced conditional discriminator should, as far as possible, judge the result of this input as 1 (true), thus minimizing the squared error relative to 1. The smaller the squared error, the smaller the loss. Adversarial loss. L D In This represents the judgment result of the Mel-derivative-enhanced conditional discriminator on the predicted generated sample. That is, the Mel-derivative-enhanced conditional discriminator should, as far as possible, judge the result of this input as 0 (false), thus minimizing the squared error relative to 0. The smaller the squared error, the smaller the loss. Adversarial loss. L G In The training objective of the speech synthesis model is to ensure that "the data generated by the generator can be misclassified as real data by the Mel derivative-enhanced conditional discriminator, that is, the Mel derivative-enhanced conditional discriminator's judgment of the predicted generated sample is 1 (true) for this input." Therefore, the goal is to minimize the squared error with respect to 1. The smaller the squared error, the smaller the loss. Finally, in addition to the traditional loss function commonly used in speech synthesis models, a new objective term for controllable and transferable emotion synthesis is introduced, making the overall loss function of the speech synthesis model and the Mel derivative-enhanced conditional discriminator... L total for: L total = L iter + L ssim + L prosody + αL emo1 + βL emo2 + γL prior + δL posterior + L G .
[0028] In the formula, L iter This is the sum of the L1 loss between the predicted and true Mel spectrograms for each Conformer block in the Mel spectrometer decoder. L ssim The structural similarity loss is used in the Mel spectrum decoder to measure the perceived similarity in the final Conformer block. L prosody From y i The L1 loss between the predicted and actual pitch, energy, and duration extracted from the data is calculated. L emo1 Cross-entropy loss for the coarse-grained sentiment modeling module. L emo2 Cross-entropy loss for the fine-grained sentiment modeling module. α , β , γ , δ These represent the weights of the corresponding items; in this embodiment, they can be set to 0.5, 0.5, 0.05, and 0.05, respectively. After obtaining the overall loss function... L totalThen, backpropagation and parameter updates are performed until the loss value converges or the preset number of training epochs is reached, completing the training. The AdamW optimizer can be used, with hyperparameters β1=0.9 and β2=0.98. The training learning rate for both the speech synthesis model and the Mel derivative-enhanced conditional discriminator is configured to 1×10⁻⁶. -4 In scenarios where the speaker is visible, the model was trained for a total of 160k steps, with a batch size of 32 on a single NVIDIA RTX 3090 GPU. In scenarios where the speaker is not visible, a batch size of 32 was achieved on three NVIDIA RTX 3090 GPUs.
[0029] To verify the effectiveness of the present invention, the end-to-end adaptive controllable emotional speech synthesis method (hereinafter referred to as Ours) in the present invention is compared with the iEmoTTS method (hereinafter referred to as iEmoTTS), METTS method (hereinafter referred to as METTS), and CMETTS method (hereinafter referred to as CMETTS).
[0030] First, the sentiment features generated by the coarse-grained sentiment modeling module are processed using the t-SNE algorithm to form, as shown below. Figure 3 The aforementioned scatter plot. Figure 3 In the diagram, each point represents the emotional feature of a speech sentence. It can be seen that speech sentences with the same emotion cluster together well, while speech sentences with different emotions have virtually no overlap and can be distinguished in a two-dimensional plane. This indicates that the coarse-grained emotion modeling module has learned the emotional information in the speech data very well. Furthermore, the features of different speech sentences with the same emotion are not completely clustered together, reflecting the stylistic differences within the same emotion. Next, the data on different methods in Chinese speech character error rate (CER) and tone error rate (TER) were tested. Lower values indicate better performance. The results are shown in the table below: method CER TER iEmoTTS 0.67 0.90 METTS 0.62 0.85 CMETTS 0.64 0.87 Ours 0.60 0.82 It is evident that the method of this invention achieves the best results compared to other methods. Subsequently, the subjective opinion scores (MOS) of different methods in terms of emotional similarity, speaker similarity, and speech naturalness were tested. Higher values indicate better performance, and the results are shown in the table below: method Emotional similarity Speaker similarity naturalness of speech iEmoTTS 3.76 3.81 3.79 METTS 3.93 3.91 3.83 CMETTS 3.58 3.67 3.72 Ours 4.13 4.15 4.10 It can also be seen that the method of this invention achieves the best results compared to other methods. Finally, the sentiment classification accuracy of different methods was tested separately, with higher values indicating better performance. The results are shown in the table below: In terms of the accuracy of sentiment classification, it can be seen that the method of this invention has achieved the best results compared with other methods.
[0031] In another embodiment, an end-to-end adaptive controllable emotional speech synthesis device is proposed, comprising: an input device, a synthesizer, and a player. The input device is used to acquire source text. x i and reference audio y j The synthesizer is based on x i and y j The synthesized emotional speech is synthesized using the end-to-end adaptive controllable emotional speech synthesis method described in the above embodiment. A player is used to play the synthesized emotional speech.
[0032] In another embodiment, a computer-readable storage medium storing a computer program is also proposed. When the computer program is executed by a processor, it implements the steps of the end-to-end adaptive controllable emotion-based speech synthesis method described in the above embodiments. The computer-readable storage medium may include, but is not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.
[0033] In another embodiment, a computer program product is also proposed, comprising computer instructions. These computer instructions are used to cause a computer to perform the steps of the end-to-end adaptive controllable emotion-based speech synthesis method described in the above embodiments. The computer program instructions may exist in a computer-readable medium in forms including, but not limited to, source files, executable files, and installation package files. Accordingly, the computer program instructions may be executed by a computer in ways including, but not limited to: the computer directly executing the instructions; the computer compiling the instructions and then executing the corresponding compiled program; the computer reading and executing the instructions; or the computer reading and installing the instructions and then executing the corresponding installed program.
[0034] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0035] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.
Claims
1. An end-to-end adaptive and controllable emotion-based speech synthesis method, characterized in that, It inputs source text into the speech synthesis model x i and reference audio y j , get and x i The content is consistent and with y j Emotionally consistent synthesized speech; It includes: Extract using a text encoder x i Context-dependent text hiding features hid text ; First, it is processed through a coarse-grained emotion encoder and emotion constraint processing. y j Obtaining emotional embedding emo align Then, it is processed through a multi-head attention mechanism. emo align Subsequently, a coarse-grained sentiment representation containing coarse-grained sentiment information and text relevance was obtained. emo c ; emo c and hid text The summation yields coarse-grained intermediate features of the sentiment text. hid text_c And it generates phoneme-level sentiment embeddings through a fine-grained sentiment predictor. F p1 Next, a bidirectional potential matching module is used for processing. F p1 Fine-grained sentiment embeddings obtained by bidirectional matching of textual and sentiment latent spaces. emo f ; Obtain speaker identity features spk i and with hid text_c , emo f The summation yields intermediate hidden features that contain text, sentiment, and speaker information. hid text_c_s ;right hid text_c_s After distribution alignment processing, a Mel spectrum spectrum is generated using a Mel spectrum decoder. The Mel spectrogram is converted into a time-domain speech signal to obtain emotional synthesized speech.
2. The end-to-end adaptive controllable emotion-based speech synthesis method according to claim 1, characterized in that, spk i The methods of obtaining it include: In a scenario where the speaker is known, the speaker's identity is encoded using a speaker lookup table to obtain... spk i ; Alternatively, in scenarios where the speaker is not visible, the speaker encoder can be used to... y i Extracted spk i .
3. The end-to-end adaptive controllable emotion-based speech synthesis method according to claim 1, characterized in that, Methods for generating Mel spectrum spectra include: Through the variational adapter hid text_c_s Prediction is performed; wherein, the variational adapter includes: a duration predictor, a pitch predictor, and an energy predictor, used to obtain the corresponding duration features, auxiliary pitch features, and energy profile features; Based on duration characteristics hid text_c_s From phoneme-level upsampling to frame-level features ftame_hid text_c_s Embedding of auxiliary pitch features and energy profile features ftame_hid text_c_s The data is then fed into a Mel spectrum decoder to generate a Mel spectrum.
4. The end-to-end adaptive controllable emotion-based speech synthesis method according to claim 1, characterized in that, The speech synthesis model is used after training; Its training methods include: A Mel derivative-enhanced conditional discriminator is introduced to conduct adversarial training with the speech synthesis model in order to constrain the generated speech output by the speech synthesis model; Among them, the adversarial loss of the Mel derivative-enhanced conditional discriminator L D for: ; Adversarial loss of speech synthesis models L G for: ; In the formula, These represent the acoustic channels corresponding to the Mel-frequency spectral characteristics, first-order derivative characteristics, and second-order derivative characteristics, respectively. Y represents the average loss within the same batch. t express t The true Mel spectrum at any given moment. express t Mel spectrum generated at time 10:00 This indicates that the Mel derivative-enhanced conditional discriminator is in t The output is the result of inputting the actual Mel-frequency spectrogram at any given time. c Expressing emotional information, This indicates that the Mel derivative-enhanced conditional discriminator is in t The output is generated after inputting the Mel spectrogram at any given time.
5. The end-to-end adaptive controllable emotion-based speech synthesis method according to claim 4, characterized in that, The speech synthesis model includes: a text encoder, a coarse-grained emotion modeling module, a fine-grained emotion modeling module, a speaker modeling module, a variational adapter, and a Mel-spectrum decoder; the coarse-grained emotion modeling module includes: a coarse-grained emotion encoder, an emotion alignment module, and an attention module; the fine-grained emotion modeling module includes: a fine-grained emotion encoder, a phoneme-level averaging module, a bidirectional latent matching module, and a fine-grained emotion predictor. In the training process of the speech synthesis model, the reference speech is used. y j Heyuan Voice y i The global sentiment representations of each entity are extracted using a coarse-grained sentiment modeling module, and then aligned using a sentiment alignment module. y j Global emotional style alignment to y i Within the global emotional style space, thus obtaining emotional embedding. emo align ; emo align After processing by the attention module, a coarse-grained sentiment representation is obtained. emo c ; Source audio y i Frame-level emotional features were extracted using a fine-grained emotion encoder. F f ;Will F f Phoneme-level emotional features are obtained by downsampling using a phoneme-level averaging module. F p2 Matching in a unified probability space via a bidirectional latent matching module. F p2 and F p1 Fine-grained emotion embeddings are obtained after latent representation. emo f ; The speaker modeling module is used to obtain speaker identity features. spk i ; spk i and hid text_c , emo f After being added together, the mixture is passed through a variational adapter and a Mel spectrum decoder to generate a predicted Mel spectrum. Predicted Mel spectrogram and source speech y i The generated Mel spectrograms are input together into the Mel derivative-enhanced conditional discriminator for evaluation.
6. The end-to-end adaptive controllable emotion-based speech synthesis method according to claim 5, characterized in that, The bidirectional latent matching module samples from priors during training and is trained using bidirectional KL divergence loss to promote consistency between the fine-grained sentiment encoder and the fine-grained sentiment predictor; the bidirectional KL divergence loss includes complementary front-end losses. L prior Rear section loss L posterior : ; ; In the formula, D KL This indicates the calculation of KL divergence loss. q ( z ' | y i ) represents the posterior distribution, derived from the source speech by a fine-grained emotion encoder. y i I learned it in school; q ( z ' | y i ) represents the intermediate hidden distribution obtained by the fine-grained emotion encoder and the phoneme-level averaging module. q ( z | y i ) represents the posterior distribution, given by q ( z ' | y i The simplified distribution obtained after passing through the bidirectional latent matching module; p ( z ' | hid text_c ) represents the intermediate hidden distribution obtained by the fine-grained sentiment predictor. p ( z | hid text_c ) represents the prior distribution, given by p ( z ' | hid text_c The complex distribution obtained after passing through the bidirectional potential matching module.
7. The end-to-end adaptive controllable emotion-based speech synthesis method according to claim 6, characterized in that, Overall loss function of speech synthesis model and Mel derivative-enhanced conditional discriminator L total for: L total = L iter + L ssim + L prosody + αL emo1 + βL emo2 + γL prior + δL posterior + L G ; In the formula, L iter Let L1 be the sum of the predicted and actual Mel spectrograms in the Mel spectrogram decoder. L ssim For the structural similarity loss of the Mel spectrum decoder, L prosody From y i The L1 loss between the predicted and actual pitch, energy, and duration extracted from the data is... L emo1 For the cross-entropy loss of the coarse-grained sentiment modeling module, L emo2 For the cross-entropy loss of the fine-grained sentiment modeling module, α , β , γ , δ These are the weights of the corresponding items; After obtaining the overall loss function L total Then backpropagation and parameter updates are performed until the loss value converges or the preset number of training rounds is reached to complete the training.
8. An end-to-end adaptive controllable emotion-based speech synthesis device, characterized in that, It includes: Input device, used to collect source text x i and reference audio y j ; synthesizer, which is based on x i and y j Emotionally synthesized speech is synthesized using the end-to-end adaptive controllable emotional speech synthesis method as described in any one of claims 1 to 7. A player used to play emotionally synthesized speech produced by a synthesizer.
9. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instruction is executed by the processor, it implements the steps of the end-to-end adaptive controllable emotional speech synthesis method as described in any one of claims 1 to 7.
10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it implements the steps of the end-to-end adaptive controllable emotional speech synthesis method as described in any one of claims 1 to 7.