A speech cloning method and system fusing an emotion enhancement mechanism and a storage medium

By employing a two-stage transfer learning framework and processing multi-source speech datasets, combined with acoustic preprocessing and emotion embedding training, the problem of insufficient emotion feature modeling in speech synthesis is solved. This achieves deep fusion and accurate modeling of emotion speech features, thereby improving the robustness and adaptability of the speech cloning system.

CN121565140BActive Publication Date: 2026-04-21HUNAN XIAOYU ZHIHE TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-23
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing technologies for speech synthesis suffer from insufficient modeling of emotional features, resulting in generated speech that lacks realism and expressiveness, has limited emotional transfer capabilities, and is difficult to adapt to diverse application scenarios.

Method used

A two-stage transfer learning framework is adopted, which combines multi-source speech datasets and acoustic preprocessing. The output of voiceprint features is achieved by constructing a directional construction mechanism, and emotion embedding is introduced for joint training. A multi-level loss function is designed to constrain the balance, and multiplicative adaptive adjustment of emotion loss weights is adopted to construct a dual-track input system for feature fusion and temporal modeling.

Benefits of technology

It achieves deep fusion and accurate modeling of emotional speech features, avoids timbre shift and speech distortion, improves the robustness and controllability of the speech cloning system, and significantly enhances the expression and adaptability of emotional features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121565140B_ABST
    Figure CN121565140B_ABST
Patent Text Reader

Abstract

This invention discloses a speech cloning method, system, and storage medium that integrates an emotion enhancement mechanism, relating to the field of speech synthesis technology. The speech cloning method, system, and storage medium that integrates an emotion enhancement mechanism include the following steps: S1, constructing a transfer learning framework and a multi-source speech dataset, and performing normalization processing; S2, based on the preprocessed multi-source speech dataset, constructing a directional construction mechanism to output voiceprint features; S3, adding emotion embeddings to a neutral model for joint training, outputting time-stamped Mel spectrum; S4, based on an emotion-annotated Changsha dialect corpus, using multiplicative adaptive adjustment of emotion loss weights; S5, constructing a dual-track input system to output emotional speech clone data. This invention effectively improves timbre consistency, emotion consistency, and speech naturalness, solving the problems of timbre and emotion feature separation, timbre drift caused by emotion transfer, and insufficient robustness across dialects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech synthesis technology, specifically to a speech cloning method, system, and storage medium that integrates emotion enhancement mechanisms. Background Technology

[0002] With the increasing application of emotional speech technology in fields such as intelligent interaction, news broadcasting, film and television dubbing, and audiobooks, its core goal is to achieve the natural transmission and realistic reproduction of emotional information in speech. Existing technologies typically rely on a single speaker encoder for voiceprint feature extraction, but this encoder is weak in capturing emotional information. This results in generated speech that, while closely resembling the target speaker's timbre, lacks emotional expressiveness and sounds stiff, failing to meet the practical needs of emotional speech synthesis. Poor emotional adaptability often requires additional manual parameter tuning or retraining to achieve appropriate tone and expression, significantly reducing the system's flexibility and real-time performance.

[0003] For example, invention patent CN120783724A discloses a method, apparatus, device, and medium for voice cloning of emotional expression, including: acquiring a user voice signal that can capture more user emotional information; preprocessing the user voice signal including noise removal; extracting the voiceprint features of the preprocessed voice signal; performing voiceprint cloning based on the voiceprint features and a voiceprint cloning model; analyzing the emotional type of the user voice signal according to the user voice signal; adjusting the cloned voiceprint according to the analysis results to obtain a target voiceprint that better expresses the user's emotions; and finally converting the target voiceprint into a target voice signal and outputting it at a volume greater than 80dB. Therefore, the voice cloning method can accurately capture and reproduce the emotional tone of the user's voice, realize complex emotional expression of the user, and make the cloned voice more natural and vivid, adaptable to scenarios that require subtle emotional expression.

[0004] For example, invention patent CN115359775A discloses an end-to-end Chinese speech cloning method for timbre and emotion transfer, including: collecting user-recorded Chinese speech as training data and extracting the required speech features; training a speech cloning synthesis model, including a timbre and emotion encoder, a synthesizer, and a vocoder; using the trained speech cloning synthesis model, generating the speech of a specified speaker based on the user-input speech or text content; or quickly cloning the timbre and emotion in the user's speech based on a short-time input speech. This invention achieves end-to-end speech synthesis and cloning by using a multi-speaker model to synthesize speech with different emotions and timbre using the same model and different speaker vector embeddings. This invention uses speaker embedding vectors generated from short speech samples, combined with a generative model trained using a large amount of corpus, to perform speech cloning, achieving speech cloning that can reflect the timbre and emotion of a specific speaker.

[0005] Existing technologies have significant shortcomings in emotional feature modeling. Speaker encoders typically only capture voiceprint information and struggle to effectively represent emotional features, resulting in generated speech that lacks realism and expressiveness, failing to meet the demands of emotional speech synthesis. Furthermore, their ability to transfer emotion is limited, often resulting in speech synthesis products that are emotionally simplistic and lack naturalness, making them unsuitable for diverse application scenarios.

[0006] Therefore, in order to address the above problems, there is an urgent need for a voice cloning method, system, and storage medium that integrates emotion enhancement mechanisms. Summary of the Invention

[0007] Technical problems to be solved

[0008] To address the shortcomings of existing technologies, this invention provides a speech cloning method, system, and storage medium that integrates emotion enhancement mechanisms, solving the problems of separation of timbre and emotional features, timbre drift caused by emotion transfer, and insufficient robustness across dialects.

[0009] Technical solution

[0010] To achieve the above objectives, this invention provides the following technical solution: a speech cloning method, system, and storage medium integrating an emotion enhancement mechanism, comprising: S1, constructing a transfer learning framework and a multi-source speech dataset, performing acoustic preprocessing and normalization on the multi-source speech dataset, and obtaining pure timbre cloned speech through a two-stage collaborative training mechanism; S2, based on the preprocessed multi-source speech dataset, constructing a directional construction mechanism to output voiceprint features, and constructing a dialect-adaptive spectrogram generation and WaveRNN vocoder conversion mechanism to output Mel spectrograms and synthesized speech; S3, adding emotion embedding joint training to the neutral model, outputting timestamped Mel spectrograms, and establishing a dialect-adaptive attention and anomaly handling mechanism through neural network adaptive learning of feature associations; S4, based on an emotion-annotated Changsha speech corpus, designing a multi-level loss function constraint balance, using multiplicative adaptive adjustment of emotion loss weights, evaluating synthesized speech, and verifying the effectiveness of weight adjustment; S5, constructing a dual-track input system, performing feature fusion and temporal modeling through deep restricted Boltzmann machine and time-delay neural network modeling, jointly driving the synthesis process to output emotional speech clone data.

[0011] Furthermore, a second aspect of the present invention provides a speech cloning system incorporating an emotion enhancement mechanism, applied to a speech cloning method incorporating an emotion enhancement mechanism, comprising: a transfer learning overall framework and a two-stage training module, used to construct the transfer learning framework and a multi-source speech dataset, perform acoustic preprocessing and normalization on the multi-source speech dataset, and obtain pure timbre cloned speech through a two-stage collaborative training mechanism; a timbre cloning and speaker encoder training module, used to construct a directional construction mechanism to output voiceprint features based on the preprocessed multi-source speech dataset, construct a dialect-adaptive spectro-frequency generation and WaveRNN vocoder conversion mechanism to output Mel spectrograms and synthesized speech; and a speech synthesis... The synthesizer and vocoder modeling module is used to jointly train the neutral model by adding emotion embeddings, outputting timestamped Mel spectra, and establishing dialect-adaptive attention and anomaly handling mechanisms through neural network adaptive learning of feature associations. The emotion cloning and emotion feature learning module is used to design a multi-level loss function constraint balance based on the emotion-annotated Changsha speech corpus, adopt multiplicative adaptive adjustment of emotion loss weights, evaluate synthesized speech, and verify the effectiveness of weight adjustment. The multi-feature fusion and temporal modeling module is used to construct a dual-track input system, and perform feature fusion and temporal modeling through deep restricted Boltzmann machine and time-delay neural network modeling, jointly driving the synthesis process to output emotion speech clone data.

[0012] Furthermore, a third aspect of the present invention provides a voice cloning storage medium incorporating an emotion enhancement mechanism, comprising: the storage medium having one or more programs, the one or more programs being executed by one or more processors to implement a voice cloning method incorporating an emotion enhancement mechanism.

[0013] Beneficial effects

[0014] The present invention has the following beneficial effects:

[0015] (1) This invention introduces a two-stage transfer learning framework in the speech cloning process, simultaneously capturing the speaker's voiceprint information and emotional features, to achieve deep fusion and accurate modeling of emotional speech features, avoiding the problems of timbre shift and speech distortion in traditional methods, realizing the effective transfer and expression of emotional features, and enabling the model to present rich emotional layers while maintaining timbre consistency.

[0016] (2) By adopting a two-stage transfer learning strategy, the present invention can stably extract the voiceprint features of the target speaker in the voice cloning stage, ensuring the individual consistency of voice; it can achieve adaptive fusion of emotional features, avoid expression imbalance caused by feature overfitting, and significantly improve the robustness and controllability of the voice cloning system in real application scenarios.

[0017] (3) This invention establishes a multi-feature fusion mechanism through the e-vector model, jointly modeling speaker identity and emotional features. This enables the establishment of a dynamic mapping relationship between the emotional space and the voiceprint space, thereby effectively improving the accuracy and robustness of emotional feature modeling. In speech synthesis tasks with different emotional categories, it exhibits higher emotional consistency and timbre fidelity, significantly outperforming traditional models.

[0018] (4) This invention introduces a joint feature weight allocation mechanism under a two-stage transfer learning framework, which weights and couples the timbre consistency index with the emotion expression index, thus avoiding the problems of emotion drift and timbre distortion. This improves the training efficiency and real-time generation of the speech cloning process, and also significantly enhances the adaptability and stability of the model in complex emotional scenarios.

[0019] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description

[0020] Figure 1 This is a flowchart of a voice cloning method that integrates emotion enhancement mechanisms according to the present invention;

[0021] Figure 2 This is a framework diagram of a voice cloning system that integrates emotion enhancement mechanisms according to the present invention;

[0022] Figure 3 This is the Mel spectrum diagram of the present invention;

[0023] Figure 4 This is a trend chart showing the change in emotional loss weight during the emotional cloning stage of this invention.

[0024] Figure 5 This is a heatmap of the emotional feature similarity matrix of the present invention;

[0025] Figure 6 This is a schematic diagram of the e-vector speaker encoder structure of the present invention;

[0026] Figure 7 This is a framework diagram of a speech cloning method that integrates emotion enhancement mechanisms according to the present invention. Detailed Implementation

[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0028] Please see Figures 1-7This invention provides a technical solution: a speech cloning method, system, and storage medium integrating an emotion enhancement mechanism, comprising: S1, constructing a transfer learning framework and a multi-source speech dataset, performing acoustic preprocessing and normalization on the multi-source speech dataset, and obtaining pure timbre cloned speech through a two-stage collaborative training mechanism; S2, based on the preprocessed multi-source speech dataset, constructing a directional construction mechanism to output voiceprint features, and constructing a dialect-adaptive spectral frequency generation and WaveRNN vocoder conversion mechanism to output Mel spectrograms and synthesized speech; S3, adding emotion embedding joint training to the neutral model, outputting timestamped Mel spectrograms, and establishing a dialect-adaptive attention and anomaly handling mechanism through neural network adaptive learning of feature association; S4, based on an emotion-annotated Changsha speech corpus, designing a multi-level loss function constraint balance, using multiplicative adaptive adjustment of emotion loss weights, evaluating synthesized speech, and verifying the effectiveness of weight adjustment; S5, constructing a dual-track input system, performing feature fusion and temporal modeling through deep restricted Boltzmann machine and time-delay neural network modeling, jointly driving the synthesis process to output emotional speech clone data.

[0029] Specifically, the process of constructing a transfer learning framework and a multi-source speech dataset, performing acoustic preprocessing and normalization on the multi-source speech dataset, and obtaining pure timbre cloned speech through a two-stage collaborative training mechanism is as follows: Constructing an overall transfer learning framework, splitting speech cloning into a progressive stage of timbre cloning and emotion cloning, and building a neutral model; acquiring and integrating publicly available speech datasets, private target speaker timbre and emotion-inducing data, and cross-scene and cross-device transfer adaptation data; collecting basic Changsha dialect corpus and constructing a multi-source speech dataset, which includes: target Changsha dialect basic... The dataset includes a corpus, a sentiment-annotated Changsha dialect corpus, and a pre-trained basic speech corpus. The target Changsha dialect basic corpus includes: neutral sentiment Changsha dialect speech from the target speaker, corresponding dialect text annotations, and voiceprint calibration samples. These were recorded in a silent environment using professional recording equipment, covering neutral scenarios such as daily conversations and news broadcasts, with consistent sampling rate, quantization bits, and single speech duration. The sentiment-annotated Changsha dialect corpus includes: Changsha dialect speech from the target speaker expressing joy, anger, sorrow, happiness, and neutral emotions, dialect text, sentiment tags, and prosodic parameter annotations. This corpus was created by manually guiding the target speaker to simulate different emotions. For each sample recorded in an emotionally resonant scene, manually extracted values ​​for fundamental frequency, speech rate, and intonation are included. The pre-trained basic speech corpus includes: general dialect speech data and speaker identity labels, which are selected from the publicly available speech database AISHELL-3 and supplemented with Changsha dialect data. Changsha dialect speech of the target speaker in a historically neutral emotional state is collected as a historically neutral speech dataset. A target-oriented acquisition and collaborative data construction mechanism is adopted. Target-oriented acquisition is tailored to the characteristics of Changsha dialect, focusing on recording speech containing dialect-specific vocabulary. The multi-source speech dataset is stored in a dedicated speech data buffer. The data is temporarily stored in a structured manner, organized hierarchically by corpus type, speaker ID, sentiment label, and speech number, preserving the original speech waveform, recording timestamp, device parameters, and unprocessed acoustic features. Acoustic preprocessing and feature Z-Score normalization are then performed on the multi-source speech dataset. Acoustic preprocessing involves noise reduction, pre-emphasis, framing, and windowing of the speech, unifying the basis for acoustic feature extraction. Feature normalization normalizes the multi-source speech dataset to a unified range using the Z-Score normalization method, preserving the bidirectional mapping relationship between the original and normalized features.

[0030] A two-stage collaborative training mechanism is established to progressively transfer timbre to emotion. The preprocessed target Changsha dialect corpus is input into the speaker encoder, and timbre feature vectors are obtained through local acoustic feature extraction and temporal voiceprint modeling. A baseline vector, generated by mean and normalization of features extracted from a historical neutral speech dataset, is used as the standard timbre vector. The standard timbre vector is obtained by subtracting the standard timbre vector from the timbre feature vector. The L1 norm of the difference vector is then squared to obtain the minimized timbre loss value. The specific formula for minimizing the timbre loss value is as follows: ;

[0031] In the formula, This represents the minimum timbre loss value, used to quantify the degree of deviation between the timbre features extracted by the encoder from the input speech and the target standard timbre features. It is the core loss indicator to ensure that the model accurately learns the Changsha dialect timbre of the target speaker during the timbre cloning stage. The square operation of the L1 norm has low sensitivity to outliers of feature deviation, can stably constrain the overall difference between timbre feature vectors, avoid the interference of local extreme deviations on the training direction, and ensure the stability of timbre reproduction. This represents the timbre feature vector, which contains unique voiceprint details of the target speaker's Changsha dialect; The standard timbre vector is used as the target anchor point for timbre cloning training to ensure that the timbre learned by the model is consistent with the authentic Changsha dialect timbre of the target speaker.

[0032] As shown in Table 1, in Sample 1, the mean feature vector of the input speech encoder is -0.123, the mean standard timbre vector is -0.118, and the minimum timbre loss is 0.015, which is considered correct; the corresponding speech scenario is Changsha dialect daily conversation (neutral). In Sample 2, the mean feature vector of the input speech encoder is 0.089, the mean standard timbre vector is 0.092, and the minimum timbre loss is 0.018, which is considered correct; the corresponding speech scenario is Changsha dialect news broadcast (neutral). In Sample 3, the mean feature vector of the input speech encoder is -0.056, the mean standard timbre vector is -0.049, and the minimum timbre loss is 0.023, which is considered incorrect; the corresponding speech scenario is Changsha dialect short sentence reading (neutral, with background noise). In Sample 4, the mean feature vector of the input speech encoder is 0.112, the mean standard timbre vector is 0.108, and the minimum timbre loss is 0.012, which is considered correct; the corresponding speech scenario is Changsha dialect daily greeting (neutral). In Sample 5, the mean value of the input speech encoder feature vector is -0.091, the mean value of the standard timbre vector is -0.087, and the calculated minimum timbre loss value is 0.019, which is determined to be correct; the corresponding speech scenario is a description of daily life in Changsha dialect (neutral).

[0033] Table 1. Minimum timbre loss values ​​during the timbre cloning stage.

[0034]

[0035] The encoder extracts the Mel spectrum from the target Changsha dialect corpus frame by frame. It captures the unique fundamental frequency, formants and pronunciation habits as timbre features through the ECAPA-TDNN model. The timbre features are input into the Tacotron2 synthesizer. A new Changsha dialect prosody adaptation layer is added to the Tacotron2 synthesizer, which incorporates the Changsha dialect tone rules, the first tone value and the second tone value, to generate the embedding vector and Mel spectrum. Then, it is converted into time-domain speech by the WaveRNN vocoder to obtain pure timbre cloned speech.

[0036] In this implementation plan, by constructing a multi-source speech dataset and performing standardized processing, and by leveraging a two-stage collaborative training mechanism and precise timbre loss constraints, the accuracy of Changsha dialect timbre cloning of the target speaker is effectively improved, and the influence of interference factors is reduced. This provides a clear optimization direction for progressive speech synthesis from timbre to emotion, and improves the overall quality and reliability of Changsha dialect speech cloning.

[0037] Specifically, based on the preprocessed multi-source speech dataset, the process of constructing a directional construction mechanism to output speakerprint features is as follows: On the basis of acoustic preprocessing and normalization of the multi-source speech dataset, a directional construction mechanism is constructed: The pre-training basic speech corpus is screened according to target focus and dialect adaptation. Target focus refers to selecting Changsha dialect speech under neutral emotion from the target speaker, eliminating samples with emotional fluctuations and non-standard pronunciation; dialect adaptation involves retaining characteristic Changsha dialect pronunciation samples, supplementing tone practice samples, and adapting to the tonal characteristics of the dialect; The screened speech is then subjected to noise reduction, pre-emphasis, framing, and windowing, and converted into Mel spectrogram features to obtain the original speech dataset as training data. A hybrid network structure is used to extract feature vectors from the original speech dataset through hierarchical feature extraction. CNN and LSTM are used for collaborative modeling. The CNN layer adopts a convolution and pooling structure. Conv1 extracts the fundamental spectral features of pitch harmonics and formants, and Pool1 compresses redundant information. Conv2 captures the spectral features of Changsha dialect, and Pool2 further compresses the frequency dimension. Conv3 deeply extracts local spectral features and outputs dimensional features to implement local spectral feature capture. The LSTM layer adopts a bidirectional LSTM. The dimensional features of the CNN output are compressed by convolution to adapt to the LSTM input format and perform bidirectional modeling. The forward LSTM learns the temporal correlation of speech frames, and the backward LSTM learns the inverse dependency. The bidirectional outputs are concatenated, and the temporal features are aggregated into a fixed-dimensional vector by global average pooling to eliminate the influence of speech length differences and perform temporal characteristic modeling.

[0038] A dual-loss constraint training optimization mechanism is established for the extracted feature vectors: minimizing the timbre loss value is used as the main loss to avoid timbre drift and implement the main loss constraint; anchor sample triples are constructed, and by minimizing the triple loss value, the feature difference between the target and non-target speakers is expanded, thereby improving the voiceprint discrimination; the triple consists of two speech samples from the same target speaker and one sample from a non-target speaker, and a margin-based triple loss is used to bring similar samples closer together and widen the embedding distance of dissimilar samples to implement auxiliary loss constraints; the AdamW optimizer is used, and the learning rate is decayed according to the cosine annealing strategy, and an early stopping mechanism is enabled to balance timbre fidelity and identity discrimination objectives.

[0039] All parameters during training are dynamically monitored, and the minimum timbre loss value, minimum triplet loss value, total loss value, and cosine similarity of sample embedding vectors for each batch are recorded in real time. GPU memory usage, temperature, and alarm triggering are also monitored. Data is stored by training task ID and epoch number to ensure traceability. Stability, discriminability, and dialect adaptability are verified. An untrained raw speech dataset is selected, and the average cosine similarity between the embedding vector and the standard timbre vector is calculated to verify timbre stability. Target and non-target inputs are fed into the encoder, and cosine similarity classification and accuracy are used to verify identity discrimination. The average similarity between the target speech and the standard timbre vector is calculated to verify the encoder's adaptability to Changsha dialect features. Finally, voiceprint features are output.

[0040] In this implementation scheme, by targeted screening and dialect adaptation processing of multi-source speech datasets, combined with hierarchical feature extraction of a hybrid CNN and bidirectional LSTM network, the distinctive voiceprint information and temporal correlation characteristics of Changsha dialect are accurately captured; the identity differentiation of voiceprint features is effectively improved; and high-quality and reliable core voiceprint feature support is provided for Changsha dialect speech cloning tasks.

[0041] Specifically, the specific process of constructing a dialect adaptation spectral frequency generation and WaveRNN vocoder conversion mechanism to output a mel spectrogram and synthesized speech is as follows: Input the embedding vector and the corresponding text in the target Changsha dialect base corpus. Map the embedding vector to a feature vector through a fully connected layer, align it with the output dimension of the Tacotron2 text encoder, and perform L2 normalization. Perform dialect adaptation processing on the corresponding text in the target Changsha dialect base corpus. First, perform word segmentation in Changsha dialect and convert it into international phoneme and dialect-specific phoneme annotations (for example, "咯" is annotated as "lo", "噻" is annotated as "sai"), and convert it into a text feature vector through the embedding layer. Add a feature fusion layer between the Tacotron2 text encoder and the attention layer, and fuse the timbre feature and the text feature in a way of element-wise addition and residual connection, and control the fusion ratio through a gating mechanism. Establish a dialect adaptation spectral frequency generation mechanism: Construct an improved text encoder at the encoding end, encode the text sequence into a high-dimensional semantic vector, and splice and fuse it with the embedding vector at the encoding end. Add a tone embedding layer to the traditional bidirectional LSTM text encoder, convert the tone labels of the high tone, rising tone, falling-rising tone, and falling tone in Changsha dialect into tone embedding vectors, and input them into the LSTM after splicing with the text feature vector, so that the encoder can capture the dialect tone differences and at the same time mask the emotion-related feature channels. In terms of the attention mechanism, introduce a position-based monotonic attention structure, superimpose the position convolution feature and the previous frame alignment probability on the original content attention score, and limit the moving step of the attention distribution through a Gaussian window constraint to avoid alignment jumps and backtracking, and ensure the sequential monotonic alignment from text to acoustic frames. In terms of the decoding mechanism, use a two-layer recurrent neural network with residual connection as the decoding core. In each decoding step, send the previous frame mel spectrogram and the fused conditional vector to the decoding end together, cooperate with the pre-network to suppress the exposure bias, add a stop marker prediction head at the output end of the decoding RNN, use binary cross-entropy loss to train and learn the sentence end feature, and determine the sentence end position through a dynamic threshold to reduce the trailing and truncation problems. In the attention layer of the improved text encoder, use position-sensitive attention and assign attention weights to the phoneme sequences of multi-syllable words in Changsha dialect. After generating the mel spectrogram, perform frame smoothing on the generated mel spectrogram to eliminate the unnatural rhythm caused by the jump between spectral frames. At the same time, perform energy normalization to adjust the spectral energy mean to be consistent with the real neutral spectrum, avoid volume fluctuations interfering with the timbre perception, and obtain the mel spectrogram. As Figure 3The shown Mel spectrogram has the time frame sequence on the horizontal axis and the Mel frequency channels on the vertical axis. The colors from dark to light represent the energy distribution from low to high. It can be seen from the figure that the energy bands in the low-frequency to mid-frequency section are continuous and have a smooth trend. The formant stripes are clear and distinct, and the spectral band fluctuations corresponding to different syllables have clear patterns, without obvious breaks, rough spots or large energy holes, indicating that the spectral structure synthesized by the Tacotron2 synthesizer and the WaveRNN vocoder is highly similar to that of real speech. At the same time, in the transition area between sentences, the energy changes harmoniously and the fundamental frequency trajectory is stable, reflecting good pronunciation coherence and prosodic naturalness, verifying the effectiveness in terms of timbre fidelity and delicate acoustic detail reconstruction. Input the Mel spectrogram to construct the WaveRNN vocoder conversion mechanism: pre-train the WaveRNN vocoder to learn the mapping relationship between the Mel spectrogram and the time-domain waveform, optimize the waveform restoration of Changsha dialect characteristic pronunciations (such as the stop sound of "呷" and the ending sound of "咯"), adapt to the target timbre, and use the Mel spectrogram and real waveform of the target speech for the loss function of the WaveRNN vocoder with waveform reconstruction loss. Fix the sampling rate, quantization bit number, number of hidden layer units and dropout rate to avoid overfitting; execute the spectrum and waveform conversion process, normalize the Mel spectrogram generated by the Tacotron2 synthesizer according to the mean and standard deviation during the pre-training of the WaveRNN vocoder. The WaveRNN vocoder uses an autoregressive sample-by-sample generation method. Starting from the spectrum, it generates waveform data for each sampling point in each step, predicts the next sampling point, generates the complete time-domain waveform, and performs DC offset removal, amplitude normalization and endpoint detection on the generated complete time-domain waveform to output the synthesized speech. Establish a mechanism for monitoring and verifying the synthesis quality of neutral timbre, record the Mel spectrogram loss of the Tacotron2 synthesizer and the waveform loss of the WaveRNN vocoder in real time, and store the monitoring data according to the synthesis task ID and training rounds; use an emotion classification model to discriminate the emotion category of the synthesized speech, with no obvious emotion bias. The emotion classification model takes the Mel spectrogram generated by the Tacotron2 synthesizer as input and minimizes the classification loss through the combined structure of a convolutional neural network and a bidirectional recurrent neural network to obtain an emotion classifier that can distinguish neutral, happy, angry and sad emotional states. Input the synthesized speech into the emotion classification model to determine that there is no obvious bias in the emotion dimension of the synthesized speech and evaluate the prosodic fluency of the synthesized speech; analyze the feature differences and dynamic change laws of the short-time energy and fundamental frequency curves of the synthesized speech and real speech, and compare the inter-frame correlation of the synthesized Mel spectrogram and the real spectrogram through the inter-frame correlation calculation method to output the neutral Mel spectrogram atlas.

[0042] This implementation scheme achieves high-fidelity synthesis of neutral timbre for Changsha dialect, resulting in a smooth and continuous Mel spectrum with details highly consistent with real speech. A neutral timbre synthesis quality monitoring mechanism is constructed to ensure that the synthesized speech possesses both timbre stability and prosodic naturalness without emotional bias, providing a reliable neutral baseline for subsequent emotion enhancement and emotion cloning.

[0043] Specifically, the process of jointly training a neutral model with emotion embedding and outputting time-stamped Mel spectrograms, and establishing a dialect-adaptive attention and anomaly handling mechanism through neural network adaptive learning of feature associations, is as follows: Emotion embedding is added to the neutral model for joint training to build the speech synthesis backbone network. Text content corresponding to the speech is extracted from the target Changsha dialect corpus and the emotion-annotated Changsha dialect corpus, preprocessed, and input into the text encoder. A high-dimensional text encoding vector aligned with the speech temporal sequence is output as a text feature for feature fusion and accurate modeling of speech-related tasks. Neutral Changsha dialect speech of the speaker is extracted from the target Changsha dialect corpus and historical neutral speech datasets, processed by the speaker encoder, and used as speaker embedding features to characterize the individualized timbre and identity attributes of the target speaker. Changsha dialect speech with labeled categories is extracted from the emotion-annotated Changsha dialect corpus. Fundamental frequency, speech rate, and intensity prosodic features are extracted using the Praat tool, and after dimensionality reduction and Z-score standardization via a lightweight fully connected network, emotion embedding features are obtained, used to adjust the emotion of the synthesized speech while maintaining timbre. The input consists of text features, speaker embedding features, and sentiment embedding features. Feature interfaces are standardized by mapping speaker embedding features and sentiment embedding features to the same dimension as the text encoding vector through fully connected layers, and mean and variance normalization is applied. A multi-source feature fusion layer is added between the text encoder and the attention layer. The standardized text encoding vector, speaker embedding vector, and sentiment embedding vector are concatenated along the feature dimension, and then linearly transformed and nonlinearly activated to form a fusion condition vector. This fusion condition vector is added to the text encoding output through residual connections and gating mechanisms. After feature dimension matching and temporal alignment, an enhanced encoding sequence is obtained. The fusion condition vector is used as the initial state of the attention RNN and the conditional input of the decoding RNN. At each decoding time step, the Mel spectrum of the previous frame and the fusion condition vector are fed into the decoder. Text encoding input, speaker feature input, and sentiment feature input are added to the multi-source feature fusion layer, and the input and output are bound. CTC forced alignment is used to obtain frame-level timestamps, and an 80-dimensional Mel spectrum is output. A forced aligner method is used to attach timestamp labels to each frame spectrum.

[0044] Feature association is performed through adaptive learning via neural networks. The adaptive compression weight matrix is ​​multiplied by the element-wise product of the text encoding vector, speaker embedding vector, emotion feature activation coefficients, and emotion embedding vector to obtain a high-dimensional feature compression term. This high-dimensional feature compression term is then added to a bias term to obtain a bias-compensated compressed feature term. A linear rectified activation function is applied to the bias-compensated compressed feature term to obtain a non-linear activation feature term. Finally, this non-linear activation feature term is added to the text encoding vector to obtain a fused feature vector. The specific calculation formula for the fused feature vector is as follows: ;

[0045] In the formula, The fusion feature vector is used to integrate multi-source features of text semantics, speaker timbre and emotional expression to form a unified feature representation that is adapted to subsequent spectrum decoding, providing comprehensive feature support for high-quality Mel spectrum generation. This represents the linear rectified activation function, which introduces a nonlinear transformation by setting the negative eigenvalues ​​to zero, thereby enhancing the nonlinear expressive power of the features and suppressing the interference of noisy features on effective information. This represents an adaptive compressed weight matrix. Through model training, it autonomously learns the correlation weights between multi-source features, realizing a non-linear mapping from high-dimensional spliced ​​features to the target dimension, ensuring that the feature dimension matches the input requirements of the subsequent decoding layer. This indicates a feature dimension concatenation operation, which concatenates the text encoding vector, speaker embedding vector, and sentiment embedding vector (after activation coefficient adjustment) in dimensional order to form a high-dimensional joint feature, fully preserving the original information of each source feature; The text encoding vector contains semantic information, phoneme features, and tone features of Changsha dialect text. It is generated by bidirectional LSTM encoding and is the core feature that ensures the semantic accuracy of synthesized speech. The speaker embedding vector contains the unique Changsha dialect timbre features of the target speaker. It is extracted and generated by the speaker encoder and is a key feature to ensure the consistency of synthesized speech timbre. This represents the activation coefficient of emotional features. It takes a value of 0 in the timbre cloning stage (suppressing the participation of emotional features) and a value of identity matrix in the emotional cloning stage (activating the participation of emotional features). It only controls whether emotional features are included in the fusion by switching on and off, and does not involve feature weight allocation. This represents element-wise multiplication, used to perform element-wise operations on activation coefficients and sentiment embedding vectors, precisely controlling the participation state of sentiment features; The emotion embedding vector contains prosodic features (fundamental frequency, speech rate, and intonation) of emotions such as joy, anger, sorrow, and happiness. It is generated from emotion-annotated corpus during the emotion cloning stage and is the core feature for achieving emotion transfer. Denoted as the bias term, it is obtained through synchronous iterative learning based on the multi-source training corpus of Changsha dialect of the target speaker (including neutral corpus and sentiment-annotated corpus) and the weight matrix It is used to adjust the overall distribution of the compressed features, so that the feature mean adapts to the input distribution requirements of the subsequent decoding layer.

[0046] In the Tacotron2 synthesizer, a three-level architecture of a position-sensitive attention layer, a dual LSTM decoder, and a frame prediction smoothing layer is adopted. The position-sensitive attention layer outputs a context vector, the dual LSTM decoder optimizes the inter-frame correlation, and the frame smoothing layer eliminates jumps through weighted averaging of adjacent frames. A dialect-adaptive attention mechanism is established: new long-sentence logical pause points and attention enhancement factors are added, a lightweight network structure is designed, a decreasing hidden layer is adopted, the convolution kernel parameters of the spectral feature mapping layer and the waveform prediction layer are shared, the Swish activation function is used in the hidden layer, and the tanh activation function is used in the output layer to constrain the waveform amplitude. The double-stage sampling of coarse sampling and fine sampling is innovated, the context of adjacent frame sampling points is cached, batch inference is supported, the speech batch generation time is reduced, and the batch synthesis requirements are met. The restoration of Changsha dialect pronunciation details is optimized: the spectra and waveform mappings of the characteristic pronunciations of the plosive "呷" and the nasalized rhyme "湘" are mainly learned, and the L1 waveform loss and perceptual loss are used for optimization, and harmonic enhancement is added to the waveform post-processing. For the joint loss function, the AdamW optimizer is adopted, the learning rate is decayed every fixed batch according to the cosine annealing strategy, and the early stopping mechanism is enabled. Monitor the quality of the backbone network. The quality of the backbone network includes: spectral quality, waveform quality, perceptual quality, and inference efficiency; spectral quality refers to the inter-frame correlation reaching a high correlation threshold, the MFCC feature distance being lower than a low difference threshold, and the spectral loss value being lower than a low loss threshold; waveform quality refers to the signal-to-noise ratio reaching a high fidelity threshold, the signal distortion ratio reaching a low distortion threshold, and the waveform loss value being lower than an extremely low loss threshold; perceptual quality refers to the naturalness score reaching a good threshold, the dialect adaptability score reaching an excellent threshold, and the recognition accuracy rate reaching a high accuracy threshold; inference efficiency refers to the single-standard-duration speech generation time being lower than the fast generation threshold and the batch generation real-time rate reaching the high-efficiency inference threshold. An exception handling mechanism is established. The exception handling mechanism refers to checking the accuracy of phoneme mapping, adjusting the attention enhancement factor, supplementing a sufficient amount of characteristic pronunciation corpus, retuning the WaveRNN vocoder, increasing the coarse sampling step size to the maximum adaptation step size, or reducing the number of hidden layer units, optimizing the parameters of the tone embedding layer, and increasing the proportion of tone samples.

[0047] In this implementation plan, by introducing emotion embedding for joint training on the basis of the neutral model, and constructing a dialect-adaptive attention mechanism, an adaptive feature fusion formula, and an exception handling mechanism, it not only ensures the objective quality of the spectrum and waveform, but also takes into account the perceptual naturalness and inference efficiency, realizing high-fidelity and controllable emotions in multi-emotion and dialect scenarios.

[0048] Specifically, based on the sentiment-annotated Changsha dialect corpus, a multi-level loss function constraint balance was designed, and the process of using multiplicative adaptive adjustment of sentiment loss weights is as follows: A structured sentiment feature representation system is constructed by inputting the sentiment-annotated Changsha dialect corpus. The fundamental frequency dynamic sequence of each corpus is extracted using a double-annotation mechanism with the Praat tool. Outliers are manually corrected to form a fundamental frequency feature annotation set. The fundamental frequency dynamic sequence includes: mean, maximum, minimum, standard deviation, and slope of change. At the bottom layer, prosodic element vectors of fundamental frequency, speech rate, intonation, and intensity are extracted. Fundamental frequency features refer to the extraction of statistical and dynamic features of the fundamental frequency dynamic sequence to form a fundamental frequency feature vector, focusing on characterizing the fluctuation pattern of fundamental frequency under different emotions. Speech rate features refer to extracting average speech rate, speech rate change rate, and speech rate values ​​of key semantic segments on a sentence-by-sentence basis to form a speech rate feature vector, capturing the regulation pattern of speech rate by emotion. Intonation features refer to converting the intonation change curve into polynomial coefficients through an intonation contour fitting algorithm, combining the position and amplitude of intonation inflection points to form an intonation feature vector, quantifying the intonation change patterns of different emotions. Intensity features refer to the extraction of statistical and dynamic features of short-term energy to form intensity feature vectors, distinguishing the influence of emotion on speech intensity. After concatenation by a high-level neural network, a lightweight fully connected network is used for nonlinear dimensionality reduction to output emotion feature vectors. Z-Score normalization is applied to normalize the emotion feature vectors. Cluster analysis is performed on the feature vectors of each emotion category to generate standard emotion feature vectors for joy, anger, sorrow, happiness, and neutrality. A multi-level loss function system is designed: For each emotion element, mean squared error, mean absolute error, mean squared error, and DTW distance are used to constrain the mean of the squared differences between the predicted and actual mean frequency feature vectors, constraining the statistical characteristics and dynamic changes of the mean frequency to match the true emotion. For speech rate loss, mean absolute error is used to calculate the mean of the absolute differences between the predicted and actual speech rate feature vectors, ensuring that the speech rate adjustment patterns under different emotions are learned. The intonation loss is a weighted sum of the mean squared error between the intonation feature vector and the true intonation feature vector, and the DTW distance. This constrains both the static features of intonation and matches the temporal patterns of intonation changes. The intensity loss uses the mean squared error to calculate the average of the squared differences between corresponding elements of the intensity feature vector and the true intensity feature vector, ensuring that the intensity of emotional expression matches the real-world scenario. This balances the overall emotional loss.

[0049] A multiplicative adaptive adjustment strategy is employed to dynamically update the emotional loss weights. After each training cycle, it no longer relies solely on fixed weights or linear adjustments based only on the current metric. Instead, a smooth and stable weight update mechanism is constructed using a nonlinear saturation function and exponential mapping, thereby enhancing emotional expression while avoiding timbre drift and training instability. It can adaptively balance the relationship between emotional consistency and timbre similarity during training, allowing for fine-tuning of the focus at different training stages. The emotional recognition rate of the (k-1)th training cycle is added to a minimum positive constant as the emotional term. The emotional accuracy threshold is added to a minimum positive constant to obtain the emotional target term. The emotional target term is divided by the emotional term to obtain the emotional ratio. The emotional ratio minus a fixed value yields the initial emotional term. The initial emotional term is then multiplied by the emotional adjustment coefficient after hyperbolic tangent calculation to obtain the emotional bias term. Similarly, the timbre similarity score of the (k-1)th training cycle is added to a minimum positive constant as the timbre term. The timbre threshold is added to a minimum positive constant to obtain the timbre target term. The timbre term is divided by the timbre target term to obtain the timbre ratio. The timbre ratio minus a fixed value yields the initial timbre term. The initial timbre term is calculated using the hyperbolic tangent and then multiplied by a timbre adjustment coefficient to obtain a timbre protection term. The (k-1)th emotion term is subtracted from the (k-2)th emotion term to obtain the emotion difference, which is then divided by the emotion target term to obtain the trend term. The emotion deviation term is subtracted from the timbre protection and trend terms to obtain an exponential term. An exponential operation is performed on the exponential term to obtain an adaptive factor. The emotion loss weight of the (k-1)th training period is multiplied by the adaptive factor to obtain the emotion loss weight of the kth training period. An exponential form is used for multiplicative updates, ensuring the weights are non-negative and scale-invariant to adjustments at different training stages. These three pieces of information are used to adaptively adjust the weights, dynamically adjusting the emotion loss weights during training. This allows for a greater penalty when emotion expression is insufficient and a smaller penalty when timbre is damaged, while avoiding training jitter. The emotion deviation term adjusts how far the current emotion recognition rate is from the target (deviation magnitude); the timbre protection term adjusts how far the current timbre similarity is from the target (whether to protect the timbre); and the trend term adjusts whether the emotion accuracy is improving or deteriorating (trend).

[0050] An adaptive weight adjustment strategy is employed to dynamically update the emotional loss weights. After each training cycle, the proportion of emotional loss in the overall loss function is adaptively adjusted based on the model's comprehensive performance in both emotional expression and timbre preservation. Instead of relying on fixed weights or simple linear updates, an exponential update mechanism based on performance bias is introduced. This ensures that the weight update process maintains continuity and stability while responding responsively to optimization needs at different training stages. By jointly modeling the deviations of emotional recognition accuracy and timbre similarity relative to the target value, a dynamic balance between emotional enhancement and timbre protection is achieved. This improves emotional expression while avoiding timbre drift and training instability. The weight update formula is as follows:

[0051]

[0052] ,

[0053] in, This represents the sentiment loss weight corresponding to the kth training period, which is used to control the relative contribution of sentiment loss in the overall loss function at the current stage. The emotional loss weight from the previous training cycle is used as the base value for weight updates, ensuring the continuity and stability of weight changes. This represents the emotion recognition accuracy of the model during the (k-1)th training period, used to measure the consistency between the generated speech and the target emotion label in terms of emotion expression; The target threshold representing the accuracy of emotion recognition is used to indicate the expected level of emotion expression intensity for the current task; This represents the emotion expression bias term, which reflects the gap between the current emotion expression and the target emotion requirement. When the emotion recognition accuracy is lower than the target threshold, this bias term is positive, thereby increasing the emotion loss weight to strengthen the learning of emotion features. This represents the timbre similarity score of the model during the (k-1)th training period, which is used to measure the degree of similarity between the generated speech and the target speaker in terms of timbre features; The target threshold representing timbre similarity is used to indicate the minimum requirement for timbre preservation. This represents the timbre protection bias term. When the current timbre similarity is lower than the target threshold, this bias term is positive and is used to suppress the further increase of the emotional loss weight, so as to avoid overemphasizing emotion and destroying timbre consistency. This represents the affective sensitivity moderating coefficient, used to control the strength of the influence of affective bias on the magnitude of weight updates; This represents the timbre protection adjustment coefficient, used to control the suppression effect of timbre deviation in weight updates. Through the above adaptive update mechanism, the emotional loss weight is dynamically adjusted according to the emotional expression effect and timbre preservation state during training, so that the model can adaptively balance emotional consistency and timbre similarity at different training stages, thereby improving the overall training stability and the naturalness of the generated speech.

[0054] As shown in Table 2, in training period 1, the previous period's emotional loss weight was 0.3, the current period's emotional recognition rate was 0.65, the current period's timbre similarity was 0.92, and the updated emotional loss weight was 0.38, with the weight adjustment state being emotional enhancement. In training period 2, the previous period's emotional loss weight was 0.38, the current period's emotional recognition rate increased to 0.72, the current period's timbre similarity decreased to 0.89, and the updated emotional loss weight was 0.36. In training period 3, the previous period's emotional loss weight was 0.36, the current period's emotional recognition rate further increased to 0.78, the current period's timbre similarity rebounded to 0.91, and the updated emotional loss weight was 0.41, with the weight adjustment state being emotional enhancement. In training period 4, the previous period's emotional loss weight was 0.41, the current period's emotional recognition rate reached 0.81, the current period's timbre similarity decreased to 0.88, and the updated emotional loss weight was 0.39, with the weight adjustment state being fine-tuning suppression. In training period 5, the emotional loss weight of the previous period was 0.39, the emotional recognition rate of the current period was stable at 0.82, the timbre similarity of the current period rebounded to 0.9, the updated emotional loss weight was 0.4, and the weight adjustment state was stable and balanced.

[0055] Table 2. Updated Data on Emotional Loss Weights During the Emotional Cloning Stage

[0056]

[0057] like Figure 4 The graph showing the trend of emotional loss weight changes during the emotional cloning stage is illustrated. The horizontal axis represents the weight adjustment state number, and the vertical axis represents the weight value of the corresponding loss term. It can be seen that the changes in each weight during training exhibit the following characteristics: the main weight remains generally stable, emotionally related weights undergo small adaptive adjustments, and secondary weights remain close to zero for a long period. Weights 2 and 4 consistently maintain a high level and exhibit only minor fluctuations under different adjustment states, indicating that the core losses corresponding to these two items (timbre consistency and spectral reconstruction accuracy) are maintained as relatively stable baseline constraints during the emotional cloning stage. Weight 3 gradually increases from a moderate value with each adjustment state and then stabilizes, reflecting the system's gradual enhancement of the contribution of a certain emotionally related loss term during the emotional cloning process. Weights 1 and 7 fluctuate slightly within the range of 0.35-0.40, with the corresponding emotional sub-losses undergoing subtle increases and decreases based on model performance at different stages, used for fine-tuning emotional expression details. Weight 5 remains close to 0, indicating that the corresponding auxiliary loss basically does not participate in optimization during the emotional cloning stage and does not interfere with the convergence of the main objective. This reflects that the weight adaptive mechanism in the emotion cloning stage, while ensuring the stability of the weights of key indicators such as timbre and spectrum, makes gradual and small-step dynamic adjustments to the emotion-related loss terms, so that the training process gradually converges to a balanced state that takes into account both timbre consistency and the naturalness of emotional expression under multiple rounds of weight adjustment.

[0058] In this implementation scheme, by combining multi-level prosodic feature constraints with a multiplicative adaptive emotional loss weight update mechanism, a dynamic balance can be automatically found between emotional expression and timbre preservation during the training process. This allows the Changsha dialect emotional clone speech to achieve more natural and delicate emotional expression while maintaining timbre consistency, and makes the training process convergence more stable and controllable.

[0059] Specifically, the process of evaluating synthesized speech and verifying the effectiveness of weight adjustment is as follows: Based on the accuracy of capturing emotional features using emotional loss weights, a quantitative evaluation of the synthesized speech is conducted: the macro-average F1 score of emotion recognition is used as the F1 index for emotion recognition, to accurately quantify the model's ability to capture emotional features; the relative errors of fundamental frequency, energy, and duration prosodic elements between the synthesized speech and the target speech are compared, and the mean square error of fundamental frequency, relative deviation of energy, and syllable duration deviation rate are statistically analyzed as speaker similarity, used to learn the variation patterns of each type of emotional prosodic element; the MOSNet model is used to score the synthesized speech to obtain a naturalness score; and the emotional content is then combined with the score. Feature accuracy requires that the F1 score for emotion recognition be no lower than the emotion threshold, the average recognition accuracy for each emotion type reach a high accuracy threshold or above, and there be no cross-emotional confusion. Emotional expressiveness scoring combines the F1 score for emotion recognition with the naturalness score, using a point-based scoring system. It requires that the average score for each emotion type reach a good threshold or above, and that there be no negative feedback from human evaluation regarding emotional stiffness or unnatural rhythm. At the same time, the thresholds for achieving the target are that the F1 score for emotion recognition be no lower than the emotion threshold, the speaker similarity score be no lower than the similarity threshold, and the naturalness score estimated by MOSNet be no lower than the scoring threshold. Timbre consistency is verified by the cosine similarity of the embedding vector. The design employs a dual-track adjustment strategy of progressive and real-time feedback, dividing the training phase and binding the weight adjustment of each phase to the target. The timbre anchoring phase monitors the speaker similarity as greater than or equal to the high consistency threshold, consolidating the target timbre features. The emotional integration phase monitors the emotion recognition F1 score as not lower than the emotion threshold, the naturalness score as not lower than the score threshold, and the speaker similarity as maintained above the high consistency threshold, improving emotional expression ability. The fine balance phase monitors the emotion recognition F1 score as steadily improving, and the speaker similarity and naturalness scores as not decreasing, optimizing the subtlety of emotional expression. To address issues such as timbre drift, insufficient emotional expression, and indicator stagnation: Timbre drift occurs when speaker similarity falls below the high consistency threshold. In this case, the emotional embedding layer parameters are frozen, and only the spectrum generation of the Tacotron2 synthesizer is updated until speaker similarity rises above the high consistency threshold for several consecutive training cycles, at which point the phase progression is restarted. Insufficient emotional expression occurs when the emotional recognition F1 score is below the emotional threshold or the naturalness score is below the scoring threshold. If speaker similarity is greater than or equal to the high consistency threshold, the emotional feature learning rate is increased to the emotional learning rate enhancement value, while the emotional loss weight is increased to strengthen the fitting strength of the emotional prosody. Indicator stagnation optimization refers to reducing the emotional loss weight when the improvement of the comprehensive indicator between two adjacent stages falls below the preset improvement threshold for several consecutive times, thus avoiding overfitting and ineffective training.The test utilizes quantitative indicators such as timbre consistency and emotional feature accuracy, along with scenario-based testing. Scenario-based testing includes: cross-emotional testing within the same text and extreme emotion testing. Cross-emotional testing generates multiple emotional voice types from the same Changsha dialect text, requiring listeners familiar with the target speaker to confidently identify the same voice, and listeners without text to accurately distinguish between different emotion types. Extreme emotion testing targets highly emotional voice, requiring timbre features to remain within the normal range of the target speaker, and that the emotion recognition F1 score and naturalness score meet the threshold requirements. The effectiveness of weight adjustment is verified, and a dynamic weight strategy file is output. Figure 5 The heatmap of the emotion feature similarity matrix shown has emotion categories on both the horizontal and vertical axes, including "joy, anger, sorrow, happiness, surprise, fear, and contemplation." The color intensity and the numerical values ​​in the squares together represent the similarity between feature vectors of different emotion categories, ranging from approximately 0.75 to 1.00. The similarity on the diagonal is 1.00, corresponding to features within the same emotion category, indicating that the emotion vectors extracted from similar samples are highly aggregated and have good internal consistency. The similarity between different emotions outside the diagonal is generally moderate to high. The similarity of the combinations "joy-happiness," "joy-surprise," "happiness-surprise," and "sorrow-contemplation" is around 0.90, reflecting that these emotions have certain commonalities in actual speech performance. For example, joy and happiness are both characterized by a higher fundamental frequency and faster speech rate, while sorrow and contemplation are characterized by a relatively gentle rhythm and energy convergence. The similarity of the combinations "sorrow-fear," "surprise-contemplation," and "fear-contemplation" is relatively low, between approximately 0.77 and 0.80, indicating that the feature space effectively separates emotional states with large differences. The heatmap shows that the emotional feature space constructed by this invention can not only ensure the tight clustering of features within the same emotional category, but also form a clear and separable similarity distribution between different emotional categories.

[0060] In this implementation plan, by combining quantitative indicators with tests on cross-emotional and extreme emotional scenarios within the same text, an observable and verifiable closed-loop basis is provided for adjusting the weight of emotional loss; a dynamic weight strategy file that can be directly reused is formed, which significantly improves the engineering controllability and application reliability of Changsha dialect emotional cloned speech.

[0061] Specifically, the process of constructing a dual-track input system, using deep restricted Boltzmann machines and time-delay neural networks for feature fusion and temporal modeling, and jointly driving the synthesis process to output emotional speech clone data, is as follows: Figure 6As shown, a dual-track input system of identity features and emotional features is constructed. In the original speech dataset, voiceprint features are extracted through static and dynamic difference features of Mel-frequency cepstral coefficients, and fundamental frequency range, formant distribution, and spectral tilt are extracted as voice characteristic features through linear predictive coding. Principal component analysis is used for dimensionality reduction to remove redundant information and extract identity features. Fundamental frequency temporal features are extracted from the fundamental frequency feature annotation set using the sliding window method. Emotional prosodic features are extracted by combining emotional loss weights. Emotional prosodic features include: prosodic dynamic features, emotional state features, and feature processing. Prosodic dynamic features are the temporal sequence of fundamental frequency, the temporal change of speech rate, and the temporal sequence of intensity. Emotional state features refer to the emotional category output by the emotional classification model, reflecting the emotional category and intensity of the current speech segment. Feature processing refers to the framing of temporal features, and the concatenation of features from each frame to form a temporally ordered emotional feature sequence to ensure that it can reflect the dynamic evolution of emotion. A depth-restricted Boltzmann machine is constructed to perform nonlinear fusion by minimizing reconstruction error. The structure consists of an input layer, hidden layer 1, hidden layer 2, and an output layer. Each layer performs feature transformation through binary units. The input layer receives the concatenated identity feature vector and single-frame emotion feature vector as the raw input for fusion. Hidden layer 1 uses the ReLU activation function to perform preliminary nonlinear transformation on the input emotion prosodic features, uncovering low-order correlations. Hidden layer 2 uses the sigmoid activation function to further aggregate high-order features, capturing the implicit correlation between identity and emotion (e.g., "the fundamental frequency of a speaker will exceed the normal range by 1.2 times when angry"). The output layer outputs an intermediate feature vector, which only reflects... This method integrates identity and emotion information from a single frame, employing a temporal delay neural network to perform convolutional temporal modeling of emotional prosodic features. This captures the local dynamic correlations of emotional prosodic features across frames. A temporal feature extraction architecture using convolutional layers, pooling layers, and fully connected layers is employed, with convolutional kernels of different receptive fields designed to cover multi-scale temporal information. The input layer receives the intermediate feature sequence output from a depth-limited Boltzmann machine. Convolutional layer 1 uses a convolutional kernel to capture short-term dynamics in adjacent frames (e.g., instantaneous changes in fundamental frequency). Convolutional layer 2 uses a convolutional kernel to capture mid-range dynamics within the range (e.g., phased changes in speech rate). The pooling layer uses max pooling to compress the temporal dimension while retaining key dynamic features. The fully connected layer outputs temporal aggregated features. The temporal aggregated features are then subjected to mean pooling and dimensionality pruning to obtain an e-vector. This e-vector is used as the conditional feature input to the speaker feature input of the Tacotron2 synthesizer in the speech synthesis backbone network. The e-vector, text features, and speaker features jointly drive the synthesis process, outputting emotional speech clone data.

[0062] In this implementation scheme, by using dual-track input of identity features and emotional prosodic features, nonlinear fusion of deep-restricted Boltzmann machines, and cross-frame temporal modeling of time-delay neural networks, the temporal variation patterns of fundamental frequency, speech rate, and intensity under different emotions are reproduced, significantly improving the accuracy of emotional expression and the naturalness of the listening experience of emotional speech clone data.

[0063] Specifically, the second aspect of this invention provides a speech cloning system integrating an emotion enhancement mechanism, applied to a speech cloning method integrating an emotion enhancement mechanism, comprising: a transfer learning overall framework and a two-stage training module, used to construct a transfer learning framework and a multi-source speech dataset, integrate publicly available speech datasets, collect basic Changsha dialect corpus and perform acoustic preprocessing and normalization on the multi-source speech dataset, perform structured temporary storage through a dedicated speech data buffer, and obtain pure timbre cloned speech through a two-stage collaborative training mechanism; a timbre cloning and speaker encoder training module, used to construct a directional construction mechanism based on the preprocessed multi-source speech dataset, extract feature vectors from the original speech dataset using a hybrid network structure hierarchical feature extraction, establish a dual-loss constraint training optimization mechanism to output voiceprint features, construct a dialect-adapted spectral frequency generation and WaveRNN vocoder conversion mechanism to output Mel spectrograms and synthesized speech, analyze the short-time energy and fundamental frequency curves of synthesized speech and real speech, compare the inter-frame correlation between the synthesized Mel spectrogram and the real spectrum, and output a neutral Mel spectrogram. The system comprises four modules: a speech synthesizer and vocoder modeling module, used to jointly train a neutral model with emotion embedding, build a speech synthesis backbone network, output timestamped Mel spectra, obtain fused feature vectors through neural network adaptive learning of feature associations, and establish dialect-adaptive attention and anomaly handling mechanisms; an emotion cloning and emotion feature learning module, used to construct a structured emotion feature representation system based on an emotion-annotated Changsha dialect corpus, design a multi-level loss function constraint balance, adopt a multiplicative adaptive adjustment method to dynamically update the emotion loss weights, evaluate synthesized speech, and verify the effectiveness of weight adjustment; and a multi-feature fusion and temporal modeling module, used to construct a dual-track input system of identity features and emotion features, perform feature fusion and temporal modeling through deep restricted Boltzmann machines and time-delay neural networks, output temporal aggregated features, and perform mean pooling and dimension pruning to obtain an e-vector. This e-vector is used as a conditional feature input to the speech synthesis backbone network, and the e-vector, text features, and voiceprint features jointly drive the synthesis process to output emotional speech clone data. Figure 7The diagram illustrates a framework for a speech cloning method incorporating an emotion enhancement mechanism. It comprises three parts: a speaker encoder, a Tacotron2 synthesizer, and a WaveRNN vocoder. The speech is converted into a Mel spectrogram and input into the speaker encoder along with an emotion label. Through a three-layer long short-term memory network and a fully connected layer, a speaker emotion feature vector integrating timbre and emotional information is extracted during dual-stage training (timbre training and emotion training). The input text is processed by the encoder, attention mechanism, and decoder in the Tacotron2 synthesizer and fused with the speaker emotion feature vector to generate a target Mel spectrogram with a time series. The WaveRNN vocoder synthesizes the corresponding waveform speech based on this Mel spectrogram, thus outputting a speech cloning result that combines the target speaker's timbre and the target emotional expression.

[0064] In this implementation plan, through the coordinated efforts of modules such as transfer learning and two-stage training for timbre cloning, emotion cloning and adaptive weight adjustment mechanism, multi-feature fusion and temporal modeling, the prosodic changes and expressive details under multiple emotional states are accurately depicted and synthesized, thereby improving the overall performance and engineering usability of speech cloning in terms of emotional naturalness, speaker similarity and cross-scene adaptability.

[0065] Specifically, a third aspect of the present invention provides a voice cloning storage medium that integrates an emotion enhancement mechanism, comprising: the storage medium having one or more programs, the one or more programs being executed by one or more processors to implement a voice cloning method that integrates an emotion enhancement mechanism.

[0066] In this implementation scheme, the speech cloning method that integrates emotion enhancement mechanisms is stored as a program on a storage medium, which can be directly loaded and executed on different hardware platforms. This enables rapid deployment and large-scale reuse of the algorithm, facilitating stable application in various speech synthesis and human-computer interaction scenarios. Based on the storage medium, remote upgrades and version management can be easily performed. Model parameters and functional modules can be iteratively optimized without replacing hardware, enabling the system to continuously adapt to new emotional corpora and application requirements, thereby improving overall maintainability and scalability.

[0067] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0068] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.

Claims

1. A voice cloning method incorporating emotion enhancement mechanisms, characterized in that, Includes the following steps: S1. Construct a transfer learning framework and a multi-source speech dataset. Perform acoustic preprocessing and normalization on the multi-source speech dataset. Obtain pure timbre cloned speech through a two-stage collaborative training mechanism. S2, based on the preprocessed multi-source speech dataset, constructs a directional construction mechanism to output speaker features, and constructs a dialect-adaptive spectral frequency generation and WaveRNN vocoder conversion mechanism to output Mel spectrogram and synthesized speech; S3 incorporates sentiment embeddings into the neutral model for joint training, outputting a time-stamped Mel spectrum. Through neural network adaptive learning of feature associations, a dialect-adaptive attention and anomaly handling mechanism is established. S4, based on the sentiment-annotated Changsha speech corpus, designs a multi-level loss function constraint balance, adopts multiplicative adaptive adjustment of sentiment loss weights, evaluates synthesized speech, and verifies the effectiveness of weight adjustment; The specific process of designing a multi-level loss function constraint balance based on the sentiment-annotated Changsha discourse corpus and adopting multiplicative adaptive adjustment of sentiment loss weights is as follows: A structured sentiment feature representation system is constructed using a sentiment-annotated Changsha dialect corpus. A fundamental frequency feature annotation set is formed using a double-annotation mechanism with the Praat tool. The prosodic element vectors of fundamental frequency, speech rate, intonation, and intensity are extracted at the bottom layer. After high-level concatenation, the vectors are reduced in dimensionality by a lightweight fully connected network, standardized using Z-scores, and then clustered to generate various sentiment standard feature vectors. A multiplicative adaptive adjustment strategy is used to dynamically update the sentiment loss weights. A smooth and stable weight update mechanism is constructed through a nonlinear saturation function and exponential mapping. The sentiment recognition rate in the (k-1)th training cycle is added to a minimum positive constant to obtain the sentiment term. The sentiment accuracy threshold is added to a minimum positive constant to obtain the sentiment target term. The sentiment target term is divided by the sentiment term to obtain the final result. The emotional ratio is calculated by subtracting a fixed value from the emotional ratio to obtain the initial emotional term. The initial emotional term is then multiplied by the emotional adjustment coefficient after applying a hyperbolic tangent to obtain the emotional deviation term. The timbre similarity score of the (k-1)th training period is first added to a minimum positive constant to obtain the timbre term. The timbre threshold is then added to a minimum positive constant to obtain the timbre target term. The timbre term is divided by the timbre target term to obtain the timbre ratio. Subtracting a fixed value from the timbre ratio yields the initial timbre term. Applying a hyperbolic tangent to the initial timbre term and multiplying by the timbre adjustment coefficient yields the timbre protection term. The emotional difference is obtained by subtracting the emotional term of the (k-2)th training period from the emotional term of the (k-1)th period. The emotional difference is then divided by the emotional target term to obtain the trend term. Finally, the timbre protection term and the trend term are subtracted from the emotional deviation term to obtain the exponential term. Perform an exponential operation on the exponential term to obtain the adaptive factor; multiply the sentiment loss weight of the (k-1)th training period by the adaptive factor to obtain the sentiment loss weight of the kth training period. S5 constructs a dual-track input system, using deep restricted Boltzmann machine and time-delay neural network modeling to perform feature fusion and temporal modeling, jointly driving the synthesis process to output emotional speech clone data; The specific process of constructing a dual-track input system, using deep restricted Boltzmann machines and time-delay neural networks for feature fusion and temporal modeling, and jointly driving the synthesis process to output emotional speech clone data, is as follows: A dual-track input system of identity and emotion features is constructed. In the original speech dataset, voiceprint features are extracted through Mel-Cepstral Coefficients and formant voice feature features are extracted through Linear Predictive Coding. Identity features are extracted through Principal Component Analysis, and fundamental frequency temporal features are extracted from the fundamental frequency feature annotation set using the sliding window method. Emotional prosodic features are extracted by combining emotion loss weights. Nonlinear fusion is performed by minimizing reconstruction error to construct a deep restricted Boltzmann machine and output an intermediate feature vector. The intermediate feature vector only reflects the fusion information of identity and emotion in a single frame. Convolutional temporal modeling of emotional prosodic features is performed through a time-delay neural network to capture the local dynamic correlation of emotional prosodic features across frames. The output temporal aggregated features are then subjected to mean pooling and dimensionality pruning to obtain an e-vector. The e-vector is used as a conditional feature input to the speech synthesis backbone network. The e-vector, text features, and voiceprint features jointly drive the synthesis process and output emotional speech clone data.

2. The speech cloning method incorporating emotion enhancement mechanisms according to claim 1, characterized in that: The specific process of constructing the transfer learning framework and multi-source speech dataset, performing acoustic preprocessing and normalization on the multi-source speech dataset, and obtaining pure timbre cloned speech through a two-stage collaborative training mechanism is as follows: A general framework for transfer learning is constructed, which breaks down speech cloning into a progressive stage of timbre cloning and emotion cloning, and builds a neutral model. Public speech datasets, private target speaker timbre and emotion-inducing data, and cross-scene and cross-device transfer adaptation data are acquired and integrated. Changsha dialect basic corpus is collected and a multi-source speech dataset is constructed, which includes: target Changsha dialect basic corpus, emotion-annotated Changsha dialect corpus, and pre-trained basic speech corpus. Changsha dialect speech of the target speaker in historical neutral emotion state is collected as historical neutral speech dataset. A target-oriented collection and collaborative data construction mechanism is adopted, and the multi-source speech dataset is temporarily stored in a structured manner through a dedicated speech data cache. Acoustic preprocessing and Z-score normalization of features were performed on the multi-source speech dataset. A two-stage collaborative training mechanism was established to progressively transfer timbre to emotion. The preprocessed target Changsha dialect corpus was input into the speaker encoder. After local acoustic feature extraction and temporal voiceprint modeling, timbre feature vectors were obtained. The baseline vector generated by the encoder after feature extraction, mean taking and normalization was used as the standard timbre vector. The timbre feature vector was subtracted from the standard timbre vector to obtain the difference vector. The L1 norm squared operation was performed on the difference vector to obtain the minimized timbre loss value. The encoder extracted Mel spectra from the target Changsha dialect corpus frame by frame. The ECAPA-TDNN model was used to capture the unique fundamental frequency, formants and pronunciation habits as timbre features, generating embedding vectors and Mel spectra to obtain pure timbre cloned speech.

3. The speech cloning method incorporating emotion enhancement mechanisms according to claim 1, characterized in that: The specific process of constructing the directional construction mechanism to output speakerprint features based on the preprocessed multi-source speech dataset is as follows: Based on acoustic preprocessing and normalization of multi-source speech datasets, a targeted construction mechanism is constructed: the pre-training basic speech corpus is processed according to target focus and dialect adaptation to obtain the original speech dataset as training data. The original speech dataset is subjected to hierarchical feature extraction using a hybrid network structure to obtain feature vectors. Local spectral feature capture and temporal characteristic modeling are implemented. A dual-loss constraint training optimization mechanism is established for the extracted feature vectors, implementing main loss constraints and auxiliary loss constraints. All parameters in the training process are dynamically monitored to verify stability, discriminability and dialect adaptation effect, and output voiceprint features.

4. The speech cloning method incorporating emotion enhancement mechanisms according to claim 1, characterized in that: The specific process of constructing the dialect-adapted spectral frequency generation and WaveRNN vocoder conversion mechanism to output the Mel spectrogram and synthesized speech is as follows: The algorithm inputs the embedded vectors and the corresponding text from the target Changsha dialect corpus, and establishes a dialect-adapted spectro-frequency generation mechanism: An improved text encoder is constructed at the encoding end to encode the text sequence into a high-dimensional semantic vector, which is then concatenated and fused with the embedded vector at the encoding end; a position-based monotonic attention structure is introduced for the attention mechanism; for the decoding mechanism, a two-layer recurrent neural network with residual connections is used as the decoding core, a stop marker prediction head is added to the output of the decoding RNN, and sentence end features are trained using binary cross-entropy loss, with a dynamic threshold used to determine the sentence end position; after generating the Mel spectrum, post-processing is performed to obtain the Mel spectrogram; the Mel spectrogram is input, and a WaveRNN vocoder conversion mechanism is constructed: the mapping relationship between the Mel spectrum and the time-domain waveform is learned, the spectrum and waveform conversion process is executed, and the synthesized speech is output; a neutral timbre synthesis quality monitoring and verification mechanism is established to evaluate the prosodic fluency of the synthesized speech; the short-time energy and fundamental frequency curves of the synthesized speech and real speech are analyzed, the inter-frame correlation between the synthesized Mel spectrum and the real spectrum is compared, and a neutral Mel spectrogram set is output.

5. The speech cloning method incorporating emotion enhancement mechanisms according to claim 1, characterized in that: The specific process of adding sentiment embeddings to the neutral model for joint training, outputting timestamped Mel spectrum, and establishing a dialect-adaptive attention and anomaly handling mechanism through neural network adaptive learning of feature associations is as follows: In a neutral model, sentiment embedding is added for joint training to build a speech synthesis backbone network: the text content of the corresponding speech is extracted from the target Changsha dialect basic corpus and the sentiment-annotated Changsha dialect corpus, and the text is preprocessed and input into the text encoder. The output is a high-dimensional text encoding vector aligned with the speech temporal sequence as the text feature. Neutral Changsha dialect speech of the speaker is extracted from the target Changsha dialect basic corpus and historical neutral speech dataset, and processed by the speaker encoder as speaker embedding feature; Changsha dialect speech with labeled categories was extracted from the sentiment-annotated Changsha dialect corpus. Based on the Praat tool, fundamental frequency, speech rate, and intensity prosodic features were extracted. After dimensionality reduction and Z-score standardization using a lightweight fully connected network, sentiment embedding features were obtained. The input text features, speaker embedding features, and sentiment embedding features are standardized through feature interface standardization. A multi-source feature fusion layer is added between the text encoder and the attention layer. The standardized text encoding vector, speaker embedding vector, and sentiment embedding vector are concatenated along the feature dimension to form a fusion condition vector. This fusion condition vector is added to the text encoding output through residual connections and gating mechanisms to obtain an enhanced encoding sequence. The fusion condition vector is used as the initial state of the attention RNN and the conditional input of the decoding RNN. At each decoding time step, the Vimel spectrum of the previous frame and the fusion condition vector are fed into the decoder. Text encoding input, speaker feature input, and sentiment feature input are added to the multi-source feature fusion layer, and the output input is bound to the output Vimel spectrum. A timestamp label is added to each frame spectrum using a forced aligner method. Feature association is achieved through adaptive learning via neural networks. The adaptive compression weight matrix is ​​multiplied by the combination of the text encoding vector, speaker embedding vector, emotion feature activation coefficients, and emotion embedding vector element-wise to obtain a high-dimensional feature compression term. The high-dimensional feature compression term is added to the bias term to obtain the offset-compensated compressed feature term. A linear rectified activation function is applied to the offset-compensated compressed feature term to obtain a non-linear activation feature term. The non-linear activation feature term is added to the text encoding vector to obtain the fused feature vector. In the Tacotron2 synthesizer, a three-level architecture of position-sensitive attention layer, dual LSTM decoder, and frame prediction smoothing layer is adopted to establish a dialect-adaptive attention mechanism. The backbone network quality is monitored, and an anomaly handling mechanism is established.

6. The speech cloning method incorporating emotion enhancement mechanisms according to claim 1, characterized in that: The specific process for evaluating the synthesized speech and verifying the effectiveness of weight adjustment is as follows: Based on the accuracy of capturing emotional features using emotional loss weights, a quantitative evaluation of synthesized speech is conducted. The macro-average F1 score of emotion recognition is used as the F1 score for emotion recognition. The relative errors of fundamental frequency, energy, and prosodic elements between synthesized speech and target speech are compared. The mean square error of fundamental frequency, relative deviation of energy, and syllable duration deviation rate are statistically analyzed as speaker similarity. The MOSNet model is used to score the synthesized speech to obtain a naturalness score. Combined with the accuracy of emotional features, the cosine similarity of the embedding vector is used to verify timbre consistency. A dual-track adjustment strategy of phased progression and real-time feedback is designed to address timbre drift, insufficient emotional expression, and index stagnation. The effectiveness of weight adjustment is verified through quantitative indicators of timbre consistency and emotional feature accuracy, as well as scenario testing, and a dynamic weight strategy file is output.

7. A speech cloning system incorporating an emotion enhancement mechanism, employing a speech cloning method incorporating an emotion enhancement mechanism as described in any one of claims 1-6, characterized in that, include: The overall framework of transfer learning and the two-stage training module are used to construct the transfer learning framework and the multi-source speech dataset. The multi-source speech dataset is subjected to acoustic preprocessing and normalization. Pure timbre cloned speech is obtained through a two-stage collaborative training mechanism. The timbre cloning and speaker encoder training module is used to construct a directional construction mechanism to output voiceprint features based on a preprocessed multi-source speech dataset, and to construct a dialect-adaptive spectro-frequency generation and WaveRNN vocoder conversion mechanism to output Mel spectrograms and synthesized speech. The speech synthesizer and vocoder modeling module is used to add emotion embeddings to the neutral model for joint training, outputting time-stamped Mel spectra, and establishing dialect-adaptive attention and anomaly handling mechanisms through neural network adaptive learning of feature associations. The emotion cloning and emotion feature learning module is used to design a multi-level loss function constraint balance based on the emotion-annotated Changsha speech corpus, adopt multiplicative adaptive adjustment of emotion loss weights, evaluate synthesized speech, and verify the effectiveness of weight adjustment. The multi-feature fusion and temporal modeling module is used to construct a dual-track input system. It performs feature fusion and temporal modeling through deep restricted Boltzmann machine and time-delay neural network modeling, which together drive the synthesis process to output emotional speech clone data.

8. A voice cloning storage medium incorporating an emotion enhancement mechanism, characterized in that, include: The storage medium has one or more programs, which are executed by one or more processors to implement a voice cloning method incorporating an emotion enhancement mechanism as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Chinese speech cloning method for end-to-end tone and emotion migration

    CN115359775A

  • Voice cloning method and device for emotion expression, equipment and medium

    CN120783724A

  • Speech emotion recognition method and system based on multi-modal feature fusion

    CN121565211A