Semi-supervised dialect emotion speech synthesis system based on hybrid experts
Through a hybrid expert architecture and a semi-supervised learning framework, data scarcity and feature modeling complexity in dialect emotional pronunciation synthesis are solved, and high-quality dialect emotional pronunciation synthesis is achieved, improving nature and adaptability.
Patent Information
- Application Number
- CN202510778709.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing dialect emotional pronunciation synthesis technology faces the problems of data scarcity and complexity of emotional expression modeling, and it is difficult to accurately capture dialect features and emotional expression under low resource conditions. The existing methods fail to effectively utilize unlabeled data, resulting in the synthetic pronunciation being stiff or inconsistent with the context.
A semi-supervised learning architecture based on hybrid experts is adopted, including text analysis module, hybrid expert module, dynamic routing module, semi-supervised learning module and acoustic parameter generation module. Through the collaborative work of dialect acoustic experts, rhythm experts, emotional experts and general feature experts, combined with supervised learning and self-supervised learning strategies, feature decoupling and collaborative optimization are achieved.
Under the conditions of limited labeled data, the naturalness and fidelity of dialect emotional pronunciation synthesis are significantly improved, the amount of labeled data is reduced, and multiple emotions are generated and cross-dial dialect knowledge transfer is supported, which is more adaptable than traditional methods.
Smart Images

Figure CN120299449A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to dialect speech synthesis, and in particular to a semi-supervised dialect emotional speech synthesis system based on mixture of experts. Background Art
[0002] The current dialect emotional speech synthesis technology faces two core challenges: data scarcity and complexity of emotional expression modeling. Existing mainstream speech synthesis methods, such as Tacotron, FastSpeech2, etc., rely on large-scale labeled data for training. However, the cost of collecting dialect speech data is high and the annotation requires the participation of linguistics experts, resulting in the model being prone to overfitting. For example, the text-to-speech alignment data of dialects such as Cantonese and Minnan dialect is less than 5% of the general language, severely restricting the effect of supervised learning. The end-to-end neural network synthesis model performs poorly in low-resource scenarios and is difficult to capture the phonetic features and emotional expressions of dialects. Although the cross-dialect adaptation framework alleviates the data scarcity problem to a certain extent, it still relies on a large number of target dialect samples for fine-tuning.
[0003] Dialect emotional expression is strongly related to regional cultural characteristics (such as intonation fluctuations, prosody habits, etc.). Traditional models are difficult to capture its uniqueness due to the lack of fine-grained emotional annotation data, and the synthesized speech often appears rigid or inconsistent with the context emotion. For low-resource scenarios, although semi-supervised learning has been introduced into the field of speech synthesis, its application in dialects still has limitations. Existing methods such as self-supervised pre-training combined with fine-tuning usually directly reuse the general speech feature extraction strategy. However, there are significant differences in phoneme distribution and acoustic characteristics between dialects and general languages, resulting in low knowledge transfer efficiency. Pre-trained models based on HuBERT, etc., are prone to pronunciation errors or prosody distortion in the generated speech due to incomplete phoneme coverage in the dialect scenario.
[0004] Existing emotional speech synthesis technologies such as the VAE-based emotion control model or GST framework perform well in processing standard speech emotions, but are difficult to accurately capture the unique emotional expression patterns of dialects. Although MoE has been used to improve model capacity and task adaptability, such as Google's Switch Transformer achieving multi-task learning through a dynamic routing mechanism, existing MoE architectures are mostly designed for tasks with sufficient annotations and do not perform collaborative optimization for labeled and unlabeled data. The traditional dynamic routing strategy only relies on the semantic information of the input text and does not combine the relevance between dialect acoustic features and emotional labels, resulting in inefficient division of labor among experts. At the same time, existing methods do not design an expert training mechanism for unlabeled data and cannot make full use of the massive low-cost dialect speech resources.
[0005] In summary, the existing technical solutions mainly have the following limitations: 1) The integrated model is difficult to balance different features. Traditional end-to-end synthesis models use a single network architecture to process each feature dimension of speech simultaneously, making it difficult to balance dialect features and emotional expression under limited data conditions. 2) The annotation efficiency is low. Existing methods usually require a large number of dialect samples with emotional annotations, which is difficult to meet in practical applications. 3) There is a lack of an effective knowledge transfer mechanism, and the common features between standard speech and dialects, as well as between different dialects, are not fully utilized for knowledge transfer. 4) Modal separation leads to fragmented features. Some methods artificially divide speech features into independent modules for processing, but lack an effective feature fusion mechanism, resulting in poor overall coordination of the synthesized speech. Summary of the Invention
[0006] In view of the above-mentioned drawbacks of the existing technology, the present invention provides a semi-supervised dialect emotional speech synthesis system based on a mixture of experts, which can effectively overcome the defect that it is difficult to accurately synthesize dialect emotional speech in the case of scarce sample resources.
[0007] To achieve the above objectives, the present invention is realized through the following technical solutions: A semi-supervised dialect emotional speech synthesis system based on a mixture of experts, comprising the following component modules: A text analysis module, which preprocesses the input dialect text and generates a text representation vector through feature extraction and feature fusion; A mixture of experts module, which obtains dialect acoustic features, prosodic features, emotional features, and general acoustic features; A dynamic routing module, which realizes intelligent cooperation between experts through a task-aware soft routing algorithm; A semi-supervised learning module, which trains supervised learning using labeled dialect emotional speech data and simultaneously trains self-supervised learning using unlabeled dialect emotional speech data; An acoustic parameter generation module, which integrates the outputs of each expert to generate a complete set of acoustic parameters; A neural vocoder, which converts the set of acoustic parameters into the final dialect emotional speech.
[0008] Preferably, the mixture of experts module includes a dialect acoustic expert, a prosody expert, an emotion expert, and a general feature expert; The dialect acoustic expert uses an improved dialect acoustic autoencoder and focuses on capturing the timbre and pronunciation features of the dialect. The phoneme sequence P first obtains an initial representation through a phoneme embedding layer and then is processed by an attention-based encoder-decoder network: , where, speakerid is the speaker identifier, used to identify and distinguish the voice characteristics of different speakers. DialectAcousticEncoder represents the improved dialect acoustic encoder, A dialect is the dialect acoustic feature; In addition, a phoneme-level alignment mechanism is designed. Through forced alignment training and attention constraints, accurate dialect pronunciation modeling is achieved. The dialect phoneme perception attention unit calculates through: , where softmax represents the softmax function, Q is the query matrix Query, representing the features of the current processed phoneme sequence P, K is the key matrix Key, used to match with the query matrix Query, K T is the transpose matrix of K, V is the value matrix Value, containing the information that actually needs to be extracted, d is the feature dimension, used to scale the score of dot product attention, M dialect is the attention mask pre-calculated based on dialect phoneme similarity, used to guide the model to focus on the pronunciation details of the dialect; The prosody expert, based on the hierarchical recurrent neural network, models the prosody features including the rhythm, pause and stress of the sentence. It adopts a three-layer cascade structure, corresponding to the prosody feature modeling at the syllable, word and sentence levels respectively: , , , , where SyllableLevelRNN represents the syllable-level recurrent neural network, used to model the prosody features at the syllable level, R syllable is the syllable-level prosody feature; WordLevelRNN represents the word-level recurrent neural network, used to model the prosody features at the word level, word boundaries is the word boundary information, identifying the start and end positions of the word, R word is the word-level prosody feature; PhraseLevelRNN represents the phrase-level recurrent neural network, used to model the prosody features at the sentence level, phrase boundaries is the phrase boundary information, identifying the demarcation points of phrases in the sentence, R phrase is the phrase-level prosody feature; ProsodyProjector represents the prosody projector, used to fuse the prosody features at different levels to generate the comprehensive prosody feature R prosody ; The hierarchical recurrent neural network introduces an autoregressive prediction mechanism that predicts the prosodic feature r at the current time t by considering the historical prosodic trajectory t , captures the prosodic patterns and long-term dependencies of dialects, and dynamically integrates prosodic features at different levels through an adaptive hierarchical gating mechanism: , Among them, f represents the autoregressive prediction function, which predicts the prosodic feature r at the current time t based on the historical prosodic trajectory and the current content content t r t , r t-1 , r t-2 , …, r t-k are the historical prosodic features at the past k time moments; The emotion expert, based on a hybrid architecture of variational autoencoder VAE and normalizing flow, realizes fine control of emotion expression. This expert first learns the emotion latent space Z from the annotated dialect emotion speech data X emo : emo , , Among them, Encoder emo represents the emotion encoder, which is used to map the input annotated dialect emotion speech data X emo emo to the emotion latent space Z emo ; Decoder emo represents the emotion decoder, which is used to obtain the reconstructed emotion features from the emotion latent space Z emo ; Then, the operability of the emotion latent space Z is enhanced through normalizing flow: emo , Among them, NormalizingFlow represents the normalizing flow model, which is used to enhance the operability of the emotion latent space Z emo , is the emotion latent space after being processed by normalizing flow; Emotion decoupling learning decomposes the emotion latent space processed by normalizing flow into three orthogonal dimensions: intensity, category, and style: , Among them, Z intensity is the emotion intensity dimension, representing the strength of emotion, and Z category is the emotion category dimension, representing the type of emotion, and Z styleThe emotional style dimension represents the way and characteristics of expressing emotions, so that each dimension can be controlled independently; Emotion experts also designed an emotion transfer mechanism based on reference audio, which extracts the emotional features of the reference audio and transfers them to the target synthesized speech: , , Among them, EmotionExtractor represents the emotion extractor, which is used to extract the emotion from the reference audio Audio ref Extract reference audio emotion feature Z ref , Reference Audio ref Contains target sentiment characteristics; EmotionTransfer represents the emotion transfer function, which is used to transfer the reference audio emotion feature Z ref Transfer to the target synthesized speech, linguistic context Provide semantic information for language context, E emotion The final generated emotional features; The universal feature expert uses a pre-training-fine-tuning paradigm to extract universal acoustic features that are independent of dialects. This expert extracts shared features across dialects through a contrastive learning method: , Among them, GeneralFeatureEncoder represents a general feature encoder used to extract the audio features from the phoneme sequence P and audio features Extract the universal acoustic features F that are independent of dialect general , audio features audio features Contains spectral characteristics, fundamental frequency and energy; The general feature expert uses the Transformer-XL architecture to support long-distance dependency modeling, and compresses the model size through knowledge distillation technology to maintain reasoning efficiency: , Among them, KnowledgeDistillation represents knowledge distillation, which is used to transfer the knowledge of a large model to a smaller model. Temperature=2.0 represents the temperature parameter used in the knowledge distillation process, which is used to control the smoothness of the soft label. A higher temperature parameter makes the probability distribution smoother, which helps to capture the subtle feature relationships in the large model. distilled It is a compressed version of the universal acoustic features obtained through knowledge distillation, which retains the key characteristics of the large model while significantly reducing the model size. This design enables universal feature experts to provide an effective acoustic foundation for low-resource dialects and significantly reduce the data requirements for dialect acoustic modeling.
[0009] Preferably, the dynamic routing module realizes intelligent collaboration among experts through a task-aware soft routing algorithm, including: S1. Calculate the correlation scores between the input features and each expert: , where CompatibilityScorer represents the correlation scorer, which is used to evaluate the matching degree between the input feature input features and the experts. The input feature input features contains the text and audio information to be processed, is the feature description of the i-th expert, which includes the expertise field of the expert, and S i is the input feature input features and the correlation score between the input feature input and the i-th expert; features The correlation scorer uses a multi-layer perceptron network, considering the text features, target emotion types, and dialect features of the input feature input , where MLP represents the multi-layer perceptron network, text features is the text feature, target emotion is the target emotion type, and dalect id is the dialect identifier; S2. Calculate the weight distribution of each expert through an adaptive gating network: , where softmax represents the softmax function, which is used to convert the correlation score S i into a probability distribution. The temperature is a dynamically adjusted temperature parameter, which is used to control the smoothness of the weight distribution, and W i is the initial weight of the i-th expert; S3. To enhance the effect of expert collaboration, introduce a task adaptability score and dynamically adjust the expert weights through historical performance: , , where TaskPerformanceEvaluator represents the task performance evaluator, which is used to evaluate the historical performance of the expert in the current task type. expert i is the i-th expert, current task is the current task, and A i is the adaptability score of the i-th expert expert i in the current task type; sum(W j *A j ) represents the weighted sum of the adaptability scores of all experts on the current task type, which is used to normalize the weights. is the weight of the i-th expert after task adaptability adjustment; S4. Designed a soft gating fusion strategy that allows multiple experts to participate simultaneously, but with different degrees of contribution: , where sum represents the sum function. is the output result of the i-th expert, and Output is the output result of the final fusion, which is obtained by weighted summation.
[0010] Preferably, the semi-supervised learning module uses the labeled dialect emotional speech data to train the supervised learning, including: The total loss function of the supervised learning is L supervised : , where L acoustic is the acoustic reconstruction loss, which measures the difference between the synthetic speech and the real speech in acoustic features, and calculates the difference of the Mel spectrogram using the mean square error or L1 loss. L prosody is the prosody matching loss, which evaluates the consistency between the synthetic speech and the target speech in the prosody pattern including the fundamental frequency contour, duration, and energy variation. L emotion is the emotion expression loss, which ensures that the synthetic speech expresses the target emotion features and is realized based on the backpropagation of the emotion classifier. and are the weight coefficients, which are used to adjust the relative importance of different loss terms and control the weights of the prosody matching loss and the emotion expression loss in the total loss respectively; The semi-supervised learning module uses the unlabeled dialect emotional speech data to train the self-supervised learning by adopting the following three training strategies, including: 1) Feature reconstruction strategy, which learns robust features by adding random perturbations and requiring the model to recover: , , , where ADDNoise represents the function of adding noise, which is used to introduce random perturbations to the unlabeled dialect emotional speech data X original , noise ratio is the noise ratio parameter, which is used to control the intensity of adding noise. noise ratio =0.2 means adding 20% random noise, Xnoisy is the data after adding noise; Model represents the speech synthesis model to be trained, X reconstructed is the output reconstructed by the model from the data X after adding noise noisy ; ReconstructionLoss represents the reconstruction loss function, and the L1 or L2 norm is used to calculate the difference between the unlabeled dialect emotional speech data X original and the reconstructed output X reconstructed , and L recon is the calculated feature reconstruction loss value; 2) Consistency regularization strategy, applying different augmentation methods to the same input and requiring the model output to be consistent: , , , , , where Augmentation1 and Augmentation2 represent two different data augmentation functions, including time stretching, pitch transformation, and volume adjustment, X aug1 , X aug2 are the data obtained after being processed by different augmentation methods; F1 and F2 are the feature representation outputs of the model for the two augmented data X aug1 , X aug2 respectively. ConsistencyLoss represents the consistency loss function, which requires the model to generate similar feature representations for the same data in different augmented versions, and is measured using cosine similarity or mean squared error. L consist is the calculated consistency loss value; 3) Pseudo-label iterative optimization strategy, using the current model to generate pseudo-labels for the unlabeled dialect emotional speech data X original and then updating the model using the high-confidence pseudo-labels: , , , where CurrentModel is the speech synthesis model in the current training state, Y pseudo is the pseudo-label generated by the model for the unlabeled dialect emotional speech data X original , including predicted acoustic parameters, prosodic features, and emotion labels; SelectHighConfidence is a high-confidence sample selection function, and threshold is the confidence threshold. is the set of selected high-confidence samples; PseudoLabelLoss represents the pseudo-label loss function, which requires the output of the model to be consistent with the pseudo-labels. is the label generated by the model for the set of selected high-confidence samples L pseudo is the calculated pseudo-label loss value; The semi-supervised learning module balances the contributions of the supervised learning and self-supervised learning paths through a dynamic weight adjustment mechanism to obtain the total loss function L of model training total : , where is the weight coefficient of supervised learning, which is dynamically adjusted according to the training progress, initially large and gradually decreasing to the balance value, realizing a smooth transition from supervised to semi-supervised. is the weight coefficient of self-supervised learning. , , are the weight coefficients of the feature reconstruction loss L recon , the consistency loss L consist , and the pseudo-label loss L pseudo respectively, which are used to balance the relative importance of the three self-supervised learning training strategies.
[0011] Preferably, the acoustic parameter generation module integrates the outputs of each expert to generate a complete set of acoustic parameters, including: The acoustic parameter generation module adopts an attention-based multi-level feature fusion strategy, and the processing process is as follows: S1. The feature alignment layer solves the temporal difference problem of the outputs of different experts, and uses the dynamic time warping algorithm for sequence alignment: , where TemporalAlignment represents the temporal alignment function, which is used to solve the temporal difference problem of different feature sequences, and F aligned is the aligned feature set, ensuring that various features are accurately aligned in the time dimension; S2. The cross-attention layer realizes deep interaction between features, and captures the dependence relationship between different features through the multi-head attention mechanism: , , , , , Among them, LinearProjection Q , LinearProjection K , LinearProjection V respectively represent the linear projection functions of query, key, and value. F aligned [i], F aligned [j] are the aligned features of the i-th class and the j-th class respectively. Q i is the query vector of the aligned features of the i-th class, K j , V j are the key vector and value vector of the aligned features of the j-th class respectively; MultiHeadAttention represents the multi-head attention calculation function, which is used to capture the associations between features. Attention ij is the attention result of the aligned features of the i-th class to the aligned features of the j-th class; ConcatenateAndProject represents the concatenation and projection function, which is used to combine and reduce the dimension of all attention results. F interaction is the interaction feature, which contains the deep association information between various features; S3, the sequence refinement layer generates the final set of acoustic parameters A through residual connection and gating mechanism: , , , Among them, Conv1D represents a one-dimensional convolutional neural network, which is used to extract the local patterns of the interaction feature F interaction . sigmoid represents the sigmoid function, and G is the gating value, which is used to control the degree of information flow; F residual is the residual feature, which is extracted from the interaction feature F interaction through convolution. LayerNorm represents the layer normalization operation, which is used to stabilize the training process. A is the set of acoustic parameters, which combines the interaction feature F interaction and the residual feature F residual , and the mixing ratio is controlled by the gating mechanism; S4. Use the parameter adjustment network for dialect speech to fine-tune the generated set of acoustic parameters A to better match the dialect features: , Among them, DialectAdaptationNetwork represents the dialect adaptation network, which fine-tunes acoustic parameters according to different dialect features. dialect id is the dialect identifier, which guides the dialect adaptation network to adjust the direction. A refined is the set of fine-tuned acoustic parameters, which is more in line with the acoustic characteristics of the target dialect; S5. The final output contains the set of fine-tuned acoustic parameters A of the mel spectrogram, fundamental frequency curve, and energy contour refined and directly passes it to the neural vocoder for waveform generation.
[0012] Preferably, the neural vocoder converts the set of acoustic parameters into the final dialect emotional speech, including: The neural vocoder adopts an improved HiFi-GAN architecture, which is optimized for dialect speech characteristics and includes a generator network and two types of discriminator networks: 1) The generator adopts a multi-scale convolutional structure to map the acoustic parameters into a speech waveform: , Among them, G’ represents the generator network, which is composed of multiple transposed convolutional layers and residual blocks, and is used to map the set of fine-tuned acoustic parameters A refined into a speech waveform x, and the speech waveform x is a one-dimensional sequence in the time domain; A dialect adaptive convolutional layer is designed in the generator to dynamically adjust the convolutional kernel parameters according to the target dialect characteristics: , Among them, DialectModulator represents the dialect modulator, which generates a convolutional kernel scaling factor for the dialect according to the dialect identifier dialect id Conv represents the standard convolutional operation, DialectBias represents the dialect bias generator, and provides unique bias parameters for the dialect according to the dialect identifier dialect id The dialect identifier dialect id is used to identify the target dialect type, * represents element multiplication, and realizes the dynamic scaling of the convolutional kernel. Conv adaptive represents the dialect adaptive convolutional layer, which can dynamically adjust the convolutional kernel parameters according to the target dialect characteristics; 2) The multi-resolution spectrogram discriminator focuses on the restoration degree of spectral details and evaluates the generation quality by analyzing the spectral Figure 1 consistency of different frequency bands: , Among them, D spec_full represents the sub-network that discriminates the full-band spectrogram and focuses on the overall spectral contour. D spec_midIt represents a sub-network for discriminating the mid-frequency band spectrogram, focusing on the details of the main frequency bands of human voices, D spec_high It represents a sub-network for discriminating the high-frequency band spectrogram, focusing on high-frequency details and voice brightness, D spec D(x) is the output of the multi-resolution spectrogram discriminator; 3) The multi-period discriminator focuses on the prosodic patterns unique to dialects and evaluates the prosodic naturalness by analyzing the signal periodicity at different time scales: , where D p1 , D p2 , …, D pn represents a series of period discriminators for different time periods, which respectively process the signal periodicity at different time scales. n is the number of period discriminators, and D period D(x) is the output of the multi-period discriminator.
[0013] Compared with the prior art, the semi-supervised dialect emotional speech synthesis system based on mixture of experts provided by the present invention has the following beneficial effects: 1) An innovative semi-supervised learning architecture based on mixture of experts is proposed. Through the collaborative work of dialect acoustic experts, prosody experts, emotion experts and general feature experts, the decoupling and collaborative optimization of dialect features and emotional expressions are realized. Compared with the traditional single-model architecture, the present invention has significantly improved in terms of the naturalness of emotional expression and the fidelity of dialect features, effectively solving the core problem of low-resource dialect emotional speech synthesis; 2) A task-aware dynamic routing mechanism is designed, which can automatically adjust the expert weight distribution according to the input features and the target emotion type, significantly improving the adaptability of the system to different dialects and emotional scenarios. Compared with the fixed-weight mixture, this mechanism improves the naturalness of dialect emotional synthesis, reduces the conflict problem between emotions and dialect features, and realizes more accurate feature control; 3) An innovative combination of supervised learning and self-supervised learning strategies is adopted. Through three strategies of feature reconstruction, consistency regularization and pseudo-label iterative optimization, the limited labeled data and a large amount of unlabeled data are efficiently utilized. The present invention greatly reduces the amount of labeled data required for dialect emotional speech synthesis, while ensuring excellent synthesis quality, significantly reducing the resource threshold of dialect emotional speech synthesis; 4) Fine-grained emotion control and cross-dialect knowledge transfer are realized. The system can independently adjust in three dimensions of emotion intensity, category and style, supporting the generation of multiple basic emotions and fine-grained emotion variants. When extended to a new dialect, only a small amount of target dialect data is required for fine-tuning, and the adaptability is much higher than that of traditional methods, providing an efficient solution for the emotional speech synthesis of minority dialects. Description of the Drawings
[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0015] Figure 1 It is the system architecture diagram of the present invention. Specific embodiments
[0016] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0017] A semi-supervised dialect emotion speech synthesis system based on mixture of experts, as Figure 1 shown, includes the following component modules: The text analysis module preprocesses the input dialect text and generates a text representation vector through feature extraction and feature fusion; The mixture of experts module obtains dialect acoustic features, prosodic features, emotion features, and general acoustic features; The dynamic routing module realizes intelligent cooperation between experts through a task-aware soft routing algorithm; The semi-supervised learning module uses labeled dialect emotion speech data to train supervised learning, and at the same time uses unlabeled dialect emotion speech data to train self-supervised learning; The acoustic parameter generation module integrates the outputs of each expert to generate a complete set of acoustic parameters; The neural vocoder converts the set of acoustic parameters into the final dialect emotion speech.
[0018] The text analysis module includes a dialect text preprocessing unit, a phoneme sequence mapping unit, a text representation extraction unit, a dialect vocabulary semantic feature extraction unit, a text emotion feature extraction unit, and a text representation vector generation unit; The dialect text preprocessing unit standardizes the input dialect text; The phoneme sequence mapping unit maps the standardized dialect text into a phoneme sequence through a text-to-phoneme converter; The text representation extraction unit extracts text representation through a multi-layer cascade network. First, the character embedding layer is used to obtain the embedding representation of the standardized dialect text. Then, a one-dimensional convolutional neural network is used to capture the dependency between local language features and adjacent characters. Then, a BiLSTM network is used to model long-distance dependencies. Finally, the global context information is integrated through the self-attention mechanism to obtain the text representation. The dialect vocabulary semantic feature extraction unit uses a combination of dictionary matching and deep learning to identify dialect vocabulary in the standardized dialect text and extract its semantic features to obtain dialect vocabulary semantic features; The text sentiment feature extraction unit uses a BERT-based sentiment classifier to extract the sentiment tendency and intensity contained in the standardized dialect text to obtain the text sentiment features; The text representation vector generation unit generates a text representation vector by fusing the phoneme sequence, text representation, dialect vocabulary semantic features and text sentiment features through a feature fusion network.
[0019] Specifically, for the input dialect text T input Perform standardization, including word segmentation, symbol normalization, and digital normalization operations: T norm =Normalize(T input ) Among them, Normalize represents the text normalization processing function, T norm It is the dialect text after standardization; The standardized dialect text T is converted into norm Mapped to phoneme sequence P: P=Text2Phoneme(T norm ) Among them, Text2Phoneme represents the text-to-phoneme converter function. The text-to-phoneme converter integrates the rule system and the statistical model, and can accurately handle the pronunciation rules and sound changes of dialects; Obtain the standardized dialect text T through the character embedding layer norm The embedding representation E char : E char =Embedding(T norm ) Among them, Embedding represents the character embedding function, which is used to convert each character in the text into a dense vector representation. char It is a character-level embedding representation, that is, each character is mapped to a vector space of fixed dimension; Capture local language features and dependencies between adjacent characters using a one-dimensional convolutional neural network: E conv =Conv1D(E char ) where Conv1D represents a one-dimensional convolutional neural network, and E conv is the feature representation after one-dimensional convolutional processing; Model long-distance dependencies using a BiLSTM network: E lstm =BiLSTM(E conv ) where BiLSTM represents a bidirectional long short-term memory network that can consider both forward and backward information of the sequence, and E lstm is the feature representation after bidirectional LSTM processing; Fuse global context information through the self-attention mechanism to obtain the text representation T repr : T repr =MultiHeadAttention(E lstm ,E lstm ,E lstm ) where MultiHeadAttention represents the multi-head self-attention mechanism, which captures the correlations between different positions by fusing global context information; Identify the dialect words in the standardized dialect text T norm through a method that combines dictionary matching and deep learning, and extract their semantic features to obtain the dialect word semantic features D feat : D feat =DialectVocabEncoder(T norm ,DialectLexicon) where DialectVocabEncoder represents the dialect word encoder, which combines dictionary matching and deep learning to identify dialect words and extract their semantic features, and DialectLexicon represents the dialect dictionary, which contains dialect words and their corresponding semantic information; Use a BERT-based sentiment classifier to extract the sentiment tendency and intensity contained in the standardized dialect text T norm to obtain the text sentiment features E emo : E emo =EmotionAnalyzer(T norm ) where EmotionAnalyzer represents the BERT-based sentiment classifier; The phoneme sequence P and the text representation T are subjected to feature fusion through a feature fusion network repr , the semantic features D of dialect vocabulary feat and the text emotion feature E emo to generate a text representation vector T: T = FeatureFusionNetwork(P, T repr , D feat , E emo ) Among them, FeatureFusionNetwork represents the feature fusion network.
[0020] In the technical solution of this application, the innovation of the text analysis module lies in designing a context-aware text analysis algorithm for the dialect context, which can accurately identify and represent the unique linguistic features and emotional expressions of dialects, and provide high-quality text representations for subsequent processing.
[0021] The mixture of experts module includes a dialect acoustics expert, a prosody expert, an emotion expert, and a general feature expert; The dialect acoustics expert uses an improved dialect acoustic autoencoder and focuses on capturing the timbre and pronunciation features of dialects. The phoneme sequence P first obtains an initial representation through a phoneme embedding layer, and then is processed through an attention-based encoder-decoder network:
[0022] Among them, speaker id is the speaker identifier, used to identify and distinguish the voice features of different speakers. DialectAcousticEncoder represents the improved dialect acoustic encoder, and A dialect is the dialect acoustic feature; In addition, a phoneme-level alignment mechanism is designed. Through forced alignment training and attention constraints, accurate dialect pronunciation modeling is achieved. The dialect phoneme perception attention unit calculates through: , Among them, softmax represents the softmax function, Q is the query matrix Query, representing the features of the currently processed phoneme sequence P, K is the key matrix Key, used to match with the query matrix Query, and K T is the transpose matrix of K, V is the value matrix Value, containing the information that actually needs to be extracted, d is the feature dimension, used to scale the scores of dot product attention, and M dialect is the attention mask pre-computed based on dialect phoneme similarity, used to guide the model to focus on the pronunciation details of dialects; The prosody expert models prosody features including the rhythm, pauses, and accents of utterances based on a hierarchical recurrent neural network. It adopts a three-level cascaded structure, corresponding to prosody feature modeling at the syllable, word, and sentence levels respectively: , , , , Among them, SyllableLevelRNN represents the syllable-level recurrent neural network, which is used to model prosody features at the syllable level, and R syllable is the syllable-level prosody feature; WordLevelRNN represents the word-level recurrent neural network, which is used to model prosody features at the word level, and word boundaries is the word boundary information, indicating the start and end positions of words, and R word is the word-level prosody feature; PhraseLevelRNN represents the phrase-level recurrent neural network, which is used to model prosody features at the sentence level, and phrase boundaries is the phrase boundary information, indicating the demarcation points of phrases in the sentence, and R phrase is the phrase-level prosody feature; ProsodyProjector represents the prosody projector, which is used to fuse prosody features at different levels to generate comprehensive prosody feature R prosody ; The hierarchical recurrent neural network introduces an autoregressive prediction mechanism to predict the prosody feature r at the current time t by considering the historical prosody trajectory t , capturing the prosody patterns and long-term temporal dependencies of dialects, and realizing the dynamic integration of prosody features at different levels through an adaptive hierarchical gating mechanism: ,
[0023] Among them, f represents the autoregressive prediction function, which predicts the prosody feature r at the current time t based on the historical prosody trajectory and the current content content t and r t , r t-1 , r t-2 , …, r t-k are the historical prosody features at the past k time moments; The emotion expert realizes fine control of emotion expression based on a hybrid architecture of variational autoencoder (VAE) and normalizing flow. This expert first learns the emotion latent space Z from the labeled dialect emotion speech data X emo : emo : , , Among them, Encoder emo represents an emotion encoder, which is used to map the input labeled dialect emotion speech data X emo to the emotion latent space Z emo ; Decoder emo represents an emotion decoder, which is used to obtain the reconstructed emotion features from the emotion latent space Z emo ; ; Then, the operability of the emotion latent space Z is enhanced through normalizing flow: emo : , Among them, NormalizingFlow represents a normalizing flow model, which is used to enhance the operability of the emotion latent space Z emo , is the emotion latent space processed by normalizing flow; Emotion decoupling learning decomposes the emotion latent space processed by normalizing flow into three orthogonal dimensions: intensity, category, and style: , Among them, Z intensity is the emotion intensity dimension, which represents the strength of emotion, Z category is the emotion category dimension, which represents the type of emotion, Z style is the emotion style dimension, which represents the way and characteristics of expressing emotion, so that each dimension can be independently controlled; The emotion expert also designed an emotion transfer mechanism based on the reference audio, by extracting the reference audio emotion features and transferring them to the target synthesized speech: , , Among them, EmotionExtractor represents an emotion extractor, which is used to extract the reference audio emotion features Z ref from the reference audio Audio ref , and the reference audio Audio ref contains the target emotion features; EmotionTransfer represents an emotion transfer function, which is used to transfer the reference audio emotion features Z ref to the target synthesized speech, linguistic context is the language context, which provides semantic information, and E emotion is the finally generated emotion feature; The General Feature Expert, adopting the pre-training - fine-tuning paradigm, extracts general acoustic features independent of dialects. This expert extracts shared features across dialects through contrastive learning: , Among them, GeneralFeatureEncoder represents the general feature encoder, which is used to extract general acoustic features F independent of dialects from the phoneme sequence P and audio feature audio features in which the audio feature audio general contains spectral features, fundamental frequency, and energy; features The General Feature Expert adopts the Transformer-XL architecture, which supports long-distance dependence modeling and compresses the model size through knowledge distillation technology to maintain inference efficiency: , Among them, KnowledgeDistillation represents knowledge distillation, which is used to transfer the knowledge of a large model to a smaller model. temperature = 2.0 represents the temperature parameter used in the knowledge distillation process, which is used to control the smoothness of soft labels. A higher temperature parameter makes the probability distribution smoother and helps to capture subtle feature relationships in the large model. F distilled is the compressed general acoustic feature obtained through knowledge distillation. While retaining the key features of the large model, it significantly reduces the model size. This design enables the General Feature Expert to provide an effective acoustic basis for low-resource dialects and significantly reduces the data requirements for dialect acoustic modeling.
[0024] The Dynamic Routing Module realizes intelligent collaboration among experts through a task-aware soft routing algorithm, including: S1. Calculate the correlation scores between the input features and each expert: , Among them, CompatibilityScorer represents the correlation scoring device, which is used to evaluate the matching degree between the input feature input features and the expert. The input feature input features contains the text and audio information to be processed, is the characteristic description of the i-th expert, including the expertise field of this expert, S i is the input feature input features and the correlation score between the i-th expert; The correlation scoring device uses a multi-layer perceptron network, considering the text features, target emotion type, and dialect characteristics of the input feature input features : , Among them, MLP represents a multi-layer perceptron network, and text features is the text feature, and target emotion is the target sentiment type, and dalect id is the dialect identifier; S2. Calculate the weight distribution of each expert through an adaptive gating network: , Among them, softmax represents the softmax function, which is used to convert the correlation score S i into a probability distribution, and temperature is a dynamically adjusted temperature parameter, which is used to control the smoothness of the weight distribution, and W i is the initial weight of the i-th expert; S3. To enhance the collaborative effect of experts, introduce a task adaptability score, and dynamically adjust the expert weights through historical performance: , , Among them, TaskPerformanceEvaluator represents a task performance evaluator, which is used to evaluate the historical performance of experts on the current task type, and expert i is the i-th expert, and current task is the current task, and A i is the adaptability score of the i-th expert expert i on the current task type; sum(W j *A j ) represents the weighted sum of the adaptability scores of all experts on the current task type, which is used to normalize the weights, is the weight of the i-th expert after task adaptability adjustment; S4. Designed a soft gating fusion strategy, allowing multiple experts to participate simultaneously, but with different degrees of contribution: , Among them, sum represents the sum function, is the output result of the i-th expert, and Output is the finally fused output result, which is obtained by weighted summation.
[0025] In the technical solution of this application, compared with the traditional fixed-weight mixing in the design of the dynamic routing module, the dynamic routing mechanism can improve the naturalness of dialect emotion synthesis and reduce the conflict problem between emotion and dialect features.
[0026] The semi-supervised learning module uses the labeled dialect emotion speech data to train the supervised learning, including: The total loss function of supervised learning is L supervised : , where L acoustic is the acoustic reconstruction loss, which measures the difference in acoustic features between the synthesized speech and the real speech, and calculates the difference in the Mel spectrogram using the mean squared error or L1 loss. L prosody is the prosody matching loss, which evaluates the consistency between the synthesized speech and the target speech in the prosody pattern including fundamental frequency contour, duration, and energy variation. L emotion is the emotion expression loss, which ensures that the synthesized speech expresses the target emotion features and is realized based on the backpropagation of the emotion classifier. and are weight coefficients used to adjust the relative importance of different loss terms and control the weights of the prosody matching loss and the emotion expression loss in the total loss respectively; The semi-supervised learning module uses the following three training strategies to train self-supervised learning with unlabeled dialect emotion speech data, including: 1) Feature reconstruction strategy, which learns robust features by adding random perturbations and requiring the model to recover: , , , where ADDNoise represents the function of adding noise, which is used to introduce random perturbations to the unlabeled dialect emotion speech data X original , noise ratio is the noise ratio parameter used to control the intensity of the added noise. noise ratio = 0.2 means adding 20% random noise, and X noisy is the data after adding noise; Model represents the speech synthesis model to be trained, and X reconstructed is the output reconstructed by the model from the data X noisy after adding noise; ReconstructionLoss represents the reconstruction loss function, which calculates the difference between the unlabeled dialect emotion speech data X original and the reconstructed output X reconstructed using the L1 or L2 norm, and L recon is the calculated feature reconstruction loss value; 2) Consistency regularization strategy, which applies different augmentation methods to the same input and requires the model output to be consistent: , , , , , Among them, Augmentation1 and Augmentation2 represent two different data augmentation functions, including time stretching, pitch transformation, and volume adjustment. X aug1 , X aug2 are the data obtained after being processed by different augmentation methods; F1 and F2 are the feature representation outputs of the model for the two augmented data X aug1 , X aug2 respectively. ConsistencyLoss represents the consistency loss function, which requires the model to generate similar feature representations for the same data in different augmented versions, measured using cosine similarity or mean squared error. L consist is the calculated consistency loss value; 3) Pseudo-label iterative optimization strategy: Use the current model to generate pseudo-labels for the unlabeled dialect emotional speech data X original , and then update the model using the high-confidence pseudo-labels: , , , Among them, CurrentModel is the speech synthesis model in the current training state, and Y pseudo is the pseudo-label generated by the model for the unlabeled dialect emotional speech data X original , including predicted acoustic parameters, prosodic features, and emotion labels; SelectHighConfidence is the high-confidence sample selection function, threshold is the confidence threshold, is the set of selected high-confidence samples; PseudoLabelLoss represents the pseudo-label loss function, which requires the output of the model to be consistent with the pseudo-label. is the label generated by the model for the set of selected high-confidence samples , and L pseudo is the calculated pseudo-label loss value; The semi-supervised learning module balances the contributions of the supervised learning and self-supervised learning paths through a dynamic weight adjustment mechanism to obtain the total loss function L total : , Among them, is the weight coefficient for supervised learning, which is dynamically adjusted according to the training progress. It is initially large and gradually decreases to a balanced value to achieve a smooth transition from supervised to semi-supervised learning. is the weight coefficient for self-supervised learning. 、 、 are the weight coefficients of the feature reconstruction loss L recon 、consistency loss L consist 、pseudo-label loss L pseudo respectively, and are used to balance the relative importance of the three self-supervised learning training strategies.
[0027] In the technical solution of this application, the design of the semi-supervised learning module reduces the amount of labeled data required for dialect emotion speech synthesis, while ensuring the synthesis quality of the speech.
[0028] The acoustic parameter generation module integrates the outputs of each expert to generate a complete set of acoustic parameters, including: The acoustic parameter generation module adopts an attention-based multi-level feature fusion strategy, and the processing process is as follows: S1. The feature alignment layer solves the problem of temporal differences in the outputs of different experts, and uses the dynamic time warping algorithm for sequence alignment: , Among them, TemporalAlignment represents the temporal alignment function, which is used to solve the problem of temporal differences in different feature sequences, and F aligned is the aligned feature set, ensuring that various features are precisely aligned in the time dimension; S2. The cross-attention layer realizes deep interaction between features, and captures the dependence relationship between different features through the multi-head attention mechanism: , , , , , Among them, LinearProjection Q 、LinearProjection K 、LinearProjection V represent the linear projection functions of query, key, and value respectively, F aligned [i]、F aligned [j] are the aligned features of the i-th class and the j-th class respectively, Q i is the query vector of the aligned feature of the i-th class, K j 、V jThe key vector and value vector of the aligned features of the j-th class respectively; MultiHeadAttention represents the multi-head attention calculation function, which is used to capture the associations between features, Attention ij is the attention result of the aligned features of the i-th class to the aligned features of the j-th class; ConcatenateAndProject represents the concatenation and projection function, which is used to combine and reduce the dimension of all attention results, F interaction is the interaction feature, which contains the deep association information between various types of features; S3. The sequence refinement layer generates the final set of acoustic parameters A through residual connection and gating mechanism: , , , Among them, Conv1D represents the one-dimensional convolutional neural network, which is used to extract the local patterns of the interaction feature F interaction , sigmoid represents the sigmoid function, G is the gating value, which is used to control the degree of information flow; F residual is the residual feature, which is extracted from the interaction feature F through convolution interaction , LayerNorm represents the layer normalization operation, which is used to stabilize the training process, A is the set of acoustic parameters, which combines the interaction feature F interaction and the residual feature F residual , and the mixing ratio is controlled by the gating mechanism; S4. Use the parameter adjustment network for dialect speech to fine-tune the generated set of acoustic parameters A to better match the dialect features: , Among them, DialectAdaptationNetwork represents the dialect adaptation network, which fine-tunes the acoustic parameters for different dialect features, dialect id is the dialect identifier, which guides the direction of adjustment of the dialect adaptation network, A refined is the fine-tuned set of acoustic parameters, which better conforms to the acoustic characteristics of the target dialect; S5. Finally, output the fine-tuned set of acoustic parameters A containing the mel spectrogram, fundamental frequency curve, and energy contour refined , and directly pass it to the neural vocoder for waveform generation.
[0029] In the technical solution of this application, the innovation of the acoustic parameter generation module lies in realizing the seamless fusion of multi-expert outputs, ensuring the coordination and consistency of various types of features.
[0030] The neural vocoder converts the set of acoustic parameters into the final dialectal emotional speech, including: The neural vocoder adopts an improved HiFi-GAN architecture, which is optimized for the characteristics of dialectal speech and includes a generator network and two types of discriminator networks: 1) The generator adopts a multi-scale convolutional structure to map the acoustic parameters into a speech waveform: , Among them, G’ represents the generator network, which is composed of multiple transposed convolutional layers and residual blocks, and is used to map the fine-tuned set of acoustic parameters A refined into the speech waveform x, and the speech waveform x is a one-dimensional sequence in the time domain; A dialect adaptation convolutional layer is designed in the generator to dynamically adjust the convolutional kernel parameters according to the characteristics of the target dialect: , Among them, DialectModulator represents the dialect modulator, which generates the convolutional kernel scaling factor for the dialect according to the dialect identifier dialect id Conv represents the standard convolution operation, DialectBias represents the dialect bias generator, and provides unique bias parameters for the dialect according to the dialect identifier dialect id The dialect identifier dialect id is used to identify the target dialect type, * represents element multiplication, and realizes the dynamic scaling of the convolutional kernel. Conv adaptive represents the dialect adaptation convolutional layer, which can dynamically adjust the convolutional kernel parameters according to the characteristics of the target dialect; 2) The multi-resolution spectrogram discriminator focuses on the restoration degree of spectral details and evaluates the generation quality by analyzing the spectral consistency of different frequency bands: Figure 1 : , Among them, D spec_full represents the sub-network that discriminates the full-band spectrogram and focuses on the overall spectral contour. D spec_mid represents the sub-network that discriminates the mid-frequency band spectrogram and focuses on the details of the main frequency band of the human voice. D spec_high represents the sub-network that discriminates the high-frequency band spectrogram and focuses on the high-frequency details and sound brightness. D spec (x) is the output of the multi-resolution spectrogram discriminator; 3) The multi-period discriminator focuses on the dialect-specific prosodic patterns and evaluates the prosodic naturalness by analyzing the signal periodicity at different time scales: , Among them, D p1 , D p2 , …, D pnRepresents a series of period discriminators for different time periods, which respectively process the signal periodicity at different time scales. n is the number of period discriminators, and D period (x) is the output of the multi-period discriminator.
[0031] In the technical solution of this application, the neural vocoder is trained by combining the adversarial loss and the feature matching loss: , , Among them, L adv is the adversarial loss function, which prompts the generator to generate real speech waveforms. D represents the set of discriminators, including the multi-resolution spectrogram discriminator and the multi-period discriminator. y is the real speech waveform, G’(A) is the speech waveform generated by the generator according to the set of acoustic parameters A, and E is the expected value, which is approximated by the batch average during training; L fm is the feature matching loss function, which requires the features of the generated samples in each layer of the discriminator to be similar to those of the real samples. D i represents the feature representation of the i-th layer of the discriminator, represents calculating the L1 norm; At the same time, a perceptual loss function unique to the dialect is added, which pays special attention to the harmonic structure and prosody pattern of the dialect speech: , Among them, L dialect is the perceptual loss function, which is used to retain the unique acoustic features of the dialect speech. DialectPerceptualLoss is the comprehensive loss function, which combines the harmonic structure loss and the prosody pattern loss and is optimized for the dialect characteristics; The total loss function of the neural vocoder is L vocoder : , Among them, is the weight coefficient of the feature matching loss, which is used to balance the contributions of the adversarial loss and the feature matching loss, is the weight coefficient of the perceptual loss, which is used to control the degree of retention of the dialect characteristics.
[0032] The operation process of this system is divided into a training stage and an inference stage: 1) The training stage includes the following key steps: Train each expert model using the labeled dialect emotional speech data (about 50 - 200h), and further optimize the model using a large amount of unlabeled dialect emotional speech data (500 - 2000h) through a self-supervised learning framework. Adopt a progressive training strategy, and gradually increase the proportion of unlabeled data used as the training progresses. Finally, fine-tune the entire system end-to-end to balance the performance of each module. In the joint optimization stage, use a small batch size and a low learning rate setting, and continuously iterate until the performance of the validation set stabilizes.
[0033] 2) The process of the inference stage is as follows: The user inputs the dialect text and the target emotion type, and the system loads the corresponding expert model and parameters; The text analysis module processes the input text to generate a phoneme sequence and a linguistic feature vector, and the calculation time is about 10 - 30ms; The dynamic routing module calculates the weight distribution of each expert according to the input features and the target emotion type, and the calculation time is about 5ms; Each expert in the mixture-of-experts module works in parallel to generate their respective feature representations, and the calculation time is about 50 - 100ms; The acoustic parameter generation module integrates the outputs of each expert to generate a complete set of acoustic parameters, and the calculation time is about 30 - 50ms; The neural vocoder converts the set of acoustic parameters into the final dialect emotional speech, and the calculation time is about 10 - 20ms / second of speech.
[0034] The end-to-end processing time of the entire inference process for 10s of speech is controlled within 300ms, meeting the real-time interaction requirements. The system supports batch processing mode, and can process up to 16 parallel synthesis requests at a time, which is suitable for large-scale speech generation scenarios.
[0035] Through the organic combination of the mixture-of-experts architecture and the semi-supervised learning framework, the present invention innovatively solves the problems of data scarcity and feature modeling complexity in dialect emotional speech synthesis, and realizes high-quality dialect emotional speech synthesis under the condition of limited labeled data.
[0036] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements will not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A semi-supervised dialect emotion speech synthesis system based on mixture of experts, characterized in that: It includes the following component modules: The text analysis module preprocesses the input dialect text and generates a text representation vector through feature extraction and feature fusion; The mixture of experts module obtains dialect acoustic features, prosodic features, emotional features, and general acoustic features; The dynamic routing module realizes intelligent collaboration among experts through a task-aware soft routing algorithm; The semi-supervised learning module uses labeled dialect emotional speech data to train supervised learning and unlabeled dialect emotional speech data to train self-supervised learning; The acoustic parameter generation module integrates the outputs of each expert to generate a complete set of acoustic parameters; The neural vocoder converts the set of acoustic parameters into the final dialect emotional speech.
2. The semi-supervised dialect emotion speech synthesis system based on a mixture of experts according to claim 1, characterized in that: The text analysis module includes a dialect text preprocessing unit, a phoneme sequence mapping unit, a text representation extraction unit, a dialect vocabulary semantic feature extraction unit, a text emotion feature extraction unit, and a text representation vector generation unit; The dialect text preprocessing unit standardizes the input dialect text; The phoneme sequence mapping unit maps the standardized dialect text to a phoneme sequence through a text-to-phoneme converter; The text representation extraction unit extracts the text representation through a multi-level cascaded network. First, it obtains the embedding representation of the standardized dialect text through a character embedding layer, then uses a one-dimensional convolutional neural network to capture the local language features and the dependencies between adjacent characters, then adopts a BiLSTM network to model the long-distance dependencies, and finally uses a self-attention mechanism to fuse the global context information to obtain the text representation; The dialect vocabulary semantic feature extraction unit identifies the dialect vocabulary in the standardized dialect text and extracts its semantic features through a method combining dictionary matching and deep learning to obtain the dialect vocabulary semantic features; The text emotion feature extraction unit uses a BERT-based emotion classifier to extract the emotional tendency and intensity contained in the standardized dialect text to obtain the text emotion features; The text representation vector generation unit performs feature fusion on the phoneme sequence, text representation, dialect vocabulary semantic features, and text emotion features through a feature fusion network to generate a text representation vector.
3. The semi-supervised dialect emotion speech synthesis system based on a mixture of experts according to claim 2, wherein: Perform normalization processing on the input dialect text T input including word segmentation, symbol normalization, and number normalization operations: T norm =Normalize(T input ), Among them, Normalize represents the text normalization processing function, and T norm is the dialect text after normalization processing; Map the standardized dialect text T through a text-to-phoneme converter norm to a phoneme sequence P: P = Text2Phoneme(T norm ) Among them, Text2Phoneme represents the text-to-phoneme converter function. The text-to-phoneme converter combines a rule system and a statistical model and can accurately handle the pronunciation rules and tone-changing phenomena of dialects; Obtain the embedded representation E of the standardized dialect text T through the character embedding layer norm char : E char =Embedding(T norm ), Among them, Embedding represents a character embedding function used to convert each character in the text into a dense vector representation. E char is a character-level embedding representation, that is, each character is mapped into a vector space of a fixed dimension; Use a one-dimensional convolutional neural network to capture the local language features and the dependencies between adjacent characters: E conv =Conv1D(E char ), Among them, Conv1D represents a one-dimensional convolutional neural network, and E conv is the feature representation after one-dimensional convolutional processing; Adopt a BiLSTM network to model the long-distance dependencies: E lstm =BiLSTM(E conv ), Among them, BiLSTM represents a bidirectional long short-term memory network, which can consider both the forward and backward information of the sequence, and E lstm is the feature representation processed by the bidirectional LSTM; Fuse global context information through the self-attention mechanism to obtain the text representation T repr : T repr =MultiHeadAttention(E lstm ,E lstm ,E lstm ), Among them, MultiHeadAttention represents the multi-head self-attention mechanism, which captures the associations between different positions by fusing the global context information; Identify the dialect words in the standardized dialect text T through a method that combines dictionary matching and deep learning, and extract their semantic features to obtain the dialect word semantic features D norm feat : D feat =DialectVocabEncoder(T norm ,DialectLexicon) Among them, DialectVocabEncoder represents the dialect vocabulary encoder, which combines dictionary matching and deep learning to identify dialect vocabulary and extract its semantic features. DialectLexicon represents the dialect dictionary, which contains dialect vocabulary and their corresponding semantic information; Extract the standardized dialect text T using a BERT-based sentiment classifier norm to obtain the sentiment tendency and intensity contained therein, and get the text sentiment feature E emo : E emo =EmotionAnalyzer(T norm ), Among them, EmotionAnalyzer represents the BERT-based emotion classifier; Through the feature fusion network, the phoneme sequence P, the text representation T repr , the semantic features D of dialect vocabulary feat and the text emotion feature E emo are fused to generate the text representation vector T: T = FeatureFusionNetwork(P, T repr , D feat , E emo ), Among them, FeatureFusionNetwork represents the feature fusion network.
4. The semi-supervised dialect emotion speech synthesis system based on a mixture of experts according to claim 1, characterized in that: The mixture-of-experts module includes a dialect acoustic expert, a prosody expert, an emotion expert, and a general feature expert; The dialect acoustic expert uses an improved dialect acoustic autoencoder, which focuses on capturing the timbre and pronunciation features of the dialect. The phoneme sequence P first obtains an initial representation through the phoneme embedding layer, and then is processed by an attention-based encoder-decoder network: , Among them, speaker id is the speaker identifier, used to identify and distinguish the voice characteristics of different speakers. DialectAcousticEncoder represents an improved dialect acoustic encoder, and A dialect is the dialect acoustic feature; In addition, a phoneme-level alignment mechanism is designed to achieve accurate dialect pronunciation modeling through forced alignment training and attention constraints. The dialect phoneme perception attention unit calculates through: , Among them, softmax represents the softmax function, Q is the query matrix Query, representing the features of the current processed phoneme sequence P, K is the key matrix Key, which is used to match with the query matrix Query, and K T is the transposed matrix of K, V is the value matrix Value, which contains the information actually needed to be extracted, d is the feature dimension, which is used to scale the scores of dot-product attention, and M dialect is the attention mask pre-computed based on the dialect phoneme similarity, which is used to guide the model to focus on the pronunciation details of the dialect; The prosody expert, based on a hierarchical recurrent neural network, models the prosody features including the rhythm, pause, and stress of the sentence, and adopts a three-layer cascade structure, corresponding to the prosody feature modeling at the syllable, word, and sentence levels respectively: , , , , Among them, SyllableLevelRNN represents a syllable-level recurrent neural network, which is used to model the prosodic features at the syllable level, and R syllable is the syllable-level prosodic feature; WordLevelRNN represents a word-level recurrent neural network, which is used to model word-level prosodic features, word boundaries is word boundary information, identifying the start and end positions of a word, R word is the word-level prosodic feature; PhraseLevelRNN represents a phrase-level recurrent neural network, which is used to model prosodic features at the sentence level, phrase boundaries is the phrase boundary information, which identifies the demarcation points of phrases in a sentence, R phrase is the phrase-level prosodic feature; The ProsodyProjector represents a prosody projector that is used to fuse prosody features at different levels to generate a comprehensive prosody feature R prosody ; The hierarchical recurrent neural network introduces an autoregressive prediction mechanism, which predicts the prosodic feature r at the current time t by considering the historical prosodic trajectory t , captures the prosodic patterns and long-term temporal dependencies of dialects, and realizes the dynamic integration of prosodic features at different levels through an adaptive hierarchical gating mechanism: , Among them, f represents the autoregressive prediction function, which predicts the prosodic feature r at the current moment t based on the historical prosodic trajectory and the current content content t The prosodic feature r at the current moment t is predicted t , r t-1 , r t-2 , …, r t-k are the historical prosodic features at the past k moments; An emotion expert, based on a hybrid architecture of variational autoencoder (VAE) and normalizing flow, realizes fine control of emotional expression. The expert first learns an emotional latent space Z from the labeled dialect emotional speech data X emo and emo : , , Among them, Encoder emo represents an emotion encoder, which is used to map the input labeled dialect emotion speech data X emo to the emotion latent space Z emo ; Decoder emo It represents an emotion decoder, which is used to obtain the reconstructed emotion features from the emotion latent space Z emo ; ; Then enhance the operability of the sentiment latent space Z emo through normalizing flow: , Among them, NormalizingFlow represents the normalizing flow model, which is used to enhance the operability of the sentiment latent space Z emo ; is the sentiment latent space after being processed by the normalizing flow; Affective decoupling learning decomposes the affective latent space processed by normalizing flow into three orthogonal dimensions: intensity, category, and style , Among them, Z intensity is the emotional intensity dimension, representing the strength of emotion. Z category is the emotional category dimension, representing the type of emotion. Z style is the emotional style dimension, representing the way and characteristics of expressing emotion, so that each dimension can be independently controlled; The emotion expert also designs an emotion transfer mechanism based on the reference audio, by extracting the emotion features of the reference audio and transferring them to the target synthetic speech: , , Among them, EmotionExtractor represents an emotion extractor, which is used to extract the reference audio emotion feature Z from the reference audio Audio ref from the reference audio Audio ref The reference audio Audio ref contains the target emotion feature; EmotionTransfer represents the emotion transfer function, which is used to transfer the emotional features Z of the reference audio ref to the target synthesized speech, and linguistic context is the language context that provides semantic information, and E emotion is the finally generated emotional feature; The general feature expert adopts a pre-training-fine-tuning paradigm to extract general acoustic features independent of the dialect. This expert extracts cross-dialect shared features through a contrastive learning method: , Among them, GeneralFeatureEncoder represents a general feature encoder, which is used to extract general acoustic features F independent of dialects from the phoneme sequence P and the audio feature audio features ; general The audio feature audio features includes spectral features, fundamental frequency, and energy; The general feature expert adopts the Transformer-XL architecture, which supports long-distance dependence modeling, and compresses the model size through knowledge distillation technology to maintain the inference efficiency: , Among them, KnowledgeDistillation represents knowledge distillation, which is used to transfer the knowledge of large models to smaller models. temperature = 2.0 represents the temperature parameter used in the knowledge distillation process, which is used to control the smoothness of the soft labels. A higher temperature parameter makes the probability distribution smoother and helps to capture the subtle feature relationships in large models, F distilled is the compressed general acoustic feature obtained through knowledge distillation. While retaining the key features of the large model, it significantly reduces the model size. This design enables the general feature expert to provide an effective acoustic basis for low-resource dialects and significantly reduces the data requirements for dialect acoustic modeling.
5. The semi-supervised dialect emotion speech synthesis system based on a mixture of experts according to claim 1, characterized in that: The dynamic routing module realizes intelligent collaboration among experts through a task-aware soft routing algorithm, including: S1. Calculate the correlation scores between the input features and each expert: , Among them, CompatibilityScorer represents a relevance scorer used to evaluate the matching degree between the input feature input features and an expert, and the input feature input features contains the text and audio information to be processed, is the characteristic description of the i-th expert, including the expertise field of this expert, S i is the input feature input features and the relevance score between the input feature input and the i-th expert; The relevance scorer uses a multi-layer perceptron network and considers the text features of the input feature input, the target sentiment type, and the dialect characteristics: features , Among them, MLP represents a multi-layer perceptron network, text features is the text feature, target emotion is the target sentiment type, dalect id is the dialect identifier; S2. Calculate the weight distribution of each expert through an adaptive gating network: , Among them, softmax represents the softmax function, which is used to convert the correlation score S i into a probability distribution. The temperature is a dynamically adjusted temperature parameter used to control the smoothness of the weight distribution. W i is the initial weight of the i-th expert; S3. To enhance the expert collaboration effect, a task adaptability score is introduced to dynamically adjust the expert weights based on historical performance: , , Among them, TaskPerformanceEvaluator represents the task performance evaluator, which is used to evaluate the historical performance of the expert in the current task type, and expert i is the i-th expert, and current task is the current task, and A i is the adaptability score of the i-th expert expert i on the current task type; sum(W j *A j ) represents the weighted sum of the adaptability scores of all experts on the current task type, which is used to normalize the weights, is the weight of the i-th expert after task adaptability adjustment; S4. A soft gating fusion strategy is designed to allow multiple experts to participate simultaneously, but with different degrees of contribution: , Among them, sum represents the sum function, is the output result of the i-th expert, and Output is the finally fused output result, which is obtained by weighted summation.
6. The semi-supervised dialect emotion speech synthesis system based on a mixture of experts according to claim 1, wherein: The semi-supervised learning module uses the labeled dialect emotion speech data to train the supervised learning, including: The total loss function of supervised learning is L supervised : , Among them, L acoustic is the acoustic reconstruction loss, which measures the difference in acoustic features between the synthesized speech and the real speech. The mean square error or L1 loss is used to calculate the difference in the mel spectrogram. L prosody is the prosody matching loss, which evaluates the consistency between the synthesized speech and the target speech in the prosody pattern including fundamental frequency contour, duration, and energy variation. L emotion is the emotion expression loss, which ensures that the synthesized speech expresses the target emotion features and is realized based on the backpropagation of the emotion classifier. and are weight coefficients, which are used to adjust the relative importance of different loss terms and respectively control the weights of the prosody matching loss and the emotion expression loss in the total loss. The semi-supervised learning module uses the unlabeled dialect emotion speech data to train the self-supervised learning using the following three training strategies, including: 1) Feature reconstruction strategy, by adding random perturbations and requiring the model to recover to learn robust features: , , , Among them, ADDNoise represents the function of adding noise, which is used to introduce random perturbations, noise, into the unlabeled dialect emotional speech data X original where noise ratio is the noise ratio parameter used to control the intensity of adding noise ratio and noise = 0.2 means adding 20% random noise to X noisy and X' is the data after adding noise; The Model represents the speech synthesis model to be trained, X reconstructed is the output reconstructed by the model from the data X after adding noise noisy ; ReconstructionLoss represents the reconstruction loss function, and the L1 or L2 norm is used to calculate the difference between the unlabeled dialect emotional speech data X original and the reconstructed output X reconstructed , and L recon is the calculated feature reconstruction loss value; 2) Consistency regularization strategy, applying different augmentation methods to the same input and requiring the model outputs to be consistent: , , , , , Among them, Augmentation1 and Augmentation2 represent two different data augmentation functions, including time stretching, pitch shifting, and volume adjustment, and X aug1 , X aug2 are the data obtained after being processed by different augmentation methods; F1 and F2 are the feature representation outputs of the model for two augmented data X aug1 , X aug2 respectively. ConsistencyLoss represents the consistency loss function, which requires the model to generate similar feature representations for the same data with different augmentation versions, measured using cosine similarity or mean squared error. L consist is the calculated consistency loss value; 3) Pseudo-label iterative optimization strategy, using the current model for unlabeled dialect emotion speech data X original to generate pseudo-labels, and then update the model using high-confidence pseudo-labels: , , , Among them, CurrentModel is the speech synthesis model in the current training state, and Y pseudo is the pseudo-label generated by the model for the unlabeled dialect emotion speech data X original and contains predicted acoustic parameters, prosodic features, and emotion labels; SelectHighConfidence is a function for selecting high-confidence samples, and threshold is the confidence threshold, which is the set of selected high-confidence samples; PseudoLabelLoss represents the pseudo-label loss function, which requires the output of the model to be consistent with the pseudo-labels. is the set of highly confident samples selected by the model The generated label, L pseudo is the calculated pseudo-label loss value; The semi-supervised learning module balances the contributions of the supervised learning and self-supervised learning paths through a dynamic weight adjustment mechanism to obtain the total loss function L for model training total : , Among them, is the weight coefficient of supervised learning, which is dynamically adjusted according to the training progress, initially large and gradually decreasing to a balanced value, realizing a smooth transition from supervised to semi-supervised. is the weight coefficient of self-supervised learning. , , are the weight coefficients of the feature reconstruction loss L recon , the consistency loss L consist , and the pseudo-label loss L pseudo respectively, which are used to balance the relative importance of the three self-supervised learning training strategies.
7. The semi-supervised dialect emotion speech synthesis system based on mixture of experts according to claim 1, wherein: The acoustic parameter generation module integrates the outputs of each expert to generate a complete set of acoustic parameters, including: The acoustic parameter generation module adopts an attention-based multi-level feature fusion strategy, and the processing process is as follows: S1. The feature alignment layer solves the problem of temporal differences in the outputs of different experts, and uses the dynamic time warping algorithm for sequence alignment: , Among them, TemporalAlignment represents a temporal alignment function, which is used to solve the temporal difference problem of different feature sequences, and F aligned is the set of aligned features to ensure that various features are precisely aligned in the time dimension; S2. The cross-attention layer realizes deep interaction between features, and captures the dependence relationships between different features through the multi-head attention mechanism: , , , , , Among them, LinearProjection Q 、LinearProjection K 、LinearProjection V respectively represent the linear projection functions of query, key, and value. F aligned [i] and F aligned [j] are the aligned features of the i-th class and the j-th class respectively. Q i is the query vector of the aligned features of the i-th class, K j , V j are the key vector and value vector of the aligned features of the j-th class respectively; MultiHeadAttention represents the multi-head attention calculation function, which is used to capture the correlations between features, and Attention ij is the attention result of the aligned feature of the i-th class to the aligned feature of the j-th class; ConcatenateAndProject represents the concatenation and projection function, which is used to combine and reduce the dimension of all attention results, F interaction is the interaction feature, which contains the deep correlation information between various features; S3. The sequence refinement layer generates the final set of acoustic parameters A through residual connections and gating mechanisms: , , , Among them, Conv1D represents a one-dimensional convolutional neural network for extracting local patterns of the interaction feature F, sigmoid represents the sigmoid function, and G is the gating value used to control the degree of information flow; interaction F residual is the residual feature, extracted from the interaction feature F through convolution. LayerNorm represents the layer normalization operation, which is used to stabilize the training process. A is the set of acoustic parameters, which combines the interaction feature F interaction and the residual feature F interaction , and the mixing ratio is controlled by the gating mechanism; residual S4. Use a parameter adjustment network for dialect speech to fine-tune the generated set of acoustic parameters A to better match the dialect features: , Among them, the Dialect Adaptation Network represents the dialect adaptation network, which fine-tunes the acoustic parameters according to different dialect features. The "dialect" id is the dialect identifier, which guides the adjustment direction of the dialect adaptation network. The "A" refined is the set of fine-tuned acoustic parameters, which is more in line with the acoustic characteristics of the target dialect; S5. The final output contains a set of fine-tuned acoustic parameters A including mel spectrograms, fundamental frequency curves, and energy contours, refined and is directly passed to a neural vocoder for waveform generation.
8. The semi-supervised dialect emotion speech synthesis system based on a mixture of experts according to claim 1, characterized in that: The neural vocoder converts the set of acoustic parameters into the final dialect emotion speech, including: The neural vocoder adopts an improved HiFi-GAN architecture, optimized for dialect speech characteristics, and includes a generator network and two types of discriminator networks: 1) The generator adopts a multi-scale convolutional structure to map acoustic parameters into speech waveforms: , Among them, G’ represents the generator network, which consists of multiple transposed convolutional layers and residual blocks and is used to map the fine-tuned acoustic parameter set A refined to the speech waveform x, which is a one-dimensional sequence in the time domain; A dialect adaptive convolutional layer is designed in the generator to dynamically adjust the convolutional kernel parameters according to the characteristics of the target dialect: , Among them, DialectModulator represents a dialect modulator that generates a convolutional kernel scaling factor for the dialect according to the dialect identifier dialect id Conv represents a standard convolution operation, and DialectBias represents a dialect bias generator that provides unique bias parameters for the dialect according to the dialect identifier dialect id The dialect identifier dialect is used to identify the target dialect type. * represents element-wise multiplication to achieve dynamic scaling of the convolutional kernel, and Conv id represents a dialect adaptive convolutional layer that can dynamically adjust the convolutional kernel parameters according to the target dialect characteristics; adaptive 2) The multi-resolution spectrogram discriminator focuses on the restoration degree of spectral details and evaluates the generation quality by analyzing the spectrogram consistency of different frequency bands: , Among them, D spec_full represents a sub-network for discriminating the full-band spectrogram, focusing on the overall spectral contour. D spec_mid represents a sub-network for discriminating the mid-band spectrogram, focusing on the details of the main frequency band of the human voice. D spec_high represents a sub-network for discriminating the high-band spectrogram, focusing on high-frequency details and sound brightness. D spec D(x) is the output of the multi-resolution spectrogram discriminator; 3) The multi-period discriminator focuses on the prosodic patterns unique to dialects and evaluates the prosodic naturalness by analyzing the signal periodicity at different time scales: , Among them, D p1 , D p2 , …, D pn represent a series of period discriminators for different time periods, respectively processing the signal periodicity at different time scales. n is the number of period discriminators, and D period (x) is the output of the multi-period discriminator.
9. The semi-supervised dialect emotion speech synthesis system based on a mixture of experts according to claim 8, characterized in that: The neural vocoder is trained by combining adversarial loss and feature matching loss: , , where L adv is the adversarial loss function that prompts the generator to generate real speech waveforms, D represents the set of discriminators, including the multi-resolution spectrogram discriminator and the multi-period discriminator, y is the real speech waveform, G’(A) is the speech waveform generated by the generator according to the set of acoustic parameters A, E is the expected value, which is approximated by the batch average during training; L fm is the feature matching loss function, which requires that the features of the generated samples in each layer of the discriminator be similar to those of the real samples. D i represents the feature representation of the i-th layer of the discriminator, represents calculating the L1 norm; At the same time, a dialect-specific perceptual loss function is added, which particularly focuses on the harmonic structure and prosodic patterns of dialect speech: , Among them, L dialect is a perceptual loss function used to preserve the unique acoustic features of dialect speech. DialectPerceptualLoss is a comprehensive loss function that combines harmonic structure loss and prosody pattern loss and is optimized for dialect characteristics; The total loss function of the neural vocoder is L vocoder : , Among them, is the weight coefficient of the feature matching loss, which is used to balance the contributions of the adversarial loss and the feature matching loss. is the weight coefficient of the perceptual loss, which is used to control the degree of retention of dialect features.
Citation Information
Patent Citations
Web attack detection method based on gating Transform
CN116527357A
Traffic flow prediction method based on de-noising attention enhancement cyclic multi-graph convolutional network
CN118506589A
Mixed expert-based multi-dialect speech recognition model and training method
CN118609545A
Multi-modal classification method fusing graph convolutional neural network and capsule graph neural network
CN119397371A
Intelligent robot speech synthesis method based on deep learning
CN119446117A
Cited By
Voice cloning method and device based on emotion enhancement and related medium
CN120599998A
Intelligent dish ordering method and device, electronic equipment and storage medium
CN121029930A
Intelligent ordering method and device, electronic equipment and storage medium
CN121029930B
Dialect speech recognition method, system and model
CN121053966A
Dialect speech recognition method, system and model
CN121053966B