Digital human emotional speech generation method based on emotional semantic modeling
By constructing emotional semantic units and an emotional contribution evaluation mechanism, and combining graph neural networks to model the interaction and evolution of emotional semantic units, the problems of coarse granularity and lack of structured relationships in existing emotional modeling technologies are solved, and high-precision and natural emotional speech generation is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING XILIANLIAN TECHNOLOGY CO LTD
- Filing Date
- 2026-04-15
- Publication Date
- 2026-07-10
AI Technical Summary
Existing methods for generating emotional speech are coarse in terms of emotion modeling granularity, ignoring the differences between different semantic segments within a sentence, making it difficult to accurately depict the impact of local semantics on the overall emotion, and lacking structured emotion relationship modeling, resulting in insufficient emotional coherence and naturalness in the generated speech.
By constructing emotional semantic units and introducing an emotional contribution evaluation mechanism, and combining graph neural networks to jointly model the interaction relationships and evolution process between emotional semantic units, an emotional interaction graph and an emotional evolution graph are constructed, enabling fine-grained emotional modeling and structured expression of dynamic emotional changes.
It significantly improves the accuracy and controllability of emotional expression in emotional speech generation, and enhances the ability of generated speech to express emotional coherence and naturalness.
Smart Images

Figure CN122369428A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, specifically to a method for generating emotional speech in digital humans based on emotion semantic modeling. Background Technology
[0002] With the rapid development of digital humans, intelligent customer service, and virtual interaction technologies, emotional speech generation, as an important technology to improve the naturalness of human-computer interaction, has received widespread attention. Existing emotional speech generation methods are mostly based on deep learning models, which automatically generate speech signals by modeling text or speech input. However, they mainly suffer from the following problems:
[0003] (1) The granularity of emotion modeling is coarse. Existing methods usually model emotions uniformly at the sentence level or speech segment level, ignoring the differences in emotional expression between different semantic segments within a sentence. It is difficult to accurately depict the influence of local semantics on the overall emotion, resulting in the problem of uniformity and coarse granularity in the generated speech in terms of emotional expression.
[0004] (2) Lack of structured emotion relationship modeling. Most existing methods use sequence models to model speech emotions, focusing on the modeling ability of the time dimension, but lack explicit modeling of complex interaction relationships between different semantic units. Especially in scenarios with multiple emotions intertwined or gradual emotion changes, it is difficult to effectively express the dynamic evolution process of emotions, resulting in insufficient emotional coherence and naturalness in the generated speech. Summary of the Invention
[0005] To address the above issues, this invention provides a digital human emotional speech generation method based on emotion semantic modeling. By constructing emotional semantic units and introducing an emotion contribution evaluation mechanism, fine-grained emotion modeling of speech semantic structure is achieved. This method can accurately identify the differential contributions of different semantic segments in a sentence to the overall emotional expression, thereby significantly improving the accuracy and controllability of emotional expression in generated emotional speech. Furthermore, by constructing a dual-graph structure of emotion interaction graph and emotion evolution graph, and combining graph neural networks to jointly model the interaction relationships and evolution process between emotional semantic units, a structured expression of complex dynamic emotional changes is achieved, thereby enhancing the generated speech's performance in terms of emotional coherence and naturalness.
[0006] The technical solution adopted by this invention is as follows: This invention provides a digital human emotional speech generation method based on emotion semantic modeling, which includes the following steps:
[0007] Step S1: Multimodal data preprocessing, obtaining multi-source input data for emotional speech generation, including text data, speech data and lip visual data corresponding to speech, and performing standardization processing, including alignment, noise reduction, segmentation and feature normalization processing, to obtain a multimodal input data sequence;
[0008] Step S2: Multimodal semantic feature extraction. Feature extraction is performed on the multimodal input data sequence to obtain textual semantic features, acoustic features, and visual features. A unified semantic representation is constructed through a cross-modal alignment mechanism.
[0009] Step S3: Emotion semantic unit segmentation. Based on the unified semantic representation, the multimodal input data sequence is structurally segmented to obtain emotion semantic units. For each emotion semantic unit, the corresponding emotion features are extracted to construct an emotion semantic association representation.
[0010] Step S4: Based on graph structure, emotional interaction is carried out by constructing an emotional relationship graph based on emotional semantic units and emotional semantic association representations. By modeling the interaction relationships between different emotional semantic units and the evolution process of individual emotional semantic units, the emotional state is structurally modeled to obtain dynamic emotional representation.
[0011] Step S5: Generating speech expression parameters. Based on the unified semantic representation and dynamic emotion representation, cross-modal fusion processing is performed to generate expression parameters for speech synthesis, including pitch, duration, energy, and prosodic information.
[0012] Step S6: Emotional speech generation. Based on the expression parameters used for speech synthesis, an emotional speech signal is generated using a speech synthesis model. The generated speech signal is then enhanced to obtain an enhanced speech signal.
[0013] Step S7: Digital human emotion collaboration drive, based on the enhanced emotional speech signal, drive the digital human to perform synchronized control of facial expressions and lip movements.
[0014] Furthermore, step S3 specifically includes the following steps:
[0015] Step S31: Semantic sequence structure parsing. The multimodal input data sequence corresponding to the unified semantic representation undergoes structure parsing and adaptive segmentation based on semantic changes to generate a candidate semantic unit set. The multimodal input data sequence is represented as follows:
[0016] ;
[0017] in, To unify the semantic representation of the corresponding multimodal feature sequences, For the first A fused feature vector of a time step or semantic segment The sequence length;
[0018] The sequence is segmented using a semantic change detection function to obtain a set of candidate semantic units, as shown below:
[0019] ;
[0020] in, For the set of candidate emotion semantic units, For the first Each of the candidate emotion semantic units Depend on The subsequences constitute the composition. This represents the total number of candidate emotion semantic units;
[0021] Step S32: Identify the dominant emotion unit. Evaluate the emotion contribution of the candidate emotion semantic unit set, set an emotion contribution threshold, and update the candidate emotion semantic units with emotion contributions higher than the threshold into the emotion semantic unit set, represented as follows: The unit with the highest emotional contribution is defined as the emotional anchor unit. The emotional contribution of each emotional semantic unit is as follows:
[0022] ;
[0023] in, For the first The emotional contribution of each candidate emotional semantic unit. A function for calculating the contribution of emotion. For the first The textual semantic feature vector of each candidate emotion semantic unit, For the first The acoustic feature vectors of candidate emotion semantic units, For the first Visual feature vectors of candidate emotion semantic units, Indicates the threshold for emotional contribution;
[0024] The emotion-dominant unit is defined as follows:
[0025] ;
[0026] in, As an emotional anchor unit, Select the candidate emotion semantic unit with the highest emotion contribution. This represents the number of candidate emotion semantic units;
[0027] Step S33: Emotional semantic unit category identification. Based on the semantic function and emotional expression role of emotional semantic units in sentences, the set of emotional semantic units is categorized, and corresponding category embedding vectors are assigned to different categories, including emotional expression units, tone regulation units, and semantic content units. Multimodal feature representations are extracted for each emotional semantic unit, and enhanced representations are performed by combining the category embedding vectors, as shown below:
[0028] ;
[0029] in, Indicates the first Embedding vectors corresponding to the categories of each emotion semantic unit Indicates the first Category identification function for each semantic unit, Indicates the first An enhanced emotional semantic unit representation;
[0030] Step S34: Emotional semantic unit relationship perception modeling. Based on the enhanced emotional semantic unit representation, the potential relationships between different units are modeled. Potential relationships include adjacency relationships based on time order, association relationships based on semantic dependency, and interaction relationships based on emotional influence. Define the... With the The relationship weights between the enhanced sentiment semantic units are represented as follows:
[0031] ;
[0032] in, express and Relationship weights This represents a similarity calculation function based on an attention mechanism. For the first An enhanced emotional semantic unit represents, ,and ;
[0033] Step S35: Construction of Structured Emotion Semantic Unit Representation. Based on the emotion semantic units and the association representations between them, a structured representation is constructed, including a set of emotion semantic unit nodes, unit enhanced feature representations, and a relationship matrix between units, forming a structured emotion semantic association representation, as shown below:
[0034] ;
[0035] ;
[0036] ;
[0037] ;
[0038] in, Represents a set of emotion semantic unit nodes. Representation of unit-enhanced feature representation, The matrix representing the relationships between elements. Represents a structured sentiment semantic graph.
[0039] Furthermore, step S4 specifically includes the following steps:
[0040] Step S41: Construction of a dual-graph structure. Based on the set of emotion semantic unit nodes, two complementary graph structures are constructed, including an emotion interaction graph and an emotion evolution graph, specifically including the following:
[0041] (1) Emotional interaction diagram, represented as Model the lateral emotional influence relationships between different emotional semantic units, with the edge set defined as follows:
[0042] ;
[0043] in, Represents the set of edges in the emotion interaction graph. Indicates the relation threshold. Indicates the temporal position of the semantic unit, corresponding to the emotion semantic unit. , Indicates time window constraints;
[0044] (2) Emotional evolution diagram, represented as The model depicts the progressive evolution of emotional semantic units over time, with the edges defined as follows:
[0045] ;
[0046] in, Represents the set of edges in the emotion evolution graph;
[0047] Step S42: Dual-graph feature initialization, initializing the feature vector for each emotion semantic unit node:
[0048] ;
[0049] in, This represents the initial representation of the node. This represents the enhanced sentiment semantic unit representation of the S3 output;
[0050] Step S43: Graph attention propagation. Perform multi-head attention propagation on the emotion interaction graph, as shown below:
[0051] ;
[0052] in, This represents the updated emotional semantic unit node representation of the emotional interaction graph. Represents the set of neighbor nodes in the interaction graph. This represents a learnable linear transformation matrix. Indicates attention weights, Represents a non-linear activation function;
[0053] Attention weights are defined as follows:
[0054] ;
[0055] in, Indicates to All adjacent nodes are normalized;
[0056] The temporal propagation on the emotion evolution diagram is represented as follows:
[0057] ;
[0058] in, This represents the updated emotional semantic unit node representation in the emotional evolution graph. Represents the time-propagation weight matrix. Indicates the bias term. Indicates the previous emotional state;
[0059] Step S44: Dual-graph information fusion. The updated emotional semantic unit node representations of the emotional interaction graph and the emotional evolution graph are fused to obtain the fused emotional semantic unit node representations. The formula used is as follows:
[0060] ;
[0061] in, This represents the fused emotion semantic unit node representation. , representing the learnable coefficient;
[0062] Step S45: Cross Figure 1 For consistency constraint optimization, a consistency constraint loss function is introduced, as follows:
[0063] ;
[0064] in, Indicates double Figure 1 Induced loss;
[0065] Step S46: Global sentiment sequence modeling. Sequence modeling is performed on the fused sentiment semantic unit node representations to obtain the global sentiment representation, as follows:
[0066] ;
[0067] in, GRU represents the global sentiment semantic representation and models long-range sentiment dependencies;
[0068] Step S47: Generate feature output for emotion-based speech, mapping the global emotion semantic representation to a dynamic emotion representation, as shown below:
[0069] ;
[0070] in, Indicates dynamic emotion expression, This represents the mapping matrix.
[0071] The beneficial effects achieved by the present invention using the above solution are as follows:
[0072] (1) By constructing emotional semantic units and introducing an emotional contribution evaluation mechanism, this invention achieves fine-grained emotion modeling of speech semantic structure, which can accurately identify the differences in the contribution of different semantic segments in a sentence to the overall emotional expression, thereby significantly improving the accuracy and controllability of emotional speech generation.
[0073] (2) This invention constructs a dual-graph structure of emotion interaction graph and emotion evolution graph, and combines graph neural network to jointly model the interaction relationship and evolution process between emotion semantic units, thereby realizing the structured expression of complex dynamic change process of emotion, thereby enhancing the performance of generated speech in terms of emotional coherence and naturalness. Attached Figure Description
[0074] Figure 1 This is a flowchart illustrating a digital human emotional speech generation method based on emotion semantic modeling provided by the present invention.
[0075] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof. Detailed Implementation
[0076] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0077] In the description of this invention, it should be understood that the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.
[0078] Example 1, see Figure 1 This invention provides a method for generating emotional speech in digital humans based on emotion semantic modeling, the method comprising the following steps:
[0079] Step S1: Multimodal data preprocessing, obtaining multi-source input data for emotional speech generation, including text data, speech data and lip visual data corresponding to speech, and performing standardization processing, including alignment, noise reduction, segmentation and feature normalization processing, to obtain a multimodal input data sequence;
[0080] Step S2: Multimodal semantic feature extraction. Feature extraction is performed on the multimodal input data sequence to obtain textual semantic features, acoustic features, and visual features. A unified semantic representation is constructed through a cross-modal alignment mechanism.
[0081] Step S3: Emotion semantic unit segmentation. Based on the unified semantic representation, the multimodal input data sequence is structurally segmented to obtain emotion semantic units. For each emotion semantic unit, the corresponding emotion features are extracted to construct an emotion semantic association representation.
[0082] Step S4: Based on graph structure, emotional interaction is carried out by constructing an emotional relationship graph based on emotional semantic units and emotional semantic association representations. By modeling the interaction relationships between different emotional semantic units and the evolution process of individual emotional semantic units, the emotional state is structurally modeled to obtain dynamic emotional representation.
[0083] Step S5: Generating speech expression parameters. Based on the unified semantic representation and dynamic emotion representation, cross-modal fusion processing is performed to generate expression parameters for speech synthesis, including pitch, duration, energy, and prosodic information.
[0084] Step S6: Emotional speech generation. Based on the expression parameters used for speech synthesis, an emotional speech signal is generated using a speech synthesis model. The generated speech signal is then enhanced to obtain an enhanced speech signal.
[0085] Step S7: Digital human emotion collaboration drive, based on the enhanced emotional speech signal, drive the digital human to perform synchronized control of facial expressions and lip movements.
[0086] Example 2, based on the above example, step S3 specifically includes the following steps:
[0087] Step S31: Semantic sequence structure parsing. The multimodal input data sequence corresponding to the unified semantic representation undergoes structure parsing and adaptive segmentation based on semantic changes to generate a candidate semantic unit set. The multimodal input data sequence is represented as follows:
[0088] ;
[0089] in, To unify the semantic representation of the corresponding multimodal feature sequences, For the first A fused feature vector of a time step or semantic segment The sequence length;
[0090] The sequence is segmented using a semantic change detection function to obtain a set of candidate semantic units, as shown below:
[0091] ;
[0092] in, For the set of candidate emotion semantic units, For the first Each of the candidate emotion semantic units Depend on The subsequences constitute the composition. This represents the total number of candidate emotion semantic units;
[0093] Step S32: Identify the dominant emotion unit. Evaluate the emotion contribution of the candidate emotion semantic unit set, set an emotion contribution threshold, and update the candidate emotion semantic units with emotion contributions higher than the threshold into the emotion semantic unit set, represented as follows: The unit with the highest emotional contribution is defined as the emotional anchor unit. The emotional contribution of each emotional semantic unit is as follows:
[0094] ;
[0095] in, For the first The emotional contribution of each candidate emotional semantic unit. A function for calculating the contribution of emotion. For the first The textual semantic feature vector of each candidate emotion semantic unit, For the first The acoustic feature vectors of candidate emotion semantic units, For the first Visual feature vectors of candidate emotion semantic units, Indicates the threshold for emotional contribution;
[0096] The emotion-dominant unit is defined as follows:
[0097] ;
[0098] in, As an emotional anchor unit, Select the candidate emotion semantic unit with the highest emotion contribution. This represents the number of candidate emotion semantic units;
[0099] Step S33: Emotional semantic unit category identification. Based on the semantic function and emotional expression role of emotional semantic units in sentences, the set of emotional semantic units is categorized, and corresponding category embedding vectors are assigned to different categories, including emotional expression units, tone regulation units, and semantic content units. Multimodal feature representations are extracted for each emotional semantic unit, and enhanced representations are performed by combining the category embedding vectors, as shown below:
[0100] ;
[0101] in, Indicates the first Embedding vectors corresponding to the categories of each emotion semantic unit Indicates the first Category identification function for each semantic unit, Indicates the first An enhanced emotional semantic unit representation;
[0102] Step S34: Emotional semantic unit relationship perception modeling. Based on the enhanced emotional semantic unit representation, the potential relationships between different units are modeled. Potential relationships include adjacency relationships based on time order, association relationships based on semantic dependency, and interaction relationships based on emotional influence. Define the... With the The relationship weights between the enhanced sentiment semantic units are represented as follows:
[0103] ;
[0104] in, express and Relationship weights This represents a similarity calculation function based on an attention mechanism. For the first An enhanced emotional semantic unit represents, ,and ;
[0105] Step S35: Construction of Structured Emotion Semantic Unit Representation. Based on the emotion semantic units and the association representations between them, a structured representation is constructed, including a set of emotion semantic unit nodes, unit enhanced feature representations, and a relationship matrix between units, forming a structured emotion semantic association representation, as shown below:
[0106] ;
[0107] ;
[0108] ;
[0109] ;
[0110] in, Represents a set of emotion semantic unit nodes. Representation of unit-enhanced feature representation, The matrix representing the relationships between elements. Represents a structured sentiment semantic graph.
[0111] In this embodiment, the code used is as follows:
[0112] import torch
[0113] import torch.nn as nn
[0114] import torch.nn.functional as F
[0115] class EmotionSemanticUnitBuilder:
[0116] def __init__(self, threshold=0.6, window_size=5):
[0117] self.threshold = threshold
[0118] self.window_size = window_size
[0119] # Embedding of Three Types of Semantic Units
[0120] self.type_embedding = {
[0121] "emotion": torch.randn(128),
[0122] "tone": torch.randn(128),
[0123] "content": torch.randn(128)
[0124] }
[0125] # -----------------------------
[0126] # 1. Semantic Change Detection
[0127] # -----------------------------
[0128] def segment_units(self, multimodal_seq):
[0129] """
[0130] multimodal_seq: [T, D]
[0131] return: list of units (each unit is list of indices)
[0132] """
[0133] T = len(multimodal_seq)[[ID=?]] [[ID=?]]
[0134] units = [][[ID=?]] [[ID=?]]
[0135] start = 0[[ID=?]] [[ID=?]]
[0136] for t in range(1, T):[[ID=?]] [[ID=?]]
[0137] diff = torch.norm(multimodal_seq[t] - multimodal_seq[t -1])[[ID=?]] [[ID=?]]
[0138] if diff > self.threshold:[[ID=?]] [[ID=?]]
[0139] units.append(list(range(start, t)))[[ID=?]] [[ID=?]]
[0140] start = t[[ID=?]] [[ID=?]]
[0141] units.append(list(range(start, T)))[[ID=?]] [[ID=?]]
[0142] return units[[ID=?]] [[ID=?]]
[0143] # -----------------------------[[ID=?]] [[ID=?]]
[0144] # 2. Emotional Contribution Calculation[[ID=?]] [[ID=?]]
[0145] # -----------------------------[[ID=?]] It seems there are some tags like ,
[0134] etc. that are not in the clear format you provided in the example. I've translated as best as possible while keeping those tags intact. If you can clarify their proper form, it would be better for a more accurate translation.
[0146] def compute_score(self, text_feat, audio_feat, visual_feat):
[0147] """
[0148] Simplified fusion scoring function
[0149] """
[0150] score = (
[0151] torch.mean(text_feat) * 0.4 +
[0152] torch.mean(audio_feat) * 0.4 +
[0153] torch.mean(visual_feat) * 0.2 )
[0155] return score
[0156] # -----------------------------
[0157] # 3. Constructing Semantic Units
[0158] # -----------------------------
[0159] def build_units(self, multimodal_seq, text, audio, visual):
[0160] segments = self.segment_units(multimodal_seq)
[0161] units = []
[0162] scores = []
[0163] for seg in segments:
[0164] t_feat = text[seg].mean(dim=0)
[0165] a_feat = audio[seg].mean(dim=0)
[0166] v_feat = visual[seg].mean(dim=0)
[0167] score = self.compute_score(t_feat, a_feat, v_feat)
[0168] units.append({
[0169] "segment": seg,
[0170] "text": t_feat,
[0171] "audio": a_feat,
[0172] "visual": v_feat,
[0173] "score": score
[0174] })
[0175] scores.append(score)
[0176] return units, scores
[0177] # -----------------------------
[0178] # 4. Screening emotional semantic units + anchor points
[0179] # -----------------------------
[0180] def select_emotion_units(self, units):
[0181] valid_units = [u for u in units if u["score"] > self.threshold]
[0182] if len(valid_units) == 0:
[0183] valid_units = units
[0184] anchor_unit = max(valid_units, key=lambda x: x["score"])
[0185] return valid_units, anchor_unit
[0186] # -----------------------------
[0187] # 5. Type Enhancement
[0188] # -----------------------------
[0189] def enhance_units(self, units):
[0190] enhanced = []
[0191] for u in units:
[0192] score = u["score"]
[0193] # Simple Rule Classification
[0194] if score > 0.75:
[0195] utype = "emotion"
[0196] elif score > 0.55:
[0197] utype = "tone"
[0198] else:
[0199] utype = "content"
[0200] embed = self.type_embedding[utype]
[0201] fused = u["text"] + u["audio"] + u["visual"] + embed
[0202] enhanced.append({
[0203] "feature": fused,
[0204] "type": utype,
[0205] "score": score <管理>})
[0207] return enhanced
[0208] # -----------------------------
[0209] # 6. Relationship Matrix
[0210] # -----------------------------
[0211] def build_relation_matrix(self, enhanced_units):
[0212] n = len(enhanced_units)
[0213] matrix = torch.zeros(n, n)
[0214] for i in range(n):
[0215] for j in range(n):
[0216] if i == j:
[0217] continue
[0218] fi = enhanced_units[i]["feature"]
[0219] fj = enhanced_units[j]["feature"]
[0220] sim = F.cosine_similarity(fi, fj, dim=0)
[0221] # Time Decay
[0222] time_decay = 1.0 / (abs(i - j) + 1)
[0223] matrix[i][j] = sim * time_decay
[0224] return matrix
[0225] # -----------------------------
[0226] # 7. Overall Process
[0227] # -----------------------------
[0228] def forward(self, multimodal_seq, text, audio, visual):
[0229] units, scores = self.build_units(multimodal_seq, text, audio,visual)
[0230] valid_units, anchor = self.select_emotion_units(units)
[0231] enhanced = self.enhance_units(valid_units)
[0232] relation_matrix = self.build_relation_matrix(enhanced)
[0233] return {
[0234] "units": valid_units,
[0235] "anchor": anchor,
[0236] "enhanced_features": enhanced,
[0237] "relation_matrix": relation_matrix
[0238] }
[0239] Example 3, based on the above examples, specifically includes the following steps in step S4:
[0240] Step S41: Construction of a dual-graph structure. Based on the set of emotion semantic unit nodes, two complementary graph structures are constructed, including an emotion interaction graph and an emotion evolution graph. Specifically, this includes the following:
[0241] (1) Emotional interaction diagram, represented as Model the lateral emotional influence relationships between different emotional semantic units, with the edge set defined as follows:
[0242] ;
[0243] in, Represents the set of edges in the emotion interaction graph. Indicates the relation threshold. Indicates the temporal position of the semantic unit, corresponding to the emotion semantic unit. , Indicates time window constraints;
[0244] (2) Emotional evolution diagram, represented as The model depicts the progressive evolution of emotional semantic units over time, with the edges defined as follows:
[0245] ;
[0246] in, Represents the set of edges in the emotion evolution graph;
[0247] Step S42: Dual-graph feature initialization, initializing the feature vector for each emotion semantic unit node:
[0248] ;
[0249] in, This represents the initial representation of the node. This represents the enhanced sentiment semantic unit representation of the S3 output;
[0250] Step S43: Graph attention propagation. Perform multi-head attention propagation on the emotion interaction graph, as shown below:
[0251] ;
[0252] in, This represents the updated emotional semantic unit node representation of the emotional interaction graph. Represents the set of neighbor nodes in the interaction graph. This represents a learnable linear transformation matrix. Indicates attention weights, Represents a non-linear activation function;
[0253] Attention weights are defined as follows:
[0254] ;
[0255] in, Indicates to All adjacent nodes are normalized;
[0256] The temporal propagation on the emotion evolution diagram is represented as follows:
[0257] ;
[0258] in, This represents the updated emotional semantic unit node representation in the emotional evolution graph. Represents the time-propagation weight matrix. Indicates the bias term. Indicates the previous emotional state;
[0259] Step S44: Dual-graph information fusion. The updated emotional semantic unit node representations of the emotional interaction graph and the emotional evolution graph are fused to obtain the fused emotional semantic unit node representations. The formula used is as follows:
[0260] ;
[0261] in, This represents the fused emotion semantic unit node representation. , representing the learnable coefficient;
[0262] Step S45: Cross Figure 1 For consistency constraint optimization, a consistency constraint loss function is introduced, as follows:
[0263] ;
[0264] in, Indicates double Figure 1 Induced loss;
[0265] Step S46: Global sentiment sequence modeling. Sequence modeling is performed on the fused sentiment semantic unit node representations to obtain the global sentiment representation, as follows:
[0266] ;
[0267] in, GRU represents the global sentiment semantic representation and models long-range sentiment dependencies;
[0268] Step S47: Generate feature output for emotion-based speech, mapping the global emotion semantic representation to a dynamic emotion representation, as shown below:
[0269] ;
[0270] in, Indicates dynamic emotion expression, This represents the mapping matrix.
[0271] In this embodiment, the code used is as follows:
[0272] import torch
[0273] import torch.nn as nn
[0274] import torch.nn.functional as F
[0275] # =========================
[0276] # Graph Attention Layer (corresponding to S43 Interaction Graph)
[0277] # =========================
[0278] class GraphAttentionLayer(nn.Module):
[0279] def __init__(self, in_dim, out_dim):
[0280] super().__init__()
[0281] self.W = nn.Linear(in_dim, out_dim, bias=False)
[0282] self.a = nn.Linear(2 * out_dim, 1, bias=False)
[0283] def forward(self, x, adj):
[0284] # x: [N, d]
[0285] h = self.W(x) # [N, out_dim]
[0286] N = h.size(0)
[0287] # Constructing Attention Input
[0288] h_i = h.unsqueeze(1).repeat(1, N, 1)
[0289] h_j = h.unsqueeze(0).repeat(N, 1, 1)
[0290] a_input = torch.cat([h_i, h_j], dim=-1) # [N, N, 2*out_dim]
[0291] e = self.a(a_input).squeeze(-1) # [N, N]
[0292] # mask non-adjacent nodes
[0293] e = e.masked_fill(adj == 0, float('-inf'))
[0294] attention = F.softmax(e, dim=1)
[0295] h_prime = torch.matmul(attention, h)
[0296] return F.relu(h_prime)
[0297] # =========================
[0298] # Emotion Evolution Module (corresponding to Evolution Diagram S43)
[0299] # =========================
[0300] class EvolutionModule(nn.Module):
[0301] def __init__(self, dim):
[0302] super().__init__()
[0303] self.W = nn.Linear(dim, dim)
[0304] def forward(self, x):
[0305] # x: [N, d]
[0306] out = []
[0307] for i in range(len(x)):
[0308] if i == 0:
[0309] out.append(x[i])
[0310] else:
[0311] out.append(F.relu(self.W(x[i-1])))
[0312] return torch.stack(out)
[0313] # =========================
[0314] # Dual-graph model (core S4)
[0315] # =========================
[0316] class DualGraphEmotionModel(nn.Module):
[0317] def __init__(self, dim=256, hidden=128):
[0318] super().__init__()
[0319] # Interactive Diagram
[0320] self.gat = GraphAttentionLayer(dim, hidden)
[0321] # Evolutionary diagram
[0322] self.evo = EvolutionModule(dim)
[0323] # Fusion parameters (λ)
[0324] self.lambda_param = nn.Parameter(torch.tensor(0.5))
[0325] # GRU (S46)
[0326] self.gru = nn.GRU(hidden, hidden, batch_first=True)
[0327] # Output layer (S47)
[0328] self.fc = nn.Linear(hidden, 64)
[0329] def forward(self, x, adj):
[0330] """
[0331] x: [N, d] (h_i')
[0332] adj: [N, N] (adjacency matrix after A_ij>τ)
[0333] """
[0334] # ===== S43: Interactive Diagram =====
[0335] x_int = self.gat(x, adj) # [N, hidden]
[0336] # ===== S43: Evolutionary Diagram =====
[0337] x_evo = self.evo(x) # [N, d]
[0338] x_evo = x_evo[:, :x_int.shape[1]] # Alignment dimension
[0339] # ===== S44: Fusion =====
[0340] lam = torch.sigmoid(self.lambda_param)
[0341] x_fused = lam * x_int + (1 - lam) * x_evo
[0342] # ===== S45: Consistency Loss =====
[0343] loss_cons = torch.mean((x_int - x_evo) ** 2)
[0344] # ===== S46: GRU =====
[0345] x_fused = x_fused.unsqueeze(0) # [1, N, hidden]
[0346] _, h_global = self.gru(x_fused)
[0347] h_global = h_global.squeeze(0) # [hidden]
[0348] # ===== S47: Output =====
[0349] z = self.fc(h_global) #
[64]
[0350] return z, loss_cons.
[0351] Here is an example of usage:
[0352] # Assume there are 5 emotional semantic units
[0353] N = 5
[0354] dim = 256
[0355] # Input: S3 Output
[0356] x = torch.randn(N, dim)
[0357] # Construct the adjacency matrix (A_ij>τ)
[0358] adj = torch.randint(0, 2, (N, N))
[0359] adj.fill_diagonal_(0)
[0360] model = DualGraphEmotionModel(dim=256, hidden=128)
[0361] z, loss = model(x, adj)
[0362] print("Emotional Control Vector:", z.shape) #
[64]
[0363] print("Consistency loss:", loss.item()).
[0364] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0365] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
[0366] The present invention and its embodiments have been described above. This description is not restrictive, and the accompanying drawings are only one embodiment of the present invention; the actual structure is not limited thereto. In conclusion, if those skilled in the art are inspired by this description and design similar structures and embodiments without departing from the spirit of the invention, such designs should fall within the protection scope of the present invention.
Claims
1. A method for generating emotional speech in digital humans based on emotion semantic modeling, characterized in that: The method includes the following steps: Step S1: Multimodal data preprocessing, obtaining multi-source input data for emotional speech generation, including text data, speech data and lip visual data corresponding to speech, and performing standardization processing, including alignment, noise reduction, segmentation and feature normalization processing, to obtain a multimodal input data sequence; Step S2: Multimodal semantic feature extraction. Feature extraction is performed on the multimodal input data sequence to obtain textual semantic features, acoustic features, and visual features. A unified semantic representation is constructed through a cross-modal alignment mechanism. Step S3: Emotion semantic unit segmentation. Based on the unified semantic representation, the multimodal input data sequence is structurally segmented to obtain emotion semantic units. For each emotion semantic unit, the corresponding emotion features are extracted to construct an emotion semantic association representation. Step S4: Based on graph structure, emotional interaction is carried out by constructing an emotional relationship graph based on emotional semantic units and emotional semantic association representations. By modeling the interaction relationships between different emotional semantic units and the evolution process of individual emotional semantic units, the emotional state is structurally modeled to obtain dynamic emotional representation. Step S5: Generating speech expression parameters. Based on the unified semantic representation and dynamic emotion representation, cross-modal fusion processing is performed to generate expression parameters for speech synthesis. Step S6: Emotional speech generation. Based on the expression parameters used for speech synthesis, an emotional speech signal is generated using a speech synthesis model. The generated speech signal is then enhanced to obtain an enhanced speech signal. Step S7: Digital human emotion collaboration drive, based on the enhanced emotional speech signal, drive the digital human to perform synchronized control of facial expressions and lip movements.
2. The digital human emotional speech generation method based on emotion semantic modeling according to claim 1, characterized in that: Step S3 specifically includes the following steps: Step S31: Semantic sequence structure parsing. The multimodal input data sequence corresponding to the unified semantic representation undergoes structure parsing and adaptive segmentation based on semantic changes to generate a candidate semantic unit set. The multimodal input data sequence is represented as follows: ; in, To unify the semantic representation of the corresponding multimodal feature sequences, For the first A fused feature vector of a time step or semantic segment The sequence length; The sequence is segmented using a semantic change detection function to obtain a set of candidate semantic units, as shown below: ; in, For the set of candidate emotion semantic units, For the first Each of the candidate emotion semantic units Depend on The subsequences constitute the composition. This represents the total number of candidate emotion semantic units; Step S32: Identify the dominant emotion unit. Evaluate the emotion contribution of the candidate emotion semantic unit set, set an emotion contribution threshold, and update the candidate emotion semantic units with emotion contributions higher than the threshold into the emotion semantic unit set, represented as follows: The unit with the highest emotional contribution is defined as the emotional anchor unit. The emotional contribution of each emotional semantic unit is as follows: ; in, For the first The emotional contribution of each candidate emotional semantic unit. A function for calculating the contribution of emotion. For the first The textual semantic feature vector of each candidate emotion semantic unit, For the first The acoustic feature vectors of candidate emotion semantic units, For the first Visual feature vectors of candidate emotion semantic units, Indicates the threshold for emotional contribution; The emotion-dominant unit is defined as follows: ; in, As an emotional anchor unit, Select the candidate emotion semantic unit with the highest emotion contribution. This represents the number of candidate emotion semantic units; Step S33: Emotional semantic unit category identification. Based on the semantic function and emotional expression role of emotional semantic units in sentences, the set of emotional semantic units is categorized, and corresponding category embedding vectors are assigned to different categories, including emotional expression units, tone regulation units, and semantic content units. Multimodal feature representations are extracted for each emotional semantic unit, and enhanced representations are performed by combining the category embedding vectors, as shown below: ; in, Indicates the first Embedding vectors corresponding to the categories of each emotion semantic unit Indicates the first Category identification function for each semantic unit, Indicates the first An enhanced emotional semantic unit representation; Step S34: Emotional semantic unit relationship perception modeling. Based on the enhanced emotional semantic unit representation, the potential relationships between different units are modeled. The potential relationships include adjacency relationships based on time order, association relationships based on semantic dependence, and interaction relationships based on emotional influence, thus obtaining the association representation between emotional semantic units. Step S35: Construction of structured emotion semantic unit representation. Based on the emotion semantic units and the association representation between emotion semantic units, a structured representation is constructed, including the set of emotion semantic unit nodes, the unit enhanced feature representation, and the relationship matrix between units, forming a structured emotion semantic association representation.
3. The digital human emotional speech generation method based on emotion semantic modeling according to claim 2, characterized in that: Step S34 specifically includes the following: Emotional semantic unit relation perception modeling: Based on the enhanced emotional semantic unit representation, the potential relationships between different units are modeled. These potential relationships include adjacency relationships based on time order, association relationships based on semantic dependency, and interaction relationships based on emotional influence. The first... With the The relationship weights between the enhanced sentiment semantic units are represented as follows: ; in, express and Relationship weights This represents a similarity calculation function based on an attention mechanism. For the first An enhanced emotional semantic unit represents, ,and .
4. The digital human emotional speech generation method based on emotion semantic modeling according to claim 2, characterized in that: Step S35 specifically includes the following: The structured representation of emotional semantic units is constructed based on the relationship representation between emotional semantic units. This structured representation includes a set of emotional semantic unit nodes, enhanced feature representations of units, and a relationship matrix between units, forming a structured emotional semantic association representation, as shown below: ; ; ; ; in, Represents a set of emotion semantic unit nodes. Representation of unit-enhanced feature representation, The matrix representing the relationships between elements. Represents a structured sentiment semantic graph.
5. The digital human emotional speech generation method based on emotion semantic modeling according to claim 1, characterized in that: Step S4 specifically includes the following steps: Step S41: Construction of a dual-graph structure. Based on the set of emotion semantic unit nodes, two complementary graph structures are constructed, including an emotion interaction graph and an emotion evolution graph, specifically including the following: (1) Emotional interaction diagram, represented as Model the lateral emotional influence relationships between different emotional semantic units, with the edge set defined as follows: ; in, Represents the set of edges in the emotion interaction graph. Indicates the relation threshold. Indicates the temporal position of the semantic unit, corresponding to the emotion semantic unit. , Indicates time window constraints; (2) Emotional evolution diagram, represented as The model depicts the progressive evolution of emotional semantic units over time, with the edges defined as follows: ; in, Represents the set of edges in the emotion evolution graph; Step S42: Dual-graph feature initialization, initializing the feature vector for each emotion semantic unit node: ; in, This represents the initial representation of the node. This represents the enhanced sentiment semantic unit representation of the S3 output; Step S43: Graph attention propagation. Perform multi-head attention propagation on the emotion interaction graph, as shown below: ; in, This represents the updated emotional semantic unit node representation of the emotional interaction graph. Represents the set of neighbor nodes in the interaction graph. This represents a learnable linear transformation matrix. Indicates attention weights, Represents a non-linear activation function; Attention weights are defined as follows: ; in, Indicates to All adjacent nodes are normalized; The temporal propagation on the emotion evolution diagram is represented as follows: ; in, This represents the updated emotional semantic unit node representation in the emotional evolution graph. Represents the time-propagation weight matrix. Indicates the bias term. Indicates the previous emotional state; Step S44: Dual-graph information fusion. The updated emotional semantic unit node representations of the emotional interaction graph and the emotional evolution graph are fused to obtain the fused emotional semantic unit node representations. The formula used is as follows: ; in, This represents the fused emotion semantic unit node representation. , representing the learnable coefficient; Step S45: Cross-graph consistency constraint optimization, introducing a consistency constraint loss function, expressed as follows: ; in, This represents the consistency loss between two graphs; Step S46: Global sentiment sequence modeling. Sequence modeling is performed on the fused sentiment semantic unit node representations to obtain the global sentiment representation, as follows: ;… in, GRU represents the global sentiment semantic representation and models long-range sentiment dependencies; Step S47: Generate feature output for emotion-based speech, mapping the global emotion semantic representation to a dynamic emotion representation, as shown below: ; in, Indicates dynamic emotion expression, This represents the mapping matrix.