An intelligent digital human customer service system based on multi-modal driving
Patent Information
- Application Number
- CN202611024867.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-10
- Publication Date
- 2026-09-29
AI Technical Summary
例如,文本回应的语速变化、语音的抑扬顿挫与面部表情的起止时序之间缺乏基于语义内容的协同约束,容易产生口型动画滞后或情绪表达脱节等问题,破坏数字人的自然感和可信度
[0056]本发明通过多模态语义融合与情感驱动的联合建模显著提升了数字人客服的交互自然度。通过模态间语义关联度的动态计算滤除了冗余噪声,情感表达需求向量与初始语义贡献度的交叉匹配机制确保数字人的表情、语调及肢体动作与用户情绪状态高度契合,避免传统客服机械式回应带来的生硬感。表达权重分布对模态间语义关联度的反向调节实现了多模态特征的自适应重校准。当语音模态承载关键情感信息时,其语义贡献度自动增强,而冗余视觉特征被有效抑制,调节后语义贡献度与调节后语义关联度的耦合计算生成了时序同步约束矩阵,严格保证音频输出、口型动画及指令动作在毫秒级时间窗口内完全同步,彻底消除多通道输出的视觉滞后或音画不同步问题。基于时序约束的数字人表达控制参数与语义校正算法共同提升了交互输出的鲁棒性。回复内容的多模态转换过程融合了语义一致性校验与情感平滑处理,使数字人的语音语调、面部微表情及手势幅度呈现类人般的连贯变化。不同模态输出内容在语义校正环节的交叉验证机制有效避免了歧义表达,支持多用户并发场景下的稳定交互,显著改善了客服系统的服务能效与用户满意度。
Smart Images

Figure CN122840957A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent customer service technology, and in particular to an intelligent digital human customer service system based on multimodal driving. Background Technology
[0002] Current multimodal driving technologies in the field of intelligent digital human customer service mainly rely on the independent processing and simple concatenation of input data from each modality. Common practices include extracting feature vectors from text, speech, and visual modalities separately, forming a unified semantic representation through linear weighting or cascaded fusion, and then using this representation to identify user intent and generate response content. At the output end, the digital human typically converts the response content into synchronous signals for each modality based on a preset temporal template or fixed weight allocation, driving lip-syncing, speech, and facial expression animations. This workflow framework has been widely implemented in industrial applications, such as simple matching based on sentiment tags or adjusting expression parameters based on statistical rules.
[0003] However, the calculation of semantic correlation between modalities often uses static weighting or preset values, lacking the ability to dynamically adapt to the user's real-time emotional state and expressive needs. When user input contains contradictory or ambiguous multimodal information (such as a smiling voice accompanied by an angry tone), fixed correlations cannot accurately reflect the differentiated contributions of different modalities to the overall semantics, leading to a mismatch between the emotional expression of the digital human's response and the actual interactive context, resulting in a stiff user experience. The temporal synchronization control of multimodal output content usually relies on local timestamp alignment or simple linear interpolation, ignoring the coupling relationship between different modalities at the semantic level. For example, the lack of semantic content-based collaborative constraints between changes in speech rate in text responses, intonation, and the start and end times of facial expressions can easily lead to problems such as lagging lip-sync animation or disconnected emotional expression, undermining the naturalness and credibility of the digital human.
[0004] Limited by the aforementioned technical bottlenecks, existing solutions struggle to achieve emotionally coherent, naturally expressive, and precisely timed digital human customer service responses when dealing with complex and dynamic multimodal interaction scenarios. Summary of the Invention
[0005] This invention provides an intelligent digital human customer service system based on multimodal driving, which can solve the problems in the prior art.
[0006] A first aspect of this invention provides a multimodal-driven intelligent digital human customer service method, comprising:
[0007] Acquire user input data from at least two of the following modalities: text, speech, and visual; and extract feature vectors for each modality.
[0008] Calculate the semantic correlation degree between modalities, fuse the feature vectors of each modality based on the semantic correlation degree to generate a unified semantic representation, and extract the initial semantic contribution of each modality from the unified semantic representation;
[0009] Based on unified semantic representation, user intent and emotional state are identified. An emotional expression demand vector is constructed according to the emotional state. The emotional expression demand vector is cross-matched with the initial semantic contribution to generate expression weight distribution and digital human expression control parameters.
[0010] The semantic correlation between modalities is adjusted inversely based on the expression weight distribution, and the adjusted semantic contribution is obtained by re-integrating the feature vectors of each modality based on the adjusted semantic correlation.
[0011] The response content is obtained based on the user's intent. The adjusted semantic contribution and adjusted semantic relevance are coupled to generate a temporal synchronization constraint matrix. Under the constraints of the digital human expression control parameters and the temporal synchronization constraint matrix, the response content is multimodal converted and the output content of each modality is generated through semantic correction.
[0012] Multimodal output instructions are generated based on the output content of each modality to drive the digital human to perform interactions.
[0013] Calculate the semantic correlation degree between modalities, fuse the feature vectors of each modality based on the semantic correlation degree to generate a unified semantic representation, and extract the initial semantic contribution of each modality from the unified semantic representation, including:
[0014] Cross-modal attention is calculated for each modal feature vector to obtain the attention weight matrix of each modal feature vector to other modal feature vectors;
[0015] The modality missing and modality quality status in the user input data are detected, and the attention weight matrix is compensated and adjusted according to the modality missing and modality quality status to obtain the compensated attention weight matrix.
[0016] The bidirectional attention score between any two modal feature vectors is calculated based on the compensated attention weight matrix, and the semantic association degree between modalities is obtained by symmetric processing of the bidirectional attention score.
[0017] Using the semantic correlation degree between modalities as the fusion coefficient, the feature vectors of each modality are weighted and fused to generate a unified semantic representation;
[0018] The unified semantic representation is decomposed, and the projection intensity of the unified semantic representation on the feature vector of each modality is calculated. The projection intensity is normalized to obtain the initial semantic contribution of each modality.
[0019] An emotion expression demand vector is constructed based on the emotional state. This vector is then cross-matched with the initial semantic contribution to generate an expression weight distribution and digital human expression control parameters, including:
[0020] Extract the modal contribution values from the initial semantic contribution values, calculate the variance of each modal contribution value to obtain the emotional intensity value, calculate the covariance of each modal contribution value to obtain the emotional type value, and combine the emotional intensity value and the emotional type value to encode an emotional expression demand vector.
[0021] Modal collaboration constraint rules are obtained based on emotional state labels. The similarity between the emotional expression demand vector and the contribution value of each modality is calculated to obtain the modal matching degree. The modal matching degree is substituted into the modal collaboration constraint rules for cross-modal weighting to generate the expression weight distribution.
[0022] The modal emotional expression weights and emotional intensity values in the expression weight distribution are fused to obtain the modal activation intensity. The information entropy of the distribution formed by the activation intensity of each modality is calculated as the distribution balance. The adjustment coefficient is determined based on the distribution balance.
[0023] The modal activation intensities are weighted based on the adjustment coefficient and the emotion type value to obtain the adjusted modal activation intensities. The temporal change rate of the adjusted modal activation intensities is calculated to obtain the modal activation intensity change rate. The modal activation intensity change rate is used as the expression control parameter of the digital human.
[0024] Modal collaborative constraint rules are obtained based on sentiment state labels. The similarity between the sentiment expression demand vector and the contribution value of each modality is calculated to obtain the modal matching degree. The initial matching degree of each modality is substituted into the modal collaborative constraint rules for cross-modal weighting to generate the expression weight distribution, including:
[0025] Based on the emotion state label, the modal co-constraint rules are obtained by querying the preset emotion-modal mapping table. The modal co-constraint rules include inter-modal mutual exclusion coefficients and co-constraint coefficients.
[0026] The emotional expression demand vector is decomposed into emotional intensity components and emotional type components. The intensity similarity between the emotional intensity component and the contribution value of each modality and the directional similarity between the emotional type component and the contribution value of each modality are calculated. The intensity similarity and directional similarity are weighted and fused to obtain the modality matching degree.
[0027] Based on the modality matching degree, the intermodality matching degree difference value is calculated. The modality matching degree and the intermodality matching degree difference value are substituted into the intermodality mutual exclusion coefficient to perform cross-modality suppression weighting to obtain the suppression weight. The intermodality synergy coefficient is substituted into the intermodality enhancement weighting to obtain the enhancement weight. The suppression weight and the enhancement weight are combined to generate the expression weight distribution.
[0028] The semantic correlation between modalities is adjusted inversely based on the expression weight distribution. The adjusted semantic contribution is obtained by re-fusing the feature vectors of each modality based on the adjusted semantic correlation, including:
[0029] Suppression weights and enhancement weights are extracted from the expression weight distribution. The numerical deviation between the suppression weight and the initial semantic contribution of the corresponding modality is calculated to obtain the suppression bias, and the numerical deviation between the enhancement weight and the initial semantic contribution of the corresponding modality is calculated to obtain the enhancement bias.
[0030] A semantic relevance adjustment matrix is constructed based on suppression bias and enhancement bias.
[0031] The adjusted semantic correlation degree is obtained by performing matrix operations on the intermodal semantic correlation degree and the semantic correlation degree adjustment matrix.
[0032] The modal feature vectors are weighted and fused according to the adjusted semantic relevance to obtain the adjusted unified semantic representation. The adjusted unified semantic representation is decomposed, and the projection intensity of the adjusted unified semantic representation on each modal feature vector is calculated. The projection intensity is normalized to obtain the adjusted semantic contribution.
[0033] Based on user intent, the response content is obtained. The adjusted semantic contribution and adjusted semantic relevance are coupled to generate a temporal synchronization constraint matrix. Under the constraints of digital human expression control parameters and the temporal synchronization constraint matrix, the response content undergoes multimodal transformation and semantic correction to generate output content for each modality, including:
[0034] The unified semantic representation is semantically fused with the user intent to obtain the response semantic representation. The response semantic representation is then decoded to generate the response content, and the modal demand vector of the response content in each modal space is extracted.
[0035] The allocation deviation of the modal demand vector and the adjusted semantic contribution degree and the coordination deviation of the adjusted semantic correlation degree are calculated. The adjusted semantic contribution degree and the adjusted semantic correlation degree are coupled based on the allocation deviation and the coordination deviation to generate the time-series synchronization constraint matrix.
[0036] The modal demand vector is coupled with the rate of change of activation intensity of each modality in the digital human expression control parameters and the allocation deviation in the temporal synchronization constraint matrix to obtain the temporal parameters of each modal transition.
[0037] Under the constraints of digital human expression control parameters and timing synchronization constraint matrix, the response content is multimodal converted according to the timing parameters of each modality conversion to obtain the initial output content of each modality;
[0038] Extract cross-modal semantic consistency features of the initial modal output content, verify the consistency of the cross-modal semantic consistency features with the coordination deviation in the temporal synchronization constraint matrix to obtain semantic correction vectors, and perform semantic correction on the initial modal output content based on the semantic correction vectors to generate the output content of each modality.
[0039] Under the constraints of digital human expression control parameters and timing synchronization constraint matrix, the response content is multimodal transformed according to the timing parameters of each modality transformation to obtain the initial output content of each modality, including:
[0040] By performing a time-series mapping between the temporal transition parameters of each modality and the rate of change of activation intensity of each modality, the starting point and duration of content generation for each modality can be obtained;
[0041] The response content is segmented into temporal segments according to the starting point and duration of each modality, and the transformation weight distribution of each modality temporal segment content is calculated based on the allocation deviation.
[0042] The pre-conversion results of each modality are obtained by weighting the temporal segments according to the conversion weight distribution;
[0043] Extract cross-modal semantic deviation features between the pre-conversion results of each modality, perform deviation matching between the cross-modal semantic deviation features and the co-conversion deviation to obtain the cross-modal adjustment vector, and perform cross-modal consistency adjustment on the pre-conversion results of each modality based on the cross-modal adjustment vector to obtain the initial output content of each modality.
[0044] A second aspect of this invention provides a multimodal driven intelligent digital human customer service system, comprising:
[0045] A multimodal feature unit is used to acquire user input data from at least two modalities, including text modality, speech modality and visual modality, and extract feature vectors for each modality;
[0046] The semantic fusion unit is used to calculate the semantic correlation degree between modalities, fuse the feature vectors of each modality based on the semantic correlation degree between modalities to generate a unified semantic representation, and extract the initial semantic contribution of each modality from the unified semantic representation.
[0047] The emotion weighting unit is used to identify user intent and emotional state based on unified semantic representation, construct an emotion expression demand vector according to the emotional state, and cross-match the emotion expression demand vector with the initial semantic contribution to generate expression weight distribution and digital human expression control parameters.
[0048] The reverse adjustment unit is used to reverse adjust the semantic correlation between modalities according to the expression weight distribution, and to obtain the adjusted semantic contribution by re-fusing the feature vectors of each modality based on the adjusted semantic correlation.
[0049] The multimodal generation unit is used to obtain response content according to user intent, couple the adjusted semantic contribution degree and the adjusted semantic relevance degree to generate a temporal synchronization constraint matrix, and perform multimodal conversion on the response content under the constraints of digital human expression control parameters and temporal synchronization constraint matrix, and generate each modal output content through semantic correction.
[0050] The drive instruction unit is used to generate multimodal output instructions based on the output content of each modality, and drive the digital human to perform interactions.
[0051] A third aspect of the present invention provides an electronic device, comprising:
[0052] processor;
[0053] Memory used to store processor-executable instructions;
[0054] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0055] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0056] This invention significantly improves the naturalness of digital human customer service interaction through joint modeling of multimodal semantic fusion and emotion-driven approaches. Dynamic calculation of intermodal semantic correlation filters out redundant noise, and a cross-matching mechanism between the emotional expression demand vector and the initial semantic contribution ensures that the digital human's facial expressions, tone of voice, and body language highly match the user's emotional state, avoiding the stiffness of traditional customer service's mechanical responses. The inverse adjustment of intermodal semantic correlation by the expression weight distribution enables adaptive recalibration of multimodal features. When the voice modality carries key emotional information, its semantic contribution is automatically enhanced, while redundant visual features are effectively suppressed. The coupled calculation of the adjusted semantic contribution and the adjusted semantic correlation generates a temporal synchronization constraint matrix, strictly ensuring that audio output, lip-sync animation, and command actions are completely synchronized within a millisecond-level time window, completely eliminating visual lag or audio-visual asynchrony issues in multi-channel output. The temporal constraint-based digital human expression control parameters and semantic correction algorithm jointly enhance the robustness of the interactive output. The multimodal conversion process of the response content integrates semantic consistency verification and emotional smoothing, enabling the digital human's voice tone, facial micro-expressions, and gesture amplitude to exhibit consistent, human-like changes. The cross-validation mechanism of different modal output content in the semantic correction stage effectively avoids ambiguous expressions, supports stable interaction in multi-user concurrent scenarios, and significantly improves the service efficiency and user satisfaction of the customer service system. Attached Figure Description
[0057] Figure 1 This is a flowchart illustrating a multimodal-driven intelligent digital human customer service approach.
[0058] Figure 2 Flowchart for multimodal digital human response generation and semantic correction. Detailed Implementation
[0059] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0060] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0061] Figure 1 This is a flowchart illustrating the intelligent digital human customer service method based on multimodal driving according to an embodiment of the present invention.
[0062] A multimodal-driven intelligent digital human customer service method includes:
[0063] Acquire user input data from at least two of the following modalities: text, speech, and visual; and extract feature vectors for each modality.
[0064] Calculate the semantic correlation degree between modalities, fuse the feature vectors of each modality based on the semantic correlation degree to generate a unified semantic representation, and extract the initial semantic contribution of each modality from the unified semantic representation;
[0065] Based on unified semantic representation, user intent and emotional state are identified. An emotional expression demand vector is constructed according to the emotional state. The emotional expression demand vector is cross-matched with the initial semantic contribution to generate expression weight distribution and digital human expression control parameters.
[0066] The semantic correlation between modalities is adjusted inversely based on the expression weight distribution, and the adjusted semantic contribution is obtained by re-integrating the feature vectors of each modality based on the adjusted semantic correlation.
[0067] The response content is obtained based on the user's intent. The adjusted semantic contribution and adjusted semantic relevance are coupled to generate a temporal synchronization constraint matrix. Under the constraints of the digital human expression control parameters and the temporal synchronization constraint matrix, the response content is multimodal converted and the output content of each modality is generated through semantic correction.
[0068] Multimodal output instructions are generated based on the output content of each modality to drive the digital human to perform interactions.
[0069] Calculate the semantic correlation degree between modalities, fuse the feature vectors of each modality based on the semantic correlation degree to generate a unified semantic representation, and extract the initial semantic contribution of each modality from the unified semantic representation, including:
[0070] Cross-modal attention is calculated for each modal feature vector to obtain the attention weight matrix of each modal feature vector to other modal feature vectors;
[0071] The modality missing and modality quality status in the user input data are detected, and the attention weight matrix is compensated and adjusted according to the modality missing and modality quality status to obtain the compensated attention weight matrix.
[0072] The bidirectional attention score between any two modal feature vectors is calculated based on the compensated attention weight matrix, and the semantic association degree between modalities is obtained by symmetric processing of the bidirectional attention score.
[0073] Using the semantic correlation degree between modalities as the fusion coefficient, the feature vectors of each modality are weighted and fused to generate a unified semantic representation;
[0074] The unified semantic representation is decomposed, and the projection intensity of the unified semantic representation on the feature vector of each modality is calculated. The projection intensity is normalized to obtain the initial semantic contribution of each modality.
[0075] When performing cross-modal attention computation on feature vectors of each modality, the feature vector of one modality is used as the query, and the feature vectors of other modalities are used as the key and value, respectively. The attention response of that modality to the other modalities is calculated through a scaled dot product attention mechanism. Specifically, let there be a total The first mode, the second The eigenvectors of the modes are Then the first The first mode for the first Types of modes ( The original attention weights are obtained as follows: Mapped to query vector via linear transformation ,Will Mapped to key vectors respectively Sum value vector ,calculate and The dot product of these values, divided by the square root of the feature dimension, is then normalized using softmax to obtain the attention weight scalar. For all modal pairs By repeating the above process, the complete attention weight matrix can be constructed. , of which elements Indicates the first The first mode for the first Attention weights for each modality.
[0076] In real-world user input scenarios, factors such as the status of the data acquisition device, network transmission quality, or differences in user behavior can lead to situations where a particular modality is completely missing or its quality is locally degraded. For example, users may not provide visual input in pure text interaction scenarios, or the signal-to-noise ratio of the speech signal may be low due to background noise. To address these two types of issues, adjustments to the attention weight matrix are necessary. Compensation adjustments are made. For cases where a modality is completely missing, the attention weights for the corresponding rows and columns are reset to zero, while the weights for the remaining modalities are proportionally normalized to ensure that information fusion does not introduce invalid components due to missing modalities. For cases where modality quality deteriorates, a quality assessment score is introduced. ( The smaller the value, the higher the value. The worse the quality of a modality, the more the elements in the attention weight matrix related to that modality are multiplied by the corresponding quality factor. This reduces the influence of low-quality modalities in subsequent fusion. After the above compensation process, the compensated attention weight matrix is obtained. , of which elements It also reflects the strength of semantic association between modalities and the credibility of the modality itself.
[0077] Based on the compensated attention weight matrix Calculate the bidirectional attention score between any two modalities. The bidirectional attention score comprehensively considers the first... The first mode for the first attention of various modalities and the The first mode for the first attention of various modalities The symmetric intermodal semantic correlation is obtained by adding the two together and taking the average. The purpose of symmetry processing is to eliminate the inconsistencies caused by different attention calculation directions, so that arbitrary modal pairs... The degree of association between them is direction-independent, thus constructing a symmetric semantic association matrix. Its elements Description of the Type 1 mode and the first The strength of bidirectional semantic coupling between modalities. In typical scenarios where text, speech, and vision modalities coexist, for The symmetric matrix has diagonal elements that can be set to 1 to represent the self-correlation of the modality itself, while off-diagonal elements reflect the degree of cross-modal semantic correlation.
[0078] semantic relevance matrix As a fusion coefficient, the feature vectors of each modality are weighted and fused to generate a unified semantic representation. Specifically, for the first A modality, based on its semantic association with all other modalities. ( ) are weights for each modal feature vector Perform a weighted summation to obtain the first... Context-enhanced representation of various modalities The context-enhanced representations of all modalities are then concatenated and passed through a nonlinear mapping layer (such as a multilayer perceptron or a Transformer encoder) to obtain the final unified semantic representation. This unified semantic representation, during the fusion process, not only preserves the semantic information of each modality itself, but also introduces complementary cross-modal information through a semantic relevance weighting mechanism, thus enabling... It can comprehensively reflect the overall semantic content of user input in a high-dimensional semantic space. When a certain modality is of poor quality or has missing parts, the contribution of that modality to the unified semantic representation will also decrease accordingly because the corresponding weights have been suppressed by the compensated attention weight matrix, thus ensuring the robustness of the fusion result.
[0079] From a unified semantic representation When extracting the initial semantic contribution of each modality, for Perform modal decomposition and calculate In the modal eigenvectors Projected intensity in the direction. Projected intensity Defined as and inner product divided by The modulus length, i.e. This value reflects the first The projection intensity represents the amount of information carried by the feature vectors of each modality in the unified semantic representation. A higher projection intensity indicates a more significant contribution of that modality to the unified semantic representation. To ensure the comparability of the contributions of each modality and to satisfy probability normalization constraints, the projection intensities of all modalities are normalized to obtain the initial semantic contribution of each modality. The negative projection is truncated to avoid semantic confusion caused by negative contributions. Initial semantic contribution vector. satisfy This intuitively reflects the relative importance of each modality to the overall semantic understanding in the current user input, providing a quantitative basis for the cross-matching of subsequent sentiment expression demand vectors and initial semantic contributions, as well as the generation of expression weight distribution.
[0080] In practical deployments, the cross-modal attention computation module can be initialized using a pre-trained multimodal alignment model to accelerate convergence and improve generalization capabilities in low-resource scenarios. Modal quality assessment score. Online prediction can be achieved through a dedicated quality estimation subnetwork, or heuristic estimation can be performed based on engineering metrics such as signal-to-noise ratio and confidence score. The dimensions of the unified semantic representation should be reasonably configured according to the actual deployed hardware resources and response latency requirements, controlling computational overhead while ensuring semantic expressiveness, in order to meet the low-latency requirements of the digital human customer service system for real-time interaction.
[0081] An emotion expression demand vector is constructed based on the emotional state. This vector is then cross-matched with the initial semantic contribution to generate an expression weight distribution and digital human expression control parameters, including:
[0082] Extract the modal contribution values from the initial semantic contribution values, calculate the variance of each modal contribution value to obtain the emotional intensity value, calculate the covariance of each modal contribution value to obtain the emotional type value, and combine the emotional intensity value and the emotional type value to encode an emotional expression demand vector.
[0083] Modal collaboration constraint rules are obtained based on emotional state labels. The similarity between the emotional expression demand vector and the contribution value of each modality is calculated to obtain the modal matching degree. The modal matching degree is substituted into the modal collaboration constraint rules for cross-modal weighting to generate the expression weight distribution.
[0084] The modal emotional expression weights and emotional intensity values in the expression weight distribution are fused to obtain the modal activation intensity. The information entropy of the distribution formed by the activation intensity of each modality is calculated as the distribution balance. The adjustment coefficient is determined based on the distribution balance.
[0085] The modal activation intensities are weighted based on the adjustment coefficient and the emotion type value to obtain the adjusted modal activation intensities. The temporal change rate of the adjusted modal activation intensities is calculated to obtain the modal activation intensity change rate. The modal activation intensity change rate is used as the expression control parameter of the digital human.
[0086] After obtaining the initial semantic contribution vectors of each modality Next, it needs to be combined with the currently identified emotional state to construct a vector representation that reflects the user's emotional expression needs. The contribution value corresponding to each modality is extracted from the initial semantic contribution value, denoted as the... The contribution value of each mode is ,in Calculate the variance of all modal contributions to obtain the emotional intensity value. The calculation method is as follows: ,in The mean of the contributions from all modalities. Emotional intensity value. It reflects the dispersion of the contribution distribution of each modality in the current interaction: when the contribution of each modality is large, the emotional intensity value is high, indicating that the user has concentrated emotional expression in a certain modality; when the contribution of each modality tends to be balanced, the emotional intensity value is low, indicating that the emotional expression is relatively mild.
[0087] Further calculation of the covariance between the contribution values of each modality yields the sentiment type value. For modal With mode Covariance between After averaging the covariances of all modality pairs and normalizing the result, a scalar form of the sentiment type value is obtained. Sentiment type values capture the coordinated direction of changes in contribution across different modalities: if multiple modal contribution values show a positive correlation, then... Positive values correspond to positively reinforcing affective expressions (such as anger, excitement, and other highly arousing emotions); negative correlations indicate... Negative values correspond to contrasting emotional expressions (such as mixed emotions like ambivalence and hesitation). The emotional intensity value... With sentiment type value After concatenation, the vector is encoded into an emotional expression demand vector via linear mapping. This vector represents the current user's emotional expression needs in the high-dimensional emotional semantic space.
[0088] In generating emotional expression demand vectors Next, the corresponding modal collaboration constraint rules need to be obtained by combining the emotion state labels. The emotion state labels are output from the intent recognition stage and include the emotion category (e.g., happy, sad, angry, neutral, etc.) and its confidence level. Different emotion categories correspond to different modal collaboration constraint rules. For example, in the happy emotion state, the collaboration weight between the facial expression modality and the speech modality should be higher than that of the text modality; in the sad emotion state, the prosodic features of the speech modality should be given higher expression priority. The modal collaboration constraint rules are expressed as a constraint matrix. Stored in the form of , its elements Indicates the modality in the current emotional state. For modes The strength of collaborative constraints.
[0089] Calculate the emotional expression demand vector Contribution values of each mode The similarity is used to obtain the modality matching degree for each mode. Similarity calculation uses cosine similarity. Expand to Vectors of the same dimension are then subjected to a dot product normalization operation. Modal matching degree. This reflects the degree to which the contribution characteristics of this modality align with the current emotional expression needs. The matching degree of each modality... Substitute into the modal collaborative constraint rule matrix Cross-modal weighting is performed, specifically, for each mode... Its weighted expression weight The following is obtained by weighting and summing all modal matching degrees according to the constraint rule row vectors: Expression weights for all modalities Normalization is performed to obtain the expression weight distribution. ,satisfy The expression weight distribution characterizes the proportion that each modality should bear in the digital human's output expression under the current emotional state.
[0090] The modality sentiment expression weights in the expression weight distribution With emotional intensity value By fusing, the activation intensity of each mode is obtained. The fusion method is as follows: ,in This is a preset emotional intensity gain coefficient, used to control the degree to which emotional intensity amplifies activation intensity. When At higher levels, the overall activation intensity of each modality is stretched, thereby driving the digital human to produce stronger emotional expressions; when When the activation strength approaches zero, it degenerates into a normalized expression weight, ensuring the stability of the output.
[0091] Calculate the activation intensity of each mode The information entropy of the resulting distribution serves as the distribution's degree of equilibrium. : ,in For the normalized probability values of activation intensity for each mode, satisfying When the activation intensity of each mode tends to be equal, the information entropy... Reaching the maximum value indicates a balanced distribution; when the activation intensity of one mode is much higher than that of other modes, the information entropy... A smaller value indicates a concentrated distribution. This is based on the degree of distribution uniformity. Determine the adjustment coefficient :when When the distribution is relatively small (concentrated), Take a larger value to suppress the uniformity of expression caused by excessive concentration; when When the distribution is large (balanced), Take the smaller value to maintain the current distribution. Specifically, and The negative correlation mapping relationship can be achieved through a preset piecewise linear function or a lookup table.
[0092] Activation intensity of each mode Based on adjustment coefficient and sentiment type value The activation intensities of each mode are obtained after weighted calculation. : .
[0093] In the above formula, when At that time, the sentiment type value has a positive reinforcing effect on the activation intensity, further amplifying the activation level of each modality; when At this time, the emotion type value has an inhibitory or contrastive modulating effect on the activation intensity, causing differentiation in activation intensity between different modalities, thus presenting the complexity and layering of emotions in the digital human's output. (Adjusted activation intensity of each modality) While retaining the original emotional weight distribution, directional information of emotional types has been incorporated, making the digital human's expression closer to real emotional semantics.
[0094] Calculate the activation intensity of each mode after adjustment. The temporal variation rate is used to obtain the variation rate of activation intensity for each mode. The temporal rate of change is calculated by differentiating the adjusted activation intensity between consecutive time frames, reflecting the dynamic evolution speed of emotional expression over time. For the first... Each time step Defined as: .
[0095] in and These represent the adjusted activation intensities at the current and previous time steps, respectively. The rate of change of activation intensity for each mode. As the output of digital human expression control parameters, it is used for the time-driven control of digital human facial movements, speech rhythm and body movements in the subsequent multimodal conversion stage, so as to ensure that digital human has natural and smooth dynamic change characteristics in the process of emotional expression, rather than static fixed expressions or tones.
[0096] Modal collaborative constraint rules are obtained based on sentiment state labels. The similarity between the sentiment expression demand vector and the contribution value of each modality is calculated to obtain the modal matching degree. The initial matching degree of each modality is substituted into the modal collaborative constraint rules for cross-modal weighting to generate the expression weight distribution, including:
[0097] Based on the emotion state label, the modal co-constraint rules are obtained by querying the preset emotion-modal mapping table. The modal co-constraint rules include inter-modal mutual exclusion coefficients and co-constraint coefficients.
[0098] The emotional expression demand vector is decomposed into emotional intensity components and emotional type components. The intensity similarity between the emotional intensity component and the contribution value of each modality and the directional similarity between the emotional type component and the contribution value of each modality are calculated. The intensity similarity and directional similarity are weighted and fused to obtain the modality matching degree.
[0099] Based on the modality matching degree, the intermodality matching degree difference value is calculated. The modality matching degree and the intermodality matching degree difference value are substituted into the intermodality mutual exclusion coefficient to perform cross-modality suppression weighting to obtain the suppression weight. The intermodality synergy coefficient is substituted into the intermodality enhancement weighting to obtain the enhancement weight. The suppression weight and the enhancement weight are combined to generate the expression weight distribution.
[0100] After obtaining the initial semantic contribution vector and sentiment state labels, it is necessary to establish a precise mapping relationship between sentiment expression needs and the output capabilities of each modality to generate a reasonable expression weight distribution. The sentiment state labels are output from the intent recognition and sentiment analysis stages, carrying discrete category information such as "anger," "anxiety," "satisfaction," and "confusion." Based on these labels, a pre-defined sentiment-modality mapping table is queried. This mapping table, constructed by domain experts using extensive interactive corpora during system initialization, stores the modal collaboration constraint rules corresponding to each sentiment state. The modal collaboration constraint rules are organized in matrix form, containing two key parameters: inter-modal mutual exclusion coefficients and inter-modal collaboration coefficients. The mutual exclusion coefficient describes the degree of expression conflict generated when two modalities are simultaneously activated at high intensity. For example, when conveying the sentiment of "calm reassurance," there is a significant mutual exclusion relationship between the high-intensity body language modality and the low-speed speech modality. The collaboration coefficient describes the expression gain generated when two modalities are jointly activated. For example, when conveying the sentiment of "warm welcome," there is a strong collaboration effect between facial smiles and positive tone of voice. By querying the sentiment-modality mapping table, we can obtain the complete modal collaborative constraint rules corresponding to the current sentiment state label, which provides a constraint basis for subsequent cross-modal weighted calculations.
[0101] Emotional expression needs vector Includes emotional intensity component And emotional type component These two components respectively encode the "intensity of emotion" and the "directional semantics of emotion." Emotion Intensity Component As a scalar, its range is normalized to [value range]. This reflects the level of activation in the current emotional state. Emotional type component. As a high-dimensional vector, it encodes the category orientation of emotion in the sentiment semantic space, such as a joint representation of multiple sentiment dimensions like the "positive-negative" axis and the "active-passive" axis. For the ... Each mode, its contribution value Similarly, it can be decomposed into an estimate of the expressive power of the intensity and orientation dimensions. Intensity similarity Through calculation and The normalized difference between the intensity of expression was obtained, specifically, ,in For the first Normalized projection of modal contributions onto the intensity dimension. Orientation similarity. By calculating the emotional type components With the The direction vector of a modality in the sentiment semantic space The cosine similarity between them is obtained, that is Intensity similarity and orientation similarity are fused using learnable weights. Perform weighted fusion to obtain the first Matching degree of each mode The calculation method is as follows ,in It is obtained through adaptive learning during the training phase based on the differences in sensitivity to intensity and direction of different emotion categories.
[0102] In obtaining the modality matching degree Next, the intermodal matching degree difference value is calculated to quantify the relative competitive relationship between modes. For modes... With mode The difference in their matching degree Defined as This value is a signed quantity; a positive value indicates a mode. relative modes Positive values indicate stronger adaptability to emotional expression, while negative values indicate the opposite. Matching degree difference matrix. elements It will be used in subsequent cross-modal weighted calculations for mutual exclusion suppression and synergistic enhancement as an important input signal for adjusting the weights of each modality.
[0103] The cross-modal suppression weighting process combines the modal matching degree of each mode. Difference in matching degree between modalities Substitute the intermodal mutual exclusion coefficients Perform the calculation. Mutual exclusion coefficient. The value is directly read from the modal collaborative constraint rules; the larger the value, the more modal... With mode The stronger the expression conflict between them. For modalities Its inhibition weights after being suppressed by other modes The result is obtained by summing the suppression amounts applied to all modes and then subtracting them from the original matching degree. The calculation method is as follows: ,in Ensure only when modality The matching degree is higher than that of the modality. Only when modality Inhibition is applied to prevent low-fit modalities from unreasonably suppressing high-fit modalities. This mechanism allows modalities with lower fit in emotion expression to be competitively suppressed by high-fit modalities, thereby reducing their expression intensity at the output end.
[0104] The cross-modal enhancement weighting process combines the matching degree of each modality. Intermodal coordination coefficient Combined calculation with weight enhancement. Synergy coefficient. Similarly, it is read from the modal collaborative constraint rules, reflecting the modal... With mode The contribution of joint activation to the overall emotional expression gain. For modality Its enhanced weight It is obtained by aggregating the contributions of all modes with which it has a cooperative relationship, and the calculation method is as follows: , where the indicator function Ensure only when modality Its adaptability is no less than that of modal Only then will it accept modal The synergistic gain prevents the synergistic effect of low-quality modalities from lowering the overall expression accuracy. The synergistic enhancement mechanism enables multiple highly adaptable modalities to reinforce each other in emotional expression, producing a "1+1>2" joint expression effect. Especially in complex emotional scenarios (such as "gratitude with a slight regret"), multimodal synergy can more accurately reproduce the subtle layers of emotion.
[0105] Suppress weights With enhanced weight Combine them to generate the first Comprehensive representation weights of various modalities The combination method uses a weighted summation. The fusion ratio coefficient Dynamically adjust based on the competitiveness of the current emotional state: for emotional scenarios with strong mutual exclusion (such as "anger"). Take a larger value to highlight the inhibitory effect; for emotional scenarios with strong synergy (such as "joy"). Smaller values are selected to amplify the collaborative gain. The combined expression weights of all modalities are softmax normalized to ensure the sum of the weights is 1, ultimately forming a complete expression weight distribution. The expression weight distribution is directly used to generate subsequent digital human expression control parameters, determining the relative intensity distribution of each output modality such as voice, facial expressions, and body movements in the final interactive response, thereby achieving multimodal expression output that is highly matched with the user's emotional state.
[0106] The semantic correlation between modalities is adjusted inversely based on the expression weight distribution. The adjusted semantic contribution is obtained by re-fusing the feature vectors of each modality based on the adjusted semantic correlation, including:
[0107] Suppression weights and enhancement weights are extracted from the expression weight distribution. The numerical deviation between the suppression weight and the initial semantic contribution of the corresponding modality is calculated to obtain the suppression bias, and the numerical deviation between the enhancement weight and the initial semantic contribution of the corresponding modality is calculated to obtain the enhancement bias.
[0108] A semantic relevance adjustment matrix is constructed based on suppression bias and enhancement bias.
[0109] The adjusted semantic correlation degree is obtained by performing matrix operations on the intermodal semantic correlation degree and the semantic correlation degree adjustment matrix.
[0110] The modal feature vectors are weighted and fused according to the adjusted semantic relevance to obtain the adjusted unified semantic representation. The adjusted unified semantic representation is decomposed, and the projection intensity of the adjusted unified semantic representation on each modal feature vector is calculated. The projection intensity is normalized to obtain the adjusted semantic contribution.
[0111] After obtaining the expression weight distribution, this distribution needs to be used to inversely adjust the semantic correlation between modalities, so that the subsequent modality fusion process can more accurately reflect the emotional expression needs. The expression weight distribution already includes inhibition weights for each modality. and enhanced weight These two types of weights represent the degree of suppression and enhancement of a modality's contribution to information, respectively. The core idea of reverse adjustment is: if the enhancement weight of a modality is much higher than its initial semantic contribution, it means that the current fusion process is underutilizing that modality and needs to be compensated at the semantic relevance level; conversely, if the suppression weight of a modality is higher than its initial semantic contribution, it means that the modality is overutilized in the current semantic fusion and its relevance strength in subsequent fusions needs to be reduced.
[0112] The method for calculating the suppression bias is as follows: for the first... Type of modality, suppress its weights Initial semantic contribution of this modality Subtraction yields the suppressed bias. .when When the expression weight distribution indicates that the mode needs to be further suppressed, that is, the contribution of the current initial fusion to the mode is still too high; when When this value is less than the suppression threshold, it indicates that the initial contribution of this mode is already below the suppression threshold, and no further suppression is needed. Similarly, the enhancement bias is calculated as follows: ,when When this occurs, it indicates that the modality's contribution to the initial fusion is insufficient, and it needs to be given higher attention weight in the adjusted semantic relevance; when When the initial contribution of the modality is sufficient to meet the enhancement requirements, no further improvement is needed. These two types of biases together characterize the gap between the current modality fusion state and the ideal sentiment expression state, and are the core basis for constructing the semantic relevance adjustment matrix.
[0113] Based on the suppression bias vector and enhanced bias vector Construct a semantic relevance adjustment matrix , of which The row corresponds to the first The modulus of correlation between a mode and all other modes. Specifically, matrix elements. The calculation comprehensively considers the modes Enhancement bias and modality The suppression bias is determined as follows: ,in To enhance the sensitivity coefficient for bias adjustment, Both the adjustment sensitivity coefficient and the variable are positive real numbers, used to suppress bias, and can be calibrated according to specific business scenarios. When this is the case, it indicates that modal enhancement is needed. For modes The semantic association strength; when When, it indicates that the strength of the association needs to be weakened; when When this value is zero, it indicates that the correlation between the modes does not need to be adjusted. Adjustment matrix. Dimensions and semantic association matrix between modalities They are the same, both are The real-valued matrix.
[0114] The intermodal semantic correlation matrix Adjustment matrix of semantic relevance Perform element-wise matrix addition to obtain the adjusted semantic relevance matrix. ,Right now To ensure the numerical reasonableness of the adjusted semantic relevance, the following steps are taken: Each element in the code is truncated to keep it within a preset valid range. Within this range, avoid negative correlations or extreme values that exceed the physical meaning. At the same time, for... Each row is normalized to ensure that the sum of the correlations of each modality with other modalities satisfies the normalization constraint, thereby guaranteeing the conservation of the total weight of each modality during the subsequent weighted fusion process. Adjusted semantic correlation matrix Compared to the original matrix At the semantic level, it is closer to the current needs of emotional expression and can guide the subsequent integration process to evolve in a direction that is more in line with the emotional state.
[0115] Based on the adjusted semantic relevance matrix For each modal feature vector Weighted fusion is performed to obtain a unified semantic representation after adjustment. During the integration process, the first Context-enhanced representation of various modalities By processing all modal feature vectors according to The Middle The weighted sum of the correlation between rows is obtained, i.e. ,in For the adjusted semantic relevance matrix The Line 1 Column elements, For the first The value vector obtained by linear transformation of each modality. Unified semantic representation after adjustment. It is obtained by concatenating or weighting the context-enhanced representations of all modalities and then passing them through a nonlinear mapping layer, so that... It can comprehensively reflect the joint expressive ability of various modalities under the modified association structure in the semantic space.
[0116] Unified semantic representation after adjustment Perform mode decomposition and calculate its projection intensity along the eigenvector directions of each mode. For the ... Types of modes, projected intensity Defined as and The scalar value of the inner product between them after modulus normalization, i.e. The projection intensity reflects the adjusted unified semantic representation in the first... The magnitude of the component along the feature direction of a modality; a higher value indicates a more significant contribution of that modality to the adjusted fusion result. Since the feature vector dimensions and units of different modalities may differ, the projection intensity... Using cosine similarity naturally eliminates the influence of dimensions, ensuring the fairness of cross-modal comparisons.
[0117] Projected intensity of all modes ( The normalization process is performed to obtain the adjusted semantic contribution. The calculation method is as follows During normalization, the negative projection intensity is truncated to zero to ensure that the semantic contribution is non-negative, consistent with the physical meaning of contribution. The normalized adjusted semantic contribution vector is shown below. The constraint that the sum of all components is 1 can be directly used as the modal weight input in the subsequent construction of the temporal synchronization constraint matrix. This is related to the initial semantic contribution vector. In comparison, the adjusted semantic contribution vector It has fully incorporated the regulatory information of emotional expression needs, and can more accurately reflect the actual contribution ratio of each modality to the final output in a specific emotional state, providing a more accurate weight basis for the generation of subsequent multimodal output content.
[0118] like Figure 2 As shown, Figure 2 This is a flowchart of the multimodal digital human response generation and semantic correction process in this embodiment.
[0119] Based on user intent, the response content is obtained. The adjusted semantic contribution and adjusted semantic relevance are coupled to generate a temporal synchronization constraint matrix. Under the constraints of digital human expression control parameters and the temporal synchronization constraint matrix, the response content undergoes multimodal transformation and semantic correction to generate output content for each modality, including:
[0120] The unified semantic representation is semantically fused with the user intent to obtain the response semantic representation. The response semantic representation is then decoded to generate the response content, and the modal demand vector of the response content in each modal space is extracted.
[0121] The allocation deviation of the modal demand vector and the adjusted semantic contribution degree and the coordination deviation of the adjusted semantic correlation degree are calculated. The adjusted semantic contribution degree and the adjusted semantic correlation degree are coupled based on the allocation deviation and the coordination deviation to generate the time-series synchronization constraint matrix.
[0122] The modal demand vector is coupled with the rate of change of activation intensity of each modality in the digital human expression control parameters and the allocation deviation in the temporal synchronization constraint matrix to obtain the temporal parameters of each modal transition.
[0123] Under the constraints of digital human expression control parameters and timing synchronization constraint matrix, the response content is multimodal converted according to the timing parameters of each modality conversion to obtain the initial output content of each modality;
[0124] Extract cross-modal semantic consistency features of the initial modal output content, verify the consistency of the cross-modal semantic consistency features with the coordination deviation in the temporal synchronization constraint matrix to obtain semantic correction vectors, and perform semantic correction on the initial modal output content based on the semantic correction vectors to generate the output content of each modality.
[0125] After obtaining the user intent, the unified semantic representation is semantically fused with the user intent to obtain the response semantic representation. Specifically, the unified semantic representation carries comprehensive semantic information from multiple modalities such as text, speech, and vision, while the user intent is a high-level abstraction of this semantic information. The two are fused through attention mechanisms or splicing mapping, so that the response semantic representation retains the fine-grained perceptual information of the multimodalities while explicitly encoding the user's core needs. When decoding the response semantic representation, an autoregressive or one-time decoding method is used to generate response content in natural language form, ensuring that the response content is highly consistent with the user intent at the semantic level. After obtaining the response content, modal demand vectors of the response content in each modal space are further extracted. The modal demand vectors describe the expected output features of the response content in various modal dimensions such as text expression, speech prosody, and facial animation. For example, the text modal demand vector encodes vocabulary selection and sentence structure preferences, the speech modal demand vector encodes speech rate, pitch, and pause rhythm, and the visual modal demand vector encodes lip-sync, facial expression amplitude, and body posture.
[0126] Calculate the modal demand vector and its adjusted semantic contribution. The allocation deviation between them, and the modal demand vector and the adjusted semantic relevance matrix The coordination bias between modalities. Allocation bias reflects the gap between the actual demand of the response content for each modality and the current modality contribution allocation. If the magnitude of the demand vector of a certain modality is significantly higher than its adjusted semantic contribution, then the modality has a positive allocation bias, meaning that more expression resources need to be allocated to it; conversely, there is a negative allocation bias. Coordination bias measures the degree of deviation between the cooperative relationship between modalities and the modal cooperative structure described by the adjusted semantic relevance matrix. If the joint demand strength of two modalities at the response content level is significantly higher than the actual demand of each modality, then the coordination bias is higher than the actual demand of each modality. Corresponding element A mismatch results in a collaborative bias. The adjusted semantic contribution vector... With the adjusted semantic association matrix Based on the above-mentioned allocation deviation and coordination deviation, a coupled calculation is performed to generate the timing synchronization constraint matrix. Timing synchronization constraint matrix elements Representing modes With mode The synchronization constraint strength in the output timing takes into account the distribution deviation of each modal contribution and the coordination deviation between modes, thus providing accurate timing alignment guidance for the subsequent multimodal conversion process.
[0127] After obtaining the temporal synchronization constraint matrix, the modal demand vector is compared with the rate of change of activation intensity of each modality in the digital human expression control parameters. and timing synchronization constraint matrix The distributed deviation components are coupled and calculated to obtain the timing parameters of each mode transition. Rate of change of activation intensity for each mode This describes the dynamic trend of the activation levels of each modality in the digital human at the current time step. Multiplying this by the required amplitude of the corresponding modality in the modality demand vector, and then adding the correction for the allocation deviation, yields the temporal offset and duration at which the modality should undergo a conversion operation in the current frame. For example, when the allocation deviation of the speech modality is positive and the rate of change in activation intensity is large, the speech modality conversion timing parameters will indicate that speech synthesis should be initiated earlier and the syllable duration appropriately extended to meet stronger speech expression needs. (Modality conversion timing parameters are listed below.) Together, they form a timing scheduling scheme to ensure that text rendering, speech synthesis, and facial animation generation proceed in a coordinated manner on the timeline.
[0128] In the expression control parameters and timing synchronization constraint matrix of digital humans Under the dual constraints, based on the timing parameters of each mode transition The response content is subjected to multimodal transformation to obtain the initial output content of each modality. The digital human's expression control parameters include the activation intensity of each modality. This parameter determines the overall expressive range of each modality in the output. For example, higher activation intensity of the speech modality will moderately increase the volume and speech rate of the synthesized speech, while higher activation intensity of the visual modality will drive richer facial expression changes. (Temporal synchronization constraint matrix) Constraints are imposed on the temporal alignment between modalities, requiring that the temporal alignment error between lip-sync animation and speech phonemes not exceed a preset threshold, and that the display timing of text subtitles be synchronized with the speech broadcast. During multimodal conversion, the text modality renders the response content into a displayable subtitle sequence through a natural language generation model, the speech modality converts the text into an audio waveform with prosodic information through neural network speech synthesis, and the visual modality maps semantic content into a sequence of facial expressions and body movements of a digital human through a facial motion encoding unit and a pose generation network. All three are scheduled by their respective modal conversion timing parameters to ensure the consistency of output frame rate and timestamp.
[0129] Cross-modal semantic consistency features are extracted from the initial output content of each modality. These features describe the degree of semantic corroboration between different modal output contents; for example, whether the emotional tendency in the speech output matches the emotional polarity of the text content, and whether the emotional category of facial expressions matches the emotion conveyed by the speech prosody. The cross-modal semantic consistency features are then compared with a temporal synchronization constraint matrix. Consistency verification is performed on the co-variance components in the model, and the degree of difference between them is calculated to obtain the semantic correction vector. Semantic correction vector Each component corresponds to a semantic correction value for a particular modality. If there is a significant difference between the cross-modal semantic consistency features and the cooperative bias of a certain modality, then the magnitude of the correction component for that modality is large; otherwise, it is close to zero. Based on the semantic correction vector... Semantic correction is performed on the initial output content of each modality. Specific operations include: replacing or supplementing words or phrases in the text output that deviate significantly from semantics; fine-tuning the prosodic parameters of the speech output (such as pitch curves and speech rate variations); and regenerating or interpolating and smoothing facial expression frames or action segments in the visual output that do not conform to semantics. After semantic correction, the output content of each modality achieves a high degree of consistency at the semantic level, ultimately generating output content that meets the requirements of multimodal collaborative expression. This provides accurate, synchronous, and emotionally natural multimodal signals for subsequently driving digital humans to perform interactions.
[0130] Under the constraints of digital human expression control parameters and timing synchronization constraint matrix, the response content is multimodal transformed according to the timing parameters of each modality transformation to obtain the initial output content of each modality, including:
[0131] By performing a time-series mapping between the temporal transition parameters of each modality and the rate of change of activation intensity of each modality, the starting point and duration of content generation for each modality can be obtained;
[0132] The response content is segmented into temporal segments according to the starting point and duration of each modality, and the transformation weight distribution of each modality temporal segment content is calculated based on the allocation deviation.
[0133] The pre-conversion results of each modality are obtained by weighting the temporal segments according to the conversion weight distribution;
[0134] Extract cross-modal semantic deviation features between the pre-conversion results of each modality, perform deviation matching between the cross-modal semantic deviation features and the co-conversion deviation to obtain the cross-modal adjustment vector, and perform cross-modal consistency adjustment on the pre-conversion results of each modality based on the cross-modal adjustment vector to obtain the initial output content of each modality.
[0135] Under the constraints of digital human expression control parameters and temporal synchronization constraint matrix, the process of multimodal conversion of the response content requires precise coordination of the time start and duration of each modality to ensure that the final output text, speech, and visual actions maintain consistency and natural fluency on the timeline. (Temporal parameter vectors for each modality conversion are also provided.) It carries the transition control information of each mode in the temporal dimension, and the rate of change of activation intensity of each mode. This reflects the dynamic activation trend of each mode at the current moment. A temporal mapping between the two is performed; specifically, for the... A modality, whose content generation starting point Depend on Corresponding components and The weighted superposition determines, i.e. ,in For the first Transition timing parameter components of the modal, This is the temporal sensitivity coefficient, used to control the impact of the activation intensity change rate on the starting point offset. Duration Then based on the adjusted activation intensity Duration compared to the baseline Determined jointly, expressed as ,in The duration extension coefficient ensures that modalities with higher emotional activation levels receive relatively more output duration, thereby maintaining a match between the expression level and the emotional state.
[0136] After obtaining the starting point and duration of each modality's content generation, the response content is segmented according to these temporal parameters. The response content itself is a semantically complete text sequence, segmented according to the... The starting point corresponding to each mode and duration It is divided into temporal segments for each modality. During segmentation, temporal windows of different modalities may overlap or have gaps. Content in overlapping areas is shared across multiple modalities, while gaps are filled through interpolation to ensure semantic continuity. After obtaining the temporal segmentation content for each modality, the transformation weight distribution for each modality is calculated. (Assignment bias) The first The difference between the actual semantic content carried by a modality in the current temporal segment and the ideal allocation is used to adjust the transformation weights based on the allocation deviation, thus obtaining the result. Transition weights of different modes The calculation method is as follows ,in This represents the total number of modalities. The smaller the allocation deviation, the higher the match between the current segment content of that modality and its expressive ability, and the greater the corresponding transformation weight, thus receiving more transformation resource allocation in subsequent weighted transformations.
[0137] The temporal segments of each modality are weighted and transformed according to the transformation weight distribution to obtain the pre-transformation results of each modality. The core of weighted transformation lies in mapping time-series segmented content into structured response text through a language generation model, specifically for text modalities. The generated confidence level is scaled; for the speech modality, the text sequence is converted into an audio feature sequence through a speech synthesis engine, and the conversion process is based on... Adjust the generation intensity of speech rate, pitch, and energy envelope; for the visual modality, map semantic content into facial expression parameter sequences and body movement keyframes, in order to... The scaling ratio of the range of motion and intensity of facial expressions is controlled. Under the constraints of the digital human's expression control parameters, the conversion process of each modality must also meet the preset limits on the range of facial expressions and movements, speech rate boundary conditions, and text output format specifications to ensure that the generated content conforms to the digital human role setting at the perceptual level.
[0138] Preconversion results of each mode There may be semantic discrepancies between them, such as inconsistencies between the emotional tone of spoken language and the emotional signals conveyed by facial expressions, or a mismatch between the semantic focus of textual expression and the directional nature of visual actions. Therefore, cross-modal semantic deviation features are extracted. It describes the modes With mode The pre-conversion results show directional and magnitude differences in semantic space. Specifically, the pre-conversion results will... and Each semantic vector is projected onto a unified semantic space, and the cosine deviation and Euclidean distance deviation between them are calculated. The concatenated vectors are then used to obtain the cross-modal semantic deviation feature vector. .
[0139] Cross-modal semantic deviation features are matched with collaborative deviations. Derived from the timing synchronization constraint matrix Corresponding element The residual between the actual modal coordination degree and the actual modal coordination degree reflects the modal coordination degree. With mode The degree of under- or over-coordination in the current output state. The deviation matching process calculates... and The joint mapping yields the cross-modal adjustment vector. Its direction points to the semantic dimension that needs to be corrected, and its magnitude represents the strength of the correction. For cases involving multiple modal pairs, all... According to modal weights Weighted aggregation yields the first... Synthetic cross-modal adjustment vector of various modes .
[0140] Cross-modal consistency adjustment is performed on the pre-conversion results of each modality based on the cross-modal adjustment vector to obtain the initial output content of each modality. The adjustment process is expressed as follows: The addition operation is performed in the feature space of the corresponding modality. For the speech modality, the adjustment is reflected in the correction of the prosodic parameter sequence; for the visual modality, the adjustment is reflected in the offset compensation of facial expression parameters and action keyframes; for the text modality, the adjustment is reflected in the semantic direction correction of word vector representations, which affects the word selection probability distribution in the subsequent decoding stage. After cross-modal consistency adjustment, the initial output content of each modality tends to be consistent in direction in the semantic space, and the intermodal expressive synergy is significantly improved, laying the foundation for subsequent semantic correction to generate the final output content of each modality. The entire multimodal conversion process operates under the dual constraints of the temporal synchronization constraint matrix and the digital human expression control parameters, ensuring that the output content meets the requirements of high-quality digital human interaction in three dimensions: temporal synchronization, emotional consistency, and semantic coherence.
[0141] A second aspect of this invention provides a multimodal driven intelligent digital human customer service system, comprising:
[0142] A multimodal feature unit is used to acquire user input data from at least two modalities, including text modality, speech modality and visual modality, and extract feature vectors for each modality;
[0143] The semantic fusion unit is used to calculate the semantic correlation degree between modalities, fuse the feature vectors of each modality based on the semantic correlation degree between modalities to generate a unified semantic representation, and extract the initial semantic contribution of each modality from the unified semantic representation.
[0144] The emotion weighting unit is used to identify user intent and emotional state based on unified semantic representation, construct an emotion expression demand vector according to the emotional state, and cross-match the emotion expression demand vector with the initial semantic contribution to generate expression weight distribution and digital human expression control parameters.
[0145] The reverse adjustment unit is used to reverse adjust the semantic correlation between modalities according to the expression weight distribution, and to obtain the adjusted semantic contribution by re-fusing the feature vectors of each modality based on the adjusted semantic correlation.
[0146] The multimodal generation unit is used to obtain response content according to user intent, couple the adjusted semantic contribution degree and the adjusted semantic relevance degree to generate a temporal synchronization constraint matrix, and perform multimodal conversion on the response content under the constraints of digital human expression control parameters and temporal synchronization constraint matrix, and generate each modal output content through semantic correction.
[0147] The drive instruction unit is used to generate multimodal output instructions based on the output content of each modality, and drive the digital human to perform interactions.
[0148] A third aspect of the present invention provides an electronic device, comprising:
[0149] processor;
[0150] Memory used to store processor-executable instructions;
[0151] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0152] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0153] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.
[0154] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multimodal driven intelligent digital human customer service method, characterized in that, include: Acquire user input data from at least two of the following modalities: text, speech, and visual; and extract feature vectors for each modality. Calculate the semantic correlation degree between modalities, fuse the feature vectors of each modality based on the semantic correlation degree to generate a unified semantic representation, and extract the initial semantic contribution of each modality from the unified semantic representation; Based on unified semantic representation, user intent and emotional state are identified. An emotional expression demand vector is constructed according to the emotional state. The emotional expression demand vector is cross-matched with the initial semantic contribution to generate expression weight distribution and digital human expression control parameters. The semantic correlation between modalities is adjusted inversely based on the expression weight distribution, and the adjusted semantic contribution is obtained by re-integrating the feature vectors of each modality based on the adjusted semantic correlation. The response content is obtained based on the user's intent. The adjusted semantic contribution and adjusted semantic relevance are coupled to generate a temporal synchronization constraint matrix. Under the constraints of the digital human expression control parameters and the temporal synchronization constraint matrix, the response content is multimodal converted and the output content of each modality is generated through semantic correction. Multimodal output instructions are generated based on the output content of each modality to drive the digital human to perform interactions.
2. The method according to claim 1, characterized in that, Calculate the semantic correlation degree between modalities, fuse the feature vectors of each modality based on the semantic correlation degree to generate a unified semantic representation, and extract the initial semantic contribution of each modality from the unified semantic representation, including: Cross-modal attention is calculated for each modal feature vector to obtain the attention weight matrix of each modal feature vector to other modal feature vectors; The modality missing and modality quality status in the user input data are detected, and the attention weight matrix is compensated and adjusted according to the modality missing and modality quality status to obtain the compensated attention weight matrix. The bidirectional attention score between any two modal feature vectors is calculated based on the compensated attention weight matrix, and the semantic association degree between modalities is obtained by symmetric processing of the bidirectional attention score. Using the semantic correlation degree between modalities as the fusion coefficient, the feature vectors of each modality are weighted and fused to generate a unified semantic representation; The unified semantic representation is decomposed, and the projection intensity of the unified semantic representation on the feature vector of each modality is calculated. The projection intensity is normalized to obtain the initial semantic contribution of each modality.
3. The method according to claim 1, characterized in that, An emotion expression demand vector is constructed based on the emotional state. This vector is then cross-matched with the initial semantic contribution to generate an expression weight distribution and digital human expression control parameters, including: Extract the modal contribution values from the initial semantic contribution values, calculate the variance of each modal contribution value to obtain the emotional intensity value, calculate the covariance of each modal contribution value to obtain the emotional type value, and combine the emotional intensity value and the emotional type value to encode an emotional expression demand vector. Modal collaboration constraint rules are obtained based on emotional state labels. The similarity between the emotional expression demand vector and the contribution value of each modality is calculated to obtain the modal matching degree. The modal matching degree is substituted into the modal collaboration constraint rules for cross-modal weighting to generate the expression weight distribution. The modal emotional expression weights and emotional intensity values in the expression weight distribution are fused to obtain the modal activation intensity. The information entropy of the distribution formed by the activation intensity of each modality is calculated as the distribution balance. The adjustment coefficient is determined based on the distribution balance. The modal activation intensities are weighted based on the adjustment coefficient and the emotion type value to obtain the adjusted modal activation intensities. The temporal change rate of the adjusted modal activation intensities is calculated to obtain the modal activation intensity change rate. The modal activation intensity change rate is used as the expression control parameter of the digital human.
4. The method according to claim 3, characterized in that, Modal collaborative constraint rules are obtained based on sentiment state labels. The similarity between the sentiment expression demand vector and the contribution value of each modality is calculated to obtain the modal matching degree. The initial matching degree of each modality is substituted into the modal collaborative constraint rules for cross-modal weighting to generate the expression weight distribution, including: Based on the emotion state label, the modal co-constraint rules are obtained by querying the preset emotion-modal mapping table. The modal co-constraint rules include inter-modal mutual exclusion coefficients and co-constraint coefficients. The emotional expression demand vector is decomposed into emotional intensity components and emotional type components. The intensity similarity between the emotional intensity component and the contribution value of each modality and the directional similarity between the emotional type component and the contribution value of each modality are calculated. The intensity similarity and directional similarity are weighted and fused to obtain the modality matching degree. Based on the modality matching degree, the intermodality matching degree difference value is calculated. The modality matching degree and the intermodality matching degree difference value are substituted into the intermodality mutual exclusion coefficient to perform cross-modality suppression weighting to obtain the suppression weight. The intermodality synergy coefficient is substituted into the intermodality enhancement weighting to obtain the enhancement weight. The suppression weight and the enhancement weight are combined to generate the expression weight distribution.
5. The method according to claim 1, characterized in that, The semantic correlation between modalities is adjusted inversely based on the expression weight distribution. The adjusted semantic contribution is obtained by re-fusing the feature vectors of each modality based on the adjusted semantic correlation, including: Suppression weights and enhancement weights are extracted from the expression weight distribution. The numerical deviation between the suppression weight and the initial semantic contribution of the corresponding modality is calculated to obtain the suppression bias, and the numerical deviation between the enhancement weight and the initial semantic contribution of the corresponding modality is calculated to obtain the enhancement bias. A semantic relevance adjustment matrix is constructed based on suppression bias and enhancement bias. The adjusted semantic correlation degree is obtained by performing matrix operations on the intermodal semantic correlation degree and the semantic correlation degree adjustment matrix. The modal feature vectors are weighted and fused according to the adjusted semantic relevance to obtain the adjusted unified semantic representation. The adjusted unified semantic representation is decomposed, and the projection intensity of the adjusted unified semantic representation on each modal feature vector is calculated. The projection intensity is normalized to obtain the adjusted semantic contribution.
6. The method according to claim 1, characterized in that, Based on user intent, the response content is obtained. The adjusted semantic contribution and adjusted semantic relevance are coupled to generate a temporal synchronization constraint matrix. Under the constraints of digital human expression control parameters and the temporal synchronization constraint matrix, the response content is subjected to multimodal transformation and semantic correction to generate output content for each modality, including: The unified semantic representation is semantically fused with the user intent to obtain the response semantic representation. The response semantic representation is then decoded to generate the response content, and the modal demand vector of the response content in each modal space is extracted. The allocation deviation of the modal demand vector and the adjusted semantic contribution degree and the coordination deviation of the adjusted semantic correlation degree are calculated. The adjusted semantic contribution degree and the adjusted semantic correlation degree are coupled based on the allocation deviation and the coordination deviation to generate the time-series synchronization constraint matrix. The modal demand vector is coupled with the rate of change of activation intensity of each modality in the digital human expression control parameters and the allocation deviation in the temporal synchronization constraint matrix to obtain the temporal parameters of each modal transition. Under the constraints of digital human expression control parameters and timing synchronization constraint matrix, the response content is multimodal converted according to the timing parameters of each modality conversion to obtain the initial output content of each modality; Extract cross-modal semantic consistency features of the initial modal output content, verify the consistency of the cross-modal semantic consistency features with the coordination deviation in the temporal synchronization constraint matrix to obtain semantic correction vectors, and perform semantic correction on the initial modal output content based on the semantic correction vectors to generate the output content of each modality.
7. The method according to claim 6, characterized in that, Under the constraints of digital human expression control parameters and timing synchronization constraint matrix, the response content is multimodal transformed according to the timing parameters of each modality transformation to obtain the initial output content of each modality, including: By performing a time-series mapping between the temporal transition parameters of each modality and the rate of change of activation intensity of each modality, the starting point and duration of content generation for each modality can be obtained; The response content is segmented into temporal segments according to the starting point and duration of each modality, and the transformation weight distribution of each modality temporal segment content is calculated based on the allocation deviation. The pre-conversion results of each modality are obtained by weighting the temporal segments according to the conversion weight distribution; Extract cross-modal semantic deviation features between the pre-conversion results of each modality, perform deviation matching between the cross-modal semantic deviation features and the co-conversion deviation to obtain the cross-modal adjustment vector, and perform cross-modal consistency adjustment on the pre-conversion results of each modality based on the cross-modal adjustment vector to obtain the initial output content of each modality.
8. A multimodal driven intelligent digital human customer service system, used to implement the method as described in any one of claims 1-7, characterized in that, include: A multimodal feature unit is used to acquire user input data from at least two modalities, including text modality, speech modality and visual modality, and extract feature vectors for each modality; The semantic fusion unit is used to calculate the semantic correlation degree between modalities, fuse the feature vectors of each modality based on the semantic correlation degree between modalities to generate a unified semantic representation, and extract the initial semantic contribution of each modality from the unified semantic representation. The emotion weighting unit is used to identify user intent and emotional state based on unified semantic representation, construct an emotion expression demand vector according to the emotional state, and cross-match the emotion expression demand vector with the initial semantic contribution to generate expression weight distribution and digital human expression control parameters. The reverse adjustment unit is used to reverse adjust the semantic correlation between modalities according to the expression weight distribution, and to obtain the adjusted semantic contribution by re-fusing the feature vectors of each modality based on the adjusted semantic correlation. The multimodal generation unit is used to obtain response content according to user intent, couple the adjusted semantic contribution degree and the adjusted semantic relevance degree to generate a temporal synchronization constraint matrix, and perform multimodal conversion on the response content under the constraints of digital human expression control parameters and temporal synchronization constraint matrix, and generate each modal output content through semantic correction. The drive instruction unit is used to generate multimodal output instructions based on the output content of each modality, and drive the digital human to perform interactions.
9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.