Multi-modal data fusion intelligent interaction method and system based on virtual digital human
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HANGZHOU DIGITAL CLOUD TECH CO LTD
- Filing Date
- 2026-05-27
- Publication Date
- 2026-08-07
AI Technical Summary
[0002]随着人工智能与数字媒体技术的快速发展,虚拟数字人作为人机交互的新型载体,被广泛应用于智能客服、在线教育、数字展陈、情感陪伴及元宇宙社交等领域;现有虚拟数字人交互系统通常以语音识别与自然语言处理为核心,通过将用户的语音输入转化为文本后进行意图解析与应答生成,驱动虚拟数字人完成基础的问答交互;部分系统在此基础上引入了面部表情识别或肢体姿态捕捉等单一视觉感知能力,但各模态数据的处理过程相互独立,通常采用简单拼接或固定规则进行信息整合,未能实现多模态信息在特征层面的深度融合;同时,现有虚拟数字人的响应生成多侧重于语义内容的准确性,其面部表情、语气语调与肢体动作之间往往缺乏协同一致性,导致交互过程中虚拟数字人的表现呈现机械感,难以达到自然流畅的拟人化交互效果
通过同步采集用户的语音信号、面部图像序列与肢体姿态数据,从语义内容、语调韵律、面部微表情与肢体运动四个维度提取多维交互特征,实现对用户交互行为的全方位感知与多层次表征,从而有效克服仅依赖单一语音模态或有限模态组合进行用户意图理解的不足;
Smart Images

Figure CN122346260B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of human-computer intelligent interaction technology, and more specifically, to a multimodal data fusion intelligent interaction method and system based on virtual digital humans. Background Technology
[0002] With the rapid development of artificial intelligence and digital media technologies, virtual digital humans, as a new type of human-computer interaction carrier, are widely used in fields such as intelligent customer service, online education, digital exhibitions, emotional companionship, and metaverse social networking. Existing virtual digital human interaction systems typically focus on speech recognition and natural language processing. They convert user voice input into text, then perform intent parsing and response generation to drive the virtual digital human to complete basic question-and-answer interactions. Some systems have introduced single visual perception capabilities such as facial expression recognition or body posture capture on this basis. However, the processing of data from each modality is independent, and information integration is usually achieved through simple splicing or fixed rules, failing to achieve deep fusion of multimodal information at the feature level. At the same time, the response generation of existing virtual digital humans focuses more on the accuracy of semantic content. Their facial expressions, tone of voice, and body movements often lack coordination and consistency, resulting in a mechanical performance of the virtual digital human during the interaction process, making it difficult to achieve a natural and smooth anthropomorphic interaction effect.
[0003] Furthermore, existing virtual digital human interaction technologies have significant shortcomings in user understanding, primarily manifested in a lack of ability to differentiate between users' explicit task intentions and implicit emotional states. Most existing systems treat user input as a uniform signal for end-to-end processing, failing to effectively decouple cognitive needs from emotional expression. This results in virtual digital human responses often prioritizing task execution over emotional empathy, or struggling to achieve a reasonable balance between emotional response and task completion. Especially in application scenarios requiring high-quality emotional interaction, such as emotional companionship, psychological counseling, and personalized education, existing technologies struggle to comprehensively assess a user's true state based on multi-dimensional signals such as semantic content, tone of voice, facial micro-expressions, and body language, thus failing to generate interactive responses that combine task execution and emotional empathy. This severely restricts the application depth and user experience of virtual digital humans in high-value scenarios.
[0004] In view of this, the present invention proposes a multimodal data fusion intelligent interaction method and system based on virtual digital humans to solve the above problems. Summary of the Invention
[0005] To overcome the aforementioned shortcomings of the prior art and achieve the above objectives, the present invention provides the following technical solution: a multimodal data fusion intelligent interaction method based on virtual digital humans, comprising: Collect users’ voice signals, facial image sequences and body posture data, and extract semantic content features, intonation and prosody features, facial micro-expression features and body movement features to form a multi-dimensional interactive feature set; Cognitive-emotional attribution determination is performed on the multidimensional interactive feature set. The task-oriented component in the semantic content feature and the task-oriented component in the body movement feature are extracted into a cognitive intention feature stream. The emotional orientation component in the semantic content feature, the emotional orientation component in the body movement feature, the intonation and prosody feature, and the facial micro-expression feature are extracted into an emotional state feature stream. Based on the cognitive intent feature flow, and combined with a pre-set domain knowledge base, intent matching and task planning are performed to generate a task response plan that includes response content, execution steps and task urgency. Emotion inference is performed on the distribution characteristics and intensity changes of each feature in the emotional state feature stream to determine the user's emotional category and emotional intensity value. Based on the emotional category and emotional intensity value, the corresponding empathy response level and empathy expression mode are matched to generate an emotional response scheme. Based on the task urgency in the task response plan and the emotional intensity in the emotional response plan, the fusion ratio of cognitive decision-making and emotional decision-making is dynamically adjusted. Based on the fusion ratio, the task response plan and the emotional response plan are synergistically fused to determine the final response content, tone style parameters and facial expression parameters of the virtual digital human, thereby driving the virtual digital human to generate interactive responses that combine task execution and emotional empathy.
[0006] Furthermore, methods for forming multidimensional interaction feature sets include: Endpoint detection is performed on the speech signal to extract the effective speech segment; speech recognition and semantic parsing are performed on the effective speech segment to obtain the text sequence and sentence-level semantic vector; task keyword extraction and sentiment word recognition are performed on the text sequence to obtain the task orientation density and sentiment orientation density, and the semantic content features are formed by combining the text sequence and sentence-level semantic vector. Frame-by-frame acoustic prosodic analysis is performed on effective speech segments to obtain the average speech fundamental frequency, fundamental frequency fluctuation, average speech energy, energy fluctuation, and average speech rate, and to form intonation prosodic features. Facial motion coding and inter-frame dynamic change analysis are performed on facial image sequences to obtain the mean vector of motion units, the fluctuation vector of motion units, and the average facial motion amplitude, and to form facial micro-expression features. The joint motion pattern and body posture characteristics were analyzed on the limb posture data to obtain the mean vector of functional joint distances, overall motion activity, average trunk extension, trunk extension fluctuation and average gesture space range, and to form limb motion characteristics. Semantic content features, intonation and prosody features, facial micro-expression features, and body movement features are combined to form a multi-dimensional interactive feature set.
[0007] Furthermore, methods for determining cognitive-emotional attribution based on multidimensional interaction feature sets include: Cognitive-emotional attribution polarity analysis is performed on semantic content features to obtain semantic cognitive attribution coefficients and semantic-emotional attribution coefficients. Based on the semantic cognitive attribution coefficients and semantic-emotional attribution coefficients, the sentence-level semantic vectors are decoupled to obtain semantic cognitive projection vectors and semantic-emotional projection vectors. The text sequence, semantic cognitive projection vectors, and task orientation density are integrated to form the task orientation component of semantic content features. The semantic-emotional projection vectors and sentiment orientation density are integrated to form the sentiment orientation component of semantic content features. The attribution attributes of each sub-feature in the limb movement characteristics are determined, and they are divided into inherent attribution characteristics and shared attribution characteristics. For shared attribution characteristics, a cross-modal attribution allocation mechanism is used for decomposition, decomposing the overall motor activity into cognitive motor activity components and emotional motor activity components, and decomposing the average gesture spatial range into cognitive gesture range components and emotional gesture range components. The functional joint distance mean vector, cognitive motor activity components and cognitive gesture range components are integrated to form the task orientation component of the limb movement characteristics. The average trunk extension degree, trunk extension fluctuation degree, emotional motor activity components and emotional gesture range components are integrated to form the emotional orientation component of the limb movement characteristics.
[0008] Furthermore, methods for extracting cognitive intention feature streams and emotional state feature streams include: Based on the preset median reference for motor activity and overall motor activity, a limb activity index is calculated; based on the task orientation density and the limb activity index, cross-modal cognitive synergy is calculated, and combined with the preset cognitive synergy enhancement coefficient, a cognitive fusion enhancement factor is calculated; based on the cognitive fusion enhancement factor, each element in the semantic cognitive projection vector and the functional joint distance mean vector is synergistically enhanced to obtain an enhanced semantic cognitive vector and an enhanced functional joint vector, and combined with the text sequence, task orientation density, cognitive motor activity component and cognitive gesture range component to form a cognitive intent feature stream; Based on the fundamental frequency fluctuation and the preset median reference for prosodic fluctuation, a prosodic activity index is calculated; based on the average facial movement amplitude and the preset median reference for facial movement, a facial activity index is calculated; based on the emotion orientation density, prosodic activity index, facial activity index, and limb activity index, a cross-modal emotion synergy is calculated, and combined with a preset emotion synergy enhancement coefficient, an emotion fusion enhancement factor is calculated; based on the emotion fusion enhancement factor, each element in the semantic emotion projection vector and the action unit mean vector is synergistically enhanced to obtain an enhanced semantic emotion vector and an enhanced action unit mean vector, and combined with the emotion orientation components of emotion orientation density, action unit fluctuation vector, average facial movement amplitude, and limb movement features, and the intonation prosodic features, an emotion state feature stream is formed.
[0009] Furthermore, methods for intent matching and task planning based on cognitive intent feature streams include: A pre-defined domain knowledge base is provided, which includes an intent template library, a task execution graph library, and a gesture intent pattern library. The intent template library contains multiple intent templates, the task execution graph library contains multiple task execution graphs, and the gesture intent pattern library contains multiple gesture intent patterns. Based on the enhanced semantic cognition vector, text sequence and enhanced function joint vector, multi-channel intent matching is performed on the intent template library to obtain the comprehensive intent matching confidence of each intent template, and the comprehensive intent matching confidence with the largest value is marked as the highest matching confidence. Based on the comprehensive intent matching confidence, the user intent recognition result is determined; based on the user intent recognition result, task planning is performed to generate response content and execution step sequence, and the corresponding task complexity level is set; based on task pointing density, cognitive motor activity component, cognitive gesture range component, highest matching confidence and task complexity level, task urgency is calculated; the response content, execution step sequence and task urgency are integrated to form a task response plan.
[0010] Furthermore, methods for determining a user's emotion category and emotion intensity value include: A pre-defined emotion prototype library is used, containing multiple emotion prototypes. Each emotion prototype includes an emotion category label, a semantic emotion prototype vector, a facial action unit prototype vector, a prosodic prototype parameter set, and a body posture prototype parameter set. For each emotion prototype, an emotion evidence score for each modality is calculated. The modalities are, in order, semantic modality, facial modality, prosodic modality, and body modality. Based on the activity level of each modality in the emotion state feature stream, a dynamic fusion weight is calculated for each modality. Based on the dynamic fusion weights of each modality, the emotion evidence scores for each modality corresponding to the same emotion prototype are weighted and summed to obtain a comprehensive emotion matching score for each emotion prototype. The emotion category label corresponding to the emotion prototype with the highest comprehensive emotion matching score is taken as the user's emotion category, and the corresponding comprehensive emotion matching score is taken as the preferred emotion matching score. The user's emotion intensity value is calculated based on the deviation of each modality from the neutral state in the emotion state feature stream and the preferred emotion matching score.
[0011] Furthermore, methods for generating emotional response plans include: A pre-defined empathy strategy library contains multiple empathy strategy entries. Each entry includes the applicable emotion category, applicable intensity range, empathy response level, and empathy expression method. Based on the user's emotion category and intensity value, the library is searched for empathy strategy entries whose applicable emotion category matches the user's and whose intensity value falls within the applicable intensity range, and these are marked as matching strategy entries. The corresponding empathy response level and empathy expression method are obtained from the matching strategy entries. The user's emotion category, intensity value, empathy response level, and empathy expression method are integrated to form an emotion response plan.
[0012] Furthermore, methods for dynamically adjusting the integration ratio of cognitive and affective decision-making include: Methods for dynamically adjusting the integration ratio of cognitive and affective decision-making include: Based on task urgency and emotional intensity, channel interaction status parameters are calculated. These parameters include channel tension index, total channel activity, and channel dominance polarity. The channel tension index is compared with a preset tension conflict threshold, and the current interaction mode is determined based on the comparison result. The current interaction mode includes coordinated fusion mode and adversarial fusion mode. If the current interaction is in a coordinated fusion mode and the total channel activity is greater than zero, the fusion ratio is calculated based on task urgency and total channel activity. The fusion ratio includes cognitive fusion ratio and emotional fusion ratio. If the current interaction is in an adversarial fusion mode, then the mutual inhibition strength factor is calculated based on the channel tension index and the tension conflict threshold; the initial cognitive proportion is calculated based on the task urgency and the total channel activity; if the channel dominance polarity value is greater than zero, then the initial emotional proportion is calculated based on the initial cognitive proportion, and the cognitive enhancement amount is calculated in combination with the mutual inhibition strength factor; the sum of the initial cognitive proportion and the cognitive enhancement amount is calculated to obtain the cognitive fusion proportion; if the channel dominance polarity value is less than zero, then the subordination retention coefficient is calculated based on the mutual inhibition strength factor, and the cognitive fusion proportion is calculated in combination with the initial cognitive proportion; the difference between the initial and cognitive fusion proportions is calculated to obtain the emotional fusion proportion.
[0013] Furthermore, methods for synergistically integrating task response plans and emotional response plans include: Based on the current fusion modality and fusion ratio of the interaction, the final response content is determined; based on the fusion ratio, tone style parameters and facial expression parameters are calculated; among them, tone style parameters include fusion speech rate coefficient, fusion pitch coefficient, and fusion tone temperature coefficient; facial expression parameters include final facial expression identifier and fusion action amplitude coefficient. Perform cross-channel expression consistency verification on tone style parameters and facial expression parameters: use the fused tone temperature coefficient as tone expression valence and the fused action amplitude coefficient as action expression valence; calculate cross-channel expression deviation based on tone expression valence and action expression valence; if the cross-channel expression deviation is greater than the preset expression consistency threshold, perform consistency correction; if the cross-channel expression deviation is less than or equal to the expression consistency threshold, do not perform consistency correction. The consistency correction process is as follows: Calculate the target concordance valence based on the fusion ratio, tone expression valence, and action expression valence; calculate the tone correction amount based on the target concordance valence, tone expression valence, and the preset consistency correction step size; update the fusion tone temperature coefficient in the tone style parameters based on the sum of the fusion tone temperature coefficient and the tone correction amount; calculate the action correction amount based on the target concordance valence, action expression valence, and the consistency correction step size; update the fusion action amplitude coefficient in the facial expression action parameters based on the sum of the fusion action amplitude coefficient and the action correction amount. The final response content, tone and style parameters, facial expression and action parameters, and execution step sequence are integrated to form a comprehensive interactive driving instruction; the virtual digital human is then driven to perform interactive responses based on the comprehensive interactive driving instruction.
[0014] The multimodal data fusion intelligent interaction system based on virtual digital humans, and the implementation of the aforementioned multimodal data fusion intelligent interaction method based on virtual digital humans, include: The multimodal parsing module is used to collect users' voice signals, facial image sequences, and body posture data, and extract semantic content features, intonation and prosody features, facial micro-expression features, and body movement features to form a multi-dimensional interactive feature set. The dual-stream decoupling module is used to perform cognitive-emotional attribution determination on the multi-dimensional interaction feature set. It extracts the task-oriented component in the semantic content features and the task-oriented component in the body movement features into a cognitive intention feature stream, and extracts the emotional orientation component in the semantic content features, the emotional orientation component in the body movement features, the intonation and prosody features, and the facial micro-expression features into an emotional state feature stream. The cognitive decision-making module is used to perform intent matching and task planning based on the cognitive intent feature flow and a preset domain knowledge base, and generate a task response plan that includes response content, execution steps and task urgency. The emotion decision module is used to infer the emotion from the distribution characteristics and intensity changes of each feature in the emotion state feature stream, determine the user's emotion category and emotion intensity value, and generate an emotion response plan by matching the corresponding empathy response level and empathy expression mode according to the emotion category and emotion intensity value. The dual-stream driving module is used to dynamically adjust the fusion ratio of cognitive decision-making and emotional decision-making based on the task urgency in the task response plan and the emotional intensity value in the emotional response plan. Based on the fusion ratio, the task response plan and the emotional response plan are collaboratively fused to determine the final response content, tone style parameters and facial expression parameters of the virtual digital human, and drive the virtual digital human to generate interactive responses that combine task execution and emotional empathy.
[0015] The technical effects and advantages of this invention based on a multimodal data fusion intelligent interaction method and system for virtual digital humans are as follows: By simultaneously collecting users' voice signals, facial image sequences, and body posture data, multi-dimensional interaction features are extracted from four dimensions: semantic content, intonation and rhythm, facial micro-expressions, and body movements. This enables comprehensive perception and multi-level representation of user interaction behavior, thereby effectively overcoming the shortcomings of relying solely on a single voice modality or a limited combination of modalities to understand user intent. By using a cognitive-emotional attribution determination and cross-modal attribution allocation mechanism based on semantic attribution polarity values, the inherent attribution features with clear directional attributes and the shared attribution features with ambiguous attribution attributes in multidimensional interaction features are systematically decoupled to form independent cognitive intent feature streams and emotional state feature streams. This enables accurate separation of users’ explicit task requirements and implicit emotional states, avoiding the ambiguity of intent determination and inaccurate emotion recognition caused by mixing cognitive and emotional information. By employing a cross-modal collaborative fusion enhancement mechanism, the geometric mean of multimodal signal activity is used to measure the degree of collaboration between modalities. Feature enhancement is performed only when multiple modalities simultaneously present active signals, effectively improving the reliability and robustness of cross-modal information complementarity verification. By fusing semantic matching, keyword coverage, and gesture-assisted matching into a multi-channel intent recognition method, and combining task complexity level and intent matching confidence to calculate task urgency, a comprehensive intent determination that balances semantic understanding depth, vocabulary matching accuracy, and limb gesture assistance is achieved. This effectively improves the accuracy of user intent recognition and the rationality of task response priority in complex interaction scenarios. By using a sentiment inference method that dynamically calculates fusion weights based on the real-time activity level of each modality, adaptive sentiment recognition is achieved under different channel qualities and user expression habits, effectively overcoming the problem of decreased recognition accuracy in noisy environments or modality-deficient scenarios when fixed-weight fusion is used. By introducing a channel tension index and a fusion modality dual-state determination mechanism, the cognitive-emotional fusion process is divided into a coordinated fusion modality and an adversarial fusion modality. In the adversarial fusion modality, a mutual inhibition adjustment mechanism is employed to achieve nonlinear enhancement of the dominant channel and nonlinear compression of the inferior channel, effectively solving the problem of ambiguous responses and unclear priorities in high-conflict situations caused by a simple linear mixture of cognitive and emotional decisions. Through a phase-separated response organization strategy, independent expression phases are assigned to different channels in the temporal structure of the response, ensuring the integrity of the dominant channel's expression and the appropriate preservation of key information in subordinate channels, achieving structured coordination of cognitive task content and emotional empathy content in the response text. Through a cross-channel expression consistency verification mechanism, the tone style parameters and facial expression parameters are bidirectionally corrected with the target coordination valence weighted by the fusion ratio as the correction direction, ensuring the consistency of the virtual digital human's voice tone and facial body expression in the emotional dimension, effectively avoiding unnatural interactive performances where tone and expression contradict each other. Ultimately, this achieves high-quality intelligent interactive responses from the virtual digital human while balancing task performance and emotional empathy. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of a multimodal data fusion intelligent interaction system based on a virtual digital human, according to Embodiment 1 of the present invention. Figure 2 This is a flowchart of the multimodal data fusion intelligent interaction method based on virtual digital humans according to Embodiment 2 of the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Example 1:
[0018] Please see Figure 1 As shown in this embodiment, the multimodal data fusion intelligent interaction system based on virtual digital humans includes a multimodal parsing module, a dual-stream decoupling module, a cognitive decision-making module, an emotional decision-making module, and a dual-stream driving module; each module is connected via wired and / or wireless means to realize data transmission between modules.
[0019] The multimodal parsing module is used to collect users' voice signals, facial image sequences, and body posture data, and extract semantic content features, intonation and prosody features, facial micro-expression features, and body movement features to form a multi-dimensional interactive feature set.
[0020] Methods for collecting users' voice signals, facial image sequences, and body posture data include: Audio acquisition devices are used to acquire the user's voice signal in real time. These devices include microphones, microphone arrays, and headsets, which are equipped with audio acquisition capabilities. Endpoint detection is performed on the voice signal to identify the start and end times of valid voice activities, and valid voice segments are extracted based on these times. The start and end times include both the end and start times. Endpoint detection is a conventional technique in this field and will not be elaborated upon here. A valid voice segment is a continuous voice fragment after removing silent parts. Image acquisition devices are used to acquire a sequence of facial images of the user in real time; these devices include, for example, RGB cameras, infrared cameras, depth cameras, and other devices with image acquisition capabilities; the facial image sequence consists of multiple frames of facial images arranged in chronological order of acquisition time. The system employs posture acquisition devices to acquire user limb posture data in real time. These devices include, for example, depth cameras and inertial measurement units (IMUs) with posture acquisition capabilities. The limb posture data consists of multiple frames of joint coordinate data arranged in chronological order of acquisition time. Each frame of joint coordinate data includes an acquisition timestamp and the three-dimensional spatial coordinates of multiple human joints. These human joints include, but are not limited to, head joints, neck joints, left and right shoulder joints, left and right elbow joints, left and right wrist joints, left and right hip joints, left and right knee joints, and left and right ankle joints.
[0021] Methods for forming multidimensional interaction feature sets include: Speech recognition processing is performed on valid speech segments to convert them into corresponding text sequences. The text sequence consists of text content composed of multiple words arranged in word order. Sentence-level semantic encoding is performed on the text sequence to obtain sentence-level semantic vectors. The sentence-level semantic vectors are fixed-dimensional vectors obtained by compressing and representing the overall semantic information of the text sequence. Sentence-level semantic encoding can be achieved using methods such as word vector aggregation and pre-trained language models. Task keywords are extracted from the text sequence to identify words related to task execution, resulting in a task keyword set. These keywords include, but are not limited to, action instruction words (such as "recommend," "explain," "demonstrate," etc.), interrogative words (such as "can," "how much," "how," etc.), and object entity nouns (such as music names, application names, product names, etc.), all related to specific task requests. The number of task keywords in the task keyword set is counted to obtain the task word count. The total word count is obtained by counting all words in the text sequence. The ratio of the task word count to the total word count is calculated to obtain the task targeting density. The task targeting density reflects the density of task-oriented words in the text sequence. Sentiment word identification is performed on the text sequence to identify words carrying emotional tendencies, resulting in a sentiment word set. Sentiment words include, but are not limited to, emotional adjectives (such as excitement, anxiety, anger, etc.), degree adverbs (such as very, slightly, especially, etc.), and interjections (such as sigh, wow, hum, etc.). The number of sentiment words in the sentiment word set is counted to obtain the sentiment word count. The ratio of the sentiment word count to the total word count is calculated to obtain the sentiment orientation density. The sentiment orientation density reflects the density of sentiment-oriented words in the text sequence. The text sequence, sentence-level semantic vector, task orientation density, and sentiment orientation density are integrated to form semantic content features. These features are used to characterize the textual semantic information, task orientation intensity, and sentiment orientation intensity of the user's speech. It should be noted that the specific implementation methods of speech recognition processing, sentence-level semantic encoding, task keyword extraction, and sentiment word recognition are all conventional techniques in this field and will not be elaborated on here.
[0022] The effective speech segment is divided into frames according to a preset frame length and frame shift to obtain multiple time frames. The frame length is the number of sampling points contained in each time frame; the frame shift is the offset of sampling points between corresponding starting positions of adjacent time frames. Both the frame length and frame shift are preset by those skilled in the art based on the speech analysis accuracy requirements. Fundamental frequency is extracted from each time frame to obtain the corresponding speech fundamental frequency value, which is then arranged in chronological order to form a speech fundamental frequency sequence. The specific implementation method of fundamental frequency extraction is a conventional technique in this field and will not be elaborated upon here. The speech fundamental frequency sequence is then subjected to mean and standard deviation calculations to obtain the average speech fundamental frequency and fundamental frequency fluctuation, respectively. The average speech fundamental frequency reflects the overall pitch of the user's speech; the fundamental frequency fluctuation reflects the amplitude of pitch variation. The mean of the squares of the audio amplitudes corresponding to each sampling point within each time frame is calculated to obtain the short-time energy value of each time frame, and these values are arranged in chronological order to form a speech energy sequence. The audio amplitude is the discrete amplitude value of each sampling point in the effective speech segment. The mean and standard deviation of the speech energy sequence are calculated to obtain the average speech energy and energy fluctuation, respectively. The average speech energy reflects the overall loudness level of the user's speech, while the energy fluctuation reflects the variation in speaking intensity. The difference between the end and start times of the effective speech segment is calculated to obtain the speech duration. The ratio of the total word count to the speech duration is calculated to obtain the average speech rate, which reflects the speed at which the user speaks. The average speech fundamental frequency, fundamental frequency fluctuation, average speech energy, energy fluctuation, and average speech rate are integrated to form intonation prosodic features; among them, intonation prosodic features are used to characterize the acoustic prosodic properties of user speech.
[0023] For each frame of a facial image sequence, facial keypoint detection is performed to obtain a set of facial keypoint coordinates for each frame. Facial keypoints are anatomically significant feature points on the face, including but not limited to eyebrow endpoints, corners of the eyes, tip of the nose, and corners of the mouth. For each frame of a facial image sequence, facial action unit recognition is performed to obtain an action unit activation vector for each frame. A facial action unit is the smallest facial muscle movement unit defined in the facial action coding system, such as inner eyebrow lift (AU1), outer eyebrow lift (AU2), corner of mouth lift (AU12), and corner of mouth pull down (AU15). Each element in the action unit activation vector corresponds to the activation intensity value of a facial action unit; the activation intensity value reflects the contraction amplitude of the facial action unit. For each facial motion unit in all motion unit activation vectors, the mean and standard deviation of the corresponding activation intensity value are calculated to obtain the motion unit mean vector and motion unit fluctuation vector. The motion unit mean vector reflects the overall activation level of each facial motion unit, while the motion unit fluctuation vector reflects the variation in activation intensity of each facial motion unit over time. For example, the facial motion unit activation vectors include two types of facial motion units: Unit 1 and Unit 2. The facial motion unit activation vectors corresponding to each frame of facial image are as follows: , , Therefore, the mean of the activation intensity value corresponding to unit 1 is 0.7, and the standard deviation is 0.082; the mean of the activation intensity value corresponding to unit 2 is 0.3, and the standard deviation is 0.082; therefore, the action unit mean vector is... The motion unit fluctuation vector is ; Each pair of adjacent facial images in the facial image sequence is considered as a set of image frames. For each set of image frames, the Euclidean distance between the coordinates of each facial key point in the corresponding set of facial key point coordinates is calculated, and the Euclidean distances of all facial key points are summed to obtain the corresponding inter-frame facial motion amount. The average inter-frame facial motion amount of all image frame pairs is calculated to obtain the average facial motion amplitude. The average facial motion amplitude is used to reflect the overall intensity of the user's facial expression changes. The mean vector of the action unit, the fluctuation vector of the action unit, and the average facial movement amplitude are integrated to form facial micro-expression features. Among them, facial micro-expression features are used to characterize the muscle movement patterns and dynamic changes of the user's facial expressions. It should be noted that the specific implementation methods of facial key point detection and facial action unit recognition are conventional technical means in this field, and will not be elaborated on here.
[0024] A pre-defined set of functional joint pairs is used, containing multiple functional joint pairs. Each functional joint pair is a combination of two human joints related to a task-performing action, such as a forearm manipulation joint pair formed by the wrist and elbow joints, an upper arm extension joint pair formed by the elbow and shoulder joints, and an arm extension joint pair formed by the wrist and shoulder joints. Each functional joint pair is pre-set by those skilled in the art based on human kinematics and common task-performing action patterns. For each frame of joint coordinate data in the limb posture data, the three-dimensional spatial distance between each functional joint pair in the set is calculated, forming a functional joint distance vector for each frame. For each functional joint pair in all functional joint distance vectors, the mean of the corresponding three-dimensional spatial distance is calculated, resulting in a mean functional joint distance vector. This mean functional joint distance vector reflects the average spatial relationship of each functional joint pair in the task-performing action. In the limb posture data, every two adjacent sets of joint coordinate data are taken as a joint frame pair. For each joint frame pair, the Euclidean distance between the three-dimensional spatial coordinates of each human joint is calculated to obtain the inter-frame displacement of each human joint. The average inter-frame displacement of the same human joint in all joint frame pairs is calculated to obtain the average motion velocity of each human joint. The average motion velocities of all human joints are summed to obtain the overall motion activity. The overall motion activity reflects the overall activity level of the user's limb movements. For each frame of joint coordinate data, the three-dimensional spatial distance between the left and right shoulder joints is calculated to obtain the shoulder spacing; the three-dimensional spatial distance between the left and right hip joints is calculated to obtain the hip spacing; the mean of the shoulder spacing and hip spacing is calculated to obtain the trunk spread; the mean and standard deviation of all trunk spread are calculated to obtain the average trunk spread and trunk spread fluctuation, respectively; the average trunk spread reflects the degree of openness or closure of the user's limb posture; the trunk spread fluctuation reflects the stability of the user's limb posture opening and closing state. For each frame of joint coordinate data, the mean of the three-dimensional spatial coordinates of the left and right shoulder joints and the left and right hip joints is calculated to obtain the torso center coordinates. The three-dimensional spatial distances from the left and right wrist joints to the torso center coordinates are calculated respectively, and the maximum value of the two is taken to obtain the gesture spatial range value. The gesture spatial range value is used to reflect the maximum unfolding distance of the hand movement relative to the torso in a single frame. The mean of all gesture spatial range values is calculated to obtain the average gesture spatial range. The average gesture spatial range is used to reflect the spatial unfolding amplitude of the user's gesture movement. The mean vector of functional joint distances, overall motion activity, average trunk extension, trunk extension fluctuation, and average gesture space range are integrated to form limb motion features. Among them, limb motion features are used to characterize the motion pattern, activity level, and posture characteristics of the user's limb movements.
[0025] Semantic content features, intonation and prosody features, facial micro-expression features, and body movement features are combined to form a multi-dimensional interactive feature set.
[0026] The dual-stream decoupling module is used to perform cognitive-emotional attribution determination on the multi-dimensional interaction feature set. It extracts the task-oriented component in the semantic content features and the task-oriented component in the body movement features into a cognitive intent feature stream, and extracts the emotional orientation component in the semantic content features, the emotional orientation component in the body movement features, the intonation and prosody features, and the facial micro-expression features into an emotional state feature stream.
[0027] Methods for determining cognitive-emotional attribution based on multidimensional interaction feature sets include: Semantic content features, intonation and prosody features, facial micro-expression features, and body movement features are obtained from a multi-dimensional interaction feature set. Cognitive-emotional attribution polarity analysis and component decoupling are performed on semantic content features. Specifically, text sequences, sentence-level semantic vectors, task-oriented density, and sentiment-oriented density are obtained from semantic content features. The sum of task-oriented density and sentiment-oriented density is calculated to obtain the total semantic orientation density. If the total semantic orientation density is greater than zero, the difference between task-oriented density and sentiment-oriented density is calculated and then divided by the total semantic orientation density to obtain the semantic attribution polarity value. If the total semantic orientation density is equal to zero, the semantic attribution polarity value is set to zero. A positive semantic attribution polarity value indicates that the semantic content is generally biased towards cognitive task orientation, a negative semantic attribution polarity value indicates that the semantic content is generally biased towards sentiment expression orientation, and a semantic attribution polarity value of zero indicates that cognitive task information and sentiment expression information in the semantic content are balanced. Based on the semantic affiliation polarity value, calculate the semantic cognitive affiliation coefficient and the semantic affective affiliation coefficient; the expression for the semantic cognitive affiliation coefficient is: The expression for the semantic sentiment attribution coefficient is: In the formula, This is the semantic cognitive attribution coefficient. The semantic sentiment attribution coefficient. The semantic attribution polarity value is denoted as follows: the semantic cognitive attribution coefficient reflects the proportion of cognitive task-oriented information in the semantic content; the semantic emotional attribution coefficient reflects the proportion of emotionally oriented information in the semantic content. Each element of the sentence-level semantic vector is multiplied by a semantic cognitive attribution coefficient to obtain a semantic cognitive projection vector; the semantic cognitive projection vector is used to represent the semantic information component in the sentence-level semantic vector that is related to the cognitive task. Each element of the sentence-level semantic vector is multiplied by a semantic sentiment attribution coefficient to obtain a semantic sentiment projection vector; the semantic sentiment projection vector is used to represent the semantic information component in the sentence-level semantic vector that is related to sentiment expression. The text sequence, the semantic cognitive projection vector, and the task orientation density are integrated to form the task orientation component of the semantic content feature. The semantic sentiment projection vector and the sentiment orientation density are integrated to form the sentiment orientation component of the semantic content feature.
[0028] The cognitive-emotional components of limb movement features are decoupled. Specifically, the mean vector of functional joint distances, overall motor activity, average trunk extension, trunk extension fluctuation, and average gesture spatial range are obtained from limb movement features. The attribution attributes of each sub-feature in the limb movement features are determined, and each sub-feature is divided into inherent attribution features and shared attribution features. Inherent attribution features are sub-features with natural and clear cognitive or emotional orientation attributes. Shared attribution features are sub-features that may carry both cognitive task information and emotional expression information, and need to be decomposed through a cross-modal attribution allocation mechanism. The mean vector of functional joint distances is identified as an inherent cognitive attribution feature. A functional joint pair is defined as a combination of two human joints related to task-performing actions. The mean vector of functional joint distances is calculated from the three-dimensional spatial distance of the functional joint pair, thus naturally possessing cognitive task-oriented attributes. Average trunk extension and trunk extension fluctuation are identified as inherent emotional attribution features. The degree of trunk opening and closing and stability primarily reflect the user's emotional state, possessing natural emotional attribution attributes. Overall motor activity and average gesture spatial range are identified as shared attribution features. Overall motor activity may reflect the frequency of actions when the user is actively performing a task, or the degree of limb agitation when the user is emotionally excited. The average gesture spatial range may reflect the spatial amplitude of the user's task-indicating gestures, or the spatial amplitude of the user's emotionally expressive gestures. For shared attribution features, a cross-modal attribution allocation mechanism is adopted for decomposition. This mechanism utilizes semantic cognitive attribution coefficients and semantic emotional attribution coefficients at the semantic level as cross-modal attribution guidance signals to allocate cognitive-emotional proportions for shared sub-features with ambiguous attribution attributes in limb movement features. Specifically, the product of overall motor activity and the semantic cognitive attribution coefficient is calculated to obtain the cognitive motor activity component; the product of overall motor activity and the semantic emotional attribution coefficient is calculated to obtain the emotional motor activity component. The cognitive motor activity component reflects the degree of limb movement activity driven by task execution; the emotional motor activity component reflects the degree of limb movement activity driven by emotional state. The product of the average gesture spatial range and the semantic cognitive attribution coefficient is calculated to obtain the cognitive gesture range component; the product of the average gesture spatial range and the semantic emotional attribution coefficient is calculated to obtain the emotional gesture range component. The cognitive gesture range component reflects the degree of gesture spatial expansion related to task execution; the emotional gesture range component reflects the degree of gesture spatial expansion related to emotional expression. The mean vector of functional joint distances, the cognitive motor activity component, and the cognitive gesture range component are integrated to form the task-oriented component of limb movement characteristics; the mean trunk extension, trunk extension fluctuation, the emotional motor activity component, and the emotional gesture range component are integrated to form the emotional-oriented component of limb movement characteristics.
[0029] Methods for extracting features into a cognitive intent feature stream include: Cross-modal cognitive synergy is achieved by fusing the task-oriented components of semantic content features and limb movement features to form a cognitive intent feature stream. Specifically, a median reference for movement activity is preset, which is pre-set by those skilled in the art based on the typical activity level of a user's limb movements. The ratio of overall movement activity to the joint scale of activity (the sum of overall movement activity and the median reference for movement activity) is calculated to obtain a limb activity index. The limb activity index is used to map overall movement activity into a normalized index reflecting the richness of limb modality signals. The task-oriented density is used as a semantic modality cognitive signal index. The product of the semantic modality cognitive signal index and the limb activity index is calculated, and the square root is taken to obtain the cross-modal cognitive synergy. The cross-modal cognitive synergy reflects the degree of synergy between the cognitive task signal of the semantic modality and the movement signal of the limb modality. The cross-modal cognitive synergy approaches a high value only when both modalities have high signal activity. A cognitive collaboration enhancement coefficient is preset, which is pre-set by those skilled in the art according to the cross-modal cognitive fusion enhancement requirements; the product of the cognitive collaboration enhancement coefficient and the cross-modal cognitive collaboration degree is calculated and then one is added to obtain the cognitive fusion enhancement factor; each element in the semantic cognitive projection vector is multiplied by the cognitive fusion enhancement factor to obtain the enhanced semantic cognitive vector; each element in the functional joint distance mean vector is multiplied by the cognitive fusion enhancement factor to obtain the enhanced functional joint vector; the enhanced semantic cognitive vector, the enhanced functional joint vector, the text sequence, the task pointing density, the cognitive motion activity component, and the cognitive gesture range component are integrated to form a cognitive intent feature stream.
[0030] Methods for extracting emotional state feature streams include: A cross-modal emotional synergy fusion is performed on the emotional orientation components of semantic content features, the emotional orientation components of body movement features, intonation prosody features, and facial micro-expression features to form an emotional state feature stream. Specifically, a median reference for prosodic fluctuation and a median reference for facial movement are preset. The median reference for prosodic fluctuation and the median reference for facial movement are respectively preset by those skilled in the art based on the typical fluctuation degree of the user's intonation prosody and the typical change amplitude of facial expressions. The ratio of the fundamental frequency fluctuation degree to the joint scale of fluctuation degree (the sum of the fundamental frequency fluctuation degree and the median reference for prosodic fluctuation) is calculated to obtain the prosodic activity index. The ratio of the average facial movement amplitude to the joint scale of movement amplitude (the sum of the average facial movement amplitude and the median reference for facial movement) is calculated to obtain the facial activity index. The prosodic activity index and the facial activity index are used to reflect the richness of emotional signals in the intonation prosody modality and the facial micro-expression modality, respectively. Emotional orientation density is used as a semantic modality emotion signal indicator. The fourth root of the product of the semantic modality emotion signal indicator, prosodic activity indicator, facial activity indicator, and limb activity indicator is calculated to obtain the cross-modal emotion synergy. The cross-modal emotion synergy reflects the overall degree of synergy among the four modalities in the direction of emotion signals. The cross-modal emotion synergy approaches a high value only when all modalities have high signal activity simultaneously. An emotion synergy enhancement coefficient is preset, which is pre-set by those skilled in the art according to the cross-modal emotion fusion enhancement requirements. The product of the emotion synergy enhancement coefficient and the cross-modal emotion synergy is calculated and then one is added to obtain the emotion fusion enhancement factor. Each element in the semantic emotion projection vector is multiplied by the emotion fusion enhancement factor to obtain the enhanced semantic emotion vector. Each element in the action unit mean vector is multiplied by the emotion fusion enhancement factor to obtain the enhanced action unit mean vector. The enhanced semantic emotion vector, emotion orientation density, enhanced action unit mean vector, action unit fluctuation vector, average facial movement amplitude, and the emotion orientation component of limb movement features are integrated with the intonation prosodic features to form an emotion state feature stream.
[0031] The cognitive decision-making module is used to perform intent matching and task planning based on the cognitive intent feature flow and a preset domain knowledge base, and generate a task response plan that includes response content, execution steps and task urgency.
[0032] Methods for intent matching and task planning based on cognitive intent feature streams include: A pre-defined domain knowledge base is used to store the intent categories that the virtual digital human can process and the corresponding task execution knowledge. Specifically, it includes an intent template library, a task execution graph library, and a gesture intent pattern library. The intent template library contains multiple intent templates, each containing an intent category identifier, a set of intent keywords, an intent prototype semantic vector, and an associated task graph encoding. The intent category identifier is a predefined unique identifier for a user intent category, including but not limited to information query, operation instructions, content recommendation, knowledge explanation, and process guidance. The intent keyword set is a collection of multiple intent keywords semantically related to the corresponding intent category, used to represent the core semantic elements of the intent category at the lexical level. The intent prototype semantic vector is a fixed-dimensional vector obtained by sentence-level semantic encoding of the typical semantic expression of the corresponding intent category, used to represent the standard semantic pattern of the intent category at the semantic space level. The associated task graph encoding is used to identify the task execution graph associated with the corresponding intent category. Each intent template is pre-constructed by those skilled in the art based on application scenarios and interaction requirements. The task execution graph library contains multiple task execution graphs. Each task execution graph includes a task graph code, a task complexity level, and multiple execution nodes. The task graph code is a predefined unique identifier for the task execution graph, corresponding to the associated task graph code in the intent template. The task complexity level reflects the logical complexity of the task execution, including simple, medium, and complex levels. Each execution node includes a step number, a step description, and a step response template. The step number is the sequential number of the execution node in the task execution flow. The step description is the operation instructions required for the corresponding execution node. The step response template is a text template with fillable slots, where slots are placeholders that need to be dynamically replaced based on the specific user request content. Each task execution graph is pre-constructed by those skilled in the art based on the task execution flow. The gesture intent pattern library contains multiple gesture intent patterns. Each gesture intent pattern includes a gesture pattern identifier, a gesture prototype joint vector, and an associated intent category identifier. The gesture pattern identifier is a predefined unique identifier for the gesture pattern. The gesture prototype joint vector is a typical functional joint distance distribution vector in the corresponding gesture pattern. The associated intent category identifier is used to identify the intent category pointed to by the corresponding gesture pattern. Each gesture intent pattern is pre-set by those skilled in the art based on common task-directing gestures in human-computer interaction.
[0033] The system extracts enhanced semantic cognitive vectors, enhanced functional joint vectors, text sequences, task pointing density, cognitive motion activity components, and cognitive gesture range components from the cognitive intent feature stream. Based on the enhanced semantic cognitive vectors, text sequences, and enhanced functional joint vectors, it performs multi-channel intent matching on the intent template library to obtain the comprehensive intent matching confidence for each intent template. Specifically, it calculates the cosine similarity between the enhanced semantic cognitive vector and the intent prototype semantic vector of each intent template in the intent template library to obtain the semantic matching degree for each intent template. The semantic matching degree reflects the semantic closeness between the user's current semantic expression and the standard semantic patterns of each intent category. Each word in the text sequence is matched against the intent keyword set of each intent template. The number of intent keywords that match the corresponding intent keyword set in the text sequence is counted to obtain the number of matched keywords. The total number of keywords is obtained by counting all intent keywords in each intent keyword set. The ratio of matched keywords to the total number of keywords is calculated to obtain the keyword coverage rate for each intent template. The keyword coverage rate is used to reflect the degree of vocabulary coverage of the core keywords of the intent template by the text sequence. The Euclidean distance between the enhanced function joint vector and the gesture prototype joint vector of each gesture intent pattern is calculated to obtain the gesture deviation distance corresponding to each gesture intent pattern. The gesture deviation distance reflects the degree of difference between the user's current limb posture and the preset gesture pattern. A preset gesture activation distance threshold is set by those skilled in the art based on the gesture recognition accuracy requirements. The gesture deviation distance corresponding to each gesture intent pattern is compared with the gesture activation distance threshold. If the gesture deviation distance is less than or equal to the gesture activation distance threshold, the corresponding gesture intent pattern is determined to be activated; otherwise, it is determined to be inactive. The associated intent category identifiers of all activated gesture intent patterns are obtained to form a gesture-associated intent set. For each intent template, if the corresponding intent category identifier exists in the gesture-associated intent set, the corresponding gesture auxiliary matching degree is set to a preset gesture confidence bonus value; if the corresponding intent category identifier does not exist in the gesture-associated intent set, the corresponding gesture auxiliary matching degree is set to zero. The gesture confidence bonus value is set by those skilled in the art based on the degree of auxiliary contribution of the gesture signal to intent determination. The gesture auxiliary matching degree is used to provide additional matching confidence supplementation for intent categories with limb gesture support. Semantic matching degree, keyword coverage and gesture-assisted matching degree are assigned corresponding matching weights. Based on the matching weights, the semantic matching degree, keyword coverage and gesture-assisted matching degree of the same intent template are weighted and summed to obtain the comprehensive intent matching confidence degree for each intent template. The comprehensive intent matching confidence degree is used to reflect the comprehensive matching degree between the user's current interaction behavior and each intent category. Each matching weight is preset by those skilled in the art based on the reliability of each matching channel.
[0034] The user intent recognition result is determined based on the comprehensive intent matching confidence score. Specifically, all intent templates are sorted from high to low according to their comprehensive intent matching confidence scores. The intent template with the highest comprehensive intent matching confidence score is selected, and its corresponding comprehensive intent matching confidence score is taken as the highest matching confidence score. The corresponding intent category is identified as the preferred intent category. A preset intent confidence score threshold is established, which is pre-set by those skilled in the art based on the intent recognition accuracy requirements. The highest matching confidence score is compared with the intent confidence score threshold. If the highest matching confidence score is greater than or equal to the intent confidence score threshold, the intent recognition is considered successful, and the preferred intent category is taken as the user intent recognition result. If the highest matching confidence score is less than the intent confidence score threshold, the intent recognition is considered uncertain, and the user intent recognition result is marked as an intent requiring clarification. The template with the highest comprehensive intent matching confidence score is ranked higher. The intent category identifiers corresponding to the intent templates of each bit form a candidate intent category set; among which, The number of candidates is a preset number, which is pre-set by those skilled in the art based on the interactive clarification strategy.
[0035] Based on the user intent recognition results, task planning is performed to generate response content and a sequence of execution steps. Specifically, if the user intent recognition result is the preferred intent category, the associated task graph code is obtained from the corresponding intent template. The corresponding task execution graph is then retrieved from the task execution graph library based on the associated task graph code and marked as the adapted execution graph. Entity information extraction is performed on the text sequence, identifying words related to the specific object requested by the user, resulting in a set of entity information. Entity information includes, but is not limited to, object name, quantity description, time description, and location description. The specific implementation method for entity information extraction is a conventional technique in this field and will not be elaborated upon here. The task complexity level and all execution nodes corresponding to the adapted execution graph are obtained, and the entity information in the entity information set is matched with each execution node. The system matches the slots in the step response template, fills the corresponding slots with the successfully matched entity information, and generates the step response text for each execution node. All execution nodes are arranged in ascending order of step number to form an execution step sequence, and the step response text corresponding to the execution node with the smallest step number is used as the response content. This response content serves as the first-round response text from the virtual digital human to the current user's request. If the user intent identification result is an intent to be clarified, a clarification query text is generated based on the intent keyword set corresponding to each intent category identifier in the candidate intent category set. This clarification query text is used as the response content, the execution step sequence is set to an empty sequence, and the task complexity level is set to simple. The clarification query text is an inquiry-based response text that guides the user to further clarify their interaction intent.
[0036] Task urgency is calculated based on task orientation density, cognitive-motor activity component, cognitive gesture range component, highest matching confidence level, and task complexity level. Specifically, a preset upper limit for motor activity and a preset upper limit for gesture range are established, each pre-set by a person skilled in the art based on the statistical characteristics of user interaction behavior. The ratio of the cognitive-motor activity component to the upper limit for motor activity is calculated to obtain a standard motor urgency index. This standard motor urgency index reflects the relative activity level of the user's task-performing limb movements. If the standard motor urgency index is greater than one, then... The standard motion urgency index is set to one; the ratio of the cognitive gesture range component to the upper limit of the gesture range reference is calculated to obtain the standard gesture urgency index; the standard gesture urgency index reflects the relative spatial expansion of the user's task-directing gestures; if the standard gesture urgency index is greater than one, it is set to one; corresponding urgency weights are assigned to the task pointing density, standard motion urgency index, and standard gesture urgency index, respectively, and a weighted sum is calculated based on the urgency weights to obtain the multimodal urgency value; each urgency weight is determined by... Those skilled in the art pre-set the degree of urgency reflected by each indicator; the multimodal urgency value is used to reflect the degree of urgency of the task conveyed by the user through semantic expression and body language; a complexity adjustment coefficient is preset for each task complexity level; among them, the complexity adjustment coefficients corresponding to the simple, medium and complex levels increase sequentially; each complexity adjustment coefficient is pre-set by those skilled in the art according to the impact of different levels of task complexity on response priority; the complexity adjustment coefficient is used to adjust the task complexity dimension of the basic urgency, and the higher the complexity level, the more steps the task involves, and the more... The system needs to prioritize resource allocation for response; based on the task complexity level, obtain the corresponding complexity adjustment coefficient; calculate the product of the multimodal urgency value and the complexity adjustment coefficient to obtain the complexity adjustment urgency; calculate the product of the complexity adjustment urgency and the highest matching confidence to obtain the task urgency; by including the highest matching confidence in the calculation, it is ensured that the system will not respond to potentially inaccurate tasks with high urgency when the intent is ambiguous; the task urgency is used to reflect the overall urgency of the user's current task request. The higher the task urgency, the more urgent the user's task demand, the clearer the intent, and the more priority the task needs to be responded to.
[0037] The response content, execution step sequence, and task urgency are integrated to form a task response plan; the task response plan is used to provide the virtual digital human with task execution instructions and response content basis at the cognitive decision-making level.
[0038] The emotion decision module is used to infer the emotion from the distribution characteristics and intensity changes of each feature in the emotion state feature stream, determine the user's emotion category and emotion intensity value, and generate an emotion response plan by matching the corresponding empathy response level and empathy expression mode according to the emotion category and emotion intensity value.
[0039] Methods for determining a user's emotion category and emotion intensity include: A pre-defined emotion prototype library is provided, containing multiple emotion prototypes. Each emotion prototype includes an emotion category label, a semantic emotion prototype vector, a facial action unit prototype vector, a prosodic prototype parameter set, and a body posture prototype parameter set. The emotion category label is a unique identifier for a predefined emotion category, including but not limited to happiness, sadness, anger, anxiety, surprise, and calmness. The semantic emotion prototype vector is a typical emotion expression vector for the corresponding emotion category in the semantic space. The facial action unit prototype vector is a typical facial action unit activation vector for the corresponding emotion category. The prosodic prototype parameter set includes typical average speech fundamental frequency, typical fundamental frequency fluctuation, typical average speech energy, typical energy fluctuation, and typical average speech rate for the corresponding emotion category. The body posture prototype parameter set includes typical average trunk extension, typical trunk extension fluctuation, typical emotional motor activity components, and typical emotional gesture range components for the corresponding emotion category. Each emotion prototype is pre-constructed by those skilled in the art based on emotion psychology theory and a multimodal emotion database.
[0040] For each emotion prototype, an emotion evidence score for each modality is calculated; the modalities are, in order, semantic modality, facial modality, prosodic modality, and somatic modality; specifically, the cosine similarity between the enhanced semantic emotion vector and the semantic emotion prototype vector of each emotion prototype is calculated to obtain the semantic emotion evidence score corresponding to each emotion prototype; the semantic emotion evidence score is used to reflect the degree of closeness between the user's semantic expression and each emotion category in the semantic space; The Euclidean distance between the mean vector of the enhanced action unit and the prototype vector of the facial action unit for each emotion prototype is calculated to obtain the facial prototype deviation distance for each emotion prototype. A facial distance benchmark is preset, which is pre-set by a person skilled in the art based on the typical range of variation of the activation intensity of the facial action unit. The ratio of the facial prototype deviation distance to the facial distance benchmark is calculated to obtain the standard facial deviation. The difference between the standard facial deviation and the standard facial deviation is calculated to obtain the facial emotion evidence score. If the facial emotion evidence score is less than zero, the facial emotion evidence score is set to zero. The facial emotion evidence score is used to reflect the degree of conformity between the user's facial expression pattern and the typical facial pattern corresponding to each emotion category. The user prosodic feature vector is constructed by sequentially combining the average fundamental frequency, fundamental frequency fluctuation, average speech energy, energy fluctuation, and average speech rate. Similarly, the prototype prosodic feature vector is constructed by sequentially combining the parameters in the prosodic prototype parameter group for each emotion prototype. The Euclidean distance between the user prosodic feature vector and each prototype prosodic feature vector is calculated to obtain the prosodic prototype deviation distance for each emotion prototype. A prosodic distance benchmark is preset, which is pre-set by a person skilled in the art based on the typical variation range of intonation prosodic parameters. The ratio of the prosodic prototype deviation distance to the prosodic distance benchmark is calculated to obtain the standard prosodic deviation. The difference between the standard prosodic deviation and the prosodic deviation is calculated to obtain the prosodic emotion evidence score. If the prosodic emotion evidence score is less than zero, it is set to zero. The prosodic emotion evidence score reflects the degree of agreement between the user's intonation prosodic pattern and the typical prosodic pattern corresponding to each emotion category. The average torso extension, torso extension fluctuation, emotional motor activity component, and emotional gesture range component are sequentially combined to form the user's body feature vector. The parameters in the body posture prototype parameter group for each emotional prototype are sequentially combined to form the prototype body feature vector. The Euclidean distance between the user's body feature vector and each prototype body feature vector is calculated to obtain the body prototype deviation distance for each emotional prototype. A body distance benchmark is preset, which is pre-set by a person skilled in the art based on the typical variation range of body posture parameters. The ratio of the body prototype deviation distance to the body distance benchmark is calculated to obtain the standard body deviation. The difference between the standard body deviation and the body deviation is calculated to obtain the body emotional evidence score. If the body emotional evidence score is less than zero, it is set to zero. The body emotional evidence score reflects the degree of conformity between the user's limb posture pattern and the typical body pattern corresponding to each emotional category.
[0041] Based on the activity level of each modality in the emotional state feature stream, dynamic fusion weights corresponding to each modality are calculated. Specifically, emotional orientation density is used as the semantic modality activity index; the ratio of the average facial motion amplitude to a preset facial motion activity benchmark is used as the facial modality activity index; if the facial modality activity index is greater than one, it is set to one; wherein, the facial motion activity benchmark is preset by those skilled in the art based on the typical statistical distribution of facial motion amplitude; the mean of fundamental frequency fluctuation and energy fluctuation is calculated to obtain the prosodic comprehensive fluctuation; the ratio of the prosodic comprehensive fluctuation to a preset prosodic fluctuation activity benchmark is used as the prosodic modality activity index; if the prosodic modality activity index is greater than one, it is set to one; wherein, the prosodic fluctuation activity benchmark is preset by those skilled in the art based on the typical statistical distribution of prosodic fluctuation. The system is configured as follows: the ratio of the emotional motor activity component to a preset somatic motor activity benchmark is used as the somatic modality activity index. If the somatic modality activity index is greater than one, it is set to one. The somatic motor activity benchmark is preset by a person skilled in the art based on the typical statistical distribution of limb movements. The sum of the semantic modality activity index, facial modality activity index, prosodic modality activity index, and somatic modality activity index is calculated to obtain the total activity index. If the total activity index is greater than zero, the ratio of each modality activity index to the total activity index is calculated to obtain the dynamic fusion weight of each modality. If the total activity index is equal to zero, the dynamic fusion weight of each modality is set to 0.25. Each dynamic fusion weight is used to adaptively adjust the contribution ratio of each modality emotional evidence in the fusion judgment according to the richness of each modality emotional signal in the current interaction.
[0042] Based on the dynamic fusion weights of each modality, the sentiment evidence scores of each modality corresponding to the same sentiment prototype are weighted and summed to obtain the comprehensive sentiment matching score for each sentiment prototype. All sentiment prototypes are sorted from high to low according to the comprehensive sentiment matching score, and the sentiment prototype with the highest comprehensive sentiment matching score is selected. The corresponding sentiment category label is taken as the user sentiment category, and the corresponding comprehensive sentiment matching score is taken as the preferred sentiment matching score. Among them, the sentiment category is used to identify the dominant sentiment type expressed by the user in the current interaction.
[0043] The user's emotional intensity value is calculated based on the deviation of each modality from the neutral state in the emotional state feature stream. Specifically, an emotional neutrality benchmark library is preset, which contains benchmark parameters for each modality in the emotional neutral state, including neutral semantic emotion vector, neutral action unit vector, neutral prosodic feature vector, and neutral body feature vector. Each benchmark parameter is preset by those skilled in the art based on typical multimodal data in the emotional neutral state. The semantic emotion deviation distance is obtained by calculating the Euclidean distance between the enhanced semantic emotion vector and the neutral semantic emotion vector; the facial emotion deviation distance is obtained by calculating the Euclidean distance between the enhanced action unit mean vector and the neutral action unit vector; the prosodic emotion deviation distance is obtained by calculating the Euclidean distance between the user prosodic feature vector and the neutral prosodic feature vector; and the bodily emotion deviation distance is obtained by calculating the Euclidean distance between the user bodily feature vector and the neutral bodily feature vector. Preset maximum deviation benchmarks for semantics, face, prosody, and body are established, each pre-set by a person skilled in the art based on the maximum reasonable deviation range of each modality's emotional signal. The ratios of semantic emotional deviation distance to the maximum semantic deviation benchmark, facial emotional deviation distance to the maximum facial deviation benchmark, prosodic emotional deviation distance to the maximum prosodic deviation benchmark, and body emotional deviation distance to the maximum body deviation benchmark are calculated to obtain the semantic emotional deviation degree, facial emotional deviation degree, prosodic emotional deviation degree, and body emotional deviation degree. If any deviation degree is greater than one, the corresponding deviation degree is set to one. Based on dynamic fusion weights, the semantic emotional deviation degree, facial emotional deviation degree, prosodic emotional deviation degree, and body emotional deviation degree are weighted and summed to obtain the comprehensive emotional deviation degree. The mean of all elements in the motion unit fluctuation vector is calculated to obtain the mean of facial motion fluctuation; the sum of the mean of facial motion fluctuation and the average facial movement amplitude is calculated and then divided by two to obtain the facial dynamic intensity index; the facial dynamic intensity index is used to reflect the temporal dynamic activity of facial expression changes; a facial dynamic intensity benchmark is preset, which is preset by those skilled in the art based on the typical distribution range of facial dynamic intensity; the ratio of the facial dynamic intensity index to the facial dynamic intensity benchmark is calculated to obtain the standard facial dynamic intensity; if the standard facial dynamic intensity is greater than one, the standard facial dynamic intensity is set to one. The system assigns corresponding intensity weights to the overall emotional deviation and the standard facial dynamic intensity, and performs a weighted summation based on these weights to obtain the original emotional intensity value. Each intensity weight is pre-set by a person skilled in the art according to the emotional intensity assessment requirements. The emotional intensity value is obtained by multiplying the original emotional intensity value by the preferred emotional matching score. By including the preferred emotional matching score in the calculation, the system ensures that it does not report excessively high emotional intensity when the emotional classification confidence is low. The emotional intensity value reflects the overall intensity of the user's current emotional expression; a higher emotional intensity value indicates a stronger emotional expression.
[0044] Methods for generating emotional response plans include: A pre-defined empathy strategy library contains multiple empathy strategy entries. Each entry includes an applicable emotion category, an applicable intensity range, an empathy response level, and an empathy expression method. The applicable emotion category is the emotion category label corresponding to the empathy strategy entry; the applicable intensity range is the range of emotion intensity values applicable to the empathy strategy entry, expressed as a closed interval; the empathy response level includes mild empathy, moderate empathy, and deep empathy; the empathy expression method includes empathy tone labels and empathy action labels. The empathy tone label indicates the tone type the virtual digital person should use when responding, including but not limited to gentle reassurance, positive encouragement, calm guidance, and enthusiastic response; the empathy action label indicates the facial and body language expressions the virtual digital person should use during interaction, including but not limited to smiling and nodding, concerned gaze, slight forward leaning, and gesture reassurance. Each empathy strategy entry is pre-constructed by those skilled in the art based on emotional interaction design theory and user experience research. Based on the user's emotion category and emotion intensity value, a matching search is performed from the empathy strategy library. Specifically, empathy strategy entries that match the user's emotion category and whose emotion intensity value falls within the applicable intensity range are retrieved from the empathy strategy library and marked as matching strategy entries. The corresponding empathy response level and empathy expression mode are obtained from the matching strategy entries. Among them, the empathy response level is used to indicate the depth level of the virtual digital human's emotional response to the user; the empathy expression mode is used to provide emotional expression guidance for the generation of the virtual digital human's tone of voice and facial body movements. The system integrates user emotion categories, emotion intensity values, empathy response levels, and empathy expression methods to form an emotion response scheme. This scheme provides virtual digital humans with empathy expression instructions and emotional response criteria at the emotional decision-making level.
[0045] The dual-stream driving module is used to dynamically adjust the fusion ratio of cognitive decision-making and emotional decision-making based on the task urgency in the task response plan and the emotional intensity value in the emotional response plan. Based on the fusion ratio, the task response plan and the emotional response plan are collaboratively fused to determine the final response content, tone style parameters and facial expression parameters of the virtual digital human, and drive the virtual digital human to generate interactive responses that combine task execution and emotional empathy.
[0046] Methods for dynamically adjusting the integration ratio of cognitive and affective decision-making include: Based on task urgency and emotional intensity values, channel interaction status parameters are calculated. These parameters include the channel tension index, total channel activity, and channel dominance polarity. Specifically, the channel tension index is calculated by multiplying the task urgency and emotional intensity values. This index quantifies the intensity of competition between the cognitive and emotional channels for response dominance; a higher index indicates a greater conflict in response needs between the two channels. The total channel activity is calculated by summing the task urgency and emotional intensity values. If the total activity is greater than zero, the channel dominance polarity is calculated by dividing the difference between the task urgency and emotional intensity values by the total activity. If the total activity is zero, the channel dominance polarity is set to zero. A positive value indicates that the cognitive channel has a response advantage, while a negative value indicates that the emotional channel has a response advantage. The channel tension index is compared with a preset tension conflict threshold, which is preset by those skilled in the art based on the typical critical level of cognitive-emotional conflict in the interaction scenario. If the channel tension index is less than or equal to the tension conflict threshold, the current interaction is determined to be in a coordinated fusion mode. If the channel tension index is greater than the tension conflict threshold, the current interaction is determined to be in an adversarial fusion mode. The coordinated fusion mode indicates that there is no significant conflict between the response needs of the two channels, and fusion can be achieved through linear proportional allocation. The adversarial fusion mode indicates that both channels have strong and competitive response needs, and nonlinear mutual inhibition adjustment is required to ensure the effectiveness and coordination of the response. If the current interaction is in a coordinated fusion mode and the total number of active channels is greater than zero, then calculate the ratio of task urgency to total number of active channels to obtain the cognitive fusion ratio, and calculate the difference between the cognitive fusion ratio and the emotional fusion ratio to obtain the emotional fusion ratio; if the current interaction is in a coordinated fusion mode and the total number of active channels is equal to zero, then set both the cognitive fusion ratio and the emotional fusion ratio to 0.5. If the current interaction is in an adversarial fusion mode, a mutual inhibition adjustment mechanism is used to calculate the fusion ratio. Specifically, the ratio of the channel tension index to the joint tension scale (the sum of the channel tension index and the tension conflict threshold) is calculated to obtain the mutual inhibition strength factor. This mutual inhibition strength factor quantifies the nonlinear suppression of the inferior channel by the dominant channel; the larger the channel tension index, the closer the mutual inhibition strength factor is to one, and the more significant the suppression effect. The ratio of task urgency to total channel activity is calculated to obtain the initial cognitive ratio. If the channel dominance polarity value is greater than zero, the difference between one and the initial cognitive ratio is calculated to obtain the initial emotional ratio. The product of the mutual inhibition strength factor and the initial emotional ratio is calculated to obtain the cognitive enhancement amount. The sum of the initial cognitive proportion and the cognitive enhancement amount is used to obtain the cognitive fusion proportion. If the channel dominance polarity value is less than zero, the difference between the first and the mutual inhibition intensity factor is calculated to obtain the subordination retention coefficient. The product of the initial cognitive proportion and the subordination retention coefficient is calculated to obtain the cognitive fusion proportion. If the channel dominance polarity value is equal to zero, the cognitive fusion proportion is set to 0.5. The difference between the first and the cognitive fusion proportion is calculated to obtain the emotional fusion proportion. Among them, the mutual inhibition regulation mechanism is different from the traditional linear weighted fusion method. By nonlinearly enhancing the dominant channel and nonlinearly compressing the inferior channel in high conflict situations, the system can autonomously determine a clear dominant response direction and avoid the fuzzy output caused by compromise fusion.
[0047] Methods for synergistically integrating task response and emotional response strategies include: Based on the current fusion modality and fusion ratio of the interaction, the final response content is determined. Specifically, a pre-defined emotional response corpus is used, containing multiple emotional response templates. Each template includes an applicable emotional category, an applicable empathy level, an emotional response statement, and an emotional closing statement. The emotional response statement is an empathic opening text matching the corresponding emotional category and empathy level; the emotional closing statement is an empathic closing text matching the corresponding emotional category and empathy level. Each emotional response template is pre-constructed by those skilled in the art based on emotional interaction corpora and empathy expression strategies. Based on the user's emotional category and empathy response level, an emotional response template matching the user's emotional category and empathy level is retrieved from the emotional response corpus, and the corresponding emotional response statement and emotional closing statement are obtained. If the current interaction is in a coordinated fusion modality, the emotional response statement is appended to the response content to form the final response content. If the current interaction is in an adversarial fusion mode, a phased response organization strategy is used to generate the final response content. Specifically, the dominant channel for the response is determined based on the channel dominance polarity value. If the channel dominance polarity value is greater than or equal to zero, the cognitive channel is the dominant channel; if the channel dominance polarity value is less than zero, the affective channel is the dominant channel. If the cognitive channel is the dominant channel, the affective response statement is semantically simplified to obtain a simplified empathy preamble, which is then appended to the response content to form the final response content. If the affective channel is the dominant channel, key points are extracted from the response content to obtain a simplified task summary, which is then concatenated with the affective response statement, the simplified task summary, and the affective closing statement in sequence. This process forms the final response content. The phase-separated response organization strategy assigns independent expression phases to different channels within the temporal structure of the response, ensuring that the expression integrity of the dominant channel is not diluted, while key information from subordinate channels is preserved in appropriate positions. This differs from the traditional fusion method of mixing the content of two channels at a single point in time. It should be noted that semantic simplification can employ methods such as text summarization, sentence compression, and key sentence extraction; key point extraction can employ methods such as key phrase extraction, named entity extraction, and TextRank-based core sentence extraction. The specific implementation methods of semantic simplification and key point extraction are conventional techniques in this field and will not be elaborated upon here.
[0048] Based on the fusion ratio, tone style parameters are calculated. Specifically, a task-standard tone parameter set is preset, which includes a task-standard speech rate coefficient, a task-standard pitch coefficient, and a task-standard tone temperature coefficient. An emotional tone parameter set is also preset for each empathy tone tag, with each set including an emotional speech rate coefficient, an emotional pitch coefficient, and an emotional tone temperature coefficient. Both the task-standard tone parameter set and each emotional tone parameter set are preset by those skilled in the art based on the tone expression characteristics of the virtual digital human in the corresponding interactive context. The speech rate coefficient is used to control the virtual digital human's speech synthesis. The speech rate coefficient is used to control the pitch of the synthesized speech; the tone temperature coefficient is used to control the gentleness of the tone, with a higher tone temperature coefficient indicating a gentler tone and a lower tone indicating a calmer and more restrained tone; the corresponding emotional tone parameter set is obtained based on the empathic tone label; each coefficient in the task standard tone parameter set and the emotional tone parameter set is weighted and summed according to the proportion of cognitive fusion and the proportion of emotional fusion to obtain the fusion speech rate coefficient, fusion pitch coefficient, and fusion tone temperature coefficient; the fusion speech rate coefficient, fusion pitch coefficient, and fusion tone temperature coefficient are integrated to form the tone style parameter.
[0049] Based on the fusion ratio, facial expression parameters are calculated. Specifically, a set of standard task action parameters is preset, which includes a standard task expression identifier and a standard task action amplitude coefficient. A set of emotional action parameters corresponding to each empathic action tag is also preset, with each set including an emotional expression identifier and an emotional action amplitude coefficient. Both the standard task action parameter set and each emotional action parameter set are pre-set by those skilled in the art based on the facial and body expression characteristics of the virtual digital human in the corresponding interactive context. The expression identifier is used to specify the type of facial expression of the virtual digital human; the action amplitude coefficient is used to control the expression amplitude of facial expressions and body movements. The corresponding emotional action parameter set is obtained based on the empathic action tag. If the cognitive fusion ratio is greater than or equal to the emotional fusion ratio, the standard task expression identifier is used as the final expression identifier; otherwise, the emotional expression identifier is used as the final expression identifier. The standard task action amplitude coefficient and the emotional action amplitude coefficient are weighted and summed according to the cognitive fusion ratio and the emotional fusion ratio to obtain the fusion action amplitude coefficient. The final expression identifier and the fusion action amplitude coefficient are integrated to form the facial expression parameters.
[0050] Cross-channel expression consistency verification is performed on tone style parameters and facial expression parameters. Specifically, the fused tone temperature coefficient is used as the tone expression valence, and the fused action amplitude coefficient is used as the action expression valence. The tone expression valence characterizes the intensity of emotional expression in the tone dimension, and the action expression valence characterizes the intensity of emotional expression in the action dimension. The absolute value of the difference between the tone expression valence and the action expression valence is calculated to obtain the cross-channel expression deviation. A preset expression consistency threshold and consistency correction step size are established, both preset by those skilled in the art according to the requirements of virtual digital human expression coordination. If the cross-channel expression deviation is greater than the expression consistency threshold, consistency correction is performed. The sum of the cognitive tone contribution value (the product of cognitive fusion ratio and tone expression valence) and the emotional action contribution value (the product of emotional fusion ratio and action expression valence) is calculated to obtain the target coordination valence. The target coordination valence represents the valence at the current fusion level. The system calculates the consistency target value for tone and action under the given proportion; it calculates the difference between the target coordination valence and the tone expression valence, multiplies it by the consistency correction step size to obtain the tone correction amount; it updates the fused tone temperature coefficient in the tone style parameters based on the sum of the fused tone temperature coefficient and the tone correction amount; it calculates the difference between the target coordination valence and the action expression valence, multiplies it by the consistency correction step size to obtain the action correction amount; it updates the fused action amplitude coefficient in the facial expression parameters based on the sum of the fused action amplitude coefficient and the action correction amount; if the cross-channel expression deviation is less than or equal to the expression consistency threshold, no consistency correction is performed; it should be noted that the cross-channel expression consistency verification uses the target coordination valence weighted by the fused proportion as the correction direction to ensure that the virtual digital human's voice tone and facial body expression remain coordinated in the emotional expression dimension, avoiding contradictory interaction between tone and expression, and the correction direction is consistent with the channel dominance relationship determined by the mutual inhibition regulation.
[0051] The final response content, tone style parameters, facial expression and action parameters, and execution step sequence are integrated to form a comprehensive interactive driving command. The virtual digital human is then driven to execute the interactive response based on the comprehensive interactive driving command. Specifically, the speech synthesis engine that drives the virtual digital human converts the final response content into interactive speech according to the tone style parameters, and the facial expression animation engine that drives the virtual digital human synchronously generates facial expressions and body movements according to the facial expression and action parameters, outputting an interactive response that combines task execution and emotional empathy.
[0052] This embodiment simultaneously collects the user's voice signal, facial image sequence, and body posture data, and extracts multi-dimensional interaction features from four dimensions: semantic content, intonation and rhythm, facial micro-expressions, and body movements. This enables comprehensive perception and multi-level representation of the user's interactive behavior, thereby effectively overcoming the shortcomings of relying solely on a single voice modality or a limited combination of modalities to understand the user's intent. This embodiment uses a cognitive-emotional attribution determination and cross-modal attribution allocation mechanism based on semantic attribution polarity values to systematically decouple inherent attribution features with clear directional attributes and shared attribution features with ambiguous attribution attributes from multidimensional interaction features, forming independent cognitive intent feature streams and emotional state feature streams. This achieves accurate separation of users' explicit task requirements and implicit emotional states, avoiding the ambiguity of intent determination and inaccurate emotion recognition caused by mixing cognitive and emotional information. This embodiment employs a cross-modal collaborative fusion enhancement mechanism, utilizing the geometric mean of multimodal signal activity to measure the degree of collaboration between modalities. Feature enhancement is performed only when multiple modalities simultaneously exhibit active signals, effectively improving the reliability and robustness of cross-modal information complementarity verification. By fusing semantic matching, keyword coverage, and gesture-assisted matching into a multi-channel intent recognition method, and combining task complexity level and intent matching confidence to calculate task urgency, it achieves comprehensive intent determination that balances semantic understanding depth, vocabulary matching accuracy, and limb gesture assistance. This effectively improves the accuracy of user intent recognition and the rationality of task response priority in complex interaction scenarios. This embodiment achieves adaptive emotion recognition under different channel qualities and user expression habits by dynamically calculating the fusion weight based on the real-time activity level of each modality. It effectively overcomes the problem of decreased recognition accuracy of fixed weight fusion in noisy environments or modality missing scenarios. This embodiment introduces a channel tension index and a fusion modality dual-state determination mechanism to distinguish between a coordinated fusion modality and an adversarial fusion modality in the cognitive-emotional fusion process. In the adversarial fusion modality, a mutual inhibition adjustment mechanism is employed to achieve nonlinear enhancement of the dominant channel and nonlinear compression of the inferior channel. This effectively solves the problem of ambiguous responses and unclear priorities in high-conflict situations caused by a simple linear mixture of cognitive and emotional decisions. Through a phase-separated response organization strategy, independent expression phases are assigned to different channels in the temporal structure of the response, ensuring the integrity of the dominant channel's expression and the appropriate preservation of key information in subordinate channels. This achieves structured coordination of cognitive task content and emotional empathy content in the response text. Through a cross-channel expression consistency verification mechanism, the tone and facial expression parameters are bidirectionally corrected using a fusion-weighted target coordination valence as the correction direction. This ensures the consistency of the virtual digital human's voice tone and facial body expression in the emotional dimension, effectively avoiding unnatural interactions where tone and expression contradict each other. Ultimately, this achieves high-quality intelligent interactive responses from the virtual digital human while balancing task performance and emotional empathy.
[0053] Example 2
[0054] Please see Figure 2As shown, parts not described in detail in this embodiment are described in Embodiment 1. A multimodal data fusion intelligent interaction method based on virtual digital humans is provided, the method including: Collect users’ voice signals, facial image sequences and body posture data, and extract semantic content features, intonation and prosody features, facial micro-expression features and body movement features to form a multi-dimensional interactive feature set; Cognitive-emotional attribution determination is performed on the multidimensional interactive feature set. The task-oriented component in the semantic content feature and the task-oriented component in the body movement feature are extracted into a cognitive intention feature stream. The emotional orientation component in the semantic content feature, the emotional orientation component in the body movement feature, the intonation and prosody feature, and the facial micro-expression feature are extracted into an emotional state feature stream. Based on the cognitive intent feature flow, and combined with a pre-set domain knowledge base, intent matching and task planning are performed to generate a task response plan that includes response content, execution steps and task urgency. Emotion inference is performed on the distribution characteristics and intensity changes of each feature in the emotional state feature stream to determine the user's emotional category and emotional intensity value. Based on the emotional category and emotional intensity value, the corresponding empathy response level and empathy expression mode are matched to generate an emotional response scheme. Based on the task urgency in the task response plan and the emotional intensity in the emotional response plan, the fusion ratio of cognitive decision-making and emotional decision-making is dynamically adjusted. Based on the fusion ratio, the task response plan and the emotional response plan are synergistically fused to determine the final response content, tone style parameters and facial expression parameters of the virtual digital human, thereby driving the virtual digital human to generate interactive responses that combine task execution and emotional empathy.
[0055] Example 3
[0056] This application also provides an electronic device. The electronic device may include one or more processors and one or more memories. The memories store computer-readable code that, when executed by the one or more processors, can perform the multimodal data fusion intelligent interaction method based on a virtual digital human as described above.
[0057] The method or system according to the embodiments of this application can also be implemented using the architecture of the electronic device shown in this application. The electronic device may include a bus, one or more CPUs, ROM, RAM, a communication port connected to a network, input / output, a hard disk, etc. The storage device in the electronic device, such as a ROM or hard disk, may store the multimodal data fusion intelligent interaction method based on virtual digital humans provided in this application. Furthermore, the electronic device may also include a user interface. Of course, the architecture shown in this application is merely exemplary; when implementing different devices, one or more components in the electronic device shown in this application may be omitted according to actual needs.
[0058] Example 4
[0059] One embodiment of this application discloses a computer-readable storage medium. The computer-readable storage medium stores computer-readable instructions. When the computer-readable instructions are executed by a processor, the multimodal data fusion intelligent interaction method based on a virtual digital human, as described in the above-described embodiments of this application, can be executed. The storage medium includes, but is not limited to, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.
[0060] Furthermore, according to embodiments of this application, the processes described in the above-referenced flowcharts can be implemented as computer software programs. For example, this application provides a non-transitory machine-readable storage medium storing machine-readable instructions that can be executed by a processor to perform instructions corresponding to the method steps provided in this application, such as a multimodal data fusion intelligent interaction method based on a virtual digital human. When this computer program is executed by a central processing unit (CPU), it performs the functions defined in the method of this application.
[0061] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
[0062] All formulas in this manual are dimensionless and calculated numerically. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters and thresholds in the formulas are set by those skilled in the art according to the actual situation.
[0063] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.
Claims
1. A multimodal data fusion intelligent interaction method based on virtual digital humans, characterized in that: include: Collect users’ voice signals, facial image sequences and body posture data, and extract semantic content features, intonation and prosody features, facial micro-expression features and body movement features to form a multi-dimensional interactive feature set; Cognitive-emotional attribution determination is performed on the multidimensional interactive feature set. The task-oriented component in the semantic content feature and the task-oriented component in the body movement feature are extracted into a cognitive intention feature stream. The emotional orientation component in the semantic content feature, the emotional orientation component in the body movement feature, the intonation and prosody feature, and the facial micro-expression feature are extracted into an emotional state feature stream. Based on the cognitive intent feature flow, and combined with a pre-set domain knowledge base, intent matching and task planning are performed to generate a task response plan that includes response content, execution steps and task urgency. Emotion inference is performed on the distribution characteristics and intensity changes of each feature in the emotional state feature stream to determine the user's emotional category and emotional intensity value. Based on the emotional category and emotional intensity value, the corresponding empathy response level and empathy expression mode are matched to generate an emotional response scheme. Based on the task urgency in the task response plan and the emotional intensity in the emotional response plan, the fusion ratio of cognitive decision-making and emotional decision-making is dynamically adjusted. Based on the fusion ratio, the task response plan and the emotional response plan are synergistically fused to determine the final response content, tone style parameters and facial expression parameters of the virtual digital human, thereby driving the virtual digital human to generate interactive responses that combine task execution and emotional empathy.
2. The multimodal data fusion intelligent interaction method based on virtual digital humans according to claim 1, characterized in that, Methods for forming multidimensional interaction feature sets include: Endpoint detection is performed on the speech signal to extract the effective speech segment; speech recognition and semantic parsing are performed on the effective speech segment to obtain the text sequence and sentence-level semantic vector; task keyword extraction and sentiment word recognition are performed on the text sequence to obtain the task orientation density and sentiment orientation density, and the semantic content features are formed by combining the text sequence and sentence-level semantic vector. Frame-by-frame acoustic prosodic analysis is performed on effective speech segments to obtain the average speech fundamental frequency, fundamental frequency fluctuation, average speech energy, energy fluctuation, and average speech rate, and to form intonation prosodic features. Facial motion coding and inter-frame dynamic change analysis are performed on facial image sequences to obtain the mean vector of motion units, the fluctuation vector of motion units, and the average facial motion amplitude, and to form facial micro-expression features. The joint motion pattern and body posture characteristics were analyzed on the limb posture data to obtain the mean vector of functional joint distances, overall motion activity, average trunk extension, trunk extension fluctuation and average gesture space range, and to form limb motion characteristics. Semantic content features, intonation and prosody features, facial micro-expression features, and body movement features are combined to form a multi-dimensional interactive feature set.
3. The multimodal data fusion intelligent interaction method based on virtual digital humans according to claim 2, characterized in that, Methods for determining cognitive-emotional attribution based on multidimensional interaction feature sets include: Cognitive-emotional attribution polarity analysis is performed on semantic content features to obtain semantic cognitive attribution coefficients and semantic-emotional attribution coefficients. Based on the semantic cognitive attribution coefficients and semantic-emotional attribution coefficients, the sentence-level semantic vectors are decoupled to obtain semantic cognitive projection vectors and semantic-emotional projection vectors. The text sequence, semantic cognitive projection vectors, and task orientation density are integrated to form the task orientation component of semantic content features. The semantic-emotional projection vectors and sentiment orientation density are integrated to form the sentiment orientation component of semantic content features. The attribution attributes of each sub-feature in the limb movement characteristics are determined, and they are divided into inherent attribution characteristics and shared attribution characteristics. For shared attribution characteristics, a cross-modal attribution allocation mechanism is used for decomposition, decomposing the overall motor activity into cognitive motor activity components and emotional motor activity components, and decomposing the average gesture spatial range into cognitive gesture range components and emotional gesture range components. The functional joint distance mean vector, cognitive motor activity components and cognitive gesture range components are integrated to form the task orientation component of the limb movement characteristics. The average trunk extension degree, trunk extension fluctuation degree, emotional motor activity components and emotional gesture range components are integrated to form the emotional orientation component of the limb movement characteristics.
4. The multimodal data fusion intelligent interaction method based on virtual digital humans according to claim 3, characterized in that, Methods for extracting cognitive intention feature streams and emotional state feature streams include: Based on the preset median reference for motor activity and overall motor activity, a limb activity index is calculated; based on the task orientation density and the limb activity index, cross-modal cognitive synergy is calculated, and combined with the preset cognitive synergy enhancement coefficient, a cognitive fusion enhancement factor is calculated; based on the cognitive fusion enhancement factor, each element in the semantic cognitive projection vector and the functional joint distance mean vector is synergistically enhanced to obtain an enhanced semantic cognitive vector and an enhanced functional joint vector, and combined with the text sequence, task orientation density, cognitive motor activity component and cognitive gesture range component to form a cognitive intent feature stream; Based on the fundamental frequency fluctuation and the preset median reference for prosodic fluctuation, a prosodic activity index is calculated; based on the average facial movement amplitude and the preset median reference for facial movement, a facial activity index is calculated; based on the emotion orientation density, prosodic activity index, facial activity index, and limb activity index, a cross-modal emotion synergy is calculated, and combined with a preset emotion synergy enhancement coefficient, an emotion fusion enhancement factor is calculated; based on the emotion fusion enhancement factor, each element in the semantic emotion projection vector and the action unit mean vector is synergistically enhanced to obtain an enhanced semantic emotion vector and an enhanced action unit mean vector, and combined with the emotion orientation components of emotion orientation density, action unit fluctuation vector, average facial movement amplitude, and limb movement features, and the intonation prosodic features, an emotion state feature stream is formed.
5. The multimodal data fusion intelligent interaction method based on virtual digital humans according to claim 4, characterized in that, Methods for intent matching and task planning based on cognitive intent feature streams include: A pre-defined domain knowledge base is provided, which includes an intent template library, a task execution graph library, and a gesture intent pattern library. The intent template library contains multiple intent templates, the task execution graph library contains multiple task execution graphs, and the gesture intent pattern library contains multiple gesture intent patterns. Based on the enhanced semantic cognition vector, text sequence and enhanced function joint vector, multi-channel intent matching is performed on the intent template library to obtain the comprehensive intent matching confidence of each intent template, and the comprehensive intent matching confidence with the largest value is marked as the highest matching confidence. Based on the comprehensive intent matching confidence, the user intent recognition result is determined; based on the user intent recognition result, task planning is performed to generate response content and execution step sequence, and the corresponding task complexity level is set; based on task pointing density, cognitive motor activity component, cognitive gesture range component, highest matching confidence and task complexity level, task urgency is calculated; the response content, execution step sequence and task urgency are integrated to form a task response plan.
6. The multimodal data fusion intelligent interaction method based on virtual digital humans according to claim 5, characterized in that, Methods for determining a user's emotion category and emotion intensity include: A pre-defined emotion prototype library is used, containing multiple emotion prototypes. Each emotion prototype includes an emotion category label, a semantic emotion prototype vector, a facial action unit prototype vector, a prosodic prototype parameter set, and a body posture prototype parameter set. For each emotion prototype, an emotion evidence score for each modality is calculated. The modalities are, in order, semantic modality, facial modality, prosodic modality, and body modality. Based on the activity level of each modality in the emotion state feature stream, a dynamic fusion weight is calculated for each modality. Based on the dynamic fusion weights of each modality, the emotion evidence scores for each modality corresponding to the same emotion prototype are weighted and summed to obtain a comprehensive emotion matching score for each emotion prototype. The emotion category label corresponding to the emotion prototype with the highest comprehensive emotion matching score is taken as the user's emotion category, and the corresponding comprehensive emotion matching score is taken as the preferred emotion matching score. The user's emotion intensity value is calculated based on the deviation of each modality from the neutral state in the emotion state feature stream and the preferred emotion matching score.
7. The multimodal data fusion intelligent interaction method based on virtual digital humans according to claim 6, characterized in that, Methods for generating emotional response plans include: A pre-defined empathy strategy library contains multiple empathy strategy entries. Each entry includes the applicable emotion category, applicable intensity range, empathy response level, and empathy expression method. Based on the user's emotion category and intensity value, the library is searched for empathy strategy entries whose applicable emotion category matches the user's and whose intensity value falls within the applicable intensity range, and these are marked as matching strategy entries. The corresponding empathy response level and empathy expression method are obtained from the matching strategy entries. The user's emotion category, intensity value, empathy response level, and empathy expression method are integrated to form an emotion response plan.
8. The multimodal data fusion intelligent interaction method based on virtual digital humans according to claim 7, characterized in that, Methods for dynamically adjusting the integration ratio of cognitive and affective decision-making include: Based on task urgency and emotional intensity, channel interaction status parameters are calculated. These parameters include channel tension index, total channel activity, and channel dominance polarity. The channel tension index is compared with a preset tension conflict threshold, and the current interaction mode is determined based on the comparison result. The current interaction mode includes coordinated fusion mode and adversarial fusion mode. If the current interaction is in a coordinated fusion mode and the total channel activity is greater than zero, the fusion ratio is calculated based on task urgency and total channel activity. The fusion ratio includes cognitive fusion ratio and emotional fusion ratio. If the current interaction is in an adversarial fusion mode, then the mutual inhibition strength factor is calculated based on the channel tension index and the tension conflict threshold; the initial cognitive proportion is calculated based on the task urgency and the total channel activity; if the channel dominance polarity value is greater than zero, then the initial emotional proportion is calculated based on the initial cognitive proportion, and the cognitive enhancement amount is calculated in combination with the mutual inhibition strength factor; the sum of the initial cognitive proportion and the cognitive enhancement amount is calculated to obtain the cognitive fusion proportion; if the channel dominance polarity value is less than zero, then the subordination retention coefficient is calculated based on the mutual inhibition strength factor, and the cognitive fusion proportion is calculated in combination with the initial cognitive proportion; the difference between the initial and cognitive fusion proportions is calculated to obtain the emotional fusion proportion.
9. The multimodal data fusion intelligent interaction method based on virtual digital humans according to claim 8, characterized in that, Methods for synergistically integrating task response and emotional response strategies include: Based on the current fusion modality and fusion ratio of the interaction, the final response content is determined; based on the fusion ratio, tone style parameters and facial expression parameters are calculated; among them, tone style parameters include fusion speech rate coefficient, fusion pitch coefficient, and fusion tone temperature coefficient; facial expression parameters include final facial expression identifier and fusion action amplitude coefficient. Perform cross-channel expression consistency verification on tone style parameters and facial expression parameters: use the fused tone temperature coefficient as tone expression valence and the fused action amplitude coefficient as action expression valence; calculate cross-channel expression deviation based on tone expression valence and action expression valence; if the cross-channel expression deviation is greater than the preset expression consistency threshold, perform consistency correction; if the cross-channel expression deviation is less than or equal to the expression consistency threshold, do not perform consistency correction. The consistency correction process is as follows: Calculate the target concordance valence based on the fusion ratio, tone expression valence, and action expression valence; calculate the tone correction amount based on the target concordance valence, tone expression valence, and the preset consistency correction step size; update the fusion tone temperature coefficient in the tone style parameters based on the sum of the fusion tone temperature coefficient and the tone correction amount; calculate the action correction amount based on the target concordance valence, action expression valence, and the consistency correction step size; update the fusion action amplitude coefficient in the facial expression action parameters based on the sum of the fusion action amplitude coefficient and the action correction amount. The final response content, tone and style parameters, facial expression and action parameters, and execution step sequence are integrated to form a comprehensive interactive driving instruction; the virtual digital human is then driven to perform interactive responses based on the comprehensive interactive driving instruction.
10. A multimodal data fusion intelligent interaction system based on virtual digital humans, implementing the multimodal data fusion intelligent interaction method based on virtual digital humans as described in any one of claims 1-9, characterized in that, include: The multimodal parsing module is used to collect users' voice signals, facial image sequences, and body posture data, and extract semantic content features, intonation and prosody features, facial micro-expression features, and body movement features to form a multi-dimensional interactive feature set. The dual-stream decoupling module is used to perform cognitive-emotional attribution determination on the multi-dimensional interaction feature set. It extracts the task-oriented component in the semantic content features and the task-oriented component in the body movement features into a cognitive intention feature stream, and extracts the emotional orientation component in the semantic content features, the emotional orientation component in the body movement features, the intonation and prosody features, and the facial micro-expression features into an emotional state feature stream. The cognitive decision-making module is used to perform intent matching and task planning based on the cognitive intent feature flow and a preset domain knowledge base, and generate a task response plan that includes response content, execution steps and task urgency. The emotion decision module is used to infer the emotion from the distribution characteristics and intensity changes of each feature in the emotion state feature stream, determine the user's emotion category and emotion intensity value, and generate an emotion response plan by matching the corresponding empathy response level and empathy expression mode according to the emotion category and emotion intensity value. The dual-stream driving module is used to dynamically adjust the fusion ratio of cognitive decision-making and emotional decision-making based on the task urgency in the task response plan and the emotional intensity value in the emotional response plan. Based on the fusion ratio, the task response plan and the emotional response plan are collaboratively fused to determine the final response content, tone style parameters and facial expression parameters of the virtual digital human, and drive the virtual digital human to generate interactive responses that combine task execution and emotional empathy.
Citation Information
Patent Citations
Multi-modal intelligent robot system and interaction method
CN120588256A
Emotional evolution method and terminal for virtual avatar in educational metaverse
US20250014470A1