Multi-modal fusion real-time digital human driving method based on unified behavior vector mapping

Through the unified fusion framework of speech, action and visual modes and honeycomb grid vector mapping, the problems of multimodal alignment and conflict recognition in digital human drivers are solved, and the multimodal fusion of naturalization, emotionality and scene-based multimodal fusion is realized, which improves the expressiveness and adaptability of digital humans.

CN120339477APending Publication Date: 2025-07-18CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510424742.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing digital human-driven technology mainly relies on single modal input, resulting in stiff behavior and single emotional expression. There are time resolution differences in the integration of multimodal data, lack of unified representation of cross-modal data, and insufficient conflict identification and correction mechanisms, making it difficult to achieve multimodal precise alignment, efficient fusion and dynamic conflict adaptation.

Method used

A unified fusion framework of three modalities: speech, action and vision is adopted. Through a cross-modal collaborative alignment strategy based on knowledge graph constraints and timing decomposition, a coordinated alignment of speech, action and visual modal features is carried out, a situation-driven-modal-dominated conflict adaptive correction mechanism is designed, and a three-dimensional spatial projection method of honeycomb grid vector mapping is used to realize multimodal fusion real-time driving.

Benefits of technology

The accuracy and stability of the multimodal fusion process are achieved, the expression and emotional richness of digital people are improved, the real-time needs of virtual reality and online interaction are met, and the robustness in complex scenarios is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339477A_ABST
    Figure CN120339477A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-modal fusion real-time digital human driving method based on unified behavior vector mapping, and belongs to the technical field of human-computer interaction. The method comprises the following steps of: 1) extracting multi-modal features of voice, action and vision, and generating input features by adopting language feature-emotion decoupling, staged action modeling and macroscopic-micro expression flow analysis; (2) cross-modal collaborative alignment and conflict correction are carried out, and high-precision fusion is realized through knowledge graph constraint time sequence decomposition, particle-by-particle interactive fusion and a situation-driven modal dominant strategy; and 3) constructing a three-dimensional behavior vector space, mapping the multi-modal features to a unified coordinate by utilizing honeycomb grid projection, and driving the digital human to output in combination with a coordinate-action mapping table. According to the method, the problems of difficulty in multi-modal time sequence alignment, feature isomerism and out-of-control conflict are solved, real-time interaction of naturalization, emotional and scenarized is realized, and the expressive force and adaptability of the digital human are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of human-computer interaction, and relates to a method for real-time driving a digital human based on unified behavior vector mapping and multimodal fusion. Background Art

[0002] As an important application in the cross-field of artificial intelligence and computer graphics, digital human technology has been widely promoted in the fields of virtual reality, intelligent interaction, online education, etc. Its core goal is to achieve natural interaction with users through highly realistic appearance, movements, and emotional expressions, thereby enhancing the experience and expanding the application boundaries of human-computer interaction.

[0003] Currently, digital human driving technology mainly relies on single-modal input (such as speech, motion capture, or visual analysis) to generate corresponding behaviors. However, single-modal driving has significant limitations:

[0004] (1) The essence of human interaction behavior is multimodal fusion (such as changes in speech intonation, facial expression adjustment, and limb movement coordination). Single-modal cannot comprehensively reflect multi-dimensional information, resulting in rigid digital human behaviors and single emotional expressions.

[0005] (2) The existing technologies face three major challenges in multimodal data integration:

[0006] (3) The time resolutions and feature forms of different modal data (speech, motion, expression) vary greatly, making it difficult to achieve precise time synchronization.

[0007] (4) There is a lack of unified representation for cross-modal data such as speech semantics, joint movements, and expression features, resulting in low fusion efficiency.

[0008] (5) Multimodal information may express contradictory semantics (such as calm speech but angry expression). Existing methods lack a dynamic conflict recognition and correction mechanism.

[0009] Although some studies have attempted multimodal fusion, there are still the following defects:

[0010] (1) Traditional time series alignment methods ignore the action phase decomposition (such as start - motion - stop) and instantaneous changes in micro-expressions, resulting in inaccurate fine-grained behaviors.

[0011] (2) Existing methods rely on simple weighting or splicing and do not model the high-order dependence relationships between modalities (such as the dynamic constraint of speech emotion on expression intensity).

[0012] (3) The conflict correction mechanism is fixed and cannot dynamically adjust the dominant modality according to the situation (such as dialogue, performance), resulting in the disconnection between the behavior output and the scene.

[0013] Therefore, there is an urgent need for a digital human driving method that can achieve multi-modal precise alignment, efficient fusion, and dynamic conflict adaptation to improve interaction naturalness, emotional authenticity, and scene adaptability. Summary of the Invention

[0014] In view of this, the purpose of the present invention is to provide a multi-modal fusion real-time driving digital human method based on unified behavior vector mapping. It adopts a unified fusion framework for three modalities of speech, action, and vision, uses a cross-modal collaborative alignment strategy based on knowledge graph constraints and temporal decomposition to perform collaborative alignment of speech, action, and visual modality features, adopts a cross-modal fine-grained interaction fusion mechanism, captures high-order dependency relationships between modalities through hierarchical modeling to obtain multi-modal global semantic features, and designs a conflict adaptive correction mechanism for context-driven - modality-dominated to solve the modality conflict problem and ensure the accuracy and stability of the multi-modal fusion process. Finally, a multi-modal fusion real-time driving digital human method based on unified behavior vector mapping is proposed. Using a three-dimensional space projection method based on honeycomb grid vector mapping, project the features of speech, action, and expression into three-dimensional space. When the coordinates are activated, the coordinate-action mapping table extracts each modality component to drive the digital human to output coordinated speech, action, and expression, realizing a natural, emotional, and scene-based multi-modal fusion real-time driving digital human process.

[0015] To achieve the above purpose, the present invention provides the following technical solutions:

[0016] A multi-modal fusion real-time driving digital human method based on unified behavior vector mapping, specifically including the following steps:

[0017] S1: The multi-modal input includes three parts: speech modality, action modality, and visual modality. The speech modality generates semantic features through a long short-term memory network and the language model BERT, and proposes a multi-language emotion recognition method based on language feature-emotion expression decoupling to capture emotion features in different languages. Finally, integrate the semantic and emotion features into the final speech text word granularity semantic feature vector. The action modality uses OpenPose to extract joint point coordinates and spatial relationships, and adopts a multi-dimensional joint feature extraction method based on stage-based action modeling to analyze joint spatial relationships and dynamic changes in action stages, and integrate multi-dimensional features into a high-dimensional action pose feature vector. The visual modality extracts facial key points based on OpenFace, and proposes a macro-micro expression flow feature extraction module to extract macro and micro expression features respectively. Finally, use the macro flow expression AU intensity as a constraint to correct the micro expression features through a gated recurrent unit GRU to obtain the facial expression local feature vector, providing comprehensive input for multi-modal fusion.

[0018] S2: Multimodal fusion. First, a cross-modal collaborative alignment strategy based on knowledge graph constraints and temporal decomposition is proposed. By constructing an emotion-action-expression triple knowledge base and using the temporal decomposition and hierarchical alignment mechanism, the collaborative alignment of speech, action, and visual modal features is carried out. Then, a cross-modal granularity-by-granularity interactive fusion mechanism is adopted. By hierarchical modeling, the high-order dependence relationships between modalities are captured to obtain multimodal global semantic features. Finally, a conflict adaptive correction mechanism oriented to context-driven-modal-dominated is designed. Through four steps of context awareness, modal-dominated strategy, confidence dynamic adjustment, and conflict correction, modal conflicts are eliminated to ensure the accuracy and stability of the multimodal fusion process.

[0019] S3: A real-time driving digital human method for multimodal fusion based on unified behavior vector mapping is proposed. First, a three-dimensional behavior vector space for speech, action, and expression modalities is created. By defining modal axes, features such as specific intonations, action types, and facial expression categories are assigned to each modality. Then, a three-dimensional space projection method based on honeycomb grid vector mapping is proposed to project the features of speech, action, and expression onto each grid cell in the three-dimensional space and determine their three-dimensional coordinate positions in the space. Finally, when the coordinates are activated, a coordinate-action mapping table is used to extract each modal component to drive the digital human to output coordinated speech, action, and expression, realizing the real-time driving digital human process of natural, emotional, and scenario-based multimodal fusion.

[0020] Furthermore, in S1, multimodal inputs of speech, action, and vision are obtained through a multilingual emotion recognition method based on language feature-emotion expression decoupling, a multi-dimensional joint feature extraction method based on stage action modeling, and a macro-micro expression flow feature extraction method. The specific steps are as follows:

[0021] S11: Speech modal input. First, the long short-term memory network (LSTM) is used to convert speech into text, and then the pre-trained language model BERT is adopted to perform semantic understanding of the speech text to generate the semantic feature vector F of the speech text. s i, then adopt the sentiment analysis network SER, and propose a multi - language sentiment recognition method based on the decoupling of language characteristics - sentiment expression. By separating and analyzing language characteristics from sentiment features, aiming at the sentiment expression characteristics of different languages, extract language characteristics such as intonation and syllable rhythm, and independently analyze the sentiment category and intensity, reducing the interference of language characteristics on sentiment recognition. That is, for languages with rich intonation changes, such as English, focus on analyzing high - frequency spectrum features; for languages with obvious syllable rhythms, such as Chinese, focus on the analysis of speech segmentation and rhythm changes; at the same time, aiming at the differences in sentiment intensity expression in different languages, use a dynamic sentiment intensity calibration mechanism. By analyzing the acoustic features of the speech signal and combining with language characteristics, dynamically adjust the calculation method of sentiment intensity. For languages with implicit sentiment expressions, such as Japanese, the model will enhance the analysis of subtle acoustic features; for languages with direct sentiment expressions, such as Spanish, focus on the extraction of significant acoustic features, so as to identify the sentiment features F in the speech e i , finally, integrate the semantic features and sentiment features of the speech text into a multi - dimensional feature space to generate the final speech text word - granularity semantic feature vector F1, as shown in Equation (1).

[0022]

[0023] Where n represents the dimension of the feature vector, represents the weighted fusion of all n - dimensional features, α represents the weighted coefficient of semantic features, controlling the importance and influence of semantic features in the final fused features, β represents the weighted coefficient of sentiment features, controlling the contribution of sentiment features to the final fusion, W s represents the weight matrix of semantic features, used to weight - adjust the semantic feature L(·), W e represents the weight matrix of sentiment features, used to weight - adjust the sentiment features, represents the fusion loss function, optimizing the fusion effect, ensuring the optimal combination of semantic and sentiment features, F s i represents the semantic feature vector of the i - th speech text, generated by the BERT model, representing the semantic information of the speech text, F e i represents the sentiment feature vector of the i - th speech text, extracted by the SER network, representing the sentiment information of the speech text.

[0024] S12: Action modality input. Obtain the original image through a monocular camera, use OpenPose to perform human pose estimation on the input image, and extract the two - dimensional coordinates of 18 key points, such as Figure 2As shown below. Next, considering the diversity and complexity of human movements and comprehensively taking into account multi-level features such as space, time, and motion patterns, a multi-dimensional pose capture method based on phased action modeling is proposed. First, the relative position relationships between joints are calculated, including the spatial relationship features of coordinate differences, angle changes, and geometric distances between key points among joints, so as to construct a multi-dimensional spatial relationship matrix M between joints to capture the spatial configuration of joints in different poses, as shown in Equation (2).

[0025]

[0026] Among them represents the angle change between joints, v i and v j are the direction vectors of the joint and the joint relative to the root joint, |v i | and |v j | are the magnitudes of the vectors, d ij represents the j-th distance-related feature of the i-th joint pair. m represents the number of joint pairs. Among 18 joint points, the number of joint pairs is the combination number of selecting 2 joint points from 18 joint points. n represents the number of features calculated for each joint pair, which at least includes coordinate difference, angle change, and geometric distance, that is, n≥3.

[0027] Then, in order to analyze the human movement process more precisely, a phased action modeling method is adopted. Each action is divided into three stages: start - motion - stop. In each stage, the motion pattern of the joint undergoes dynamic changes such as static, acceleration, and deceleration, providing features in different "time dimensions", so as to deeply analyze the comprehensive arrangement pattern, motion trajectory, and dynamic changes among multiple joint points, identify the detailed features of joints in different motion stages, and capture the global and local multi-dimensional features of human postures. Finally, these multi-dimensional joint space features and phased dynamic features are integrated into a high-dimensional action pose feature vector F2, which is one of the core inputs for subsequent multi-modal fusion, providing a reliable data basis for the precise action generation of digital humans, as shown in Equation (3).

[0028]

[0029] Among them, vec(·) represents the matrix vectorization operation, expanding the 153×3 matrix into a 459-dimensional vector. α i represents the stage weight, and the significance of the spatial relationship in each stage is measured by the Frobenius norm, represents the multi-dimensional spatial relationship matrix within stage T i (i = 1, 2, 3), which is composed of joint pairs and feature dimensions (coordinate difference, angle, distance). k j represents stage Tj The number of time points (for example, the startup phase T1 includes k1 sampling points), and Δt represents the time interval, which is used to normalize the acceleration dimension. denotes the vector concatenation operation, where the first part is the spatial feature and the second part is the temporal dynamic feature. The acceleration of the j-th joint at time t.

[0030] S13: Visual modality input. First, preprocess the collected facial images, such as cropping, adjusting illumination and contrast. Then, use OpenFace to ensure the standardized alignment of the face region, and extract 68 key points from the aligned facial images. These key points cover the main facial structures (such as eyebrows, eyes, nose, mouth, and chin), as Figure 3 shown. Next, to address the problem of insufficient capture of instantaneous micro-actions by traditional OpenFace, a macro-micro expression flow feature extraction module is proposed. The macro-expression flow extracts the spatial distribution and changes of 68 key points based on OpenFace, further analyzes and identifies facial action units (such as the corners of the mouth rising, eyebrows lifting, etc.), and obtains the collaborative movement trajectory of facial muscle groups through optical flow tracking; the micro-expression flow uses the TV-L1 dense optical flow algorithm to amplify local movement details (such as eyelid fluttering, slight tremors at the corners of the mouth), and constructs a 16×16 pixel-level motion energy map. Finally, use the AU intensity of the macro-expression flow as a prior constraint, and correct the micro-expression feature vector through a gated recurrent unit (GRU) to obtain the facial expression feature vector F3, as shown in Equation (4).

[0031]

[0032] where Φ1 represents the macro-expression flow feature, Φ2 represents the micro-expression flow feature, AU represents the facial action unit intensity vector, which describes the muscle activity intensity such as the corners of the mouth rising and eyebrows lifting, and O m represents the 68×2 key point coordinate matrix extracted based on OpenFace. Each row represents the position of a key point, and vec(·) represents the matrix vectorization operation, which flattens the 68×2×68 matrix into a 136-dimensional vector. E m represents the 16×16 pixel-level motion energy map generated by the TV-L1 dense optical flow algorithm, which amplifies details such as eyelid fluttering and slight tremors at the corners of the mouth. W TV-L1 represents the optical flow convolution kernel, which is used to extract local motion patterns. Pool(·) represents the max pooling operation, which reduces the dimension to 64. W g represents the learnable weights and biases, which are used to adjust the constraint of AU intensity on micro-expressions. σ(·) represents the Sigmoid function, which generates an attention weight between 0 and 1 to suppress noisy micro-features. ⊙ represents element-wise multiplication to achieve feature weighting.

[0033] Furthermore, multimodal fusion is performed in S2. First, a cross-modal collaborative alignment strategy based on knowledge graph constraints and temporal decomposition is proposed to perform collaborative alignment of speech, action, and visual modal features. Then, a cross-modal granular interactive fusion mechanism is adopted to obtain multimodal global semantic features. Finally, a conflict adaptive correction mechanism for context-driven and modality-dominated is designed to eliminate modal conflicts and ensure the accuracy and stability of the multimodal fusion process. Specifically, the following steps are included:

[0034] S21: A multimodal collaborative alignment strategy based on knowledge graph constraints and temporal decomposition is proposed. First, semantic association modeling driven by knowledge graph is adopted to extract a set of three elements, namely, emotion category, facial action unit, and action trajectory, from sentiment dictionaries, psychological research, and multilingual film and television subtitles to construct an emotion-action-expression triple knowledge base, such as (happy, upturned corners of the mouth, waving action). Then, a knowledge graph K is constructed, and the semantic associations between different modalities (emotion, action, expression) are represented by a graph structure. The speech emotion, action trajectory features, and facial expressions representing the same state are connected to ensure the semantic consistency between the modalities, as shown in formula (5).

[0035] K={(e1,a1,f1),(e2,a2,f2),...,(e n ,a n ,f n )} (5)

[0036] The knowledge graph represented by K contains multiple triples, e i represents the i-th emotion category (such as happy, angry, etc.), a i represents the feature of the ith action trajectory (such as waving, jumping, etc.), f i Represents the i-th facial expression feature (such as raised corners of the mouth, frowning, etc.).

[0037] Then, the temporal features of speech, action, and visual modalities are decomposed into shallow units and deep units using the temporal decomposition and hierarchical alignment mechanism to achieve hierarchical alignment in the time dimension. In the shallow units, the entire sentence in speech, the complete action sequence in action, and the general changes in expression in vision are preliminarily aligned to capture the overall outline between the modalities, as shown in formula (6); in the deep units, the word level in speech, the joint frame level in action, and the micro-expression changes in vision are further aligned based on the knowledge graph of triples to further align finer details, as shown in formula (7).

[0038]

[0039] Where T s and T d represents shallow and deep multimodal alignment, T a , Tm , T v respectively represent the temporal features of speech, action, and vision, and f i (T a , T m , T v ) represents the temporal feature function extracted from modality i (speech, action, vision), and g i (K) represents the mapping constraint between modalities established based on the knowledge graph K and the emotion-action-expression triple, is the L2 norm, representing the similarity measure between features, and minimizing the distance to achieve the initial alignment between modalities. h i (T a , T m , T v ) represents the fine-grained feature function in modality i (e.g., word-level, joint-frame level, micro-expression), and K i (K) is the deep constraint based on the knowledge graph K, reflecting the more precise relationship between modalities, represents the Frobenius norm, measuring the similarity between matrices, and A i j represents the cross-attention matrix between the i-th modality and the j-th modality, and λ is the regularization parameter used to balance the weights between shallow and deep alignments.

[0040] S22: Propose a cross-modal fine-grained interaction fusion mechanism analysis to obtain multi-modal global semantic features. First, divide the aligned features of each modality in S21 into the following three layers from fine to coarse in terms of granularity, and the processing objectives and contents of each layer are as follows:

[0041] (1) Fine-grained layer: Capture the local dynamic dependencies between modalities

[0042] In the speech modality, based on the local attention mechanism, calculate the dynamic feature weights of word-level time steps within a short-time window to strengthen the temporal consistency; in the action modality, model the short-time displacement changes of joint points and the collaborative relationship between adjacent joints through local graph convolution; in the visual modality, divide the facial region and extract the dynamic sub-flows of local key points to enhance the expression detail representation.

[0043] (2) Medium-grained layer: Extract the combined patterns and medium-range dependencies of each modality

[0044] In the speech modality, use intermediate window temporal convolution to capture the context semantic patterns of word groups or sentence fragments; in the action modality, analyze the collaborative movement patterns of joint groups and model the dynamic features of action fragments; in the visual modality, model the global expression joint changes of the facial region (eyebrow-eye-mouth) through convolution operations.

[0045] (3) Coarse-grained layer: Capture the global semantic relationships of the features of each modality

[0046] Extract the global emotional and semantic expression features of the entire speech in the speech modality; in the action modality, analyze the dynamic start and end points and overall coordination of the human motion trajectory; in the visual modality, extract the global dynamic change trend of the complete expression sequence.

[0047] After completing the per-granularity layering, the cross-modal interaction and fusion mechanism IAM performs bidirectional dynamic interaction modeling on the modal feature streams of each granularity layer, simultaneously capturing the high-order dependency relationships and dynamic collaboration information between modalities, as shown in Equation (8).

[0048]

[0049] where φ(Q (k) , V (k) ) = MLP(Q (k) ⊙V (k) ) represents the cross-modal semantic enhancement term, which is used to enhance the inter-modal dependency modeling. MLP(·) is a multi-layer perceptron, which is used to extract non-linear high-dimensional features from the inter-modal interaction relationships. k represents the granularity layer. Q (k) , K (k) , V (k) represent the query, key-value, and value matrices, respectively, which perform linear transformations on each modal feature stream and are used to capture the dynamic semantic interaction relationships inside and outside the modality in the granularity layer k. represents the modal mask matrix, which is used to control the inter-modal interaction weights, represents the time step marker of modalities i and j, and controls the interaction intensity of different modal time steps through the time distance. σ(·) represents the normalization activation function, which is used to extract the non-linear feature expression after the interaction.

[0050] S23: Propose a context-driven and modality-dominated conflict adaptive correction mechanism. Aiming at the problem of modal conflicts in the multi-modal fusion process, that is, the semantic contradiction of inconsistent emotions or intentions transmitted between different modalities, design a dynamic conflict correction mechanism based on context-driven and modality-dominated. Through four steps of context awareness, modality-dominated strategy, confidence dynamic adjustment, and conflict correction, evaluate and adjust the weights of each modality in real time in different scenarios to eliminate the conflicts between modalities.

[0051] First, the modality-dominated strategy driven by the context monitors the current scene in real time and dynamically selects the dominant modality according to the context and situational requirements. The main contents include:

[0052] (I) Context awareness

[0053] According to the information of the three modalities of speech, action, and vision, and combined with the scene type, emotion type, and relationship between roles, judge which modality expression is the most critical in the current context. Specifically:

[0054] (1) Use term frequency - inverse document frequency (TF-IDF) to extract important keywords or semantics from the speech and match them with the scenario dictionary. For example, the scenarios matched by keywords such as "speech" and "discussion" are conversations, meetings, or debates, and the scenarios matched by keywords such as "dance" and "performance" are stage performances or celebrations. Then, adopt a semantic analysis model (Sentence-BERT) to calculate the semantic embedding vector of the speech text, calculate its similarity with the preset scenario categories, and infer the current scenario in combination with the context before and after.

[0055] (2) Use Openpose to calculate joint angles, limb link lengths, and displacement information to determine the human action category, and combine the acceleration, speed, and angular velocity of the joint movement trajectory to judge the exercise intensity. Then, calculate the change trend of the body center of gravity in different frames to determine whether the person is in a stable state (such as standing) or a strenuous exercise state (such as jumping or running). Finally, infer the current scenario by synthesizing all the results. For example, steady standing + slight hand movement infers that the scenario is "conversation", alternating legs + forward movement infers that the scenario is "walking", and large limb swings + continuous rotation infers that the scenario is "dance".

[0056] (3) Use Openface to obtain the direction and focus of the eyes and analyze the area pointed by the eyes. For example, eyes focused forward represent a scenario of concentrating on a conversation or performance, and wandering eyes indicate distracted attention or no specific target. By combining the eye direction and facial expression, the emotional understanding of the scenario is further enhanced. That is, in a conversation scenario, focused eyes and a peaceful facial expression may represent attention and participation in the discussion, while in a drama performance, the eyes may often shift towards the audience or a specific direction. Finally, jointly infer the current situation based on the changes in facial expression and eye movement. In a dance scenario, the facial expression is usually pleasant or excited, and accompanied by large movements. The system will infer that the action modality is dominant. In a meeting scenario, a calm facial expression and slow or static movements may be consistent with the content of the debate or discussion in the speech modality, thus inferring that the conversation scenario is dominant.

[0057] Finally, infer the most appropriate dominant modality at present through three steps: mainly judging by the speech modality information, assisted by the action information, and verified and confirmed by the visual information. That is, if the speech content contains clear keywords and semantic analysis shows that the keyword highly matches a specific scenario, then give priority to choosing the speech modality information as the dominant. If the speech does not play a decisive role, the action modality is used for auxiliary judgment. If the user's action amplitude is large and the rhythm is fast, the priority of the action modality information is increased. The visual modality is used for final verification, combined with facial expression, eye direction, and attention distribution, to confirm or adjust the dominant modality. For example,

[0058] Scenario 1: If words such as "debate" and "discussion" appear in the speech, and the voice emotion is pleasant, but there is no obvious change in facial expression and the movements are smooth, it is inferred that this is a dialogue scene and the speech mode is dominant.

[0059] Scenario 2: If keywords such as "dance" and "performance" appear in the speech, and the action modality information shows dance movements, and the facial expression also shows happy emotions, it is inferred that the current scene is a performance and the action modality is dominant.

[0060] Scenario 3: If the words "I'm fine" appear in the voice, the core keywords are not included and the tone is flat, the body movements are steady and the standing or sitting posture remains normal, but the visual modal information expresses strong facial emotional characteristics such as evasive eyes, not daring to look directly at others, slightly drooping corners of the mouth and tense facial muscles, indicating that the actual state of the person may not be consistent with the voice and action content, then it is judged that the current scene is emotion concealment and the visual modality is dominant.

[0061] 2. Determination of dominant mode

[0062] The system calculates the dominant confidence of each modality based on contextual information Determine which mode is dominant in the current situation, as shown in formula (9).

[0063]

[0064] Where m∈{A,V,M} represents the modality type identifier (A=speech, V=vision, M=action), represents the current timestamp, k represents the modality traversal index, M=3 represents the total number of modalities, represents the dominant confidence of mode m at time t, λ∈[0.6,0.7] is the dynamic scene weight, W c ∈R 192×3 represents the trainable weight matrix (192=128+64-dimensional feature concatenation), Represents the modal depth features extracted by ResNet, ctx∈R 64 represents the scene coding vector (including position coding and time segment coding), γ = 0.1 represents the time decay factor, Δt = tt last Indicates the number of seconds since the last switch. represents the original confidence at the previous moment, softmax(·) represents the normalized exponential function to ensure the probability distribution, [;] represents the feature concatenation operation, and exp(·) represents the exponential function to strengthen the impact of the event.

[0065] (III) Leading Switching

[0066] As scenes and situations change, the weights between modalities need to be further adjusted. Confidence scoring and dynamic dominant switching are adopted. According to the current situation perception, the confidence score of each modality is calculated in real time, and the modality with a higher score gets a higher weight in the current scene. When the scene changes, the adaptive switching trigger condition S w , the confidence of the dominant modality will be adjusted accordingly, and the system will dynamically switch the dominance of the modality to ensure the coordination of emotional expression. After the dominant switch, the specific weights of each modality are adjusted to ensure that the role of the dominant modality in the overall fusion is maximized, as shown in formula (10).

[0067]

[0068] 1 means triggering mode switching, and 0 means keeping the current mode configuration. represents the confidence difference between adjacent moments, represents the 30-second sliding mean, represents the time domain gradient of the weight matrix, ∈=0.05 represents the gradient stability coefficient, and τ=0.5 represents the adaptive threshold.

[0069] Finally, the dynamic conflict correction mechanism is used to eliminate inconsistencies by adjusting the time step between modes and optimizing modal features. The modal features are optimized by strengthening and weakening to correct the conflicting parts. That is, in the dialogue scene, for the conflict between expression and voice, the inharmonious expression is "covered up" by strengthening the emotional intensity of the voice and weakening the facial expression movements.

[0070] Furthermore, a multimodal fusion real-time driving digital human method based on unified behavior vector mapping is proposed in S3. First, a three-dimensional behavior vector space of speech, action and expression modes is created by defining the modal axis. Then, a three-dimensional space projection method using honeycomb grid vector mapping is proposed to project the features of speech, action and expression to each grid unit in the three-dimensional space. Finally, when the coordinates are activated, the coordinate-action mapping table is used to extract the modal components to drive the digital human to perform coordinated speech, action and expression output, realizing a natural, emotional and scenario-based multimodal fusion real-time driving digital human process, which specifically includes the following steps:

[0071] S31: A multimodal fusion-driven digital human method based on unified behavior vector mapping is proposed. First, a unified three-dimensional behavior vector space UBVS is constructed to uniformly reflect the speech, action and visual modal information into a high-dimensional "coordinate system". In this space, the feature information of each modality in the global semantic features obtained in S2 is encoded into the corresponding "behavior vector", and the overall behavior state of the digital human at a specific time is fully expressed through the position of the behavior vector in the three-dimensional space.

[0072] The three-dimensional coordinate axes of the behavior vector include a voice modality axis, an action modality axis, and a visual modality axis. The voice modality axis represents the emotional category (happy, sad, angry, etc.), emotional intensity (pitch, volume), and speech rate (fast, medium, slow) of the voice input; the action modality axis represents the body movement characteristics captured by action or visually extracted, including action type (gesture, body posture), action amplitude (slight, intense), and action direction (based on joint angles and movement trajectories); the visual modality axis represents the emotions and dynamic characteristics extracted from facial expressions, including emotion category (smile, frown, surprise, etc.), expression intensity (degree based on key point displacement), and expression change speed (speed of dynamic change). Each three-dimensional coordinate point represents the overall behavior state that the digital human needs to execute at the current moment. The dynamic change of the coordinate point reflects the real-time adjustment of the modality characteristics. The vector length represents the overall behavior intensity, and the vector direction represents the modality dominant relationship, as shown in formula (11).

[0073]

[0074] Among them, in three-dimensional space, a vector point is represented by three coordinate values (x, y, z). x = S represents the voice modality characteristics, y = A represents the action modality characteristics, and z = E represents the expression modality characteristics. e s represents the emotion category, t s represents the intonation intensity, p s represents the speech rate, t a represents the action type, m a represents the action amplitude, d a represents the action direction, t e represents the expression category, s e represents the expression intensity, v e represents the change speed.

[0075] S32: Three-dimensional space projection method based on honeycomb grid behavior vector mapping. In order to map the behavior vector in S31 into three-dimensional space and ensure the accurate and coordinated behavior performance of the digital human in the space, a behavior vector mapping rule based on honeycomb grid is proposed, providing a uniform and fine hexagonal grid division in three-dimensional space, making the projection position of each behavior vector more reasonable and avoiding the uneven distribution of traditional rectangular grids in high-dimensional space.

[0076] (1) Create a honeycomb grid structure

[0077] According to the characteristic range of each modality in the behavior vector, first divide the three-dimensional space into multiple small grid units, and each unit represents a possible behavior pattern state. The honeycomb grid structure has higher space utilization and uniformity, so the distance between each grid unit is more balanced, and it can avoid the problem of uneven distribution of traditional rectangular grids in high-dimensional space.

[0078] (2) Behavioral vector projection

[0079] Projection rule: The features of each modality will be mapped to specific positions in three-dimensional space according to their magnitudes and weights on their respective coordinate axes. The honeycomb grid defines the positions of each behavioral vector in three-dimensional space, thus ensuring that the fusion of multiple modalities is orderly and uniform within the space.

[0080] Determination of behavioral vector position: Through the projection method of the honeycomb grid, the features of speech, action, and expression are projected onto a certain hexagonal grid cell in three-dimensional space, and the coordinate points within each hexagonal region will represent a set of similar behavioral states. For example, the speech emotion "happy" may be mapped to a certain high-value region on the speech modality axis, while "angry" may correspond to another region. Actions will be mapped to the corresponding regions on the action modality axis according to their amplitudes, types, and directions, and vision will be mapped according to the intensity and change speed of the expression, as shown in formula (12).

[0081]

[0082] where α s , α a , α e represent the weight coefficients of each modality, used to control the contribution size of the corresponding modality in three-dimensional space. S i , A j , E k represent the feature components of the speech, action, and expression modalities respectively. The speech modality component includes emotion category, intonation intensity, and speech rate. The action modality component includes action type, action amplitude, and action direction. The expression modality component includes expression category, expression intensity, and change speed. w si , w aj , w ek represent the elements of the weight matrix, used to adjust the contribution of each component to the overall projection.

[0083] (3) Unified behavior pattern region

[0084] Through the projection method of the honeycomb grid, multiple similar coordinate points will be aggregated into the same hexagonal region. Each hexagonal region represents a specific behavior pattern region, and the behavioral vectors within these regions are similar in features, thus avoiding duplicate mapping of coordinate points.

[0085] S33: Coordinate-action mapping table. In order to directly associate the behavioral vectors mapped to three-dimensional space with the actual behaviors of the digital human, a coordinate-action mapping table is designed to directly correspond each point in the three-dimensional coordinates with the actual speech, action, and expression behavior features of the digital human, ensuring the precise execution of the digital human's behaviors.

[0086] (1) Construct a mapping table

[0087] Each three-dimensional coordinate point corresponds to a specific behavior pattern of the digital human. Through the mapping table, the behavior vector in the three-dimensional coordinate space can be mapped to the behavior characteristics required by the digital human, as shown in formula (13). Then, based on the mapping rules, the modal feature intensity F s , F a , F e is inverted into specific semantic feature parameters, as shown in formula (14).

[0088]

[0089] Among them, F s , F a , F e respectively represent the features of the speech, action, and expression modalities. W s , W a , W e represent the weight parameters of the modal features, controlling the normalization range of the feature intensity to ensure that the values of each modality are within a reasonable range. C s , T s , R s respectively represent the speech emotion category, speech adjustment parameter, and speech rate rhythm. C a , M a , D a respectively represent the action category, action amplitude, and action direction. C e , S e , V e respectively represent the expression category, expression intensity, and expression change speed.

[0090] (2) Digital human behavior generation

[0091] When a certain coordinate point is activated, the mapping table determines the specific behavior pattern of the digital human according to the position of the coordinate point and drives the digital human to perform speech, action, and expression outputs. Adjust the emotional tone of the digital human's speech through the speech features (C s , T s , R s ), control the timbre level of the speech, and ensure that the speech matches the scene; use the action features (C a , M a , D a ) to control the limb movements of the digital human, determine the strength of the action, and specify the action direction; according to the expression features (C e , S e , V e ) specify the expression type, determine the amplitude of the digital human's expression change, and determine the speed of expression transition, thereby driving the digital human to perform a real-time "language-action-expression" behavior process that is emotional, natural, and scene-based.

[0092] The beneficial effects of the present invention are as follows:

[0093] The beneficial effects of the present invention

[0094] (1) By integrating a unified framework of speech, motion, and vision modalities, the limitations of single-modal driving in expressiveness and emotional expression are overcome. The speech modality adopts a language feature-emotion decoupling method to accurately identify multi-language emotions; the motion modality captures joint spatio-temporal features through stage modeling; the vision modality extracts detailed dynamics by combining macro-micro expression flows. The multi-modal complementarity significantly improves the realism and emotional richness of the digital human behavior.

[0095] (2) Based on the knowledge graph constraint and temporal decomposition strategy, cross-modal shallow (whole sentence / complete action) and deep (word level / micro-expression) alignment are achieved to solve the problem of multi-modal data time resolution differences. A context-driven-modal-dominated conflict correction mechanism is designed to dynamically adjust the dominant modality weights (such as strengthening speech in a dialogue scenario and strengthening motion in a performance scenario), eliminate semantic contradictions, and ensure the logical consistency of multi-modal outputs.

[0096] (3) By projecting multi-modal features into a three-dimensional space through a unified behavior vector mapping, and using the uniform distribution characteristics of the honeycomb grid, the problems of heterogeneous feature representation and low fusion efficiency in traditional methods are solved. The coordinate-action mapping table enables the rapid inversion of behavior vectors, combined with real-time projection and activation mechanisms, supports low-latency driving of digital human output, and meets the real-time requirements of scenarios such as virtual reality and online interaction.

[0097] (4) The conflict correction mechanism supports dynamic scene switching (such as from a meeting to a performance), adaptively adjusts the fusion strategy through confidence scoring, and enhances the robustness in complex scenarios. The scalable design of the three-dimensional behavior vector space allows for the rapid integration of new modalities (such as touch and environmental perception), providing a compatibility foundation for future upgrades of multi-modal interaction technologies.

[0098] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following specification. BRIEF DESCRIPTION OF THE DRAWINGS

[0099] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail with reference to the accompanying drawings, where:

[0100] Figure 1 is the main flowchart of the multi-modal fusion-driven digital human method according to the embodiment of the present invention;

[0101] Figure 2 The 17 skeletal joint points of the human body in the invention embodiment;

[0102] Figure 3 The 68 facial key points of the human body in the invention embodiment.

[0103] Reference numerals: 1 to 17: Facial contour one to facial contour seventeen;

[0104] 18 to 22: Left eyebrow eighteen to left eyebrow twenty-two;

[0105] 23 to 27: Right eyebrow twenty-three to right eyebrow twenty-seven;

[0106] 28 to 36: Nose twenty-eight to nose thirty-six;

[0107] 37 to 42: Left eye thirty-seven to left eye forty-two;

[0108] 43 to 48: Right eye forty-three to right eye forty-eight;

[0109] 49 to 68: Mouth forty-nine to mouth sixty-eight;

[0110] 90: Nose;

[0111] 91: Left eye;

[0112] 92: Right eye;

[0113] 93: Left ear;

[0114] 94: Right ear;

[0115] 95: Left shoulder;

[0116] 96: Right shoulder;

[0117] 97: Left elbow;

[0118] 98: Right elbow;

[0119] 99: Left wrist;

[0120] 910: Right wrist;

[0121] 911: Left hip;

[0122] 912: Right hip;

[0123] 913: Left knee;

[0124] 914: Right knee;

[0125] 915: Left ankle;

[0126] 916: Right ankle;

[0127] 917: Chest. Detailed implementation manners

[0128] The embodiments of the present invention will be described below through specific examples. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the drawings provided in the following examples only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following examples and the features in the examples can be combined with each other.

[0129] Among them, the drawings are only for illustrative purposes, showing only schematic diagrams rather than physical diagrams, and should not be construed as a limitation to the present invention; in order to better illustrate the embodiments of the present invention, some components in the drawings will be omitted, enlarged or reduced, which do not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0130] In the drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the drawings. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the drawings are only for illustrative purposes and should not be construed as a limitation to the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.

[0131] As Figure 1 shown, first, obtain the features of each modality according to the unified fusion framework of the three modalities of voice, action, and vision. Then, adopt the cross-modal fine-grained interactive fusion mechanism to extract the multi-modal global semantic features. At the same time, design a conflict adaptive correction mechanism oriented to situation-driven modality-dominated to eliminate the conflicts between modalities in the multi-modal fusion process. Finally, propose a multi-modal fusion-driven digital human method based on unified behavior vector mapping, project the features of each modality in the three-dimensional space through the honeycomb grid vector mapping method. When the coordinate points are activated, use the coordinate-action mapping table to extract each modality component, and drive the digital human to output coordinated voice, actions, and expressions, so as to realize the process of comprehensively, accurately, and naturally driving the digital human in real-time with multi-modal fusion, specifically including the following steps:

[0132] Step 1: The multi-modal input includes three parts: speech modality, action modality, and visual modality. The speech modality generates semantic features through a long short-term memory network and the language model BERT, and proposes a multi-language emotion recognition method based on the decoupling of language characteristics and emotion expression to capture emotion features in different languages. Finally, the semantic and emotion features are integrated into the final semantic feature vector of speech text word granularity. The action modality uses OpenPose to extract joint coordinates and spatial relationships, and adopts a multi-dimensional joint feature extraction method based on stage-based action modeling to analyze joint spatial relationships and dynamic changes in action stages, and integrates multi-dimensional features into a high-dimensional action pose feature vector. The visual modality extracts facial key points based on OpenFace, and proposes a macro-micro expression flow feature extraction module to extract macro and micro expression features respectively. Finally, using the macro-expression AU intensity as a constraint, the micro-expression features are corrected by GRU to obtain the local facial expression feature vector, providing comprehensive input for multi-modal fusion.

[0133] 101: Speech modality input. First, use the long short-term memory network LSTM to convert speech into text, and then adopt the pre-trained language model BERT for semantic understanding of speech text to generate the semantic feature vector F of speech text. s i Next, adopt the sentiment analysis network SER, and propose a multi-language emotion recognition method based on the decoupling of language characteristics and emotion expression. By separating and analyzing language characteristics and emotion features, aiming at the emotion expression characteristics of different languages, extract language characteristics such as intonation and syllable rhythm, and independently analyze emotion categories and intensities, reducing the interference of language characteristics on emotion recognition. That is, for languages with rich intonation changes, such as English, focus on analyzing high-frequency spectrum features; for languages with obvious syllable rhythms, such as Chinese, focus on the analysis of speech segmentation and rhythm changes; at the same time, aiming at the differences in emotion intensity expressions in different languages, use a dynamic emotion intensity calibration mechanism to dynamically adjust the calculation method of emotion intensity by analyzing the acoustic features of speech signals and combining language characteristics. For languages with implicit emotion expressions, such as Japanese, the model will enhance the analysis of subtle acoustic features; for languages with direct emotion expressions, such as Spanish, focus on the extraction of significant acoustic features, so as to identify the emotion features F in speech. e i Finally, integrate the semantic features and emotion features of speech text into a multi-dimensional feature space to generate the final semantic feature vector F1 of speech text word granularity, as shown in Equation (1).

[0134]

[0135] Among them, n represents the dimension of the feature vector, which means weighted fusion of all n-dimensional features. α represents the weighting coefficient of semantic features, controlling the importance and influence of semantic features in the final fused features. β represents the weighting coefficient of emotional features, controlling the contribution of emotional features to the final fusion. W s represents the weight matrix of semantic features, used for weighted adjustment of semantic features L(·), W e represents the weight matrix of emotional features, used for weighted adjustment of emotional features. represents the fusion loss function, optimizing the fusion effect and ensuring the optimal combination of semantic and emotional features. F s i represents the semantic feature vector of the i-th speech text, generated by the BERT model, representing the semantic information of the speech text. F e i represents the emotional feature vector of the i-th speech text, extracted by the SER network, representing the emotional information of the speech text.

[0136] 102: Action modality input. The original image is obtained through a monocular camera, and the OpenPose is used to estimate the human pose of the input image, extracting the two-dimensional coordinates of 18 key points, as Figure 2 shown. 90 is the nose; 91 is the left eye; 92 is the right eye; 93 is the left ear; 94 is the right ear; 95 is the left shoulder; 96 is the right shoulder; 97 is the left elbow; 98 is the right elbow; 99 is the left wrist; 910 is the right wrist; 911 is the left hip; 912 is the right hip; 913 is the left knee; 914 is the right knee; 915 is the left ankle; 916 is the right ankle; 917 is the chest.

[0137] Next, aiming at the diversity and complexity of human actions, considering multi-level features such as space, time, and motion patterns comprehensively, a multi-dimensional pose capture method based on staged action modeling is proposed. First, calculate the relative position relationships between joints, including the spatial relationship features of coordinate differences, angle changes between joints, and geometric distances between key points, so as to construct a multi-dimensional spatial relationship matrix M between joints, capturing the spatial configuration of joints in different poses, as shown in Equation (2).

[0138]

[0139] Among them represents the angle change between joints, v i and v j are the direction vectors of the joint and the joint relative to the root joint, |v i | and |v j | are the magnitudes of the vectors, d ij represents the j-th distance-related feature of the i-th joint pair. m represents the number of joint pairs. Among 18 key points, the number of joint pairs That is, it is the combination number of selecting 2 joint points from 18 joint points. n represents the number of features calculated for each joint pair, which includes at least coordinate difference, angle change, and geometric distance, that is, n≥3.

[0140] Next, in order to analyze the human motion process more precisely, a phased motion modeling method is adopted, and each motion is divided into three stages: start - motion - stop. In each stage, the motion mode of the joints will undergo dynamic changes such as stillness, acceleration, and deceleration, providing features in different "time dimensions", so as to deeply analyze the comprehensive arrangement pattern, motion trajectory, and dynamic changes among multiple joint points, identify the detailed features of the joints in different motion stages, and capture the global and local multi - dimensional features of the human body posture. Finally, these multi - dimensional joint space features and phased dynamic features are integrated into a high - dimensional motion posture feature vector F2, which is one of the core inputs for subsequent multi - modal fusion, providing a reliable data basis for the accurate motion generation of digital humans, as shown in Equation (3).

[0141]

[0142] Among them, vec(·) represents the matrix vectorization operation, which expands the 153×3 matrix into a 459 - dimensional vector. α i represents the stage weight, which measures the significance of the spatial relationship of each stage through the Frobenius norm. represents the multi - dimensional spatial relationship matrix within stage T i (i = 1, 2, 3), which is composed of joint pairs and feature dimensions (coordinate difference, angle, distance). k j represents the number of time points in stage T j (for example, the start stage T1 includes k1 sampling points), and Δt represents the time interval, which is used to normalize the acceleration dimension. represents the vector concatenation operation, where the first part is the spatial feature and the second part is the time - dynamic feature. The acceleration of the j - th joint at time t.

[0143] 103: Visual modality input. First, pre - process the collected facial images, such as cropping, adjusting lighting and contrast, then use OpenFace to ensure the standardized alignment of the face area, and extract 68 key points from the aligned facial images. These key points cover the main facial structures (such as eyebrows, eyes, nose, mouth, and chin, etc.), as Figure 3 shown. 1 - 17 are face contour one - face contour seventeen; 18 - 22 are left eyebrow eighteen - left eyebrow twenty - two; 23 - 27 are right eyebrow twenty - three - right eyebrow twenty - seven; 28 - 36 are nose twenty - eight - nose thirty - six; 37 - 42 are left eye thirty - seven - left eye forty - two; 43 - 48 are right eye forty - three - right eye forty - eight; 49 - 68 are mouth forty - nine - mouth sixty - eight;

[0144] Next, aiming at the problem of insufficient capture of instantaneous micro-actions by traditional OpenFace, a macro-micro expression flow feature extraction module is proposed. The macro-expression flow extracts the spatial distribution and changes of 68 key points based on OpenFace, further analyzes and identifies facial action units (such as the upward curl of the mouth corners, the raising of the eyebrows, etc.), and obtains the collaborative movement trajectories of facial muscle groups through optical flow tracking; the micro-expression flow uses the TV-L1 dense optical flow algorithm to amplify local movement details (such as eyelid fluttering, slight tremors at the mouth corners), and constructs a 16×16 pixel-level motion energy map. Finally, using the macro-expression flow AU intensity as a prior constraint, the micro-expression feature vector is corrected through a gated recurrent unit (GRU) to obtain the facial expression feature vector F3, as shown in Equation (4).

[0145]

[0146] Among them, Φ1 represents the macro-expression flow feature, Φ2 represents the micro-expression flow feature, AU represents the facial action unit intensity vector, which describes the muscle activity intensity such as the upward curl of the mouth corners and the raising of the eyebrows, O m represents the 68×2 key point coordinate matrix extracted based on OpenFace. Each row represents the position of a key point. vec(·) represents the matrix vectorization operation, which flattens the 68×2×68 matrix into a 136-dimensional vector, E m represents the 16×16 pixel-level motion energy map generated by the TV-L1 dense optical flow algorithm, which amplifies details such as eyelid fluttering and slight tremors at the mouth corners, W TV-L1 represents the optical flow convolution kernel, which is used to extract local motion patterns. Pool(·) represents the max pooling operation, which reduces the dimension to 64 dimensions, W g represents the learnable weights and biases, which are used to adjust the constraint of AU intensity on micro-expressions. σ(·) represents the Sigmoid function, which generates an attention weight of 0-1 to suppress noisy micro-features. ⊙ represents element-wise multiplication to achieve feature weighting.

[0147] Step 2: Multimodal fusion. First, a cross-modal collaborative alignment strategy based on knowledge graph constraints and temporal decomposition is proposed. By constructing an emotion-action-expression triple knowledge base and using the temporal decomposition and hierarchical alignment mechanism, the collaborative alignment of speech, action, and visual modality features is carried out. Then, a cross-modal fine-grained interactive fusion mechanism is adopted. By hierarchical modeling, the high-order dependence relationships between modalities are captured to obtain the multimodal global semantic features. Finally, a conflict adaptive correction mechanism oriented to context-driven and modality-dominated is designed. Through four steps of context awareness, modality-dominated strategy, confidence dynamic adjustment, and conflict correction, modality conflicts are eliminated to ensure the accuracy and stability of the multimodal fusion process.

[0148] 201: Propose a multi-modal collaborative alignment strategy based on knowledge graph constraints and temporal decomposition. First, adopt knowledge graph-driven semantic association modeling, extract a three-element set from the sentiment dictionary, psychological research, and multilingual movie subtitles, namely sentiment categories, facial action units, and action trajectories, to construct an emotion-action-expression triple knowledge base, such as (happy, upturned corners of the mouth, waving action), and then construct a knowledge graph K, and represent the semantic associations between different modalities (emotion, action, expression) through a graph structure, connecting the speech emotion, action trajectory features, and facial expressions representing the same state to ensure semantic consistency between modalities, as shown in Equation (5).

[0149] K = {(e1, a1, f1), (e2, a2, f2),..., (e n , a n , f n )} (5)

[0150] Among them, the knowledge graph represented by K contains multiple triples, e i represents the i-th sentiment category (such as happy, angry, etc.), a i represents the i-th action trajectory feature (such as waving, jumping, etc.), f i represents the i-th facial expression feature (such as upturned corners of the mouth, frowning, etc.).

[0151] Then, use the temporal decomposition and hierarchical alignment mechanism to decompose the temporal features of speech, action, and visual modalities into shallow units and deep units to achieve hierarchical alignment in the time dimension. In the shallow units, initially align the whole sentence in speech, the complete action sequence in action, and the general change in expression in vision to capture the overall contour between modalities, as shown in Equation (6); in the deep units, based on the triple knowledge graph, further perform deep alignment on more refined details for the word level in speech, the joint frame level in action, and the micro-expression change in vision, as shown in Equation (7).

[0152]

[0153] Among them, T s and T d represent shallow and deep multi-modal alignment, T a , T m , T v respectively represent the temporal features of speech, action, and vision, f i (T a , T m , T v ) represents the temporal feature function extracted from modality i (speech, action, vision), and g i (K) represents the mapping constraint between modalities established based on the knowledge graph K and the emotion-action-expression triples. is the L2 norm, representing the similarity measure between features, and minimizing the distance to achieve the initial alignment between modalities. h i (T a ,T m ,T v ), represents the fine-grained feature function in modality i (e.g., word-level, joint-frame-level, micro-expression), K i (K) is the deep constraint based on the knowledge graph K, reflecting the more precise relationship between modalities, represents the Frobenius norm, measuring the similarity between matrices, represents the cross-attention matrix between the i-th modality and the j-th modality, and λ is the regularization parameter used to balance the weights between shallow and deep alignments.

[0154] 202: A cross-modal fine-grained interaction fusion mechanism analysis is proposed to obtain multi-modal global semantic features. First, the aligned features of each modality in S21 are divided into the following three layers from fine to coarse in terms of granularity, and the processing objectives and contents of each layer are as follows:

[0155] (1) Fine-grained layer: Capture the local dynamic dependencies between modalities

[0156] In the speech modality, based on the local attention mechanism, calculate the dynamic feature weights of word-level time steps within a short-time window to strengthen the temporal consistency; in the action modality, model the short-time displacement changes of joint points and the collaborative relationship between adjacent joints through local graph convolution; in the visual modality, divide the facial region and extract the dynamic sub-stream of local key points to enhance the expression details.

[0157] (2) Medium-grained layer: Extract the combined patterns and medium-range dependencies of each modality

[0158] In the speech modality, use the intermediate window temporal convolution to capture the context semantic patterns of word groups or sentence fragments; in the action modality, analyze the collaborative motion patterns of joint groups and model the dynamic features of action segments; in the visual modality, model the global expression joint changes of the facial region (eyebrow-eye-mouth) through convolution operations.

[0159] (3) Coarse-grained layer: Capture the global semantic relationships of the features of each modality

[0160] Extract the global emotional and semantic expression features of the entire speech in the speech modality; in the action modality, analyze the dynamic start and end points and overall coordination of the human motion trajectory; in the visual modality, extract the global dynamic change trend of the complete expression sequence.

[0161] After the per - granularity layering is completed, the cross - modal interaction and fusion mechanism IAM performs bidirectional dynamic interaction modeling on the modal feature streams of each granularity layer, simultaneously capturing the high - order dependency relationships and dynamic collaboration information between modalities, as shown in Equation (8).

[0162]

[0163] Among them, φ(Q (k) , V (k) ) = MLP(Q (k) ⊙V (k) ) represents the cross - modal semantic enhancement term, which is used to enhance the inter - modal dependency modeling. MLP(·) is a multi - layer perceptron, which is used to extract non - linear high - dimensional features from the inter - modal interaction relationship. k represents the granularity layer. Q (k) , K (k) , V (k) represent the query, key, and value matrices respectively, which perform linear transformations on each modal feature stream and are used to capture the dynamic semantic interaction relationships inside and outside the modalities in the granularity layer k. represents the modal mask matrix, which is used to control the inter - modal interaction weights, represents the time - step marker of modalities i and j, which controls the interaction intensity of different modal time - steps through the time distance. σ(·) represents the normalization activation function, which is used to extract the non - linear feature expression after interaction.

[0164] 203: Propose a context - driven and modality - dominant conflict adaptive correction mechanism. Aiming at the problem of modal conflicts in the multi - modal fusion process, that is, the semantic contradiction of inconsistent emotions or intentions transmitted between different modalities, design a dynamic conflict correction mechanism based on context - driven and modality - dominant. Through four steps: context awareness, modality - dominant strategy, confidence dynamic adjustment, and conflict correction, the weights of each modality are evaluated and adjusted in real - time in different scenarios to eliminate the conflicts between modalities.

[0165] First, the modality - dominant strategy driven by context is used to monitor the current scene in real - time, and the dominant modality is dynamically selected according to the context and context requirements. The main contents include:

[0166] (1) Context awareness

[0167] According to the information of three modalities: speech, action, and vision, and combined with the scene type, emotion type, and relationship between roles, judge which modality expression is the most critical in the current context. Specifically:

[0168] (1) Use term frequency - inverse document frequency (TF-IDF) to extract important keywords or semantics from the speech and match them with the scene dictionary. For example, the situations matched by keywords such as "speech" and "discussion" are conversations, meetings or debates, and the situations matched by keywords such as "dance" and "performance" are stage performances or celebrations. Then, use the semantic analysis model (Sentence-BERT) to calculate the semantic embedding vector of the speech text, calculate its similarity with the preset scene categories, and infer the current scene in combination with the context before and after.

[0169] (2) Use Openpose to calculate joint angles, limb link lengths, and displacement information to determine the human action category, and combine the acceleration, speed, and angular velocity of the joint movement trajectory to judge the exercise intensity. Then, calculate the change trend of the center of gravity of the body in different frames to determine whether the person is in a stable state (such as standing) or a strenuous exercise state (such as jumping or running). Finally, infer the current scene by synthesizing all the results. For example, steady standing + slight hand movement infers that the scene is "conversation", alternating legs + forward movement infers that the scene is "walking", and large limb swings + continuous rotation infers that the scene is "dance".

[0170] (3) Use Openface to obtain the direction and focus of the eyes and analyze the area pointed by the eyes. For example, the eyes focused forward represent a scene of concentrating on a conversation or performance, and wandering eyes indicate distracted attention or no specific target. By combining the eye direction and facial expression, the emotional understanding of the scene is further enhanced. That is, in a conversation scene, the eyes are focused and the facial expression is calm, which may represent attention and participation in the discussion, while in a drama performance, the eyes may often shift to the audience or a specific direction. Finally, the changes in facial expression and eye movement jointly infer the current situation. In a dance scene, the facial expression is usually happy or excited and accompanied by large movements, and the system will infer that the action modality is dominant. In a meeting scene, a calm facial expression and slow or static movements may be consistent with the content of the debate or discussion in the speech modality, thus inferring that the conversation scene is dominant.

[0171] Finally, infer the most suitable dominant modality at present through three steps: mainly judging by the speech modality information, assisted by the action information, and verified and confirmed by the visual information. That is, if the speech content contains clear keywords and semantic analysis shows that the keyword highly matches a specific scene, then the speech modality information is preferentially selected as the dominant. If the speech does not play a decisive role, the action modality is used for auxiliary judgment. If the user's action amplitude is large and the rhythm is fast, the priority of the action modality information is increased. The visual modality is used for final verification, and in combination with facial expression, eye direction and attention distribution, the dominant modality is confirmed or adjusted. For example,

[0172] Scenario 1: If words like "debate" or "discussion" appear in the speech, and the speech emotion is pleasant, but there is no obvious change in the facial expression and the movement is steady, then it is inferred that this is a conversation scenario and the speech modality is dominant.

[0173] Scenario 2: If keywords such as "dance" or "performance" appear in the speech, and the movement modality information shows dance movements and the facial expression also shows a happy emotion, then it is speculated that the current is a performance scenario and the movement modality is dominant.

[0174] Scenario 3: If words like "I'm fine" appear in the speech, no core keywords appear and the intonation is flat, the body movement is steady and the standing or sitting posture remains normal, but the visual modality information expresses strong facial emotion features such as avoiding eye contact, not daring to look directly at others, the corners of the mouth drooping slightly and the facial muscles tensing, indicating that the actual state of the person may not match the speech and movement content, then it is judged that the current is an emotion concealment scenario and the visual modality is dominant.

[0175] (2) Determination of the Dominant Modality

[0176] The system calculates the dominant confidence of each modality through context information to determine which modality dominates in the current scenario, as shown in Equation (9).

[0177]

[0178] where m ∈ {A, V, M} represents the modality type identifier (A = speech, V = vision, M = movement), represents the current timestamp, k represents the modality traversal index, M = 3 represents the total number of modalities, represents the dominant confidence of modality m at time t, λ ∈ [0.6, 0.7] is the dynamic scenario weight, W c ∈ R 192×3 represents the trainable weight matrix (192 = 128 + 64 - dimensional feature concatenation), represents the modality - depth feature extracted by ResNet, ctx ∈ R 64 represents the scene - encoding vector (including position encoding and time - period encoding), γ = 0.1 represents the time decay factor, Δt = t - t last represents the number of seconds since the last switch, represents the original confidence at the previous moment, softmax(·) represents the normalized exponential function used to ensure the probability distribution, [;] represents the feature - concatenation operation, and exp(·) represents the exponential function to strengthen the influence degree of the event.

[0179] (3) Dominant Switching

[0180] As the scene and situation change, the weights between modalities need to be further adjusted. Confidence scoring and dynamic dominance switching are adopted. According to the current situation perception, the confidence scores of each modality are calculated in real time, and the modality with a higher score obtains a higher weight in the current scene. When the scene changes, the adaptive switching trigger condition S w , the confidence of the dominant modality will be adjusted accordingly, and the system dynamically switches the dominance of the modality to ensure the coordination of emotional expression. After the dominant switch, the specific weights of each modality are adjusted to ensure that the role of the dominant modality is maximized in the overall fusion, as shown in Equation (10).

[0181]

[0182] where 1 indicates triggering modality switching, 0 indicates maintaining the current modality configuration, represents the confidence difference between adjacent moments, represents the 30-second sliding mean, represents the time-domain gradient of the weight matrix, ∈ = 0.05 represents the gradient stability coefficient, and τ = 0.5 represents the adaptive threshold.

[0183] Finally, using the dynamic conflict correction mechanism, the inconsistency is eliminated through the time step adjustment between modalities and the optimization of modality features. The modality features are optimized by the method of strengthening + weakening, and the conflicting parts are corrected. That is, in the dialogue scene, for the conflict between expressions and speech, the emotional intensity of the speech is strengthened and the facial expression actions are weakened to "cover up" the inconsistent expressions.

[0184] Step 3: A real-time driving digital human method based on unified behavior vector mapping for multi-modal fusion is proposed. First, a three-dimensional behavior vector space for speech, action, and expression modalities is created. By defining the modality axes, specific features such as intonation, action types, and facial expression categories are assigned to each modality. Then, a three-dimensional space projection method based on honeycomb grid vector mapping is proposed to project the features of speech, action, and expression onto each grid cell in the three-dimensional space and determine their three-dimensional coordinate positions in the space. Finally, when the coordinates are activated, a coordinate-action mapping table is used to extract each modality component to drive the digital human to output coordinated speech, action, and expression, realizing the process of real-time driving the digital human with natural, emotional, and scene-based multi-modal fusion.

[0185] 301: A multi-modal fusion driving digital human method based on unified behavior vector mapping is proposed. First, a unified three-dimensional behavior vector space UBVS is constructed to uniformly reflect the information of speech, action, and visual modalities in a high-dimensional "coordinate system". In this space, the feature information of each modality in the global semantic features obtained in S2 is encoded as the corresponding "behavior vector", and the overall behavior state of the digital human at a specific time is completely expressed through the position of the behavior vector in the three-dimensional space.

[0186] The three - dimensional coordinate axes of the behavior vector include a speech modality axis, an action modality axis, and a visual modality axis. The speech modality axis represents the emotional category (happy, sad, angry, etc.), emotional intensity (pitch, volume), and speech rate (fast, medium, slow) of the speech input; the action modality axis represents the body movement characteristics captured by action or visually extracted, including action type (gesture, body posture), action amplitude (slight, intense), and action direction (based on joint angles and movement trajectories); the visual modality axis represents the emotional and dynamic characteristics extracted from facial expressions, including emotion category (smile, frown, surprise, etc.), expression intensity (degree based on key - point displacement), and expression change speed (speed of dynamic change). Each three - dimensional coordinate point represents the overall behavior state that the digital human needs to execute at the current moment. The dynamic change of the coordinate point reflects the real - time adjustment of the modality characteristics. The vector length represents the overall behavior intensity, and the vector direction represents the modality dominant relationship, as shown in formula (11).

[0187]

[0188] Among them, in three - dimensional space, a point (vector) is represented by three coordinate values (x, y, z). x = S represents the speech modality feature, y = A represents the action modality feature, z = E represents the expression modality feature, e s represents the emotion category, t s represents the intonation intensity, p s represents the speech rate, t a represents the action type, m a represents the action amplitude, d a represents the action direction, t e represents the expression category, s e represents the expression intensity, v e represents the change speed.

[0189] 302: Three - dimensional space projection method based on honeycomb grid behavior vector mapping. In order to map the behavior vector in S31 into three - dimensional space and ensure the accurate and coordinated behavior performance of the digital human in the space, a behavior vector mapping rule based on the honeycomb grid is proposed, providing a uniform and fine hexagonal grid division in three - dimensional space, making the projection position of each behavior vector more reasonable and avoiding the uneven distribution of traditional rectangular grids in high - dimensional space.

[0190] (1) Create a honeycomb grid structure

[0191] According to the feature ranges of each modality in the behavior vector, the three-dimensional space is first divided into multiple small grid cells, and each cell represents a possible behavior pattern state. The honeycomb grid structure has higher spatial utilization and uniformity, so the distances between each grid cell are more balanced, which can avoid the problem of uneven distribution of traditional rectangular grids in high-dimensional space.

[0192] (2) Projection of the behavior vector

[0193] Projection rule: The features of each modality will be mapped to specific positions in the three-dimensional space according to their magnitudes and weights on their respective coordinate axes. The honeycomb grid defines the positions of each behavior vector in the three-dimensional space, thus ensuring the orderly and uniform fusion of multiple modalities within the space.

[0194] Determination of the behavior vector position: Through the projection method of the honeycomb grid, the features of speech, action, and expression are projected onto a certain hexagonal grid cell in the three-dimensional space, and the coordinate points within each hexagonal area will represent a set of similar behavior states. For example, the speech emotion "happy" may be mapped to a certain high-value area on the speech modality axis, while "angry" may correspond to another area. Actions will be mapped to the corresponding areas on the action modality axis according to their amplitudes, types, and directions, and vision will be mapped according to the intensity and change speed of the expression, as shown in formula (12).

[0195]

[0196] where α s , α a , α e represent the weight coefficients of each modality, used to control the contribution size of the corresponding modality in the three-dimensional space. S i , A j , E k represent the feature components of the speech, action, and expression modalities respectively. The speech modality component includes the emotion category, intonation intensity, and speech rate. The action modality component includes the action type, action amplitude, and action direction. The expression modality component includes the expression category, expression intensity, and change speed. w si , w aj , w ek represent the elements of the weight matrix, used to adjust the contribution of each component to the overall projection.

[0197] (3) Unified behavior pattern region

[0198] Through the projection method of the honeycomb grid, multiple similar coordinate points will be aggregated into the same hexagonal area. Each hexagonal area represents a specific behavior pattern region, and the behavior vectors within these regions are similar in features, thus avoiding the repeated mapping of coordinate points.

[0199] 303: Coordinate-Action Mapping Table. In order to directly associate the behavior vector mapped to the three-dimensional space with the actual behavior of the digital human, a coordinate-action mapping table is designed to directly correspond each point in the three-dimensional coordinates with the actual voice, action, and expression behavior characteristics of the digital human, ensuring the precise execution of the digital human's behavior.

[0200] (1) Construct the mapping table

[0201] Each three-dimensional coordinate point corresponds to a specific behavior pattern of the digital human. Through the mapping table, the behavior vector in the three-dimensional coordinate space can be mapped to the required behavior characteristics of the digital human, as shown in formula (13). Then, based on the mapping rules, the modal feature intensity F s , F a , F e is inverted into specific semantic feature parameters, as shown in formula (14).

[0202]

[0203] where F s , F a , F e represent the features of voice, action, and expression modalities respectively, and W s , W a , W e represent the weight parameters of the modal features, controlling the normalization range of the feature intensity to ensure that the values of each modality are within a reasonable range. C s , T s , R s represent the voice emotion category, voice adjustment parameter, and speech rate rhythm respectively. C a , M a , D a represent the action category, action amplitude, and action direction respectively. C e , S e , V e represent the expression category, expression intensity, and expression change speed respectively.

[0204] (2) Digital human behavior generation

[0205] When a certain coordinate point is activated, the mapping table determines the specific behavior pattern of the digital human according to the position of this coordinate point and drives the digital human to perform voice, action, and expression outputs. Adjust the emotional tone of the digital human's voice through the voice features (C s , T s , R s ), control the timbre level of the voice, and ensure that the voice matches the scene; utilize the action features (C a , M a , D a)Control the limb movements of the digital human, determine the intensity of the movements and specify the movement directions; according to the expression features (C e , S e , V e )Specify the types of expressions, determine the amplitude of the digital human's expression changes and determine the speed of expression transitions, so as to drive the digital human to perform a real-time "language-action-expression" behavioral process of emotionalization, naturalization and sceneization.

[0206] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered by the scope of the claims of the present invention.

Claims

1. A real-time driving digital human method based on unified behavior vector mapping, characterized in that: It includes the following steps: S1: Obtain multimodal input data, including speech modality, action modality, and visual modality; S2: Perform multimodal fusion processing on the multimodal input data, including cross-modal collaborative alignment, fine-grained interactive fusion, and conflict correction; S3: Construct a three-dimensional behavior vector space, map the fused multimodal features to the three-dimensional behavior vector space, and drive the speech, action, and expression outputs of the digital human according to the activation coordinates.

2. The method for real-time driving a digital human based on unified behavior vector mapping according to claim 1, wherein: Specifically, S1 includes: S11: Extract semantic features and perform emotion analysis on the speech modality to generate a speech text word-level semantic feature vector; S12: Extract the joint space relationship and stage dynamic features of the action modality to generate a high-dimensional action pose feature vector; S13: Extract macro-expression and micro-expression features of the visual modality to generate a facial expression local feature vector.

3. The method for real-time driving a digital human based on unified behavior vector mapping according to claim 2, wherein: In S11, a multilingual emotion recognition method based on language feature-emotion expression decoupling is adopted, including: Separate the language features and emotion features of the speech, dynamically adjust the emotion intensity calculation method, and generate the speech text word-level semantic feature vector through weighted fusion.

4. The real-time driving digital human method based on unified behavior vector mapping and multimodal fusion according to claim 2, characterized in that: In S12, the action is decomposed into three stages of start-motion-stop through stage-based action modeling, calculate the joint space relationship matrix, and integrate spatio-temporal dynamic features to generate the high-dimensional action pose feature vector.

5. The method for real-time driving a digital human based on unified behavior vector mapping according to claim 2, wherein: In S13, through a macro-micro expression flow feature extraction module, combined with optical flow tracking and TV-L1 dense optical flow algorithm, use a gated recurrent unit to correct the micro-expression features to generate the facial expression local feature vector.

6. The real-time driving digital human method based on unified behavior vector mapping and multi-modal fusion according to claim 1, characterized in that: Specifically, S2 includes: S21: Achieve shallow and deep temporal alignment of speech, action, and visual modalities based on knowledge graph constraints and temporal decomposition strategy; S22: Capture the inter-modal dependence relationship layer by layer through a cross-modal fine-grained interactive fusion mechanism to generate multimodal global semantic features; S23: Dynamically adjust the confidence based on the context-driven-modal dominant strategy, and complete conflict correction by strengthening the dominant modal features and weakening the conflict features.

7. The method for real-time driving a digital human based on unified behavior vector mapping according to claim 6, characterized in that: In S21, construct an emotion-action-expression triple knowledge base, associate semantic consistency through a graph structure, and adopt a hierarchical alignment mechanism to decompose the temporal features into shallow alignment of the whole sentence, complete action sequence, and expression change, as well as deep alignment at the word level, joint frame level, and micro-expression.

8. The method for real-time driving a digital human based on unified behavior vector mapping according to claim 6, wherein: In S23, the context awareness module judges the scene type, combines speech keyword matching, action motion trajectory analysis, and visual expression verification, dynamically calculates the dominant confidence of each modality, and triggers the dominant switch.

9. The real-time driven digital human method based on unified behavior vector mapping and multi-modal fusion according to claim 1, wherein: Specifically, S3 includes: S31: Construct a three-dimensional behavior vector space including a speech modality axis, an action modality axis, and a visual modality axis, and each coordinate point represents the overall behavior state of the digital human; S32: Adopt a honeycomb grid vector mapping method to project the multimodal features into the hexagonal grid cells in the three-dimensional behavior vector space; S33: Convert the activation coordinate points into speech adjustment parameters, action control parameters, and expression dynamic parameters through a coordinate-action mapping table to drive the digital human output.

10. The method for real-time driving a digital human based on unified behavior vector mapping according to claim 9, wherein: The projection rule of the honeycomb grid in S32 is as follows: Determine the three-dimensional space coordinate positions based on the weighted combination of speech emotion categories, intonation intensity, speech rate, action types, amplitudes, directions, and expression categories, intensities, and change speeds, and cluster similar behavior patterns through hexagonal regions.

Citation Information

Cited By

  • Decision-making method and device guided by multi-modal semantic map, equipment and medium

    CN120952167A

  • Digital human interaction method and system based on multi-mode sensing intelligent action switching

    CN121050590A

  • Automatic video generation system and method based on AI Agent multi-mode cooperative control

    CN121126084A

  • An automated video generation system and method with AIAgentic multimodal cooperative control

    CN121126084B

  • Interaction method and system based on multi-modal data

    CN121255027A