An ai digital human expression and facial feature migration method and system
By using a dual-branch encoder network with a built-in gradient inversion layer and a multimodal cultural expression knowledge base, the cultural specificity and general emotional state of virtual digital humans are separated and reconstructed to generate multimodal animations that conform to the target culture and individual characteristics. This solves the problems of animation distortion and cultural misalignment in the cross-cultural deployment of virtual digital humans and improves the user experience.
Patent Information
- Application Number
- CN202511266975.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2045-09-05
AI Technical Summary
Existing technologies struggle to reliably and effectively address the issues of animation distortion and cultural misalignment during the cross-cultural deployment of virtual digital humans, particularly the visual unnaturalness and cultural misunderstandings caused by differences in the topological structure of 3D models and different paradigms of emotional expression.
A dual-branch encoder network with a built-in gradient inversion layer is used for non-adversarial decoupling, separating culture-related expression paradigms from culture-independent intrinsic emotional states. Combined with a multimodal cultural expression knowledge base and personalized micro-expression calibration, multimodal animations that conform to the target culture and individual characteristics are generated.
It achieves the natural presentation of virtual digital human expressions and body movements in different cultural contexts, enhances user immersion and acceptance, and solves the problems of animation distortion and cultural misalignment in cross-cultural deployment.
Smart Images

Figure CN120997351B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and more particularly to the fields of computer graphics and virtual reality technology. Specifically, it relates to a method and system for transferring facial expressions and features in AI digital humans. Background Technology
[0002] In the context of current globalization, virtual digital humans, as an emerging medium, are being deployed and applied more and more across regions and cultures. However, introducing a virtual digital human that has achieved success in a specific cultural region (e.g., the source region) to a target region with a different cultural background usually faces dual technical difficulties.
[0003] The first challenge stems from animation distortion caused by differences in the physiological morphology of virtual digital human models. To cater to the aesthetic preferences of users in the target region, entirely new facial models are typically designed for virtual digital humans. This results in the target model's underlying 3D topology (actually reflecting the differences in facial structure among residents in remote geographical locations), namely the skeletal binding and mesh structure (Rig) controlling the model's movement, being completely different from the source model's topology. In this situation, directly applying the source model's facial animation data to the target model will produce distorted and uncoordinated expressions due to the mismatch between key control points (e.g., anchor points corresponding to facial muscle groups) and mesh deformation logic, presenting an unnatural visual effect and severely impacting the user's immersion and acceptance.
[0004] A second, deeper challenge stems from differences in emotional expression paradigms across cultures. Social psychology research indicates that the external expression of human emotions is not universally applicable but rather deeply shaped by the social display rules of specific cultural environments. For the same inner emotion, the socially accepted expression can vary significantly across different cultural contexts. For example, regarding the emotion of "happiness," in cultures that encourage outward expression, a warm, cheerful, toothy laugh is seen as a positive social signal; however, in cultures that value subtlety and restraint, such an expression might be interpreted as exaggerated or inappropriate, while a more reserved, pursed-lip smile might be more appropriate. Furthermore, emotional expression is often multimodal, involving the coordinated use of facial expressions and body language (especially gestures). For instance, in some cultures, people habitually cover their mouths with their hands when expressing joy or shyness; this "facial-gesture" combination constitutes a complete and inseparable unit of emotional expression within that culture.
[0005] In existing technologies, many style transfer methods based on Generative Adversarial Networks (GANs) treat facial expressions as a kind of "texture" or "style" in a two-dimensional image and attempt to perform pixel-level transformations. These methods struggle to understand the emotional core and cultural connotations behind expressions, failing to fundamentally solve the aforementioned problems. Furthermore, these adversarial training-based methods generally suffer from technical drawbacks such as unstable training processes, susceptibility to pattern collapse, and the need for extensive fine-tuning of parameters. Therefore, existing technological solutions cannot stably and effectively overcome the dual obstacles of "physiological morphology" and "cultural paradigm," leading to either physical distortion or cultural misalignment in the transferred virtual digital human's expressions. This can confuse and bewilder audiences in the target region, even generating negative perceptions, and ultimately potentially causing the failure of cross-cultural deployment of virtual digital humans.
[0006] Therefore, there is an urgent need for a new technical solution that can operate stably and efficiently, extract universal core emotional states from source animations without cultural bias, and use these as a driving force to generate a new multimodal animation for target virtual digital humans with different facial features. This animation should conform to the target culture's expression habits and be naturally presented on its specific model, including facial expressions and coordinated body movements. Summary of the Invention
[0007] The purpose of this invention is to provide an AI digital human expression and facial feature transfer method and system to overcome the animation distortion and cultural misalignment problems caused by differences in the topological structure of virtual digital human 3D models and different cross-cultural emotional expression paradigms in the prior art. In particular, it aims to provide a stable and efficient technical solution that does not require adversarial training.
[0008] To achieve the above objectives, one aspect of the present invention provides an AI digital human expression and facial feature transfer method, which transfers the core emotions expressed by a digital virtual human in a source region through the cultural rules of the source region into another expression in a target region that can also evoke the same emotions in the viewer. The method includes the following steps:
[0009] Step 1: Receive the multimodal animation data stream of the virtual human from the source region. The data stream contains facial animation data and body animation data that carry a specific expression (hereinafter referred to as "expression X") within the cultural context of the source region.
[0010] Preferably, the facial animation data is a displacement sequence of facial vertices or a blendshape weight sequence; the limb animation data is a three-dimensional coordinate sequence of key points of the body skeleton.
[0011] Step two involves using a dual-branch encoder network with a built-in gradient inversion layer to non-adversarially decouple the multimodal animation data stream. The core of this step is to separate the culturally relevant expression paradigms and culturally unrelated intrinsic emotional states contained in "expression X" to obtain a source culture-specific expression latent vector and a general emotional latent vector, respectively.
[0012] Preferably, the dual-branch encoder network includes a general emotion encoder and a culture-specific expression encoder. The general emotion encoder is used to extract the general emotion latent vector from "expression X", which represents the intrinsic core emotional state in the animation sequence that does not change with cultural background; the culture-specific expression encoder is used to extract the source culture-specific expression latent vector, which encodes the cultural expression paradigm of the source region and the facial physiological features of the source virtual human that are unique to "expression X".
[0013] Furthermore, a gradient reversal layer (GRL) is inserted between the output of the general emotion encoder and an auxiliary culture classifier. The gradient reversal layer performs an identity transformation during forward propagation and multiplies the gradient from the culture classifier by a negative constant during backward propagation. This drives the general emotion encoder to generate features that the culture classifier cannot distinguish as originating from, thus achieving a non-adversarial separation between general emotion and cultural characteristics.
[0014] Step 3: Based on the general emotional latent vector and the input target region cultural identifier, a condition generator is used to reconstruct and generate a target culture-specific expression latent vector that conforms to the target region cultural expression paradigm, with the intrinsic emotional state defined by the general emotional latent vector as the core. This vector defines a new expression (hereinafter referred to as "expression Y") that will be presented on the virtual human in the target region.
[0015] Preferably, the condition generator retrieves a target cultural expression paradigm that matches the intrinsic emotional state and the target region's cultural identifier by querying a pre-built multimodal cultural expression knowledge base. The knowledge base is a structured database that stores data entries that associate emotional tags, cultural identifiers, and multimodal expression primitives (such as facial movement units and coordinated gestures). Essentially, this step matches a new expression of "expression Y" that conforms to the target culture for the same intrinsic emotion.
[0016] Furthermore, the construction of the multimodal cultural expression knowledge base utilizes large language models (LLM) for assisted generation, expanding the cultural expression paradigm data in the knowledge base by analyzing multi-source texts and multimedia materials.
[0017] Step four involves performing personalized online micro-expression calibration on the target virtual human. This is achieved by using a few-shot learning encoder to extract unique micro-expression features from a small amount of emotionally charged reference data of the target individual and encoding them into a personalized offset vector.
[0018] Preferably, the reference data is one or more reference images of the target individual or a short video. The few-shot learning encoder is a lightweight convolutional neural network specifically designed to extract subtle individual facial expression features that cannot be summarized by a macro-level cultural knowledge base, used for fine-tuning the "expression Y".
[0019] Step 5: Using a cross-topology multimodal decoder, based on the target culture-specific expression latent vector and the neutral pose 3D mesh model of the target virtual human, multimodal animation data "tailor-made" for the target virtual human is synthesized on the unique facial topology of the target virtual human, thereby ultimately generating a specific "expression Y".
[0020] Preferably, the multimodal decoder employs a Graph Convolutional Network (GCN) or an attention-based Transformer structure, enabling it to perform computations directly on the vertices and edges of the target 3D mesh. By perceiving and adapting to the unique facial topology of the target virtual human, it transforms the abstract instruction of "expression Y" into physically correct vertex displacements on a specific model.
[0021] Furthermore, the target culture-specific expression potential vector is fused with the personalized offset vector, and the fused vector is provided as input to the multimodal decoder to generate a highly customized animation that conforms to the target culture paradigm and incorporates individual characteristics.
[0022] Furthermore, the output of the multimodal decoder is two synchronized data streams: one is facial animation data generated for the target virtual human facial model; the other is limb animation data that drives the target virtual human body skeleton to perform in coordination with the "Y expression" and conforms to the target cultural expression definition.
[0023] Another aspect of the present invention provides an AI digital human expression and facial feature transfer system, the system comprising:
[0024] An emotion decoupling module is used to perform steps one and two above to decouple the general emotion latent vector and the source culture-specific expression latent vector from the source animation data;
[0025] An expression reconstruction module is used to perform step three above, generating a target culture-specific expression potential vector based on the general emotional potential vector and the target culture identifier;
[0026] An animation synthesis module is used to perform step five above, synthesizing target multimodal animation based on the target culture-specific expression latent vector and the target virtual human model.
[0027] Furthermore, the system may also include a personalized micro-expression calibration module for performing step four, extracting personalized offset vectors from a small number of reference samples of the target individual, and providing them to the animation synthesis module to generate an animation incorporating the individual's micro-expression features. These modules may be software programs, firmware, hardware circuits, or any combination thereof, working collaboratively to complete the entire migration process. Attached Figure Description
[0028] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0029] Figure 1 This is a flowchart illustrating an AI digital human expression and facial feature transfer method provided in an embodiment of the present invention;
[0030] Figure 2 This is a schematic diagram of the dual-branch encoder network structure and training process for non-adversarial emotion decoupling in an embodiment of the present invention;
[0031] Figure 3 This is a structural block diagram of an AI digital human expression and facial feature transfer system provided in an embodiment of the present invention;
[0032] Figure 4 This is a schematic diagram of the potential spatial distribution of the emotional decoupling effect in an embodiment of the present invention;
[0033] Figure 5 This is a comprehensive comparative diagram of the facial expression transfer effects in the embodiments of the present invention. Detailed Implementation
[0034] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0035] Example 1
[0036] This invention provides a method for transferring facial expressions and features in AI digital humans, referring to... Figure 1 The execution flow of this method may include the following steps:
[0037] Step S100: Receive the multimodal animation data stream of the virtual human in the source region.
[0038] In this step, the data stream contains facial and body animation data carrying a specific expression (hereinafter referred to as "Expression X") within a source region's cultural context. The data stream can be acquired from various sources, such as real-time capture of live actors' performance data using optical or inertial motion capture equipment, or directly reading existing digital animation files stored in standard formats (such as FBX, BVH, USD). The animation data is typically organized in the form of frame sequences, for example, at a frame rate of 30 or 60 frames per second.
[0039] Preferably, the facial animation data is a sequence of facial vertex displacements or a blendshape weight sequence. Taking a blendshape weight sequence as an example, it can be a set of floating-point numbers conforming to mainstream industry standards (such as the 52 blendshape coefficients of Apple ARKit or the action unit (AU) defined by FACS (Facial Behavior Coding System)). Each time step (e.g., 1 / 60 second) corresponds to a weight vector, and each dimension of the vector represents the activation level of a basic expression (such as "left eye wide open" or "corner of the mouth turned up"), with values typically ranging from 0.0 to 1.0. The limb animation data is a sequence of three-dimensional coordinates of key points on the body skeleton (e.g., shoulder, elbow, wrist, finger joints, etc.). These sequences record the skeletal movement trajectory of the virtual person when performing coordinated postures.
[0040] Step S200: A dual-branch encoder network with a built-in gradient inversion layer is used to non-adversarially decouple the multimodal animation data stream.
[0041] The core of this step lies in separating the culturally relevant expressive paradigms inherent in "expression X" from the culturally unrelated intrinsic emotional states, thereby obtaining a source culture-specific latent vector and a universal emotional latent vector, respectively. This process avoids the problems of unstable training and easy pattern collapse inherent in traditional Generative Adversarial Networks (GANs).
[0042] Reference Figure 2The dual-branch encoder network 200 includes a general emotion encoder 210 and a culture-specific expression encoder 220. The general emotion encoder 210 extracts the general emotion latent vector from "expression X," which, for example, is a 128-dimensional floating-point vector designed to characterize the intrinsic core emotional states in the animation sequence that do not change with cultural background, such as the pure emotional essence of "joy," "sadness," and "anger." The culture-specific expression encoder 220 extracts the source culture-specific expression latent vector, which, for example, is a 256-dimensional floating-point vector, encoding the source region's cultural expression paradigm and the source virtual human's facial physiological features specific to "expression X." For example, the emotion of "joy" might be expressed as a hearty laugh in one culture, with its expression latent vector encoding features like "open mouth" and "squinting eyes"; while in another culture it might be expressed as a pursed-lip smile, with its vector encoding features like "slightly upturned corners of the mouth" and "tense cheek muscles."
[0043] As one of the core innovations of this invention, to achieve non-adversarial separation, a gradient reversal layer (GRL) 240 is inserted between the output of the general emotion encoder 210 and an auxiliary culture classifier 230. During the forward propagation of the network, the GRL 240 performs an identity transformation, directly passing the features extracted by the general emotion encoder 210 to the culture classifier 230. However, during the backpropagation to update the network parameters, the GRL 240 multiplies the gradient received from the culture classifier 230 by a negative constant (e.g., λ = -1). To train this network, a composite loss function L_total = L_reconstruction + γ is typically defined. L_classification, where L_reconstruction is a reconstruction loss to ensure that the two latent vectors can perfectly reconstruct the original input data through a decoder (used during training), guaranteeing the integrity of the information; L_classification is the classification loss (e.g., cross-entropy loss) of the culture classifier 230, used to distinguish cultural origins; γ is a hyperparameter used to balance the two loss terms. Due to the presence of GRL 240, when minimizing L_total, the gradient of L_classification is inverted when updating the parameters of the general sentiment encoder 210, effectively maximizing the classification error of the classifier, thereby driving the general sentiment encoder 210 to generate features that confuse the culture classifier 230 and prevent it from accurately determining the cultural origin. In this way, the encoder is effectively trained to strip away the culture-specific information from the input data, thereby extracting a purer and more general sentiment representation. As an alternative, generative models such as variational autoencoders (VAEs) can also be used to achieve decoupling by imposing constraints on the distribution of the latent space, but the GRL method used in this invention has significant advantages in training stability and the directness of decoupling.
[0044] Step S300: Based on the general emotional latent vector and the input target region cultural identifier, a target culture-specific expression latent vector that conforms to the target region cultural expression paradigm is reconstructed and generated through a condition generator.
[0045] The essence of this step is to use the inherent emotional state defined by the universal emotional potential vector extracted in step S200 as the core, and to match and generate a completely new expression mode that conforms to the target culture's habits. The target region cultural identifier can be an integer, a string, or a one-hot code, used to specify the target culture, such as "East Asia" or "North America".
[0046] Preferably, the condition generator retrieves a target cultural expression paradigm that matches the intrinsic emotional state and the target region's cultural identifier by querying a pre-built multimodal cultural expression knowledge base. This knowledge base is a structured database (e.g., a JSON file, SQLite database, or key-value store like Redis), whose data entries associate emotional tags (e.g., "embarrassment"), cultural identifiers (e.g., "K country") with specific multimodal expression primitives. For example, a data entry might be recorded as: {emotion: "embarrassment", culture: "K", expression: {facial_AUs: ["AU4: brow lowerer", "AU12: lip corner puller (slight)"], gesture: "hand covering mouth"}}. After receiving the general emotional vector, the condition generator first maps it to a discrete emotional tag classification using a small fully connected network. Then, it combines this with the target cultural identifier and searches the knowledge base to obtain a description of the target expression paradigm. This description is then converted into a target culture-specific expression latent vector with the same dimension as the source culture-specific expression latent vector for use by subsequent modules.
[0047] Furthermore, to enrich and improve this knowledge base, large language models (LLMs) can be used for assisted generation. By inputting specially designed prompts into the LLM (such as GPT-4 or similar models), such as "Please analyze and list the common facial expressions and gesture differences in expressing the emotion of 'surprise' in Italian and Korean cultures, and output them in JSON format," structured expression paradigm data can be extracted semi-automatically from the large amount of text descriptions returned by the LLM. This efficiently expands the knowledge base to cover a wider range of cultures and emotions.
[0048] Step S400: Perform personalized online micro-expression calibration on the target virtual human.
[0049] This step is an optional optimization step, designed to ensure that the generated expressions not only conform to the macro-cultural paradigm but also reflect the unique temperament of the target individual. A few-shot learned encoder is used to extract the unique micro-expression features of the target individual from a small amount of emotionally charged reference data and encode them into a personalized offset vector.
[0050] Preferably, the reference data can be one or more reference images of the target individual (e.g., a smiling photo of the target virtual person or its prototype real person), or a short video containing facial expression changes. The few-shot learning encoder is a lightweight convolutional neural network (CNN), such as a simplified version of MobileNetV3 or SqueezeNet, trained to extract subtle individual facial features that cannot be generalized by a macro-cultural knowledge base, used for fine-tuning "expression Y," such as a person's unique eyebrow raising habit or asymmetrical smile. The personalized offset vector output by the encoder, such as a vector with the same dimension as the culturally specific expression latent vector, will be used to fine-tune the generated expression. This process is "online" calibration, meaning that for a new target virtual person, only a small number of samples are needed for rapid adaptation without the need for large-scale retraining.
[0051] Step S500: Use a cross-topology multimodal decoder to synthesize target multimodal animation data.
[0052] This step synthesizes "tailor-made" multimodal animation data on the unique facial topology of the target virtual human based on the target culture-specific expression potential vector generated in step S300 (and the optional personalized offset vector generated in step S400) and the neutral pose 3D mesh model of the target virtual human, thereby ultimately generating a specific "expression Y".
[0053] Preferably, the multimodal decoder employs a Graph Convolutional Network (GCN) structure. GCN can directly operate on the mesh graph structure of the 3D model, treating each vertex as a node in the graph and the edges connecting vertices as edges. This allows the decoder to perceive and adapt to the unique facial topology of the target virtual human (i.e., different rigging and mesh densities), transforming abstract representation vector instructions into physically correct and visually natural vertex displacements on a specific model. This feature fundamentally solves the animation distortion problem caused by the inconsistency between the source and target model topologies. Alternatively, a Transformer structure based on an attention mechanism can be used, treating each vertex of the mesh as a "token," and learning the interactions between different facial regions through a self-attention mechanism, thus achieving cross-topology animation generation as well.
[0054] Furthermore, before generating the final animation, the target culture-specific latent vector obtained in step S300 is fused with the personalized offset vector obtained in step S400. The fusion method can be simple vector addition or concatenation, or a weighted summation, such as `FusedVector = α`. TargetCultureVector + β `PersonalOffsetVector`, where α and β are adjustable hyperparameters, for example, they can be set to 0.8 and 0.2 respectively. The fused vector is provided as input to the multimodal decoder to generate highly customized animations that conform to the target cultural paradigm and incorporate individual characteristics.
[0055] Ultimately, the output of the multimodal decoder is two synchronized data streams: one is facial animation data (such as Blendshape weight sequences or vertex displacement sequences) generated for the target virtual human facial model; the other is limb animation data (such as three-dimensional coordinate sequences of skeletal joints) that drives the target virtual human body skeleton to perform limb animation data (such as three-dimensional coordinate sequences of skeletal joints) that are coordinated with the "Y expression" and conform to the target cultural expression definition.
[0056] Example 2
[0057] This invention also provides an AI digital human expression and facial feature transfer system, referring to... Figure 3 This system is a hardware-based and modular implementation of the above-mentioned method and process, and may include:
[0058] An emotion decoupling module 310 is used to perform the above steps S100 and S200. It integrates a dual-branch encoder network, which is responsible for decoupling the general emotion latent vector and the source culture-specific expression latent vector from the input source animation data.
[0059] An expression reconstruction module 320, used to perform the above step S300, has a condition generator at its core and is connected to a multimodal cultural expression knowledge base. This module generates a target culture-specific expression latent vector based on the general sentiment latent vector output by the sentiment decoupling module 310 and the target culture identifier from external input.
[0060] An animation synthesis module 340, used to perform the above step S500, is at its core a multimodal decoder across topologies (such as a GCN decoder). This module receives the output of the expression reconstruction module 320 and the 3D model of the target virtual human, and synthesizes the final target multimodal animation.
[0061] Furthermore, the system may also include a personalized micro-expression calibration module 330 for performing step S400. This module receives a small number of reference samples from the target individual, extracts a personalized offset vector, and provides it to the animation synthesis module 340 to generate an animation incorporating the individual's micro-expression features.
[0062] The modules can be software programs, firmware, hardware circuits (such as FPGAs or ASICs), or any combination thereof, deployed on cloud servers or local high-performance computing platforms. In a typical embodiment, the system can be deployed on one or more servers equipped with high-performance graphics processing units (GPUs, such as NVIDIA A100 or RTX 4090), running Linux, and using deep learning frameworks such as PyTorch or TensorFlow. The modules interact with each other through pre-defined application programming interfaces (APIs), such as those based on RESTful or gRPC protocols, to collaboratively complete the entire expression and facial feature transfer process.
[0063] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for transferring facial expressions in AI digital humans, used to convert a source culture expression representing a certain inner emotion into a target culture expression representing the inner emotion on a target virtual human, characterized in that, The method includes the following steps: Step 1: Receive source multimodal animation data of a virtual human containing cultural expressions from a single source; Step 2: A dual-branch encoder network with a built-in gradient inversion layer is used to perform non-adversarial feature decoupling on the source multimodal animation data. The gradient inversion layer is used to invert the gradient of an auxiliary culture classifier to penalize the cultural information contained in the feature vector during network training, thereby purifying it into a culture-independent general emotional latent vector representing an intrinsic emotion. Step 3: Based on the general emotional latent vector and an input target cultural identifier, reconstruct the expression using a condition generator. The general emotional latent vector is used as a query index to retrieve and match the target cultural expression paradigm used to constitute a target cultural expression in the multimodal cultural expression knowledge base corresponding to the target cultural identifier. Based on this, a target cultural specific expression latent vector representing the target cultural expression is generated. Step 4: Learn the encoder using a few samples to extract and encode a personalized offset vector from a small amount of reference data of a target virtual human; Step 5: Using a multimodal decoder with a cross-topology structure, the target culture-specific expression latent vector is fused with the personalized offset vector, and then synthesized on the three-dimensional mesh model provided by the target virtual human to form a target multimodal animation that represents the expression of the target culture.
2. The method according to claim 1, characterized in that, The dual-branch encoder network in step two includes a general sentiment encoder and a culture-specific expression encoder. The general sentiment encoder is used to extract the general sentiment latent vector, and the culture-specific expression encoder is used to extract a source culture-specific expression latent vector.
3. The method according to claim 2, characterized in that, Between the output of the general emotion encoder and the auxiliary culture classifier, the gradient inversion layer is inserted. The gradient inversion layer performs an identity transformation during forward propagation and multiplies the gradient from the culture classifier by a negative constant during backward propagation, thereby driving the general emotion encoder to generate features that the culture classifier cannot distinguish as cultural origins.
4. The method according to claim 1, characterized in that, Step three uses the general sentiment latent vector as a query index, specifically including: first, mapping the general sentiment latent vector to a discrete sentiment label through a fully connected network, and then using the sentiment label and the target cultural identifier as a combined index to perform a retrieval in the multimodal cultural expression knowledge base.
5. The method according to claim 1 or 4, characterized in that, The construction of the multimodal cultural expression knowledge base utilizes a large-scale language model for generation. By analyzing multi-source texts and multimedia materials, the data entries in the knowledge base that associate emotional tags, cultural identifiers, and multimodal expression primitives are expanded.
6. The method according to claim 1, characterized in that, The multimodal decoder across topology in step five is a graph convolutional network. The graph convolutional network adapts to the unique facial topology of the target virtual human by performing graph convolution operations directly on the vertices and edges of the 3D mesh model provided by the target virtual human.
7. The method according to claim 1, characterized in that, The personalized offset vector in step four is used to encode the unique micro-expression features of the target virtual human that cannot be summarized by the multimodal cultural expression knowledge base; the small amount of reference data is one or more reference images of the target virtual human or a short video.
8. The method according to claim 1, characterized in that, The target multimodal animation output in step five consists of two synchronized data streams, specifically including: facial animation data generated for the target virtual human facial model, and limb animation data that drives the target virtual human body skeleton to perform in coordination with the facial expressions.
9. A system utilizing the AI digital human facial expression transfer method as described in claim 1, characterized in that, include: An emotion decoupling module is used to receive source multimodal animation data and decouple it into a general emotion latent vector using a dual-branch encoder network with a built-in gradient inversion layer. An expression reconstruction module is used to reconstruct and generate a target culture-specific expression potential vector based on the general emotional potential vector and a target culture identifier through a condition generator; An animation synthesis module is used to synthesize a target multimodal animation by employing a multimodal decoder across a topology, based on the target culture-specific expression latent vector and a 3D mesh model of a target virtual human.
10. The system according to claim 9, characterized in that, It also includes a personalized micro-expression calibration module for extracting a personalized offset vector from a small amount of reference data of the target virtual human; and the animation synthesis module is configured to fuse the target culture-specific expression latent vector with the personalized offset vector before synthesizing the target multimodal animation.
Citation Information
Patent Citations
Digital doctor expression, action and emotion interactive simulation system based on GAN
CN120495486A
Method for real-time generation of empathy expression of virtual human based on multimodal emotion recognition and artificial intelligence system using the method
US20250200855A1