Virtual human emotion expression system driven by adaptive neural network

The virtual human emotion expression system driven by an adaptive neural network achieves unified perception and rendering of multimodal data, solves the cross-modal perception and privacy issues of existing virtual human emotion systems, and improves the realism of virtual human emotion expression and user experience.

CN121120885APending Publication Date: 2025-12-12JIANGXI INST OF FASHION TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511669407.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-14
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Existing virtual human emotion systems lack a unified system framework, making it difficult to achieve a closed loop of cross-modal perception, semantic understanding, and rendering output. The emotion models are too abstract, failing to achieve parameterized driving of emotional states and virtual human expressions, voices, and postures, and there are issues of privacy versus utility trade-offs.

Method used

The virtual human emotion expression system driven by an adaptive neural network collects data such as EEG signals through a cognitively enhanced multimodal perception module, establishes a virtual human digital twin through an embodied digital twin modeling module, updates strategies through a federated meta-reinforcement learning optimization module, conducts risk assessments through an ethical and safety governance module, performs adaptive rendering through a cross-platform collaborative rendering module, and performs personalized fine-tuning through an emotional memory and personality development module, thus achieving closed-loop processing of perception, modeling, generation, optimization, rendering, and memory.

Benefits of technology

It enhances the realism and immersion of virtual human emotional expression, ensures user data security and privacy, achieves dynamic adaptation and high-quality rendering of individualized emotional expression, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121120885A_ABST
    Figure CN121120885A_ABST
Patent Text Reader

Abstract

The invention discloses a virtual human emotion expression system driven by a self-adaptive neural network, and relates to the technical field of artificial intelligence and virtual reality crossing, and the system comprises a cognitive enhanced multi-mode sensing module which is used for collecting an electroencephalogram signal, a near-infrared brain region blood oxygen signal, eye movement data, a voice signal, an image, a text, a physiological parameter and an environment signal; and executing cross-modal causal alignment and feature fusion, and outputting unified multi-modal emotion representation and cognitive state indexes. According to the method, neural cognitive physiological indexes such as electroencephalogram signals, near-infrared brain region blood oxygen signals and eye movement data are introduced, multi-modal information such as voice, images and texts is combined, deep modeling of the emotional state and the cognitive state of a user is achieved, non-causal-related artifacts are effectively eliminated through cross-modal causal alignment and feature fusion technologies, and the accuracy of the emotion state and the cognitive state of the user is improved. Uniformity and accuracy of multi-modal emotion representation are ensured, and the virtual human can more accurately understand the real emotion and intention of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of interdisciplinary technology of artificial intelligence and virtual reality, and in particular to an adaptive neural network-driven virtual human emotion expression system. Background Technology

[0002] Existing virtual human emotion systems largely rely on unimodal emotion recognition and rule mapping, making it difficult to uniformly utilize speech, facial expressions, posture, text, and physiological signals. Traditional muscle-driven models neglect the combined effects of fascia and skeleton, easily leading to stiff facial expressions. Model personalities are mostly statically configured, lacking long-term evolution and scenario-based adaptation. Federated learning involves a trade-off between privacy and utility, and adversarial examples can easily disrupt system stability. Cross-device rendering also struggles to maintain consistent emotion delivery quality across different devices and environments.

[0003] A published patent, "An Interaction Method and System Based on Virtual Humans" (Publication No.: CN113946209A), includes: virtual human image design, virtual human emotion model, emotion lexicon analysis, and interaction software design. The virtual human image design is relevant and attractive, and can be two-dimensional or three-dimensional. Key technologies include photo-based virtual human reconstruction methods, model-based three-dimensional reconstruction methods, OpenGL-based virtual human image technology, and agent-based virtual human image technology. The virtual human emotion model includes dimensional expression in emotion classification and dimensional analysis of the emotion model. The emotion lexicon analysis uses the hierarchical analysis method. The interaction software design is a machine-computer interaction software platform with smart home application background. This invention provides an interaction method and system based on virtual humans capable of emotional interaction.

[0004] The aforementioned patents have the following defects: they suffer from multiple technical deficiencies, such as fragmented structure, outdated algorithms, and insufficient implementation. In terms of overall architecture, they lack a unified system framework and module coupling mechanism, and are merely a functional listing design. They have not established a closed loop between cross-modal perception, semantic understanding, and rendering output. In terms of emotion models, the models are too abstract and do not provide specific feature extraction and mapping algorithms. Emotion recognition is limited to the static text level and lacks temporal modeling and contextual dependency capabilities. It also fails to achieve parameterized driving of emotional states and virtual human expressions, voice, and posture. Summary of the Invention

[0005] The purpose of this invention is to address the shortcomings of existing technologies by proposing an adaptive neural network-driven virtual human emotion expression system.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: an adaptive neural network-driven virtual human emotion expression system, comprising a cognitively enhanced multimodal perception module for acquiring EEG signals, near-infrared brain region blood oxygenation signals, eye-tracking data, speech signals, images, text, physiological parameters, and environmental signals, performing cross-modal causal alignment and feature fusion, and outputting a unified multimodal emotion representation and cognitive state index; an embodied digital twin modeling module for establishing a virtual human digital twin based on a coupled dynamic model of fascia, muscles, and bones, and generating an emotion state vector, intention distribution, and expression constraint set by combining a joint causal graph of emotion and intention; and a physical-level emotion generation module for generating facial and body posture control parameters and speech synthesis based on the emotion state vector, intention distribution, and expression constraint set. The system includes: control parameters and rendering material update instructions; a federated meta-reinforcement learning optimization module for performing policy updates and structural optimizations based on user interaction feedback under privacy protection conditions, outputting individualized weight matrices and model update instructions; an ethics and safety governance module for risk assessment of emotional intensity and duration and triggering emotion buffering, expression degradation, or topic switching, while allocating privacy budgets and implementing adversarial example defense; a cross-platform collaborative rendering module for adaptively distributing rendering tasks based on device performance and user visual sensitivity, and performing ambient lighting and audio compensation; and an emotional memory and personality development module for hierarchical storage and retrieval of emotional memories and scene-specific fine-tuning of personality parameters. All modules interact through a data bus and an event bus to achieve closed-loop processing of perception, modeling, generation, optimization, governance, rendering, and memory.

[0007] As a further description of the above technical solution: The cognitive-enhanced multimodal perception module includes: a neural signal fusion layer for jointly estimating cognitive load, emotional arousal, and attention allocation by combining EEG and brain region blood oxygenation features; a context-adaptive sampling unit for dynamically adjusting the sampling rate and feature extraction strategy according to environmental complexity and cognitive load; and a cross-modal emotion causal inference network for removing non-causal artifacts and outputting a time-aligned multimodal feature sequence.

[0008] As a further description of the above technical solution: The embodied digital twin modeling module includes: a fascia, muscle and bone coupled dynamics model to constrain the biomechanical consistency of virtual human expressions and movements; an individual skeletal morphology adaptation unit to calibrate the virtual skeletal structure based on the user's facial and skeletal parameters; a dynamic emotion and intention causal graph network to achieve joint inference of emotional state and intention distribution; and an empathy mapping engine to identify empathy-sensitive dimensions and adjust mapping weights based on the user's historical interactions.

[0009] As a further description of the above technical solution: The physical-level emotion generation module includes: a skin microstructure dynamic response unit for adjusting skin elasticity and optical parameters according to physiological state; a physiological-optical linkage rendering unit for updating diffuse reflection, specular reflection and subsurface scattering parameters of the skin; a posture control unit for outputting control parameters of bones and muscle groups; and an emotion-driven speech synthesis unit for generating speech output according to emotional state and individual vocal tract parameters.

[0010] As a further description of the above technical solution: The physical-level emotion generation module includes a cross-modal synchronization control unit for synchronizing speech, lip movements and actions through temporal prediction, and driving lip and jaw micro-movements based on phoneme, muscle tone and airflow intensity signals.

[0011] As a further description of the above technical solution: The federated meta-reinforcement learning optimization module includes: a cross-domain knowledge transfer unit for performing domain adaptability filtering within the federated distillation framework to avoid negative transfer; a zero-shot and single-shot adaptation engine for accelerating the adaptation of sentiment models for new users; a multi-objective reward modeling unit for dynamically adjusting reward weights based on users' implicit needs; and a model structure evolution unit for performing pruning, quantization, and operator fusion operations based on terminal computing power.

[0012] As a further description of the above technical solution: The ethical and security governance module includes: an emotional boundary control system for triggering emotional buffering or topic switching based on emotional intensity, duration, and user psychological resilience; a privacy utility balancing unit for allocating differential privacy budgets according to data sensitivity and recording audit information; and an adversarial example detection unit for detecting forged signals and switching to backup features to ensure system security.

[0013] As a further description of the above technical solution: The cross-platform collaborative rendering module includes: a distribution engine driven by both user visual sensitivity and device capabilities to adaptively adjust between facial expression detail accuracy and rendering frame rate; and an environment-aware compensation system to perform lighting reconstruction, shadow correction, and volume balancing under different lighting and noise conditions.

[0014] As a further description of the above technical solution: The Emotional Memory and Personality Development module includes: a hierarchical emotional memory bank for configuring core, regular and temporary layers according to importance and setting decay periods and retrieval weights for each layer; a personality evolution engine for performing scenario-based fine-tuning on a long-term stable personality axis and maintaining consistency in emotional expression; and a contextual recall mechanism for waking up relevant memories and linking them with user preferences through multi-anchor triggering and fuzzy matching.

[0015] As a further description of the above technical solution: A method for a virtual human emotion expression system driven by an adaptive neural network includes the following steps: S1, acquiring multimodal signals and generating unified multimodal emotion representations and cognitive state indicators; S2, inferring emotion state vectors and intention distributions based on a digital twin modeling module; S3, generating facial expressions, postures, and speech outputs based on emotion states; S4, completing environmental compensation and real-time rendering output in a cross-platform collaborative rendering module; S5, updating policy parameters and structure using a federated meta-reinforcement learning optimization module; S6, performing risk assessment and privacy control through an ethical and security governance module; and S7, updating emotion memory and personality parameters by an emotion memory and personality development module, thereby forming a continuously evolving closed loop for virtual human emotion expression.

[0016] The present invention has the following beneficial effects: 1. In this invention, neurocognitive physiological indicators such as electroencephalogram (EEG) signals, near-infrared brain oxygenation signals, and eye-tracking data are first introduced. Combined with multimodal information such as speech, images, and text, this enables deep modeling of the user's emotional and cognitive states. Cross-modal causal alignment and feature fusion techniques effectively eliminate non-causal artifacts, ensuring the uniformity and accuracy of multimodal emotional representation. This allows the virtual human to more accurately understand the user's true emotions and intentions. A virtual human digital twin is established based on a coupled dynamic model of fascia, muscles, and bones, ensuring biomechanical consistency of expressions and movements. Combined with a joint causal graph of emotions and intentions, this allows emotional expression to extend beyond... Instead of simple facial expressions and movements, the virtual human can simulate more subtle and natural physiological changes through technologies such as dynamic response of skin microstructure and physiological optics linkage rendering. This significantly enhances the realism and immersion of the virtual human's emotional expression. The federated meta-reinforcement learning optimization module can continuously update and optimize the virtual human's emotional expression strategy by utilizing user interaction feedback while protecting privacy. This enables the generation of individualized weight matrices. The zero-shot and single-shot adaptation engine accelerates model adaptation for new users, allowing the virtual human to quickly adapt to the preferences and interaction patterns of different users. This achieves dynamic and personalized evolution of emotional expression, overcoming the limitation of the traditional virtual human's "one-size-fits-all" emotional expression.

[0017] 2. In this invention, the ethical and security governance module conducts risk assessments on the intensity and duration of emotions and can trigger emotional buffering, expression degradation, or topic switching, effectively avoiding potential user discomfort or negative impacts. By allocating privacy budgets and implementing adversarial sample defenses, it ensures the security and privacy of user data, constructing a safe and reliable virtual human interaction environment. The cross-platform collaborative rendering module adaptively distributes rendering tasks based on device performance and user visual sensitivity, and dynamically adjusts between facial expression detail accuracy and rendering frame rate. The environmental perception compensation system performs lighting reconstruction, shadow correction, and volume balancing under different lighting and noise conditions, ensuring high-quality and smooth performance of the virtual human under different hardware and environments, improving user experience, emotional memory, and individual... The personality development module stores and retrieves emotional memories in a hierarchical manner and uses the personality evolution engine to fine-tune the personality on a long-term stable personality axis in a scenario-based manner, maintaining the consistency of emotional expression. The contextual recall mechanism can awaken relevant memories through multi-anchor triggering and fuzzy matching, and link with user preferences, so that the virtual human can form a more stable and profound personality in long-term interaction with users, and express emotions more intelligently based on historical experience and user preferences. The modules interact through data bus and event bus to realize closed-loop processing of perception, modeling, generation, optimization, governance, rendering and memory, so that the system can continuously learn and optimize based on user feedback, forming an adaptive ability of dynamic personality evolution, and providing users with a highly personalized and immersive virtual human interaction experience. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of the system architecture of the present invention; Figure 2 This is a schematic diagram of the system usage method of the present invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] Reference Figure 1-2One embodiment of the present invention provides an adaptive neural network-driven virtual human emotion expression system, comprising: a cognitively enhanced multimodal perception module for acquiring EEG signals, near-infrared brain blood oxygenation signals, eye-tracking data, speech signals, images, text, physiological parameters, and environmental signals; performing cross-modal causal alignment and feature fusion; and outputting a unified multimodal emotion representation and cognitive state index. An embodied digital twin modeling module is used to establish a virtual human digital twin based on a coupled dynamic model of fascia, muscles, and bones, and to generate an emotion state vector, intention distribution, and expression constraint set by combining a joint causal graph of emotion and intention. A physical-level emotion generation module is used to generate facial and body posture control parameters and speech synthesis control parameters based on the emotion state vector, intention distribution, and expression constraint set. The system includes: a data bus and an event bus; a federated meta-reinforcement learning optimization module; a cross-platform collaborative rendering module; an ethical and safety governance module; an ethical and safety governance module; a data bus and an event bus; and a data bus and event bus. The data bus is responsible for transmitting large-scale data, such as raw sensor data, feature vectors, and model parameters. The event bus is used to transmit control signals and triggers between modules. For example, when the perception module detects a high-intensity emotion, it can immediately notify the ethical and safety governance module to conduct a risk assessment. The federated meta-reinforcement learning optimization module performs risk assessment on the intensity and duration of emotions and triggers emotion buffering, expression degradation, or topic switching, while allocating a privacy budget and performing adversarial example defense. The cross-platform collaborative rendering module adaptively distributes rendering tasks based on device performance and user visual sensitivity, and performs ambient lighting and audio compensation. The emotional memory and personality development module is used to hierarchically store and retrieve emotional memories and fine-tune personality parameters in a scenario-based manner. All modules interact through the data bus and event bus to achieve closed-loop processing of perception, modeling, generation, optimization, governance, rendering, and memory. The data bus and event bus are the communication backbone connecting all the modules. The data bus is responsible for transmitting large-scale data, such as raw sensor data, feature vectors, and model parameters. The event bus is used to transmit control signals and triggers between modules. For example, when the perception module detects a high-intensity emotion, it can immediately notify the ethical and safety governance module to conduct a risk assessment via the event bus. This modular design and clear communication mechanism ensure the system's scalability, robustness, and efficient collaboration.

[0021] The cognitive-enhanced multimodal perception module includes: a neural signal fusion layer for jointly estimating cognitive load, emotional arousal, and attention allocation using EEG and brain region oxygenation features; a context-adaptive sampling unit for dynamically adjusting the sampling rate and feature extraction strategy based on environmental complexity and cognitive load; and a cross-modal emotion causal inference network for eliminating non-causal artifacts and outputting a time-aligned multimodal feature sequence. As the system's perception front-end, the cognitive-enhanced multimodal perception module comprehensively collects multimodal signals related to user emotions and cognition. Specifically, it can collect EEG signals, near-infrared cerebral oxygenation (fNIRS) signals, eye-tracking data (such as pupil size and fixation point), speech signals, images (such as facial expressions and body posture), text (such as user input and dialogue content), physiological parameters (such as heart rate and skin conductance), and environmental signals (such as background noise and illumination). The EEG (for cognitive load and arousal) band power density (Welch) formula is as follows: , :frequency band average power density, Number of frequency points within the frequency band Spectral power density. Differential entropy formula (DE): , :frequency band The differential entropy, : Bandpass segment variance. Linear inversion formula for near-infrared fNIRS (used for cognitive load and arousal) blood oxygen concentration changes: , , Changes in oxygenated and deoxygenated hemoglobin concentrations. :wavelength The change in absorbance. : A constant determined by the extinction coefficient and the effective path factor. Eye movement and pupil (for attention) gaze concentration index formula: , Concentration Index The cumulative gaze duration falling on the area of ​​interest. Total gaze duration. Mel-frequency cepstral coefficients (MFCC) formula for speech (used for emotion and arousal): , : No. cepstral coefficients, : No. One Mel filter energy, Number of filters. Fundamental frequency statistics: within the window. Mean, standard deviation, rate of change. AU intensity vector of video face and upper body (for emotion and movement): Skeletal joint angles: and angular velocity Text (for sentiment and intent association input) sentence vectors: pre-trained language model output sentence embeddings Physiological and environmental factors (for quality assessment and compensation): Heart rate variability RMSSD: , The root mean square of the difference between adjacent heartbeats. : No. During the cardiac interval, : Number of heartbeat intervals, environmental quantity: illuminance ,noise Time window statistics were performed. After collecting this raw data, cross-modal causal alignment and feature fusion were executed. This means that it is not simply a matter of concatenating features from different modalities, but rather using advanced machine learning and deep learning algorithms, such as fusion models based on graph neural networks (GNNs) or attention mechanisms, to identify causal relationships between different modalities and eliminate non-causal artifacts. Lag estimation and artifact suppression formulas: , Modality To mode The optimal causal lag, : Lag search set, in seconds Mutual information value :right The time scrambling operation is used as a non-causal control sequence. : Forgery penalty coefficient, based on Time shifting of subordinate modes aligns the output sequence. Causal mask attention (used for subsequent fusion): , The query, key, and value matrix is ​​obtained by concatenating the alignment features of each modal encoder at the same time. : Scaling constant, equal to the dimension of the key vector. Causal mask matrix; when the time difference between the key source and the query source does not meet the requirement. When a relationship is determined to be a spurious dependency, a negative constant is added to that pair of positions to suppress attention weights. This ensures that the fused features accurately reflect the user's real-time emotional and cognitive states (e.g., cognitive load, emotional arousal, attention allocation). A unified multimodal emotional representation and cognitive state index are output for use by the backend modeling module. The quality weight formula is: Robust fusion formula: , : No. Modal quality weights This mode at time The normalized reconstruction error, This mode at time Signal-to-noise ratio estimation, Error and noise penalty coefficient, :Depend on Attention triplets obtained by linear mapping The linear mapping matrix from attention output to latent representation. Unified multimodal latent representation. Two-dimensional formula for sentiment: Parallel header of categories: , Valence and arousal regression output Distribution of sentiment categories For a fixed number of categories, Linear mapping parameters. Cognitive ternary indicator head formula: , , , Cognitive load normalized score Normalized score of emotional arousal Attention allocation normalized score : The corresponding weight vector, : The corresponding bias scalar, : Sigmoid activation function.

[0022] The embodied digital twin modeling module includes: a fascia, muscle, and skeleton coupled dynamics model to constrain the biomechanical consistency of the virtual human's expressions and movements; an individual skeletal morphology adaptation unit to calibrate the virtual skeletal structure based on the user's facial and skeletal parameters; a dynamic emotion and intention causal graph network to achieve joint inference of emotional states and intention distributions; and an empathy mapping engine to identify empathy-sensitive dimensions and adjust mapping weights based on the user's historical interactions. The embodied digital twin modeling module receives multimodal emotion representations and cognitive state indicators from the perception module. Its core lies in building a virtual human digital twin based on the fascia, muscle, and skeleton coupled dynamics model. This model is not merely a digital clone of appearance, but more deeply simulates the intrinsic mechanisms of human biomechanics, ensuring the biomechanical realism and consistency of the virtual human's expressions and movements. Combined with the joint causal graph of emotion and intention, this module can achieve joint inference of the user's emotional state and potential intentions. This graph captures the complex correlation between different emotional states and specific intentions; for example, happiness may correspond to the intention to share, and anxiety may correspond to the intention to seek help. The embodied digital twin modeling module generates an emotion state vector, an intent distribution, and a set of expression constraints. The emotion state vector describes the current emotion type and intensity, the intent distribution indicates the user's possible behavioral tendencies, and the set of expression constraints can limit the scope and intensity of emotion expression based on the virtual human's personality, scene context, or ethical considerations to prevent over-rendering or inappropriate expression.

[0023] The physical-level emotion generation module includes: a skin microstructure dynamic response unit to adjust skin elasticity and optical parameters according to physiological state; a physiological-optical linkage rendering unit to update diffuse reflection, specular reflection, and subsurface scattering parameters of the skin; a posture control unit to output control parameters for bones and muscle groups; and an emotion-driven speech synthesis unit to generate speech output based on emotional state and individual vocal tract parameters. The physical-level emotion generation module also includes: a cross-modal synchronization control unit to achieve synchronization of speech, lip movements, and actions through temporal prediction, and to drive micro-movements of the lips and jaw based on phoneme, muscle tension, and airflow intensity signals. Taking emotional state vectors, intent distributions, and expression constraint sets as input, the physical-level emotion generation module is responsible for transforming abstract emotions and intentions into executable, embodied expressions for the virtual human. It generates facial and body posture control parameters, speech synthesis control parameters, and rendering material update instructions. Facial and body posture control parameters directly drive the virtual human's skeletal, muscular, and fascial models, producing realistic and biomechanically consistent expressions and movements. Speech synthesis control parameters guide the speech synthesis engine to generate emotionally charged speech output, including speech rate, tone, volume, and timbre. The rendering material update command is used to adjust the visual effects of virtual human body parts, such as skin color, gloss, or the visibility of microvessels, to simulate physiological changes. The skin microstructure dynamic response unit adjusts skin elasticity and optical parameters based on physiological states (such as the impact of heart rate changes on blood flow); the physiological optics linkage rendering unit updates the diffuse, specular, and subsurface scattering parameters of the skin to simulate physiological optical changes such as facial flushing and pallor; the posture control unit outputs control parameters for bones and muscles, precisely controlling the virtual human's posture and expressions; and the emotion-driven speech synthesis unit generates emotionally rich speech output based on emotional states and individual vocal tract parameters. The cross-modal synchronization control unit achieves synchronization of speech, lip movements, and actions through temporal prediction, and drives lip and jaw micro-movements based on phoneme, muscle tension, and airflow intensity signals, ensuring multimodal consistency and realistic details in emotional expression.

[0024] The Federated Meta-Reinforcement Learning Optimization Module includes: a cross-domain knowledge transfer unit for performing domain-adaptive filtering within the federated distillation framework to avoid negative transfer and ensure effective knowledge sharing among different users; a zero-shot and single-shot adaptation engine for accelerating the adaptation of sentiment models for new users through meta-learning, enabling rapid personalization even with little or no historical data; a multi-objective reward modeling unit for dynamically adjusting reward weights based on users' implicit needs, making the virtual human more aligned with users' deep expectations; and a model structure evolution unit for performing pruning, quantization, and operator fusion operations based on terminal computing power to optimize model efficiency and enable efficient operation on different devices. The Federated Meta-Reinforcement Learning Optimization Module is the core of the system's adaptive capability. Under privacy-preserving conditions, it performs policy updates and structural optimization based on user interaction feedback. When users interact with the virtual human, this module collects user behavioral data, physiological feedback, or explicit ratings as reward signals. Through the Federated Learning mechanism, local models from multiple users can learn and aggregate parameters without sharing raw data, thus protecting user privacy. Meta-reinforcement learning enables models to quickly learn from small amounts of new user data and adapt to new environments, outputting individualized weight matrices and model update instructions. The individualized weight matrix is ​​used to adjust the emotional expression preferences of virtual avatars for specific users, while the model update instructions guide the evolution of the overall system model.

[0025] The ethical and security governance module includes: an emotion boundary control system to trigger emotion buffering or topic switching based on emotion intensity, duration, and user psychological resilience to avoid overstimulation; a privacy utility balancing unit to allocate differential privacy budgets based on data sensitivity and record audit information to ensure transparent and traceable data use; and an adversarial example detection unit to detect forged signals and switch to backup features to ensure system security and prevent external interference from maliciously manipulating virtual human expressions. The ethical and security governance module is responsible for ensuring the ethical compliance and security of virtual human emotional expressions. It conducts risk assessments on emotion intensity and duration, and immediately triggers interventions such as emotion buffering (reducing emotion intensity), expression downgrading (simplifying expression), or topic switching once it detects situations that may cause user discomfort or abuse. The federated meta-reinforcement learning optimization module also allocates privacy budgets and implements adversarial example defense. The privacy budget mechanism imposes restrictions on access to and use of user data, ensuring it remains within a privacy protection framework. Adversarial example defense identifies and filters potential malicious input signals to prevent the system from being attacked or misled, thereby ensuring system security and user experience.

[0026] The cross-platform collaborative rendering module includes: a dual-driven distribution engine based on user visual sensitivity and device capabilities, used to adaptively adjust between facial expression detail accuracy and rendering frame rate, balancing visual quality and performance; and an environment-aware compensation system used to perform lighting reconstruction, shadow correction, and volume balancing under different lighting and noise conditions, providing a more immersive interactive experience. The cross-platform collaborative rendering module is responsible for providing high-quality rendering output across different devices and environments. It adaptively distributes rendering tasks based on device performance (such as GPU computing power and memory) and user visual sensitivity (such as preference for detail and tolerance for stuttering). For example, higher-end devices can render more detailed facial expressions, while lower-end devices prioritize frame rate. This module also performs ambient lighting and audio compensation. Ambient lighting compensation analyzes actual ambient lighting conditions and adjusts the virtual human's lighting model to make it look natural under different lighting conditions. Audio compensation dynamically adjusts the virtual human's voice output volume and spectrum based on ambient noise levels, ensuring clear intelligibility even in noisy environments.

[0027] The Emotional Memory and Personality Development module includes: a hierarchical emotional memory bank configured with core, regular, and temporary layers based on importance, each with its own decay period and retrieval weight; a personality evolution engine that performs scenario-based fine-tuning on a long-term stable personality axis while maintaining consistency in emotional expression; and a contextual recall mechanism that uses multi-anchor triggering and fuzzy matching to awaken relevant memories and link them to user preferences. The Emotional Memory and Personality Development module is responsible for storing and retrieving the virtual human's emotional history and dynamically evolving personality parameters. It employs a hierarchical storage structure to store and retrieve emotional memories, such as configuring core, regular, and temporary layers based on importance, each with different decay periods and retrieval weights, ensuring that important memories are preserved long-term and easily recalled. The personality evolution engine performs scenario-based fine-tuning on a long-term stable personality axis, ensuring that the virtual human develops a unique and consistent personality over long-term interactions, rather than changing randomly. For example, an introverted virtual human might be slightly more active in a specific scenario, but their core introverted characteristics remain unchanged. The contextual memory mechanism awakens relevant historical memories through multi-anchor triggering and fuzzy matching, and links with user preferences, enabling virtual humans to make more intelligent and appropriate emotional responses based on past experiences and user preferences.

[0028] A method for a virtual human emotion expression system driven by an adaptive neural network includes the following steps: S1, acquiring multimodal signals and generating a unified multimodal emotion representation and cognitive state index; in this step, a cognitively enhanced multimodal perception module is responsible for acquiring various physiological and behavioral signals of the user during interaction, including EEG signals, near-infrared brain region blood oxygenation signals, eye-tracking data, speech signals, images, text, physiological parameters, and environmental signals. These heterogeneous signals undergo cross-modal causal alignment and feature fusion to eliminate non-causal artifacts and ensure temporal consistency and information purity of different modal data sources. Finally, a unified multimodal emotion representation and cognitive state index are generated, such as the user's real-time cognitive load, emotional arousal, and attention allocation state, providing high-quality input for subsequent emotion modeling. Causal constraints ensure that the alignment between signals from different modalities accurately reflects the emotional state, rather than accidental artifacts. Through joint encoding of event-related features and brain region blood oxygenation features, the system can extract core cognitive state indices such as cognitive load, emotional arousal, and attention allocation, which help eliminate noise in cross-modal data. Assuming there are two modalities... and The corresponding time series is and ,in Indicates a point in time. The core of cross-modal alignment is to ensure synchronicity and causal dependence between two signals through causal relationship modeling. Formula: The formula represents (e.g., EEG signals) The causal influence of (such as facial expression signals) requires cross-modal causal alignment networks to ensure that this causal relationship is correctly modeled, thereby avoiding artifacts that are not causally related. The first modal signal in time The value (e.g., EEG signal). The second mode signal in time The value (e.g., facial expression signals). : Represents a cause-and-effect relationship, that is right The impact. Time point. The above model can eliminate non-causal artifacts, ensure temporal consistency of different signals, and retain effective information during feature fusion. Through a neural signal fusion layer, cognitive state indicators such as cognitive load, emotional arousal, and attention allocation are estimated jointly using EEG and blood oxygenation signals. Cross-modal alignment enables unified emotional representation and cognitive state indicator output for data from different modalities. A unified emotional representation vector is defined. It combines data from different modes. Assuming each mode signal... It can be done through a specific transformation function Mapped onto the emotion representation space, the final unified emotion representation is: ,in, This represents the feature extraction function for each modality. It is modal In time The signal. A unified emotion representation vector, in time The value of . :No. A modal signal in time The values ​​(e.g., EEG, speech, video, etc.). Feature extraction function for each modality. Time Point. This method ensures the effective integration of emotional information from different modalities, thereby generating a comprehensive and accurate emotional representation. Context-Adaptive Sampling: During the acquisition process, the sampling rate is dynamically adjusted according to environmental complexity and cognitive load. In high-load scenarios, the sampling density is increased and wavelet denoising is performed; in low-load scenarios, the sampling frequency is reduced and rate-of-change features are extracted to ensure information purity. Denoising and Feature Compression: Unnecessary noise is removed through wavelet denoising and feature compression techniques while retaining key emotion-related information.

[0029] S2. The digital twin modeling module infers the emotional state vector and intention distribution. After acquiring multimodal emotional representations and cognitive state indicators, the embodied digital twin modeling module receives this information. This module first ensures the biomechanical realism of the virtual human digital twin based on a coupled dynamics model of fascia, muscles, and bones. Then, combining the joint causal graph of emotions and intentions, it infers the emotional state vector that the current virtual human should express (such as the AU (ActionUnit) weight of facial expressions and the joint angle range of body posture), the user's potential intention distribution (such as asking questions, affirmation, negation, and seeking help), and the set of expression constraints (such as cultural etiquette in the current context and user personalized preferences). For example, when the system perceives that the user is showing confused emotions and seeking explanations, it infers that the virtual human should express an emotional state of understanding and patient explanation. Then, by combining the joint causal graph of emotion and intention, the system infers the emotional state vector that the virtual human should express (such as the AU (ActionUnit) weight of facial expressions and the joint angle range of body posture), the distribution of the user's potential intentions (such as asking questions, affirming, denying, and seeking help), and the set of expression constraints (such as cultural etiquette in the current context and the user's personalized preferences). For example, when the system perceives that the user is showing confusion and seeking explanation, it infers that the virtual human should express an emotional state of understanding and patient explanation. To ensure the biomechanical realism of the virtual human digital twin, a coupled dynamic model between fascia, muscles, and bones needs to be established. This model can simulate the virtual human's body movements under various emotional expressions, including facial expressions and limb postures. This model not only considers muscle movement but also the role of fascia (the connective tissue covering the muscles) and bones to better simulate the real movement and mechanical response of the human body. Fascia-muscle-skeleton coupled dynamic equations: , Mass matrix, representing the mass distribution of muscles, fascia, and bones. Damping matrix: simulates the friction between muscles, fascia, and bones. : Stiffness matrix, representing the stiffness of each joint and muscle connection. : Displacement vectors of the joints and bones of the virtual human. : Acceleration vector, describing the acceleration of joint movement. : Velocity vector, describing the velocity of joint movement. External forces, including gravity, external propulsion, etc. Muscle strength, the force generated by muscle activity. Fascial force, considering the tensile and deformation characteristics of fascia. This equation models the motion behavior of the virtual human as the interaction of mass, damping, and stiffness. Through this equation, the virtual human's expressions and postures will be simulated based on the interaction of muscles, fascia, and bones, thus ensuring its biomechanical realism. When interacting, the virtual human needs to express emotions (such as happiness, sadness, anger, etc.) and accurately convey the user's intentions (such as asking questions, affirming, negating, etc.). This requires combining the joint causal graph of emotions and intentions to infer the emotional state vector that the virtual human should express (such as the AU weight of facial expressions, the joint angle range of body postures), as well as the distribution of the user's potential intentions. Joint causal graph model for emotion and intention inference: , : The emotional state vector of a virtual human, including facial expressions, body posture, etc. The distribution of users' potential intentions, representing intentions such as asking questions, affirming, denying, and seeking help. The user's multimodal input signals, including EEG, voice, facial expressions, etc. Cognitive state indicators of virtual humans, such as cognitive load and emotional arousal. The conditional probability of an emotional state represents the emotion that a virtual human should express given a user signal and the virtual human's cognitive state. : Conditional probability of intent, representing the user's potential intent given the user signal and the virtual human's cognitive state. The joint probability of emotional state and intention represents the association between emotion and intention. This formula infers the user's current emotional state and potential intention through a joint causal graph. This model combines multimodal signals (such as voice, facial expressions, eye movements, etc.) and the virtual human's cognitive state to accurately predict how the virtual human should express emotions and intentions. For example, when a user exhibits confusion and seeks explanation, the system can infer that the virtual human needs to express understanding and patient explanation. In inferring the virtual human's emotional state, it is necessary to consider not only its biomechanics and emotional fit but also the user's cultural background and personalized preferences. Different cultures and personalized needs may impose different constraints on emotional expression, therefore these factors need to be considered when inferring emotional state. Emotional Expression Constraint Generation Model: , The final emotion expression vector consists of constrained facial expression and posture parameters. : Preliminary inference of emotional state vectors (such as facial AU weights and posture joint angles). User intent distribution, representing their potential needs (such as asking questions, expressing affirmation, etc.). Cultural constraints refer to the user's cultural background (such as tolerance for emotional expression). Personalized preferences refer to a user's individual emotional expression preferences (such as liking humor or seriousness). The mapping function, combining emotion, intention, culture, and personality preferences, adjusts the final emotional expression. This model ensures that when virtual humans express emotions, they not only consider biomechanics and emotional matching but also adjust their emotional expression according to the user's cultural and personalized needs. For example, in some cultures, directly expressing emotions may be impolite, while personalized preferences may require virtual humans to display more humorous or more serious emotions. Through this model, virtual humans can better meet the user's needs and provide a personalized emotional interaction experience. When the system perceives that a user is exhibiting confused emotions (such as furrowed brows or shifty eyes) and an intention to seek help, the system infers the emotional state that the virtual human should express based on the causal graph and expression constraint model above. This usually means that the virtual human should show understanding and patience, for example, through a gentle tone of voice, soft facial expressions, and relaxed body language to express "understanding" and "patient explanation." Through these technologies, virtual humans can express emotions and intentions more naturally, realistically, and personally, thereby enhancing the interactive experience with users.

[0030] S3. Generate facial expressions, postures, and speech output based on emotional state. The physical-level emotion generation module generates control parameters to drive the virtual human's expression based on the emotional state vector, intent distribution, and expression constraint set generated in step S2. These parameters include facial expression control parameters, body posture control parameters, and speech synthesis control parameters. Specifically, the skin microstructure dynamic response unit adjusts skin elasticity and optical parameters; the physiological optics linkage rendering unit updates the diffuse, specular, and subsurface scattering parameters of the skin to simulate subtle physiological changes in the skin. The posture control unit outputs control parameters for bones and muscle groups, thereby driving the virtual human to generate biomechanically realistic expressions and movements. The pitch, rate, timbre, and emotional prosody of the speech are adjusted to match the emotional state; and rendering material update instructions are generated, for example, adjusting the diffuse, specular, and subsurface scattering parameters of the skin based on the virtual human's physiological state (such as blushing or sweating) to enhance realism. The emotion-driven speech synthesis unit generates emotionally rich speech output based on the emotional state and individual vocal tract parameters. The cross-modal synchronization control unit ensures the temporal synchronization and detail consistency of speech, lip movements, and actions. The physical-level emotion generation module generates various control parameters for the virtual human's performance by utilizing emotion state vectors, intention distributions, and a set of expression constraints. These control parameters refer to those that control the virtual human's facial expressions, body posture, and speech output. The emotion state vector is generated by fusing data from multiple modalities (such as EEG signals, speech signals, and facial expression data), reflecting the virtual human's emotional state. The intention distribution reflects the virtual human's intentions, such as whether it is expressing emotion, and whether the intention is understanding, helping, or being sarcastic. The set of expression constraints ensures that the virtual human's expression is biomechanically sound and conforms to ethical, safety, and cultural norms. Using this information, the system can generate a series of control parameters: Facial expression control parameters: These include control parameters for each action unit (AU) of the facial muscles, used to precisely control the virtual human's facial expressions. For example, facial expressions may include smiling, frowning, etc. These control parameters are generated through the tension of the corresponding muscles. Facial expression control parameter formula: , Facial expressions over time The control parameters, :No. Individual action units (such as frowning, smiling) in time The strength, The weights of the corresponding action units represent their impact on emotional expression. Body posture control parameters involve the angles of various joints in the virtual human skeleton. Through these control parameters, the virtual human can perform body movements consistent with emotional states, such as opening arms to welcome or bowing the head to apologize. Speech synthesis control parameters include parameters such as pitch, speech rate, timbre, and intonation to ensure that the emotional expression of the speech is consistent with the emotional state. For example, when the virtual human expresses sadness, the speech has a lower pitch, a slower speech rate, and a somber tone. The skin microstructure dynamic response unit is responsible for dynamically adjusting the physical properties of the skin according to the virtual human's emotional state. These properties include skin elasticity and optical parameters, which are crucial for simulating natural skin changes, especially during emotional expression. Skin elasticity: As emotions change (such as anger, shame, etc.), the tension and elasticity of the skin may change. For example, when angry, the facial skin may show more tension. Optical parameters include skin gloss and transparency, which have a significant impact on the visual effect of the skin during the virtual human's emotional expression. Skin may appear glossy when experiencing strong emotions, such as the reflected light from a flushed face. The physiological optics-linked rendering unit updates the skin's optical parameters, especially those for diffuse reflection, specular reflection, and subsurface scattering, to simulate subtle physiological changes in the skin. Diffuse reflection: The degree to which the skin scatters light. The skin's reflective properties differ with emotional changes; for example, when tense, the skin reflects more light. Specular reflection: The glossy effect produced when light is directly reflected onto the eyes or other facial areas. When a virtual human is angry or happy, specular reflection may be enhanced, resulting in a higher glossiness. Subsurface scattering: The process of refraction and scattering of light after entering the skin surface. Emotional changes affect blood flow to the skin, thus influencing this process. For example, increased facial blood circulation makes the skin appear more flushed. The posture control unit generates skeletal and muscle control parameters based on the virtual human's emotional state, driving the virtual human to make biomechanically reasonable expressions and movements. Skeletal control parameters: Simulating the virtual human's dynamic posture by adjusting the angles of the skeletal joints. For example, when expressing joy, the virtual human may adjust the position of its shoulders and arms, appearing relaxed and natural. Muscle control parameters: Controlling muscle tension makes the virtual human's movements more natural. For example, when angry, the virtual human's facial muscles may be more tense, and their body posture may appear aggressive. The speech synthesis control unit generates speech output based on emotional state and individual vocal tract parameters. To ensure that the virtual human's speech is consistent with its emotional state, this unit needs to adjust the following parameters: Pitch: Higher pitch when happy, lower pitch when sad. Speech rate: Faster speech rate when happy, slower speech rate when anxious. Timbre: Bright timbre when happy, dull timbre when sad. Emotional prosody: By controlling the prosody of the speech, it is made more consistent with the expressed emotion. The formula for adjusting the emotional prosody of speech synthesis is as follows: , Voice output in time audio signal, Emotional state vector (e.g., joy, sadness). :pitch, Speech rate Phono. To enhance the realism and immersion of the virtual human's performance, rendering material update commands control the optical properties of the virtual human's skin, face, and other body parts, such as gloss and transparency, ensuring that the virtual human's appearance is consistent with its emotional expression. The cross-modal synchronization control unit is responsible for ensuring the temporal synchronization and detail consistency between speech, lip movements, and actions. This unit can adjust the virtual human's lip movements in real time through analysis of the speech signal to match the timing of speech pronunciation, while ensuring coordination between body movements and facial expressions. For example, when the virtual human says "hello," the lip movements and the timing of speech pronunciation must be strictly synchronized to ensure no delays or unnatural phenomena.

[0031] S4. Environmental compensation and real-time rendering output are completed in the cross-platform collaborative rendering module. The cross-platform collaborative rendering module receives the expression instructions generated in step S3 and adaptively distributes rendering tasks based on the current device's performance (such as GPU computing power and memory) and the user's visual sensitivity (such as preference for detail and tolerance for frame rate). This module dynamically adjusts the balance between facial detail accuracy and rendering frame rate to provide the best user experience on different devices. Simultaneously, this module also performs ambient lighting and audio compensation. For example, by analyzing the lighting conditions of the user's environment in real time, it adjusts the rendering lighting parameters of the virtual human to blend it with the real environment; by analyzing ambient noise, it adjusts the volume and clarity of the virtual human's voice to ensure a high-quality, highly immersive interactive experience in different physical environments. The cross-platform collaborative rendering module first needs to evaluate the current device's performance and the user's visual sensitivity, and then adaptively allocates rendering tasks based on this information. Device performance evaluation includes GPU computing power: referring to the computing power of the device's graphics processing unit (GPU). Devices with higher GPU computing power can support more complex graphics rendering, such as high-resolution textures, higher levels of detail, and faster rendering speeds. Memory: The size of a device's memory determines the number and quality of objects that can be loaded and rendered. Devices with more memory can support more complex scene rendering and more real-time dynamic effects. User visual sensitivity assessment includes preferences for detail: Some users prefer to see higher quality images and like richly detailed rendering effects, while others prefer a smoother experience and may not pay much attention to rendering details. Tolerance for frame rate: Frame rate refers to the number of frames displayed per second. Users' requirements for frame rate vary; some users have a higher tolerance for lower frame rates, while others have a higher demand for higher frame rates. Based on device performance and user preferences, the cross-platform collaborative rendering module dynamically adjusts according to the following two aspects: Facial detail accuracy: Facial expressions and movements of virtual humans are an important component of rendering quality. Higher detail accuracy means more facial muscle action units and more complex expression control. However, this also consumes more GPU computing power and memory. Therefore, when device performance is poor or users are less sensitive to detail, the module will reduce the detail of facial expressions to reduce rendering complexity. Rendering frame rate: A higher frame rate provides a smoother user experience, which is especially crucial in fast-paced action scenes. When the device has strong performance, the system can increase the rendering frame rate to ensure smoother and more natural movements of the virtual human; conversely, when the device has weak performance, the system may reduce the frame rate to maintain system stability and avoid stuttering caused by excessively low frame rates. Formula: , Rendering quality (including detail accuracy and frame rate). The device's GPU computing power. : Device memory size. User's preference for details (high, low). User's tolerance for frame rate (high, low). This formula determines the optimal allocation of rendering quality by comprehensively considering GPU computing power, memory, detail preference, and frame rate tolerance. The cross-platform collaborative rendering module not only needs to adjust rendering quality but also needs to ensure that the virtual human can adapt to changing ambient lighting and noise conditions in different physical environments, thereby providing the best immersion. Ambient lighting compensation includes lighting conditions: In a real physical environment, the intensity and direction of light directly affect the rendering effect of the virtual human's appearance. For example, strong backlighting or low-light environments may cause the virtual human to appear too dark or too bright. To ensure consistent rendering of the virtual human in different environments, the system analyzes real-time lighting conditions and dynamically adjusts the lighting parameters during virtual human rendering. Compensation methods: If the user's ambient lighting is dark, the rendering module will enhance the virtual human's light reflection to improve the visibility of details; if the lighting is too strong, the module will reduce the light intensity to prevent the virtual human from being overexposed. Ambient noise compensation includes noise analysis: Ambient noise (such as wind noise, traffic noise, background voices, etc.) will affect the clarity of speech and immersion. The system needs to analyze and compensate for ambient noise in real time to ensure the virtual human's voice is clear and easy to understand. Compensation method: Based on the intensity of ambient noise, the system dynamically adjusts the volume and clarity of the virtual human's voice. For example, when the ambient noise is high, the voice volume will automatically increase, and the voice clarity will be enhanced to ensure that the user can hear the virtual human clearly; while in a quiet environment, the volume will be appropriately reduced to avoid excessive volume differences. Compensation formula: , The audio signal after voice compensation. Current ambient light intensity, Ambient noise intensity The original volume of the virtual human's voice. This formula indicates that the voice volume and rendering effects are adjusted based on real-time ambient light and noise intensity information.

[0032] S5. The Federated Meta-Reinforcement Learning (FMR) optimization module updates policy parameters and structure. During user interaction, the FMR optimization module continuously receives user feedback (including implicit feedback such as behavioral patterns and physiological indicators, and explicit feedback such as ratings or evaluations). Under privacy protection (through a federated learning mechanism), this module updates the emotion expression strategy and optimizes the model structure based on this feedback. The cross-domain knowledge transfer unit avoids negative transfer and accelerates the learning process; the zero-shot and single-shot adaptation engine ensures that new users can quickly obtain personalized experiences; the multi-objective reward modeling unit adjusts the reward function according to the user's implicit needs; and the model structure evolution unit performs pruning, quantization, and operator fusion based on the terminal's computing power to continuously improve the model's efficiency and effectiveness. Finally, it outputs an individualized weight matrix and model update instructions to adjust the virtual human's next emotion expression. Each user's device independently collects feedback data and trains the model using local data. After training, the device sends the updated model parameters (not the original data) to the central server. The server aggregates the model updates from all devices and then feeds them back to each device, completing the global update of the model. Through this mechanism, the virtual human's emotional expression strategy can be personalized on each user's device without leaking the user's private data. The task of the multi-objective reward modeling unit is to dynamically adjust the reward function based on the user's implicit needs, thereby guiding the virtual human to produce emotional expressions that better meet the user's expectations during interaction. The virtual human's emotional expression needs to be optimized not only based on the user's explicit feedback (such as likes, ratings, etc.) but also based on implicit needs (such as the user's preference for specific emotions, emotional changes during interaction, etc.). Implicit needs can be obtained through long-term interaction data analysis and influence the design of the reward function. Explicit feedback: includes opinions directly expressed by users, such as through ratings, selections, etc. Implicit needs: the emotions or preferences that users potentially express through interaction with the virtual human. For example, some users may prefer the virtual human to express humor or comforting emotions, while others may prefer a more rational and direct response. Assume that the reward the virtual human receives based on user feedback during interaction is... This reward can be modeled using the following multi-objective reward function: , The reward value at the current moment. Rewards from explicit feedback, such as ratings or likes. Rewards derived from implicit needs, such as behavioral patterns based on changes in user emotions. Rewards derived from emotional expression strategies, such as user satisfaction with the emotional expression of virtual avatars. Weighting coefficients represent the contributions of different feedback sources to the total reward. Based on these reward signals, the virtual human can adaptively adjust its emotional expression and optimize its interaction with the user. The role of the model structure evolution unit is to continuously optimize the efficiency and effectiveness of the virtual human model under limited device computing resources. To this end, the system uses techniques such as pruning, quantization, and operator fusion to optimize the model's computational performance and ensure that the model can run efficiently on different devices. Pruning: This refers to removing redundant neural network connections or neurons in the model, reducing computational load, and thus accelerating the inference process. The goal of pruning is to remove parts that have little impact on model performance and retain only those that contribute significantly to the output. Quantization: This converts floating-point calculations in the model into low-precision integer calculations to reduce memory usage and computational complexity. Quantization can significantly improve the running speed on low-performance devices while reducing power consumption. Operator fusion: This combines multiple computational operations into one operation to reduce intermediate steps and computational load. For example, multiple convolution operations and activation function operations can be combined into an optimized operator to accelerate the computation process. These optimization strategies help the virtual human provide high-quality emotional expression even with limited hardware resources (such as mobile phones and low-end devices). Assuming the computational resource limitations of the model can be mitigated by adjusting parameters. Let represent the efficiency of the optimized model. It can be calculated using the following formula: , The efficiency of the model is represented by the degree of reduction in computational resource consumption. Limitations of computing resources (such as memory, computing power, etc.). :Pruning ratio, indicating the proportion of neurons or connections pruned. Quantization ratio indicates the degree of reduction in parameter precision. The degree of operator fusion represents the optimization effect of merging computational operations. This formula indicates that by adjusting different optimization strategies, the computational efficiency of the virtual human emotion expression model on low-resource devices can be improved. During federated learning and model structure evolution, the virtual human needs to generate a weight matrix based on each user's personalized needs and output model update instructions to ensure that the virtual human's emotion expression can be adjusted according to each user's specific feedback. Individualized Weight Matrix: This matrix contains the emotion expression characteristics and preferences for each user, dynamically generated based on user feedback, emotion changes, and other factors. Each user's weight matrix reflects the user's preference for different emotion expressions of the virtual human. Model Update Instructions: These instructions contain information on adjusting model parameters, such as weights and biases updated during training, and changes to the model structure (such as pruned network structures or quantized models). Individualized Weight Matrix It can be represented as: , User-specific weight matrix Weighting coefficients represent the contribution of different emotional features to user preferences. In time The The influence of individual emotional features (such as smiling, frowning, etc.) on users' emotional expression. Model update instructions. It can be represented as: , The model update command indicates parameter adjustment or structural change. Individualized weight matrix : Update parameters, indicating the weights or biases that need to be adjusted. Resource constraints are used to adjust update strategies based on the device's computing power. Through these instructions, the virtual human's emotional expression model can be optimized based on user feedback after each interaction.

[0033] S6. Risk assessment and privacy control are performed through the ethics and security governance module. This module conducts real-time risk assessments of the intensity and duration of the virtual human's emotional expressions. If the assessment results indicate potential discomfort or ethical risks to the user (e.g., overly intense or prolonged emotional expression), the module will trigger emotion buffering (e.g., reducing the intensity of emotional expression), expression degradation (e.g., downgrading from strong facial expressions to micro-expressions), or topic switching. Simultaneously, this module allocates a differential privacy budget to ensure the privacy and security of sensitive user data during model training and optimization, and implements adversarial example defense to detect and resist malicious attacks, ensuring system security. For example, when the virtual human detects a user's low mood, it will avoid expressing overly positive or stimulating emotions and may proactively switch to a lighter topic. Emotional Expression Intensity: The virtual human's emotional expressions can have varying intensities. For example, a joyful emotional expression may range from a slight smile to a loud laugh, while a sad emotional expression may range from a somber tone to an extremely sorrowful demeanor. The intensity of the virtual human's emotions must be adjusted according to the user's emotional state and the interaction environment. Emotional Expression Duration: The duration of emotions also needs to be controlled. Prolonged emotional expression can cause discomfort or fatigue for users. For example, prolonged expressions of joy might seem overly affected to some users. The ethics and safety governance module analyzes the intensity and duration of the virtual human's emotional expressions in real time to determine whether they will cause discomfort or potential ethical risks to users. If the assessment results indicate a potential adverse effect on users, the module will take the following measures: Emotional buffering: When the assessment results show that the virtual human's emotional expression is too intense, the system will reduce the intensity of the emotional expression. For example, it will reduce a joyful expression from a loud laugh to a smile, or a sad emotion from a somber tone to a slight sadness. Expression degradation: If the virtual human's emotional expression lasts too long, the system will downgrade the expression, converting strong expressions into micro-expressions. For example, if the virtual human is in a state of constant joy, it may downgrade a large smile to a slight upturn of the corners of the mouth. Topic switching: If the current emotional expression cannot be adjusted to an appropriate level over a long period of time, the system will proactively switch topics to a more neutral emotional theme. Topic switching can effectively avoid discomfort caused by emotional expression. For example, if a virtual human senses that a user is feeling down and has been expressing positive emotions for an extended period, it might automatically switch to a more comforting tone to avoid making the user feel pressured. This assumes an assessment of the intensity of emotional expression. and duration Risk can be assessed using the following function: , Risk assessment value: This indicates the degree of ethical risk or user discomfort that the current emotional expression may bring. The intensity of emotional expression, quantifying the strength of virtual human's emotional expression. The duration of emotional expression indicates how long the virtual human maintains that emotion. If the risk assessment value... If the emotion exceeds a predetermined threshold, the system will trigger emotion buffering, expression degradation, or topic switching. During federated learning and model training, the virtual human's emotional expression strategy relies on a large amount of user interaction data. However, this data may contain users' private information, such as emotional changes, voice data, and behavioral patterns. To protect user privacy, differential privacy technology is used to ensure that users' personal information is not leaked even during analysis and training. Differential privacy technology protects data privacy by adding noise, ensuring that any analysis result of a single user's data does not significantly affect the overall result, thus effectively preventing privacy leaks. Differential privacy budget: Each time data is shared or the model is updated, the system allocates a privacy budget. A higher privacy budget maintains higher accuracy during model updates but offers lower privacy protection; a lower privacy budget offers stronger privacy protection but may affect model updates. (The text then abruptly shifts to a different topic: privacy budget.) The risk of data breach can be calculated using the following formula: , Privacy budget: the smaller the budget, the stronger the privacy protection. Sensitivity to data changes indicates the degree to which the model output changes when the data changes. During model training and updates, by applying differential privacy mechanisms, virtual human models can be optimized without leaking user data, while protecting user privacy. Malicious attackers may use adversarial examples (e.g., modified input data) to interfere with the emotional expression of virtual humans, causing them to produce incorrect reactions or behaviors. Adversarial example detection involves identifying potential malicious inputs through algorithms that may lead to errors in the emotional expression or model judgment of the virtual human system. Common detection methods include: Gradient detection: detecting malicious modifications by analyzing gradient changes in the input data; Adversarial training: adding adversarial examples to the training dataset to enhance the model's robustness against malicious attacks; Defense strategy formula: the defensive effect against adversarial examples. This can be expressed by the following formula: , Defense effectiveness refers to the ability to detect and defend against adversarial examples. The input data may be adversarial examples or normal data. The change between the input data and the model output represents the impact of adversarial examples on the model.

[0034] S7. The Emotional Memory and Personality Development module updates emotional memories and personality parameters, forming a continuously evolving closed loop for virtual human emotional expression. Based on user interaction history and feedback, the Emotional Memory and Personality Development module stores and retrieves emotional memories in layers and fine-tunes the virtual human's personality parameters according to specific scenarios. This includes updating core memories, regular memories, and temporary memories, and adjusting their decay cycles and retrieval weights. The personality evolution engine performs fine-tuning on a long-term stable personality axis, ensuring the consistency of the virtual human's emotional expression and allowing its personality to dynamically evolve with user interactions. The contextual recall mechanism awakens relevant memories and links them to user preferences through multi-anchor triggers and fuzzy matching, enabling the virtual human to better understand users and provide a more personalized interactive experience. For example, if a user consistently shows a preference for a certain style of humor, this module will strengthen the virtual human's personality parameters in this area, causing it to display this humor more frequently in subsequent interactions. Through the closed loop of the above steps, the virtual human can continuously learn from interactions, optimizing its emotional perception, understanding, expression, and adaptability, achieving dynamic personalized evolution. The Emotional Memory and Personality Development module stores and manages emotional memories by analyzing the user's interaction history and feedback with the virtual human. Memory is categorized into different levels and stored and retrieved based on its importance and frequency of use. There are three main memory levels: Core Memory: This includes the user's long-term preferences, emotional needs, and behavioral patterns. These memories are generally stable and reflect the user's basic emotional needs and interaction patterns. Core memories have a longer decay period and a higher retrieval weight. Regular Memory: This records the user's daily interactions and preferences, which may change over time. Regular memories have a lower retrieval weight but are updated based on the user's interaction frequency. Temporary Memory: This includes the user's recent interactions and emotional feedback, which typically have a significant impact on the virtual human's emotional expression in the short term. Temporary memories have a shorter decay period and a lower retrieval weight. Based on the user's interaction history and feedback, the virtual human regularly updates the memories at different levels and adjusts their decay period and retrieval weight according to the memory's importance and frequency of use. For example, the longer decay period of core memories means that the user's long-term emotional tendencies are not easily changed; while the shorter decay period of temporary memories means that recent interactions will affect the virtual human's emotional expression. (The text then abruptly shifts to a different topic: Assuming the weight of a certain memory...) Over time The decay process can be represented by the exponential decay formula: , In time The memory weight, Initial memory weights represent the initial weights of the memory. Attenuation coefficient: controls the rate at which memory decays. Time represents the time difference from the moment the memory was recorded. This formula shows that the weight of a memory gradually decreases over time, especially for temporary memories, which decay more rapidly. The role of the personality evolution engine is to continuously adjust the virtual human's personality traits based on user feedback and interaction history, ensuring that the virtual human's emotional expression and interaction style remain highly consistent with the user, and gradually evolve as the interaction deepens. The virtual human's personality is built on a long-term stable personality axis, which includes several dimensions, such as extroversion-introversion, humor-seriousness, etc. The personality evolution engine fine-tunes the virtual human's personality traits (such as humor, intensity of emotional response, etc.) to meet the user's needs by analyzing user feedback. For example, if the user prefers a humorous interaction style, the virtual human may add more humorous elements to the dialogue; if the user prefers rational analysis, the virtual human will reduce the humor and provide more logical and rigorous answers. The adjustment of personality evolution can be modeled using the following formula: , User personality traits over time The value, User personality traits over time The value, : Fine-tuning coefficient, indicating the degree of influence of feedback on individual adjustments. The formula represents the personality changes triggered by user feedback. It indicates that the virtual human's personality is continuously updated based on user feedback, ensuring that the virtual human's emotional expression gradually adapts to the user's personalized needs. The contextual recall mechanism uses multi-anchor triggering and fuzzy matching techniques to awaken relevant memories, enabling the virtual human to better understand the user's emotional needs and preferences and provide personalized responses. Multi-anchor triggering refers to the virtual human identifying the user's current context through multiple key events or emotional signals (such as changes in tone, topic shifts, and emotional feedback), thereby activating memories related to that context. Through the contextual recall mechanism, the virtual human can match and trigger the most relevant emotional memories by analyzing multiple signals from the user's current context. For example, the virtual human might judge the user's emotional state through changes in tone and topic, thereby evoking related historical memories and providing a more personalized emotional response. The activation of contextual recall can be represented by the following formula: , The intensity of activated emotional memories, The weight of each contextual signal (such as tone, topic, etc.) represents the contribution of that signal to memory activation. Regarding the first Contextual signals Activated Memories. This formula shows that virtual humans activate associated emotional memories through multiple contextual signals, thereby providing personalized responses to users. Virtual humans can continuously learn and optimize their emotional perception, understanding, and expression abilities from interactions with users. Each interaction with the user influences the virtual human's personality adjustments, which in turn affect subsequent emotional expression and behavioral responses. This closed-loop process achieves the dynamic and personalized evolution of the virtual human's emotional expression, enabling it to better meet the user's individual needs over time. Assuming that the virtual human updates its emotional expression and personality traits through feedback after each interaction, the closed-loop learning process can be represented by the following formula: , Users' emotional expression and personality traits over time The value, Users' emotional expression and personality traits over time The value, The learning coefficient represents the degree to which each interaction influences emotional expression and personality fine-tuning. User feedback leads to changes in emotional expression and personality. Each instance of user feedback drives the virtual human's emotions and personality to evolve in a direction that better suits the user, thus achieving long-term personalized evolution.

[0035] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. An adaptive neural network driven virtual human affective expression system, characterized in that: The cognitive enhancement multi-modal perception module is used for collecting electroencephalogram signals, near-infrared brain region blood oxygen signals, eye movement data, voice signals, images, texts, physiological parameters and environmental signals, performing cross-modal causal alignment and feature fusion, and outputting unified multi-modal emotional representation and cognitive state indicators; The embodiment digital twin modeling module is used for establishing a virtual human digital twin body based on a coupling dynamics model of fascia, muscle and skeleton, generating an emotional state vector, an intention distribution and an expression constraint set in combination with a joint causal graph of emotions and intentions; The physical level emotional generation module is used for generating facial and body posture control parameters, speech synthesis control parameters and rendering material update instructions according to the emotional state vector, the intention distribution and the expression constraint set; The federal meta-reinforcement learning optimization module is used for performing policy updating and structure optimization based on user interaction feedback under privacy protection conditions, and outputting individualized weight matrix and model update instructions; The ethical safety governance module is used for risk assessment of emotional intensity and duration and triggering emotional buffering, expression degradation or topic switching, while allocating privacy budget and performing adversarial sample defense; The cross-platform collaborative rendering module is used for adaptive distribution of rendering tasks according to device performance and user visual sensitivity, and performs environmental lighting and audio compensation; The emotional memory and personality development module is used for hierarchical storage and retrieval of emotional memory and scene-based fine-tuning of personality parameters. The modules interact through data bus and event bus to realize closed-loop processing of perception, modeling, generation, optimization, governance, rendering and memory.

2. The adaptive neural network driven virtual human emotional expression system according to claim 1, wherein: The cognitive enhancement multi-modal perception module includes: a neural signal fusion layer for estimating cognitive load, emotional arousal and attention allocation by combining electroencephalogram and brain region blood oxygen features; a context adaptive sampling unit for dynamically adjusting sampling rate and feature extraction strategy according to environmental complexity and cognitive load; a cross-modal emotional causal inference network for eliminating non-causal related artifacts and outputting time-aligned multi-modal feature sequences.

3. The adaptive neural network driven virtual human emotional expression system according to claim 1, wherein: The embodiment digital twin modeling module includes: a fascia, muscle and skeleton coupling dynamics model for constraining the biomechanical consistency of virtual human expressions and actions; an individual bone morphology adaptation unit for calibrating virtual bone structure according to user facial and skeletal parameters; a dynamic emotional and intention causal graph network for joint inference of emotional state and intention distribution; an empathy mapping engine for identifying empathy sensitive dimensions and adjusting mapping weights based on user historical interaction.

4. The adaptive neural network driven virtual human emotional expression system of claim 1, wherein: The physical level emotional generation module includes: a skin microstructure dynamic response unit for adjusting skin elasticity and optical parameters according to physiological state; a physiological optical linkage rendering unit for updating diffuse reflection, specular reflection and subsurface scattering parameters of skin; a posture control unit for outputting control parameters of skeleton and muscle groups; an emotion-driven speech synthesis unit for generating voice output according to emotional state and individual vocal tract parameters.

5. The adaptive neural network driven virtual human affective expression system according to claim 1, wherein: The physical level emotional generation module includes: a cross-modal synchronization control unit for realizing synchronization of voice, lip shape and action through time series prediction, and driving lip and jaw micro-motions based on phoneme, muscle tension and airflow intensity signals.

6. The adaptive neural network driven virtual human affective expression system of claim 1, wherein: The federal meta-reinforcement learning optimization module includes: a cross-domain knowledge transfer unit for performing domain adaptability filtering under a federal distillation framework to avoid negative transfer; a zero-shot and single-shot adaptation engine for accelerating emotional model adaptation for new users; a multi-objective reward modeling unit for dynamically adjusting reward weights according to user implicit needs; and a model structure evolution unit for performing pruning, quantization and operator fusion operations according to terminal computing power.

7. The adaptive neural network driven virtual human affective expression system of claim 1, wherein: The ethical safety governance module includes: an emotional boundary control system for triggering emotional buffering or topic switching according to emotional intensity, duration and user psychological resilience; a privacy utility balance unit for allocating differential privacy budgets according to data sensitivity and recording audit information; and an adversarial sample detection unit for detecting fake signals and switching backup features to ensure system safety.

8. The adaptive neural network driven virtual human affective expression system of claim 1, wherein: The cross-platform collaborative rendering module includes: a user visual sensitivity and device capability dual-driven distribution engine for adaptive adjustment between expression detail precision and rendering frame rate; and an environment perception compensation system for performing light reconstruction, shadow correction and volume balance processing under different light and noise conditions.

9. The adaptive neural network driven virtual human affective expression system of claim 1, wherein: The emotional memory and personality development module includes: a hierarchical emotional memory bank for configuring three layers of core, regular and temporary according to importance and setting decay periods and retrieval weights; a personality evolution engine for performing scenario-based fine-tuning on long-term stable personality axes and maintaining consistency in emotional expression; and a context recall mechanism for waking up related memories through multi-anchor triggering and fuzzy matching and linking user preferences.

10. A method applied to the adaptive neural network driven virtual human emotional expression system of claims 1-9, characterized in that: The method comprises the following steps: S1, collecting multi-modal signals and generating unified multi-modal emotional representation and cognitive state indicators; S2, inferring emotional state vectors and intention distribution based on a digital twin modeling module; S3, generating expression, posture and speech output according to emotional state; S4, completing environment compensation and real-time rendering output in a cross-platform collaborative rendering module; S5, updating policy parameters and structure using a federal meta-reinforcement learning optimization module; S6, performing risk assessment and privacy control through an ethical safety governance module; S7, updating emotional memory and personality parameters by an emotional memory and personality development module, thereby forming a continuously evolving virtual human emotional expression closed loop.

Citation Information

Patent Citations

  • Interaction method and system based on virtual human

    CN113946209A