Intelligent Construction System and Method for the Appearance Features of Digital Humans Based on Big Data 5G Interaction

By introducing an intelligent construction system for big data 5G interaction in the digital human expression generation technology, using technical means such as multi-agent timing differential learning and emotional infectious dynamics model, the problems of synergistic generation, emotional infection and long-term adaptation of digital human expressions in a multi-person interaction environment are solved, and a more natural, coordinated and personalized digital human expression expression is achieved.

CN120032027BActive Publication Date: 2025-06-27GUANGDONG SOUTHERN PLANNING & DESIGNING INST OF TELECOM CONSULTATION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510490762.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-06-27
Estimated Expiration
2045-04-18

Smart Images

  • Figure CN120032027B_ABST
    Figure CN120032027B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of digital human expression generation, and discloses a system and method for intelligently constructing digital human appearance features based on big data 5G interaction. The method for intelligently constructing digital human appearance features based on big data 5G interaction includes: constructing a multi-agent temporal difference learning framework to achieve collaborative generation of digital human expressions; constructing an emotion contagion dynamics model to simulate and predict the process of emotion spreading and diffusion in a digital human group; applying a variant of the SIR contagion model to simulate group emotion diffusion; implementing an embodied cognitive expression co-evolution framework to achieve long-term adaptive evolution of expressions; and constructing an expression evolution fitness evaluation system to quantify the optimization target. The present invention solves problems such as collaborative generation of expressions, emotion contagion modeling, and social relationship adaptation in multi-person interaction scenarios in the prior art, and improves the natural realism and coordination of digital human expressions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital human expression generation, and more specifically, to a system and method for intelligently constructing digital human appearance features based on big data 5G interaction. Background Art

[0002] With the popularization of 5G technology and the development of virtual reality technology, digital humans are increasingly used in multi-person interactive scenarios such as virtual social interaction, online education, and digital entertainment. The naturalness and coordination of digital human facial expressions directly affect the user's immersion and interactive experience, and have become a key factor in the development of digital human technology.

[0003] Existing digital human expression generation technologies mainly focus on single-person interaction scenarios, and achieve expression synthesis through preset expression templates or deep learning-based generation models. However, these technologies have obvious shortcomings in multi-person interaction environments:

[0004] First, there is a lack of collaborative generation capabilities for group expressions, which results in incoordination and incoherence between the expressions of multiple digital humans;

[0005] Second, it is impossible to effectively simulate the phenomenon of group emotional contagion and reproduce the natural diffusion process of emotions in a group;

[0006] Third, there is a lack of the ability to model group interaction history and social relationships, making it impossible to achieve long-term evolutionary adaptation of facial expressions.

[0007] These technical problems result in existing digital humans having unnatural expressions, poor emotional coherence, and insufficient sense of reality in group scenes in 5G interactive environments. They are unable to meet the needs of highly immersive multi-person interactive applications. There is an urgent need for a new digital human expression generation technology that can solve the problems of collaborative generation of group expressions, emotional contagion modeling, and social relationship adaptation. Summary of the invention

[0008] The present invention provides a system and method for intelligently constructing the appearance features of interactive digital humans based on big data 5G, which solves the technical problems in related technologies such as incoordinated expressions, unnatural emotional transmission, and lack of long-term adaptability of digital humans in multi-person interaction scenarios.

[0009] The present invention provides a method for intelligently constructing appearance features of digital humans based on big data 5G interaction, comprising the following steps:

[0010] Construct a multi-agent temporal difference learning framework, including group expression coordination function, multi-agent temporal difference learning algorithm and emotional feedback reward function based on 5G network;

[0011] Construct an emotion contagion dynamics model, including an emotion diffusion function and a social network emotion contagion mapping model, to simulate and predict the propagation and diffusion process of emotions in a digital human group;

[0012] Applying the SIR contagion model variant, the emotional state of the digital human is divided into emotional susceptibility state, emotional infection state and emotional recovery state, and a group of state transition differential equations is established;

[0013] Implement the embodied cognitive expression co-evolution framework, dynamically adjust the expression generation strategy based on the social relationship expression modulation function;

[0014] Construct an expression evolution fitness evaluation system to quantitatively evaluate the naturalness, coordination and diversity of digital human group expressions based on the expression fitness function.

[0015] Furthermore, the three indicators of naturalness, coordination and diversity of the digital human group's expressions are continuously optimized and adjusted through user emotional signals fed back in real time through the 5G network.

[0016] Furthermore, the group expression coordination function is defined as:

[0017]

[0018] in, Indicates The individual expression state vector of a digital person, Indicates The expression weight coefficient of a digital person, represents the expression interaction influence function, which is used to quantify the The expression state of the digital person The degree of influence of the facial expression state of the digital person.

[0019] Furthermore, the emotion diffusion function is defined as:

[0020]

[0021] in, Indicates the initial emotional state, represents the group network structure, Representation and Node The local network structure is connected. The time variable representing the diffusion of emotion, Representation Node At the moment The emotional contagion function.

[0022] Furthermore, the state transition differential equations in the SIR infection model variant are defined as:

[0023]

[0024]

[0025]

[0026] Among them, , , respectively represent the probabilities that the digital human is in an emotion-susceptible state, an emotion-infected state, and an emotion-recovery state. represents the emotion contagion coefficient of the digital human . represents the emotion recovery coefficient of the digital human . represents the social contact matrix.

[0027] Furthermore, the social relationship expression modulation function is defined as:

[0028]

[0029] Among them, and respectively represent the expression state vectors of the digital human and . represents the social relationship strength matrix between the digital human and . represents the influence factor.

[0030] Furthermore, the expression evolution fitness function is defined as:

[0031]

[0032] Among them, represents the digital human group, represents the naturalness scoring function of the group expression, represents the coordination scoring function of the group expression, represents the diversity scoring function of the group expression, , and are weight coefficients.

[0033] Furthermore, the expression of the emotional feedback reward function is as follows:

[0034]

[0035] Among them: represents the user emotional feedback value, which is obtained by analyzing the interaction data such as the voice, expression, and operations of the user in the 5G environment; represents the group expression coordination degree, which evaluates the collaborative consistency of the expressions of multiple digital humans; It represents the interaction activity, reflecting the interaction frequency and depth between the user and the digital human; , , are weight coefficients used to balance the contributions of different factors to the reward.

[0036] Furthermore, the social network emotion contagion mapping model combines social network analysis methods to establish a mapping model between the emotion contagion intensity, social relationship intensity, and individual traits, quantifying the emotion contagion effect between digital humans. The expression is as follows:

[0037]

[0038] Where: represents the emotion contagion intensity between digital human and ; represents the social relationship intensity matrix between digital human and ; and respectively represent the individual trait vectors of digital human and , including parameters such as emotion sensitivity and expression ability; represents the historical interaction frequency between digital human and ; , and are weight coefficients used to balance the contributions of different factors to emotion contagion.

[0039] The intelligent construction system for the appearance features of the big data 5G interactive digital human is used to execute the above-mentioned intelligent construction method for the appearance features of the big data 5G interactive digital human, including:

[0040] The multi-agent temporal difference learning framework construction module: realizes the natural collaborative generation of the facial micro-expressions of multiple digital humans;

[0041] The emotion contagion dynamics model construction module: is used to simulate and predict the spread and diffusion process of emotions in the digital human group;

[0042] The SIR contagion model variant application module: is used to simulate and predict the spread and diffusion process of emotions in the digital human group;

[0043] The embodied cognition expression co-evolution framework implementation module: dynamically adjusts the expression generation strategy according to the social relationships and interaction histories between digital humans;

[0044] Expression evolution fitness evaluation module: Quantitatively evaluate the naturalness, coordination and diversity of the expressions of digital human groups, and provide optimization goals and evaluation criteria for expression co-evolution.

[0045] The beneficial effects of the present invention are:

[0046] Through the multi-agent temporal difference learning framework, the natural collaborative generation of facial micro-expressions of multiple digital humans is achieved, which improves the naturalness score of digital human expressions by 72% and increases the number of expression types to 348, greatly improving the realism and coordination of group expressions.

[0047] By applying the emotion contagion dynamics model and the SIR contagion model variant, the process of emotion propagation and diffusion in a group of digital humans is accurately simulated, enabling digital humans to naturally adjust their expressions according to social scenarios and the emotional states of surrounding digital humans, and user interaction satisfaction is improved by 63%.

[0048] Through the embodied cognitive expression co-evolution framework, the long-term adaptive evolution of digital human expressions is achieved, enabling digital humans to continuously optimize expression generation strategies based on interaction history and social relationships. The consistency of expressions in long-term interactions is improved by 58% and the degree of personalization is improved by 79%.

[0049] It breaks through the traditional expression generation method of preset templates and realizes dynamic expression generation based on reinforcement learning and emotional contagion mechanism, so that the digital human group can present more natural, coordinated and emotionally coherent expression changes in multi-person interaction scenarios. The coordination of group expressions is improved by 67%, which significantly enhances the user immersion and experience quality in the 5G interactive environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 It is a flow chart of the method for intelligently constructing the appearance features of a digital human based on big data 5G interaction of the present invention. DETAILED DESCRIPTION

[0051] The subject matter described herein will now be discussed with reference to example implementations. It should be understood that the discussion of these implementations is only to enable those skilled in the art to better understand and implement the subject matter described herein, and the functions and arrangements of the elements discussed may be changed without departing from the scope of protection of the present specification. Various examples may omit, replace, or add various processes or components as needed. In addition, the features described in some examples may also be combined in other examples.

[0052] Implementation method 1: intelligent construction method of digital human appearance features based on big data 5G interaction, such as Figure 1 As shown, the following steps are included:

[0053] Step 1: Build a multi-agent temporal difference learning framework;

[0054] This step constructs a multi-agent temporal difference learning framework for realizing the natural collaborative generation of micro-facial expressions of multiple digital humans. Through the reinforcement learning method, this framework learns the propagation and influence patterns of digital human expressions in different social scenarios, and uses the feedback information generated by real-time interaction of the 5G network as the reward signal to continuously optimize the expression generation strategy. It specifically includes the following sub-steps:

[0055] Sub-step 1.1: Construct a group expression collaboration function;

[0056] Construct a group expression collaboration function , which is used to model the mutual influence relationship between the expression states of multiple digital humans:

[0057]

[0058] Where: represents the individual expression state vector of the -th digital human, which contains a parameter set of facial features such as eyebrows, eyes, and mouth; represents the expression weight coefficient of the -th digital human, which depends on the importance of this digital human in the current social network; represents the expression interaction influence function, which is used to quantify the influence degree of the expression state of the -th digital human on the expression state of the -th digital human.

[0059] Expression interaction influence function is defined as:

[0060]

[0061] Where: is the influence coefficient, which represents the influence intensity of the -th digital human on the -th digital human; represents the difference between the expression states of two digital humans; is the distance attenuation function, represents the distance between digital humans and in the virtual space. The farther the distance, the smaller the influence.

[0062] Sub-step 1.2: Construct a multi-agent temporal difference learning algorithm;

[0063] Based on the temporal difference learning method, a Q-learning algorithm is constructed for each digital human, enabling it to optimize the expression generation strategy according to the environmental state and the reward signal feedback by the 5G network. Each digital human is a learning agent, and its Q-value update formula is:

[0064] ;

[0065] Wherein: represents the environmental state at a moment, including information such as the current user's emotion, the expression states of other digital humans, and the type of social scenario; represents the expression actions taken by the digital human at moment represents the expression actions taken by the digital human at moment i.e., generating specific facial micro-expressions; represents the reward obtained after executing the action which is converted from the emotional signal feedback by the user through the 5G network; is the learning rate, controlling the update degree of newly acquired information to the existing knowledge; is the discount factor, determining the importance of future rewards.

[0066] Sub-step 1.3: Construct an emotion feedback reward function based on the 5G network;

[0067] Construct an emotion feedback reward function, converting the emotional signal feedback by the user through the 5G network into the reward value in reinforcement learning:

[0068]

[0069] Wherein: represents the user's emotion feedback value, which is obtained by analyzing the interaction data such as the user's voice, expression, and operations in the 5G environment; represents the group expression coordination degree, evaluating the collaborative consistency of the expressions of multiple digital humans; represents the interaction activity, reflecting the interaction frequency and depth between the user and the digital human; , , are weight coefficients, used to balance the contributions of different factors to the reward.

[0070] Sub-step 1.4: Train a multi-agent collaborative learning model;

[0071] In a multi-person interaction environment, through repeated iterative training, use the collaborative learning method to optimize the expression generation strategies of multiple digital humans, realizing the natural collaborative change of expressions:

[0072] Collect the historical data of the user's interaction with the digital human in different social scenarios, including the user's emotion feedback, the expression state of the digital human, the interaction result, etc.;

[0073] Based on the collected data, initialize the Q-value table or Q-network for each digital human;

[0074] In a simulated multi-person interaction environment, let multiple digital humans perform reinforcement learning simultaneously and update their respective Q-values;

[0075] Regularly calculate the output of the group expression coordination function to evaluate the group coordination effect of the current expression generation strategy;

[0076] According to the evaluation results, adjust the learning parameters and continue to optimize the expression generation strategy until the preset performance indicators are achieved.

[0077] The specific implementation method of the multi-agent temporal difference learning framework is as follows:

[0078] Build an expression generation model based on the Deep Q-Network (DQN). The input layer receives environmental state information, including the expression states of surrounding digital humans, user feedback signals, and the current scene type; the hidden layer uses a three-layer fully connected network, with 256 neurons in each layer, and the ReLU activation function is used; the output layer corresponds to different expression action combinations, and estimates the Q value of each action.

[0079] Implement an experience replay buffer to store state transition samples (state, action, reward, next state) during the interaction of digital humans. The buffer size is set to 10,000, and 64 samples are randomly sampled in each batch during training.

[0080] Construct a target network to stabilize the Q-learning process. Update the target network parameters every 100 steps, and adopt a soft update strategy:

[0081]

[0082] where represents the network parameters.

[0083] Implement prioritized experience replay. According to the temporal difference error:

[0084]

[0085] Assign priorities to each sample, and preferentially sample samples with large TD errors to improve the learning efficiency.

[0086] Step 2: Construct an emotion contagion dynamics model;

[0087] In this step, an emotion contagion dynamics model is constructed to simulate and predict the spread and diffusion process of emotions in a digital human group, and solve the problem that the existing technology lacks the simulation of group emotion contagion phenomena. Through mathematical modeling methods, this model accurately describes the contagion mechanism of emotions in the group network and realizes the simulation of the natural diffusion process of emotions in the digital human group. Specifically, it includes the following sub-steps:

[0088] Sub-step 2.1: Construct an emotion diffusion function;

[0089] Construct an emotion diffusion function , which is used to describe the initial emotion state In the group network structure over time diffusion process:

[0090]

[0091] Where: represents the initial emotional state, including parameters such as emotional type and intensity; represents the group network structure, describing the connection relationship between digital humans; represents the local network structure connected to node ; represents the time variable of emotional diffusion; represents node at time emotional contagion function.

[0092] Emotional contagion function is defined as:

[0093]

[0094] Where: represents the emotional sensitivity coefficient of node ; represents the emotional state at time ; represents the set of neighbor nodes of node ; represents node and connection weight between; represents node at time emotional influence function.

[0095] Sub-step 2.2: Establish a social network emotional contagion mapping model;

[0096] Combining social network analysis methods, establish a mapping model of emotional contagion intensity, social relationship intensity, and individual traits to quantify the emotional contagion effect between digital humans:

[0097]

[0098] Where: represents the emotional contagion intensity between digital humans and ; represents the social relationship intensity matrix between digital humans and ; and respectively represent digital humans and The individual trait vector, including parameters such as emotional sensitivity and expression ability; Represents the digital human and The historical interaction frequency between 、 and Are weight coefficients used to balance the contributions of different factors to emotional contagion.

[0099] Sub-step 2.3: Construct the state transition matrix of emotional diffusion;

[0100] Based on the Markov process theory, construct the emotional state transition matrix to calculate the probability distribution of the emotional states of digital humans at different times:

[0101]

[0102] Where: Represents the probability distribution vector of the emotional states of the digital human group at time ; Represents the emotional state transition matrix at time , and its element Represents the probability of transitioning from the emotional state to the state ; Represents the time step.

[0103] The element calculation formula of the state transition matrix is:

[0104]

[0105] Where, Is the transition probability function, which calculates the state transition probability according to the current emotional contagion intensity , emotional state and the group network structure .

[0106] Through the above sub-steps, a mathematical model that accurately describes the dynamic process of emotional contagion is constructed. This model can effectively simulate and predict the spread and diffusion process of emotions in the digital human group, solves the problem that the existing technology lacks the simulation of group emotional contagion phenomena, and improves the realism and coherence of the emotional performance of digital humans in group scenarios.

[0107] The specific implementation method of the emotional contagion dynamics model is as follows:

[0108] Construct an emotional propagation model based on the graph neural network (GNN), representing the digital human group as a graph structure , where the node represents the digital human, and the edge Represents the social connection between digital humans. Each node contains a feature vector, representing the emotional state, individual traits, and social attributes of the digital human.

[0109] The graph convolutional network (GCN) is adopted to achieve the propagation of emotions among nodes, and its mathematical expression is:

[0110]

[0111] Where, is the adjacency matrix with self-connections added, is the corresponding degree matrix, is the node feature matrix of the th layer, is the learnable weight matrix, is the activation function.

[0112] Combined with the recurrent neural network (RNN) to capture the temporal dependence of emotion propagation. At each time step, the update formula for the emotional state of the node is:

[0113]

[0114] Where, is the hidden state of node at time , is the external input of node at time , is the edge weight from node to , is the neighbor set of node .

[0115] Step 3: Apply a variant of the SIR infection model;

[0116] In this step, a variant of the SIR infectious disease model in epidemiology is applied to simulate the digital human emotion contagion process, further enhancing the ability to accurately model the dynamics of group emotion diffusion. The SIR model classifies individuals in the group into three categories: Susceptible, Infected, and Recovered. By improving this model, accurate propagation simulation of emotions in the digital human group is achieved. Specifically, it includes the following sub-steps:

[0117] Sub-step 3.1: Construct an SIR model for emotion contagion;

[0118] The emotional states of digital humans are divided into three categories: emotional susceptible state (S), emotional infected state (I), and emotional recovered state (R), and a system of differential equations for state transitions is established:

[0119] ;

[0120] ;

[0121] ;

[0122] Wherein: 、 、 respectively represent the probabilities of the digital human being in an emotion-susceptible state, an emotion-infected state, and an emotion-recovery state; represents the emotion contagion coefficient of the digital human and reflects its ability to receive emotional influence; represents the emotion recovery coefficient of the digital human and reflects the rate at which it recovers from a certain emotional state to a neutral state; represents the social contact matrix, which describes the social contact frequency and intensity between the digital humans and .

[0123] Sub-step 3.2: Define the social contact matrix;

[0124] Construct the social contact matrix for describing the social contact pattern and intensity between digital humans:

[0125]

[0126] Wherein: represents the social relationship intensity between the digital humans and ; represents the interaction frequency factor, which reflects the interaction frequency between the digital humans and ; represents the distance between the digital humans and in the virtual space; is the distance attenuation parameter, which controls the influence degree of distance on social contact.

[0127] Sub-step 3.3: Introduce the emotion type difference coefficient;

[0128] Considering the contagion differences between different emotion types, introduce the emotion type difference coefficient :

[0129]

[0130] Wherein: represents the digital human The basic emotional contagion coefficient; and respectively represent the emotional types of digital human and ; represents the emotional type difference coefficient, which takes a value close to 1 when the emotional types are similar and close to 0 when the emotional types are very different.

[0131] The calculation formula of the emotional type difference coefficient is:

[0132]

[0133] where and are represented as vectors in the emotional space, and the difference coefficient is obtained by calculating the cosine similarity of the vectors.

[0134] Sub-step 3.4: Implement numerical simulation of emotional contagion;

[0135] Based on the constructed SIR emotional contagion model, use numerical methods to simulate the spread process of emotions in the digital human population:

[0136] Initialize the emotional states of the digital human population and set the initial emotional infection source;

[0137] Calculate the social contact matrix and the emotional type difference coefficient;

[0138] Use numerical integration methods such as Runge-Kutta to solve the SIR differential equations;

[0139] According to the numerical solution, update the emotional state probabilities of each digital human at different times;

[0140] Based on the probability distribution, generate the change sequence of the digital human facial expressions to achieve the natural spread of emotions.

[0141] Through the above sub-steps, the variant of the SIR infectious disease model is successfully applied to the simulation of the emotional contagion process in the digital human population, greatly improving the authenticity and accuracy of the simulation of the emotional spread process, making the group changes of the digital human facial expressions more in line with the emotional contagion law in the real human population, and enhancing the natural coordination of the digital human expression changes in the multi-person interaction scenario.

[0142] The specific implementation method of the SIR contagion model variant is as follows:

[0143] Construct an emotional contagion calculation framework based on multi-layer differential equations, use the fourth-order Runge-Kutta (RK4) numerical integration method to solve the SIR differential equations, and set the time step to 0.05 seconds to ensure the time continuity and stability of the simulation.

[0144] Combined with deep learning methods, train a neural differential equation model to express the change of the emotional state of the digital human as:

[0145]

[0146] where, represents the emotional state vector at time , is a vector field function parameterized as a neural network.

[0147] Implement an adaptive social contact matrix update algorithm to dynamically adjust the contact matrix according to the interaction frequency and intensity between digital humans :

[0148]

[0149] where, is the update rate, is the interaction intensity observed at time .

[0150] Construct a parameter estimation module based on variational inference to optimize the model parameters and using real-time observation data, and minimize the KL divergence between the observed data and the model prediction.

[0151] Step 4: Implement an embodied cognitive expression co-evolution framework;

[0152] In this step, an expression co-evolution framework based on the theory of embodied cognition is implemented to solve the problem that the existing technology lacks the modeling of group interaction history and social relationships, and to achieve the long-term evolutionary adaptation of digital human expressions. The theory of embodied cognition believes that the cognitive process not only depends on the brain, but is also affected by the interaction between the body and the environment. In this step, this theory is applied to the generation of digital human expressions to achieve the co-evolution of expressions and social relationship interactions. Specifically, it includes the following sub-steps:

[0153] Sub-step 4.1: Construct a social relationship expression modulation function;

[0154] Construct a social relationship expression modulation function for dynamically adjusting the expression state according to the social relationship between digital humans:

[0155]

[0156] where: and represent the expression state vectors of digital humans and respectively; represents digital human and The social relationship strength matrix between; Represents the influence factor, controlling the strength of the social relationship on the expression modulation; Represents the difference in the expression states of two digital humans.

[0157] Social relationship strength matrix The element calculation formula of is:

[0158]

[0159] Where: Represents the strength value of the th relationship type in the social relationship matrix; Represents at time Digital human And The interaction strength of the th relationship type between; Represents the time decay weight, making the influence of recent interactions on social relationships greater; Represents the total time length of historical interaction records.

[0160] Sub-step 4.2: Construct an expression evolution learning algorithm;

[0161] Based on the genetic algorithm and reinforcement learning theory, construct an expression evolution learning algorithm so that digital humans can continuously optimize the expression generation strategy through long-term interaction:

[0162] Express the expression generation strategy of each digital human as a gene encoding, including expression parameters, triggering conditions, reaction modes, etc.;

[0163] Define gene mutation and crossover operations for generating new variants of the expression generation strategy;

[0164] Design a fitness evaluation function to evaluate the quality of the expression generation strategy according to user feedback and group coordination;

[0165] Execute the selection operation, retain the expression generation strategies with high fitness, and eliminate the strategies with low fitness;

[0166] Through multiple iterations, achieve the evolutionary optimization of the expression generation strategy.

[0167] The iterative update formula of the expression evolution learning is:

[0168] ;

[0169] Where: Represents at time Digital human 's expression generation strategy; Represents the gene mutation operation; Represents the gene crossover operation; Represents the selection operation; Represents the expression generation strategy of the fitness function value.

[0170] Sub-step 4.3: Establish a social context expression memory library;

[0171] Construct a social context expression memory library to store the digital human expression states and their effect evaluations in different social contexts, which are used to guide the expression generation in future similar contexts:

[0172]

[0173] Where: Represents the social context expression memory library; Represents the th social context description, including information such as participant relationships and environmental states; Represents the expression state generated in the context ; Represents the expression state of the effect evaluation value; Represents the number of context-expression pairs stored in the memory library.

[0174] In the new social context , the most similar context and its corresponding expression can be retrieved from the memory library through similarity matching:

[0175]

[0176] Wherein, Represents the similarity function between the context and .

[0177] Sub-step 4.4: Construct a feedback learning mechanism for expression co-evolution;

[0178] Construct a feedback learning mechanism to dynamically adjust the expression generation strategy and the social relationship model through user feedback and group interaction effects:

[0179] Collect user feedback information on the digital human expression, including satisfaction scores, interaction durations, etc.;

[0180] Analyze the group expression coordination effect and evaluate the expression consistency, contagion, and naturalness;

[0181] According to the feedback information, update the social relationship strength matrix and the expression modulation parameter ;

[0182] Adjust the fitness function weights in the expression evolution learning algorithm to strengthen effective expression generation strategies.

[0183] The parameter update formula for feedback learning is:

[0184] ;

[0185] ;

[0186] Where: represents the learning rate; and respectively represent the gradients of the fitness function with respect to the parameters and respectively.

[0187] Through the above sub-steps, an expression co-evolution framework based on the theory of embodied cognition is realized. This framework can dynamically adjust the expression generation strategy according to the social relationship and interaction history among digital humans, achieve long-term evolutionary adaptation of expressions, solve the problem that the existing technology lacks modeling of group interaction history and social relationships, and improve the natural evolution ability of digital human expressions in long-term group interactions.

[0188] The specific implementation method of the embodied cognition expression co-evolution framework is as follows:

[0189] Construct an expression generation network based on a multi-layer encoder-decoder architecture, including:

[0190] Social relationship encoder: A 3-layer graph attention network (GAT) that inputs social relationship graph data and outputs relationship embedding vectors;

[0191] Interaction history encoder: A bidirectional LSTM network that processes sequential interaction data and captures long-term and short-term dependencies;

[0192] Expression state encoder: A 6-layer convolutional neural network that extracts facial key point features;

[0193] Expression generation decoder: A generative model based on a variational autoencoder (VAE) that maps the encoded features to facial expression parameters.

[0194] Implement a co-evolution-based genetic algorithm to optimize the expression generation strategy:

[0195] Chromosome encoding: The expression generation strategy of each digital human is represented as a real number vector of length 128;

[0196] Selection operation: Adopt roulette wheel selection and elitist retention strategy to retain individuals with high fitness;

[0197] Crossover operation: realize adaptive single-point crossover with a crossover probability of 0.75;

[0198] Mutation operation: Gaussian mutation is used, and the mutation intensity decreases as the number of generations increases;

[0199] Evolution parameters: The population size is set to 50 and the maximum number of iterations is 1000 generations.

[0200] Construct an update mechanism for the social relationship strength matrix based on the Bayesian inference framework:

[0201]

[0202] in, is the prior distribution of social relationship strength, is the likelihood function of the observed interaction data, is the updated posterior distribution.

[0203] Optimizing expression modulation parameters using deep reinforcement learning , using the asynchronous advantage actor-critic (A3C) algorithm, the network structure includes:

[0204] Shared feature extraction layer: 2 layers of fully connected network, 256 neurons in each layer;

[0205] Policy network: 2-layer fully connected network, output modulation parameters The probability distribution of

[0206] Value network: a 2-layer fully connected network that estimates the value function of the current state;

[0207] Step 5: Construct an expression evolution fitness evaluation system.

[0208] This step constructs an expression evolution fitness evaluation system to evaluate the naturalness, coordination, and diversity of digital human group expressions, and to provide quantitative optimization goals and evaluation criteria for expression co-evolution. The system evaluates the quality of digital human expression generation in multiple dimensions and guides the evolutionary optimization of expression generation strategies. It specifically includes the following sub-steps:

[0209] Sub-step 5.1: construct expression fitness function;

[0210] Designing a comprehensive expression fitness function , used to evaluate the overall quality of digital human group expressions:

[0211]

[0212] in: Represents a group of digital people; A naturalness scoring function representing group expressions; Function for evaluating the coordination of group expressions; Function for evaluating the diversity of group expressions; 、 and are weight coefficients used to balance the importance of different dimensions in fitness evaluation.

[0213] Sub-step 5.2: Implement the evaluation of expression naturalness;

[0214] Based on the deep learning model, implement the function for evaluating expression naturalness , and evaluate the realism and natural smoothness of the digital human's facial expression:

[0215]

[0216] Among them: represents the number of digital humans in the group; represents the digital human 's expression state; represents the natural expression distribution model learned from real human expression data; represents the expression state and the natural expression distribution model 's similarity score.

[0217] The specific calculation process of the naturalness score is as follows:

[0218] Extract the key features of the digital human's facial expression, including the shape and position parameters of the eyebrows, eyes, mouth and other regions;

[0219] Use the pre-trained neural network for evaluating expression naturalness to map the extracted features to the naturalness score;

[0220] Consider the temporal continuity and smoothness of the expression change, and perform temporal adjustment on the score;

[0221] According to the adaptability of the scene context, perform context-related weighting on the score.

[0222] Sub-step 5.3: Implement the evaluation of expression coordination;

[0223] Construct the function for evaluating expression coordination , and evaluate the coordination and consistency among the expressions of multiple digital humans in the group:

[0224]

[0225] Among them: represents the coordination difference degree between the expression states of the digital human and after considering the social relationship ;

[0226] Coordination difference degree The calculation formula is as follows:

[0227]

[0228] Where: Represents the Euclidean distance of the expression state vector; Represents the digital human and The relationship strength between; Is the relationship influence coefficient, controlling the influence of social relationships on the coordination degree evaluation.

[0229] Sub-step 5.4: Implement expression diversity scoring;

[0230] Construct an expression diversity scoring function , evaluating the richness and variability of the group expressions:

[0231] ;

[0232] Where: Represents the number of time points within the evaluation time period; Represents the moment Group Set of expression states; Represents the moment Spatial variability of the group expression state; Represents the temporal entropy of the group expression state over the entire time period, measuring the uncertainty of the expression over time; Is the temporal diversity weight coefficient.

[0233] The formula for calculating the spatial variability is:

[0234]

[0235] Where: Represents the moment Digital human Expression state; Represents the moment Average value of the group expression state.

[0236] The formula for calculating the temporal entropy is:

[0237]

[0238] Where: Represents the number of state categories after discretizing the expression state space; Represents the probability that the expression state belongs to the category ;

[0239] Through the above sub-steps, a comprehensive expression evolution fitness evaluation system is constructed. This system can quantitatively evaluate the naturalness, coordination, and diversity of the expressions of the digital population, providing clear optimization goals and evaluation criteria for expression co-evolution, enabling the long-term adaptive evolution of the expressions of the digital population to develop in a higher-quality direction.

[0240] The specific implementation method of the expression evolution fitness evaluation system is as follows:

[0241] Construct a deep learning model for expression naturalness scoring, adopting a two-stream architecture based on the EfficientNet-B3 backbone network:

[0242] Static stream: Analyze the anatomical rationality and emotional expression clarity of a single-frame expression. The input is an RGB expression image of 224×224×3.

[0243] Dynamic stream: Evaluate the continuity and smoothness of expression changes. The input is a 16-frame temporal expression sequence.

[0244] Feature fusion layer: The attention mechanism fuses the features of the two streams to generate a comprehensive score.

[0245] Model training: Use 100,000 sets of expression data annotated by humans for supervised learning, adopting the mean squared error loss function.

[0246] Implement an expression coordination scoring system based on the graph convolutional network (GCN) architecture:

[0247] Input: The expression states of the digital population and the social relationship graph.

[0248] Graph construction: Each digital person is a node, and the edge weight between nodes is the social relationship strength.

[0249] Feature extraction: 3 graph convolutional layers to extract the structured features of the group expression.

[0250] Coordination degree calculation: Global pooling followed by a fully connected layer to output the overall coordination score.

[0251] Key optimization: Introduce a social relationship-aware attention mechanism to dynamically adjust the information transmission weight between nodes according to the relationship strength.

[0252] Construct an expression diversity scoring system, combining information theory and computer vision methods:

[0253] Spatial diversity calculation: Use a variational autoencoder (VAE) to map the expression to the latent space, and calculate the determinant of the covariance matrix of the latent vectors as the spatial diversity index.

[0254] Temporal Diversity Calculation: Apply the Dynamic Time Warping (DTW) algorithm to measure the self-similarity of the facial expression sequence, and combine Empirical Mode Decomposition (EMD) to extract the spectral features of facial expression changes;

[0255] Multi-scale Integration: Analyze the diversity at different time scales through wavelet transform, from micro-expressions (millisecond level) to sustained emotional expressions (minute level).

[0256] Balance Calculation: Introduce an adaptive weight mechanism to balance the trade-off between diversity and coherence;

[0257] Develop a fitness function weight optimizer based on reinforcement learning:

[0258] State Space: Includes the current weight coefficients 、 、 and their historical effect evaluations;

[0259] Action Space: The fine-tuning range of the weight coefficients;

[0260] Reward Function: The weighted sum of user satisfaction scores and interaction durations;

[0261] Algorithm Selection: Adopt the Proximal Policy Optimization (PPO) algorithm to ensure the stability of weight adjustment.

[0262] Technical Effects of this Embodiment:

[0263] The intelligent construction method for the appearance features of digital humans based on big data 5G interaction provided by this embodiment, through the collaborative action of the multi-agent temporal difference learning framework, the emotional contagion dynamics model, the variant application of the SIR contagion model, the embodied cognitive expression co-evolution framework, and the expression evolution fitness evaluation system, has achieved the following remarkable technical effects:

[0264] Greatly improved the group facial expression realism of digital humans in a multi-person 5G interaction environment. Through the established multi-agent temporal difference learning framework and emotional contagion dynamics model, the naturalness score of digital human facial micro-expressions has increased by 72%, and the number of facial expression types has increased to 348, making the facial expression changes of digital humans more rich and natural.

[0265] Significantly enhanced the natural expressiveness of digital humans in complex social scenarios. The variant application of the SIR contagion model realizes the accurate diffusion simulation of emotions in the group, enabling digital humans to naturally adjust their facial expressions according to the social scenario and the emotional states of surrounding digital humans, and the user interaction satisfaction has increased by 63%.

[0266] Successfully achieved the ability of digital humans to naturally evolve their expressions in long-term group interactions. Through the embodied cognitive expression co-evolution framework, digital humans can continuously optimize their expression generation strategies based on interaction history and social relationships, showing adaptive changes over time. The expression coherence in long-term interactions has increased by 58%.

[0267] Effectively enhanced the authenticity of emotion contagion of 5G interactive digital humans in group scenarios. The adopted emotion contagion dynamics model and variant of the SIR contagion model enable emotions to spread among digital human groups according to the real social rules of humans, enhancing the emotional coherence and interactivity of digital human appearance features during group interactions. The group expression coordination degree has increased by 67%.

[0268] Achieved a fundamental breakthrough in the expression generation method. Different from the method of using a preset template library in conventional technologies, this solution dynamically generates micro-expressions through reinforcement learning and an emotion contagion model, significantly improving the naturalness and personalization of expressions. Each digital human can display unique and natural expression changes, and the personalization degree of digital humans has increased by 79%.

[0269] An application example of Embodiment 1 is as follows:

[0270] The intelligent construction method for the appearance features of 5G interactive digital humans based on big data in this embodiment has been applied in a large 5G virtual reality social platform, "MetaVerse Social Space". This platform allows users to connect in real time through 5G networks and conduct social interactions in the virtual world in the form of digital humans. The main challenge faced by the platform is how to enable digital human groups to display natural, coherent, and emotionally contagious facial expressions in multi-person interaction scenarios to enhance user immersion and social authenticity.

[0271] The platform includes the following main scenarios, which pose different requirements for digital human facial expression generation technologies:

[0272] Large virtual conference / speech scenario: A single speaker and multiple audiences participate together, and it is necessary to simulate the reaction of the audience group to the speaker's emotions and the spread of emotions among the audience group.

[0273] Social gathering scenario: A casual social activity with the mixed participation of multiple users and AI digital humans, which requires simulating complex group expression interactions and emotion contagion.

[0274] Long-term community interaction scenario: Users interact with AI digital human residents in a virtual community for a long time, and it is necessary for digital human expressions to evolve as social relationships develop.

[0275] Education and training scenario: A classroom environment including teacher digital humans and student digital humans, which requires expressions to convey professional attitudes and maintain an appropriate emotional atmosphere.

[0276] Before applying the intelligent construction method for the appearance features of digital humans based on big data 5G interaction provided in this embodiment, the platform faced the following problems:

[0277] The expressions of digital humans lack group synergy. When multiple digital humans appear simultaneously, the generation of their respective expressions is independent of each other, and a natural group reaction cannot be formed.

[0278] The emotional contagion effect is not natural, and the gradualness and difference of emotional diffusion in real crowds cannot be simulated.

[0279] The expressions of digital humans lack long-term adaptability and cannot adjust the expression generation strategy according to the user interaction history.

[0280] The generation of expressions depends on a limited number of preset templates and lacks natural variations and personalized features.

[0281] Implementation process example:

[0282] The platform adopts the intelligent construction method for the appearance features of digital humans based on big data 5G interaction provided in this embodiment, and the specific implementation process is as follows:

[0283] System architecture implementation:

[0284] According to the technical solution of this embodiment, the platform constructs a complete digital human expression generation system architecture, including the core components in Table 1:

[0285] Table 1: Core components of the digital human expression generation system

[0286]

[0287] The data flow between the components of the system is as follows: The user interacts with the platform through the 5G network, and the data collection module collects the user's expressions and emotional feedback in real time; the multi-agent learning engine and the emotional contagion simulator generate the expression parameters of the digital human according to the current interaction scenario; the social relationship modeler adjusts the expression generation strategy based on the historical interaction records; the expression quality evaluator scores the generated expressions to guide learning and optimization; finally, the expression renderer converts the expression parameters into visual facial expressions.

[0288] Implementation of multi-agent temporal difference learning:

[0289] In a virtual meeting scenario, the platform implements a multi-agent expression learning system based on the deep Q-network (DQN). The system trains specific expression coordination models for different types of meetings, as shown in Table 2:

[0290] Table 2: Expression coordination strategy parameters for different meeting types

[0291]

[0292] In practical applications, when a certain user A shows an excited expression in a creative brainstorming meeting, the system first invokes the multi-agent temporal difference learning framework and calculates the expression responses of other digital humans according to the expression coordination function. According to the parameters in the above table, the system applies a relatively high expression contagion coefficient (0.75) to enable the surrounding digital humans to quickly respond and show corresponding levels of excited expressions. At the same time, the system maintains an appropriate expression independence (0.45) to ensure that there are personalized differences in the expressions of different digital humans.

[0293] Implementation of emotional contagion dynamics:

[0294] In a social gathering scenario, the platform has implemented a group emotion diffusion system based on the emotional contagion dynamics model. Taking a certain virtual New Year's party as an example, when a user expresses strong surprise emotions, the system records the spread process of emotions in different social circles, as shown in Table 3:

[0295] Table 3: Measured data of emotional contagion and diffusion

[0296]

[0297] The data shows that emotions first spread rapidly in the intimate relationship circle, then gradually spread to the general relationship circle and the stranger relationship circle, and finally reach a stable state. This diffusion pattern conforms to the emotional contagion law in real social groups, greatly improving the natural coordination of group expressions.

[0298] Implementation of embodied cognitive expression co-evolution:

[0299] In a long-term community interaction scenario, the platform has implemented a digital human expression long-term adaptation system based on the embodied cognitive expression co-evolution framework. The system has tracked and recorded the interaction history between a certain AI digital human resident and different users in a virtual community and the evolution process of the expression generation strategy, as shown in Table 4:

[0300] Table 4: Evolution tracking of digital human expression generation strategy

[0301]

[0302] The data shows that as the interaction cycle increases, the digital human expression generation strategy is continuously optimized, the naturalness and diversity of expressions are significantly improved, the strength of social relationships continues to increase, and the user satisfaction score also increases accordingly. This long-term evolution ability enables digital humans to adapt to the interaction preferences of different users and establish more natural social relationships.

[0303] Verification of technical effects:

[0304] To verify the application effect of this embodiment in the "MetaVerse social space" platform, the platform carried out a series of comparative tests and user experience evaluations. Two core technical effects, namely the natural authenticity of expressions and the group emotion synergy, were verified with emphasis.

[0305] Verification of the improvement of natural authenticity of expressions:

[0306] The platform evaluated the natural authenticity of expressions from three dimensions: the richness of expression types, the naturalness of microscopic details of expressions, and the smoothness of expression time series, and compared the effects before and after applying this embodiment, as shown in Table 5:

[0307] Table 5: Comparative test results of the natural authenticity of expressions

[0308]

[0309] The test results show that after applying this embodiment, all dimensions of the natural authenticity of expressions have been significantly improved, and the comprehensive improvement rate reaches 62.1%. In particular, the improvement of the smoothness of expression time series is the most obvious, reaching 69.1%, which is mainly due to the accurate modeling of the continuous changes of expressions by the multi-agent time series differential learning framework and the emotion contagion dynamics model.

[0310] Further analyzing the expression type data, before applying this embodiment, there were only 85 preset templates for the digital human expressions on the platform; after application, the system can generate 348 different expression types, and each type has subtle variants, and the actual distinguishable expression combinations exceed 2000. This greatly enhances the naturalness and personalization of digital human expressions.

[0311] Verification of group emotion synergy and improvement of user experience:

[0312] The platform comprehensively evaluated the application effect by comparing the emotion synergy effects of groups of different scales and large-scale user satisfaction surveys, as shown in Table 6:

[0313] Table 6: Comprehensive evaluation of group emotion synergy effect and user experience

[0314]

[0315] The data shows that after applying this embodiment, the emotion synergy of groups of all scales has been significantly improved, and the improvement rate of emotion synergy in large groups even reaches 105.7%. This is mainly due to the accurate modeling of group emotion transmission and social relationships by the variant of the SIR contagion model and the embodied cognitive expression co-evolution framework.

[0316] In terms of user experience, the satisfaction with the authenticity of emotional interaction has increased most significantly, reaching 42%, fully verifying the effectiveness of this implementation method in enhancing the intelligent construction of the appearance features of 5G interactive digital humans. In the user survey, many users also specifically mentioned that the "memory ability" and "emotional coherence" of the digital human expressions have been greatly improved, which are exactly the core advantages of the embodied cognitive expression co-evolution framework in this implementation method.

[0317] In summary, through the actual application verification on the "MetaVerse social space" platform, the intelligent construction method of 5G interactive digital human appearance features provided by this implementation method has successfully solved the main problems faced by the existing technologies, significantly improved the natural authenticity of the digital human facial expressions and the group emotion coordination, greatly improved the user experience, and provided an effective solution for the development of interactive digital human technology in the 5G environment.

[0318] The above describes the embodiments of the present invention, but these embodiments are not limited to the above specific implementation manners. The above specific implementation manners are only illustrative and not restrictive. Under the inspiration of this embodiment, those of ordinary skill in the art can also make more equivalent embodiments in various forms, all of which fall within the protection scope of this embodiment.

Claims

1. A method for intelligently constructing 5G interactive digital human appearance features based on big data, characterized in that: The following steps are involved: Construct a multi-agent temporal difference learning framework, including a group expression coordination function, a multi-agent temporal difference learning algorithm, and an emotional feedback reward function based on a 5G network; the group expression coordination function is used to model the mutual influence relationship between the expression states of multiple digital humans; based on the temporal difference learning method, a Q learning algorithm is constructed for each digital human, so that it can optimize the expression generation strategy according to the environmental state and the reward signal fed back by the 5G network; the emotional feedback reward function converts the emotional signal fed back by the user through the 5G network into a reward value in reinforcement learning; Construct an emotion contagion dynamics model, including an emotion diffusion function and a social network emotion contagion mapping model, to simulate and predict the propagation and diffusion process of emotions in a digital human group; Applying the SIR contagion model variant, the emotional state of the digital human is divided into emotional susceptibility state, emotional infection state and emotional recovery state, and a group of state transition differential equations is established; The embodied cognitive expression co-evolution framework is implemented to dynamically adjust the expression generation strategy based on the social relationship expression modulation function, which is defined as: R(E i ,E j ,S ij )=E i +β mod ·S ij ·(E j -E i ); Among them, E i and E j Respectively represent the expression state vectors of digital human i and j, S ij represents the social relationship strength matrix between digital humans i and j, β mod represents the impact factor; Construct an expression evolution fitness evaluation system to quantitatively evaluate the naturalness, coordination and diversity of the expressions of digital human groups based on the expression fitness function. The expression fitness function is defined as: F adapt (G)=ω1·N(G)+ω2·C(G)+ω3·D(G); Among them, G represents the digital human group, N(G) represents the naturalness scoring function of the group expression, C(G) represents the coordination scoring function of the group expression, D(G) represents the diversity scoring function of the group expression, and ω1, ω2 and ω3 are weight coefficients.

2. The method for intelligently constructing 5G interactive digital human appearance features based on big data according to claim 1 is characterized in that: The three indicators of naturalness, coordination and diversity of the digital human group's expressions are continuously optimized and adjusted through user emotional signals fed back in real time through the 5G network.

3. The method for intelligently constructing 5G interactive digital human appearance features based on big data according to claim 2 is characterized in that: The group expression coordination function is defined as: Among them, E i represents the individual expression state vector of the i-th digital person, w i represents the expression weight coefficient of the i-th digital person, I(E i , E j ) represents the expression interaction influence function, which is used to quantify the influence of the j-th digital human’s expression state on the i-th digital human’s expression state.

4. The method for intelligently constructing 5G interactive digital human appearance features based on big data according to claim 3 is characterized in that: The emotion diffusion function is defined as: Among them, E represents the initial emotional state, G represents the group network structure, and G i represents the local network structure connected to node i, t represents the time variable of emotion diffusion, φ i (E,G i ,τ) represents the sentiment contagion function of node i at time τ.

5. The method for intelligently constructing 5G interactive digital human appearance features based on big data according to claim 4 is characterized in that: The state transition differential equations in the SIR infection model variant are defined as: Among them, S i ,I i , R i They represent the probabilities of digital person i being in the emotionally susceptible state, emotionally infected state, and emotionally recovered state, respectively. represents the emotional contagion coefficient of digital person i, represents the emotional recovery coefficient of digital person i, A ij represents the social contact matrix, I j Represents the probability that digital person j is in an emotional infection state.

6. The method for intelligently constructing 5G interactive digital human appearance features based on big data according to claim 5 is characterized in that: The emotional feedback reward function expression is as follows: Where: E user Represents the user's emotional feedback value, which is obtained by analyzing the user's voice, expression and operation data in the 5G environment; C group Indicates the coordination of group expressions and evaluates the collaborative consistency of multiple digital human expressions; V interaction Indicates the interactive activity, reflecting the frequency and depth of interaction between users and digital humans; is the weight coefficient, which is used to balance the contribution of different factors to the reward.

7. The method for intelligently constructing 5G interactive digital human appearance features based on big data according to claim 6 is characterized in that: The social network emotion contagion mapping model combines the social network analysis method to establish a mapping model between emotion contagion intensity, social relationship intensity, and individual characteristics, and quantifies the emotion contagion effect between digital humans. The expression is as follows: Where: Ψ(i, j) represents the intensity of emotional contagion between digital persons i and j; R ij represents the social relationship strength matrix between digital humans i and j; P i and P j Respectively represent the individual trait vectors of digital people i and j, including emotional sensitivity and expressive ability; H ij represents the historical interaction frequency between digital humans i and j; and is the weight coefficient, which is used to balance the contribution of different factors to emotional contagion.

8. A system for intelligently constructing digital human appearance features based on big data 5G interaction, characterized in that: It is used to execute the method for intelligently constructing the appearance features of a 5G interactive digital human based on big data as described in any one of claims 1 to 7, comprising: Multi-agent temporal difference learning framework building blocks: achieving the natural collaborative generation of facial micro-expressions of multiple digital humans; Emotional contagion dynamics model building module: used to simulate and predict the spread and diffusion of emotions in a group of digital humans; SIR contagion model variant application module: used to simulate and predict the spread and diffusion of emotions in digital human groups; Embodied cognitive expression co-evolution framework implementation module: dynamically adjust the expression generation strategy according to the social relationship and interaction history between digital humans; Expression evolution fitness evaluation module: Quantitatively evaluate the naturalness, coordination and diversity of the expressions of digital human groups, and provide optimization goals and evaluation criteria for expression co-evolution.

Citation Information

Patent Citations

  • Interactive digital human generation method and system based on artificial intelligence

    CN119600159A

  • Virtual human interaction generating device and method therof

    KR102601159B1