A method and system for generating a meta-universe digital person based on deep learning

By constructing a neural radiation field human body model and using deep learning technology, combined with sparse sensor data and physics engine optimization, personalized full-body poses are generated, solving the problems of expensive device dependence and unnatural interaction in digital human pose generation in the metaverse, and improving the user experience.

CN120852604BActive Publication Date: 2025-12-30IVIDEA CULTURAL CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511373945.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-25
Publication Date
2025-12-30
Estimated Expiration
2045-09-25

AI Technical Summary

Technical Problem

In existing technologies, the posture generation of metaverse digital humans lacks personalized features, requires expensive professional equipment, and the interaction with the virtual environment is unnatural, resulting in a reduction in user immersion and the realism of the interaction.

Method used

A neural radiation field human body model is constructed using multi-view video data to learn user behavior sequence patterns. Combined with sparse sensor data and behavioral intent recognition, deep learning technology is used to generate full-body skeletal poses. The poses are then optimized using a differentiable physics engine and reinforcement learning algorithms to meet physical constraints.

Benefits of technology

It enables the generation of high-quality, personalized full-body poses using consumer-grade devices. The interaction between the digital human and the environment is natural and in accordance with physical laws, enhancing the user's immersive experience and the realism of the interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120852604B_ABST
    Figure CN120852604B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of computer graphics and artificial intelligence, and discloses a meta-universe digital person generation method and system based on deep learning, wherein the meta-universe digital person generation method based on deep learning comprises the following steps: constructing a neural radiation field human body model through multi-view video data and learning a user behavior sequence mode; collecting sparse sensor data of a user's head and hands and generating a preliminary full-body skeleton posture; calculating contact force feedback when a digital person interacts with a virtual environment through a differentiable physical engine; optimizing the posture to meet physical constraints by using a reinforcement learning algorithm; and performing multi-objective optimization and quality evaluation on the generated posture; the application can generate a digital person posture with personalized features and capable of naturally interacting with the environment by using a consumer-level device, and improves the meta-universe interaction experience.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer graphics and artificial intelligence, more particularly, to a meta-universe digital human generation method and system based on deep learning. BACKGROUND

[0002] In the meta-universe environment, especially in social and virtual interaction scenarios, the realism of digital human posture is a key factor affecting user immersion experience. In current technology, the posture generation of meta-universe digital humans mainly relies on preset animation libraries or general motion libraries, which cannot fully represent user individual characteristics and behavior habits.

[0003] Existing high-quality posture capture usually requires expensive multi-view video equipment or professional motion capture systems, which not only have high cost but also are complex to operate, making it difficult to popularize to ordinary users. The consumer-level devices on the market (such as VR headsets and handsets) are relatively low in price, but can only provide limited tracking point data (usually only head and hand position information), making it difficult to directly generate natural and complete full-body posture.

[0004] In addition, existing technologies are insufficient in the interaction between digital humans and virtual environment objects, lacking effective force feedback mechanisms, resulting in stiff posture performance and lack of natural mechanical feedback when digital humans push open doors, pick up objects, and other interactive behaviors, greatly reducing user immersion and interactive realism.

[0005] Therefore, how to generate a digital human posture with high individuality and natural interaction with the environment only by relying on limited sensor data collected by consumer-level devices combined with deep learning technology is a technical problem to be solved in the current development of meta-universe technology. SUMMARY

[0006] The present application provides a meta-universe digital human generation method and system based on deep learning, which solves the technical problems of lack of individuality in digital human posture, unnatural interaction with the environment, and the need for expensive professional equipment in related technologies.

[0007] The present application provides a meta-universe digital human generation method based on deep learning, comprising the following steps:

[0008] Construct a neural radiance field human model through multi-view video data and learn user behavior sequence patterns to establish an individualized posture prior model;

[0009] Based on the established individualized posture prior model, collect sparse sensor data of the user's head and hands, combine behavior intention recognition, and use a motion inference network based on Transformer to generate a preliminary full-body skeletal posture;

[0010] When a digital human interacts with objects in a virtual environment, the initial full-body skeletal posture is analyzed by a differentiable physics engine to calculate contact points, contact forces, and torques, and to generate force feedback feature vectors.

[0011] Using reinforcement learning algorithms, the initial posture is optimized based on the force feedback feature vector to meet physical constraints, thus generating the final full-body posture.

[0012] The generated final full-body pose is optimized using a multi-objective loss function and its quality is evaluated to ensure the naturalness and stability of the pose.

[0013] In a preferred embodiment, the constructed neural radiation field human body model adopts a multilayer perceptron network structure, taking the position coordinates and viewpoint direction in three-dimensional space as input, and outputting the color and density of the corresponding points to form a complete 3D human body representation.

[0014] In a preferred embodiment, the learning of user behavior sequence patterns employs a long short-term memory network. The input layer receives user posture sequence data, captures the temporal dependencies of the sequence through a bidirectional LSTM layer, and the output layer generates a behavioral pattern feature representation.

[0015] In a preferred embodiment, the personalized pose prior model employs a hybrid structure of variational autoencoder and graph neural network to encode user features into latent space vectors and establish a mapping relationship between pose and user features.

[0016] In a preferred embodiment, the Transformer-based action inference network includes an encoder and a decoder. The encoder processes sparse sensor data, and the decoder combines behavioral intention features and prior pose information to generate a full-body skeletal pose and introduces an attention mechanism for body structure awareness.

[0017] In a preferred embodiment, the differentiable physics engine includes a differentiable collision detection module, a differentiable constraint solver, and a differentiable integrator, enabling the physics interaction process to be accessed and adjusted by optimization algorithms.

[0018] In a preferred embodiment, the reinforcement learning algorithm uses the superior actor critic algorithm to train the action policy network, directly integrating physical constraints into the policy gradient calculation, and optimizing the posture to meet balance, gripping adaptability, joint range and other physical constraints.

[0019] In a preferred embodiment, the multi-objective loss function includes sparse sensor data matching loss, consistency loss with prior model, physical constraint loss, and action timing smoothing loss, and the weight of each loss term can be dynamically adjusted according to the interaction scenario.

[0020] In a preferred embodiment, the quality assessment employs a multi-channel convolutional neural network structure, comprising three parallel branches: pose naturalness assessment, physical plausibility assessment, and personalized feature similarity assessment.

[0021] In a preferred embodiment, a deep learning-based metaverse digital human generation system is used to execute a deep learning-based metaverse digital human generation method, comprising:

[0022] A personalized posture data acquisition and pre-training module is used to construct a neural radiation field human body model and a personalized posture prior model.

[0023] The sparse sensor data acquisition and fusion module is used to collect data from the user's head and hands and generate a preliminary full-body posture.

[0024] The physical environment interaction force perception and modeling module is used to calculate contact force feedback characteristics through a differentiable physics engine;

[0025] The force feedback attitude dynamic optimization and generation module is based on reinforcement learning to achieve attitude optimization under physical constraints.

[0026] The pose optimization and quality assessment module is used to improve the quality of pose generation through multi-objective optimization and quality assessment.

[0027] The beneficial effects of this invention are as follows:

[0028] This invention uses only a few sensors, such as those in consumer-grade VR headsets and controllers, to capture user movements and generate a high-quality full-body motion capture system. This reduces hardware costs and broadens the technology's accessibility and application scope.

[0029] This invention accurately captures the user's unique posture features and behavioral habits through neural radiation field human body modeling and behavioral sequence pattern learning, resulting in a digital human with improved posture personalization similarity, making the digital human more recognizable and immersive.

[0030] This invention enables digital humans to interact naturally with objects in the environment through a differentiable physics engine and force feedback posture optimization. It presents postures that conform to physical laws, such as body tilting when pushing a heavy object or arm bending under force when picking up an object. This improves the realism of the interaction and enhances the user's immersive experience in the metaverse.

[0031] Balancing computational efficiency and real-time performance: This invention employs a lightweight network structure and optimized algorithms, which improves real-time processing speed on ordinary computing devices, reduces end-to-end latency control, meets the low-latency requirements of metaverse interaction, and ensures a smooth user experience.

[0032] This invention generates poses with high naturalness and coherence through multi-objective loss function optimization and quality assessment feedback, effectively reducing the unnaturalness and mechanical feel of digital human movements and improving users' acceptance and identification with digital humans.

[0033] The method of this invention can be flexibly adapted to various application scenarios in the metaverse, such as social interaction, virtual meetings, education and training, and entertainment games, and can be extended to multi-person collaborative interaction, providing key technical support for building a realistic digital human interactive experience in the metaverse. Attached Figure Description

[0034] Figure 1 This is a flowchart of a metaverse digital human generation method based on deep learning according to the present invention. Detailed Implementation

[0035] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, some features described in the examples may be combined in other examples.

[0036] The subject name is disclosed in at least one embodiment of the present invention, such as Figure 1 As shown, it includes the following steps:

[0037] Step 1: Construct a neural radiation field human body model using multi-view video data, learn user behavior sequence patterns, and establish a personalized posture prior model.

[0038] Specifically, it includes the following sub-steps:

[0039] Step 1.1, Multi-view video data acquisition and processing;

[0040] In this embodiment of the application, multiple cameras are used to simultaneously capture short video sequences of users performing various daily actions from different angles. The captured video data is processed by a video preprocessing module, including frame synchronization, image denoising, and illumination normalization, to obtain a processed multi-view video dataset. .

[0041] In some implementations, different combinations of the number and layout of cameras can be used, such as a spherical or hemispherical distribution of 4 to 16 cameras. In resource-constrained scenarios, a minimum configuration of 3 cameras can also be used, with subsequent algorithms compensating for insufficient field of view coverage.

[0042] In addition, depth cameras (such as structured light or time-of-flight sensors) can be combined with RGB cameras to acquire more accurate 3D spatial information and improve the geometric accuracy of human body models.

[0043] Step 1.2, Construction of the human body model of neural radiation field;

[0044] Based on the processed multi-view video dataset This application uses Neural Radiance Fields (NeRF) technology to construct a high-fidelity 3D human body model of the user. The model takes the position coordinates and viewpoint direction in three-dimensional space as input and outputs the color and density of the corresponding points to form a complete 3D human body representation.

[0045] The neural radiation field human model in this application employs a multi-layer perceptron (MLP) network structure, containing eight fully connected layers, each with 128 neurons, and using the ReLU activation function. The input to this model is a 5-dimensional vector. ,in, Represents position coordinates in three-dimensional space. Indicates the viewpoint direction; the output is a 4-dimensional vector. ,in Represents RGB color values. This represents the spatial density value.

[0046] To improve the expressive power of the model, this application improves the standard NeRF model by introducing human body prior constraints and hierarchical representation. Human body prior constraints guide the NeRF model to more accurately represent human body structure by using the SMPL human body parameterized model as conditional input; hierarchical representation decomposes human body appearance into three independent branch networks: geometric structure, material properties, and lighting effects, which are modeled separately and then fused together, improving the model's ability to generalize human body appearance under different lighting and viewing angles.

[0047] When applied in metaverse social scenarios, this model can reconstruct a highly realistic 3D digital human appearance based on multi-angle videos of users under different ambient lighting conditions. Even under new perspectives not present in the training data, it can generate high-quality rendering results, providing an accurate representation of the human appearance for subsequent pose generation.

[0048] Optionally, when computational resources are sufficient, this application can also employ Dynamic Neural Radiation Field (DynamicNeRF) technology, using the time dimension as an additional input to model the continuous deformation of the human body during movement. The expression can be modified as follows:

[0049] ;

[0050] in, This represents the mapping relationship of dynamic neural radiation field functions, that is, functions that map spatial location, viewpoint direction, and temporal information into color and density values. Represents position coordinates in three-dimensional space. Indicates the direction of view. Represents a time variable. Represents RGB color values. It represents spatial density values, thus enabling the capture of detailed features such as muscle deformation and changes in clothing folds.

[0051] In certain application scenarios, such as virtual try-on, this application can combine neural radiation field technology with deformable template models to achieve separate modeling of clothing and human body, which facilitates subsequent independent control and modification of the digital human's clothing style.

[0052] Step 1.3, learning user behavior sequence patterns;

[0053] By analyzing user interaction logs in the metaverse or motion capture data in reality, this application collects user behavior sequence samples and uses a Long Short-Term Memory (LSTM) network to learn user behavior sequence patterns. It captures the temporal dependencies and habitual characteristics of user behavior.

[0054] In terms of specific implementation, the behavior sequence pattern learning network in this application has a three-layer structure: an input layer, an LSTM layer, and an output layer.

[0055] The input layer receives user pose sequence data. The feature vector input at each time step contains joint angle, position and velocity information, with a dimension of 72.

[0056] The LSTM layer contains two bidirectional LSTM layers, each with 256 hidden units, which can simultaneously capture the forward and backward dependencies of a sequence.

[0057] The output layer maps the hidden states of the LSTM to the probability distribution of behavior categories and the pose prediction for the next time step.

[0058] To improve the model's ability to learn individual behavioral habits, this application adopts an attention-enhanced LSTM structure (Attention-LSTM) and introduces a temporal attention module and a joint attention module.

[0059] The time attention module calculates the importance weights of different time steps, paying particular attention to action segments with typical personal characteristics;

[0060] The joint attention module learns the importance distribution of different body joints, highlighting key joints that express individual characteristics.

[0061] In practical applications, this model can learn personalized behavioral characteristics such as walking gait features, sitting habits, and gesture expression patterns from users' daily interaction data. For example, in virtual meeting scenarios, the model can identify users' unique speaking gesture habits, or in game scenarios, it can capture users' unique running posture features, providing key personalized behavioral constraints for subsequent posture generation.

[0062] When historical user data is insufficient, this application employs a few-shot learning method, utilizing prototype-based meta-learning to quickly adapt and extract key features from limited user behavior samples. This approach is particularly suitable for the initial stage of new users joining the Metaverse platform, enabling the rapid construction of behavioral models with preliminary personalized characteristics.

[0063] Furthermore, in scenarios that require capturing more complex emotions and social behaviors, this application can also introduce a context-aware attention mechanism to integrate information about the user's social environment (such as relationships with people around them, environment type, etc.) into the behavior learning process, generating more personalized behavior patterns that are more in line with specific social scenarios.

[0064] Step 1.4, Construction of personalized NeRF pose prior model;

[0065] Based on the above data, this application combines a 3D human body model. and behavioral sequence patterns Construct a personalized NeRF pose prior model ,in, This represents a prior probability model of posture based on neural radiation field technology. Represents human posture parameters. The latent space encoding represents user features, which is a low-dimensional vector representing the compressed personalized features of the user. This prior model can capture the user's unique pose features and behavioral habits, providing personalized constraints for subsequent pose generation.

[0066] The pose prior model in this application adopts a hybrid structure that combines variational autoencoder (VAE) and graph neural network (GNN).

[0067] The encoder part uses a graph convolutional network (GCN) to process the human skeletal structure, encoding the input pose sequence into a latent space vector. ;

[0068] The decoder part is based on a conditional VAE structure, which generates the pose distribution based on the latent vector and the current context conditions.

[0069] The model expression is:

[0070] ;

[0071] in, This indicates the latent encoding of a given user feature. Generate specific poses under certain conditions The probability distribution; It is an energy function that measures the degree of matching between posture and user features. The smaller the value, the higher the degree of matching. It is implemented through a neural network. It is an exponential function that converts energy values ​​into non-negative probability values; Represents human posture parameters. The latent space encoding of user features is a low-dimensional vector that compresses and represents the personalized features of users. This is a normalization constant, ensuring that the integral of the entire probability distribution is 1.

[0072] This application innovatively improves upon traditional pose prior models by introducing multimodal conditional constraints and hierarchical decomposition representation. Multimodal conditional constraints integrate high-level semantic information such as environmental interaction goals and emotional states into the pose generation process; hierarchical decomposition representation decomposes whole-body pose into three levels: global motion, local pose, and fine motor skills, which are modeled separately and then integrated, improving the model's ability to capture individual features at different granularities.

[0073] In metaverse game scenarios, this prior model can learn a user's unique operating style and reaction patterns based on the user's past game interaction data. For example, for the same "pick up items" task, the model can generate a unique picking posture that matches the user's habits, such as the degree of bending over and the angle of arm extension, making the digital human's behavior more personalized and enhancing the user's immersion and sense of identification.

[0074] In applications that require higher precision, this application can replace the VAE structure with an energy-based generative model, which can more accurately describe the effective distribution in the attitude space by defining a more refined energy function.

[0075] In some special scenarios, such as professional sports simulation, this application can also combine domain expert knowledge to incorporate professional posture constraints of specific sports fields into the prior model, such as dance movement norms and sports technical standards, so that the generated digital human postures can maintain personal characteristics while conforming to professional technical specifications.

[0076] Step 2: Based on the established personalized posture prior model, collect sparse sensor data of the user's head and hands, combine it with behavioral intention recognition, and use a Transformer-based action inference network to generate a preliminary full-body skeletal posture.

[0077] Specifically, it includes the following sub-steps:

[0078] Step 2.1, Sparse sensor data acquisition;

[0079] When used online, only consumer-grade devices such as VR headsets and controllers are used to collect 6 degrees of freedom (6DoF) data of the user's head and hands. Including position coordinates and rotation angle Compared to traditional motion capture systems that require multiple sensors across the entire body, this method reduces hardware requirements.

[0080] In some applications that require high motion accuracy, additional tracking points can be added to the waist and feet, for example, through portable trackers or smart insoles, to collect more motion data from key areas and further improve the accuracy of posture inference.

[0081] Furthermore, in other scenarios with more limited resources, this application also supports using only a smartphone or webcam as an alternative, extracting limited pose information from 2D images through a monocular pose estimation algorithm, as... Although the accuracy of the input is reduced, it can still maintain the basic personalized features.

[0082] Step 2.2, Behavioral Intent Recognition;

[0083] Based on user behavior sequence patterns Based on the current interaction context, this application uses the Conditional Random Field (CRF) algorithm to identify the user's current behavioral intent. This provides semantic guidance for subsequent pose generation.

[0084] In scenarios with high interaction complexity, this application can adopt a hierarchical intent recognition model to decompose user behavior into task-level intent (such as "move object"), action-level intent (such as "grab"), and detail-level intent (such as "precise grab"), forming a multi-level intent structure tree to more comprehensively guide the posture generation process.

[0085] In some application scenarios, such as metaverse social networking, this application can also integrate multimodal intent recognition, combined with additional information such as the user's voice and facial expressions, to infer more complex social intents, such as expressing friendliness or doubt, making the behavior of digital humans more diverse.

[0086] Step 2.3, Action inference based on Transformer;

[0087] This application will use sparse sensor data and behavioral intention As input, a preliminary full-body skeletal pose is generated through a Transformer-based action inference network. .

[0088] This network uses a self-attention mechanism to capture the dependencies between different body parts, while also utilizing the NeRF pose prior model. As a regularization constraint, it ensures that the generated poses conform to the user's behavioral habits.

[0089] The core formula of the action inference network is:

[0090] ;

[0091] in, This represents a neural network model based on a self-attention mechanism, used to process sequential data and capture dependencies between elements; This represents sparse sensor data, including the position coordinates of the user's head and hands. and rotation angle information; Indicating behavioral intent is the user's current interaction purpose identified through a conditional random field algorithm; The NeRF posture prior model is a probabilistic model that captures a user's unique posture features and behavioral habits. This represents the initial full-body skeletal posture generated, containing the angles and positions of all joints in the complete human skeleton.

[0092] In terms of specific implementation, the action inference network of this application adopts an improved Transformer architecture, which includes two parts: an encoder and a decoder.

[0093] The encoder part contains three self-attention layers, each with eight attention heads, and has an input dimension of 256, processing sparse sensor data sequences.

[0094] The decoder consists of three cross-attention layers that fuse the encoder output with behavioral intent features and previous pose information to generate the full-body pose parameters at the current moment.

[0095] This application innovatively improves the standard Transformer model by introducing a Body-Structure-AwareAttention (BSAA) mechanism. By adjusting the attention distribution through a predefined human joint affinity matrix, the model can better capture the physical connections and interrelationships between human joints. Furthermore, a progressive refinement generation strategy is introduced, first generating a coarse pose skeleton and then gradually refining the joint angles of each part, improving the model's pose inference accuracy under limited sensor data conditions.

[0096] In practical applications, such as VR collaborative creation environments, this network can infer a complete full-body posture using only 6DoF data from the user's headset and controllers. Even when the user only makes simple hand gestures, it can accurately infer the user's habitual standing posture, arm posture, and body center of gravity shift based on learned behavioral habits. This allows other participants to see a complete and personalized digital human image, rather than a floating head and hand controllers, greatly enhancing the realism of social interaction.

[0097] In high-speed interactive scenarios requiring real-time response, this application can employ a lightweight version of the action inference network. Through knowledge distillation technology, the capabilities of the complete model are compressed into a small network, achieving pose inference with lower latency and meeting the needs of time-sensitive applications such as virtual sports competitions.

[0098] Furthermore, for different types of interaction scenarios, this application can also adopt a scenario-adaptive action inference strategy to train a dedicated action inference model for different scenarios (such as social chat, sports, fine motor operations, etc.) and dynamically switch according to the currently detected scenario type to further improve the accuracy and naturalness of pose generation.

[0099] Step 3: When the digital human interacts with objects in the virtual environment, the initial full-body skeletal posture is analyzed through a differentiable physics engine to calculate the contact points, contact forces, and torques, and to generate force feedback feature vectors.

[0100] Specifically, it includes the following sub-steps:

[0101] Step 3.1, Define the physical attributes of the virtual environment;

[0102] In this embodiment of the application, physical properties, including mass, are defined for interactive objects in a virtual environment. Stiffness coefficient of friction Parameters such as these are used to establish a database of object properties that conform to the physical laws of the real world. .

[0103] In some complex physical scenarios, this application can define objects with multi-level physical properties. It can not only define the overall physical parameters, but also specify different physical properties for different components of the object. For example, the armrests and cushions of a sofa can have different elastic coefficients, making the interactive response more detailed and realistic.

[0104] In other specific application scenarios, such as materials science education, this application can also define more specialized physical properties, such as nonlinear deformation characteristics, anisotropy, viscoelasticity and other complex physical properties, so that the interaction between digital humans and special materials can exhibit a professional level of physical realism.

[0105] Step 3.2, Construction of the differentiable physics engine;

[0106] This application introduces a differentiable physics engine module. This engine can calculate the physical interactions between objects in real time and supports gradient backpropagation, which facilitates subsequent pose optimization. Built on the principles of Lagrange mechanics, and combined with neural networks to achieve differentiability, the physical interaction process can be accessed and adjusted by optimization algorithms.

[0107] The differentiable physics engine of this application comprises three core modules: a differentiable collision detection module, a differentiable constraint solver, and a differentiable integrator.

[0108] The differentiable collision detection module represents the geometry of the object through an implicit surface function (Signed Distance Function, SDF), making the collision detection process smooth and differentiable. The differentiable constraint solver realizes the differentiability of the constraint conditions by transforming rigid constraints into soft constraint energy functions. The differentiable integrator uses numerical solutions of differential equations to ensure the differentiability of state updates during the physical simulation.

[0109] To address the non-differentiability problem in collision detection and contact handling of traditional physics engines, this application proposes a smooth collision model based on a symbolic implicit function. This model uses a neural network to approximate the distance field between objects, achieving a smooth transition in the collision response.

[0110] This model transforms the original binary (contact / non-contact) collision state into a continuous contact intensity value, calculated using the following expression:

[0111] ;

[0112] in, Represents objects and objects The contact strength value between them ranges from 0 to 1, with a larger value indicating a stronger contact. Represents objects The location or contact point; Represents objects The location or contact point; Represents objects and objects The shortest distance between them, i.e., the value of the signed distance field function; This represents the Sigmoid function, used to smoothly map distance values ​​to the range of 0-1. This is a smoothness coefficient that controls the width of the contact transition range; the larger the value, the smoother the transition. This indicates that the contact strength increases as the distance between two objects decreases.

[0113] In metaverse collaborative scenarios, this differentiable physics engine can realistically simulate the interaction between digital humans and objects in the virtual environment. For example, when a digital human pushes a virtual table, the table will move accordingly based on the magnitude and direction of the force applied. At the same time, the digital human's posture will also be adjusted based on the feedback of the contact force, presenting the feeling of the body leaning forward and applying force when pushing a heavy object.

[0114] In performance-sensitive application scenarios, this application can adopt a hybrid physics simulation strategy, using high-precision differentiable physics simulation for objects that users directly interact with, while using simplified physics models for distant or unimportant objects, thereby achieving a reasonable allocation of computing resources, ensuring the accuracy of important interactions while maintaining real-time performance.

[0115] Furthermore, in specific professional simulation scenarios, such as virtual medical training, this application can also integrate specialized biomechanical models to accurately simulate the physical properties and reactions of human tissues, enabling medical operation training to have highly realistic tactile feedback.

[0116] Step 3.3, Contact point detection and force feedback calculation;

[0117] When a digital human comes into contact with an interactive object in its environment, this application uses a collision detection algorithm to determine the precise location of the contact point. ,in, , , They represent the first , , The three-dimensional coordinates of each contact point This indicates the total number of contact points.

[0118] The contact force at each contact point is calculated using a physics engine. ,in, , , They respectively represent the effects on the first , , The three-dimensional force vectors at each contact point describe the magnitude and direction of the contact force. This indicates the total number of contact points.

[0119] Simultaneously calculate torque ,in, , , They respectively represent the effects on the first , , The three-dimensional moment vector at each contact point describes the magnitude and direction of the rotational effect generated by the contact. This indicates the total number of contact points.

[0120] For contact detection of objects with complex shapes, this application can adopt a hierarchical collision model. First, a coarse bounding volume is used for rapid screening, and then a fine mesh model is applied to the areas where collisions may occur for accurate detection, balancing computational efficiency and detection accuracy.

[0121] In some high-precision haptic feedback applications, such as virtual sculpture, this application can also simulate the microscopic friction and texture characteristics of the contact surface, calculate more detailed haptic feedback, and make the interaction between digital humans and objects present subtle differences in materials.

[0122] Step 3.4, force feedback feature encoding;

[0123] This application encodes the calculated contact point location, contact force, and torque data through a feature encoding network to generate a force feedback feature vector. This serves as the conditional input for subsequent attitude optimization.

[0124] This feature vector contains key information such as the location, direction, and intensity of the contact, which can guide the digital human to make posture adjustments that conform to the laws of physics.

[0125] The formula for a feature coding network is:

[0126] ;

[0127] in, This represents the encoder network, used to convert contact information into feature representations; This represents the generated force feedback feature vector, which contains the encoded contact mechanics information; This represents the set of contact point locations, with each element containing three-dimensional coordinates. ; This represents the set of contact forces acting at each contact point, with each element being a three-dimensional force vector describing the magnitude and direction of the force. This represents the set of torques acting at each contact point, with each element being a three-dimensional torque vector that describes the magnitude and direction of the rotational effect.

[0128] In terms of specific implementation, the force feedback feature encoding network of this application adopts a hybrid architecture of Point Cloud Transformation Network (PointNet++) and Graph Convolutional Network (GCN).

[0129] The contact points and their corresponding force and torque data are treated as three-dimensional point clouds with mechanical properties, and features are extracted using the PointNet++ network.

[0130] Construct a connection diagram between contact points and various joints of the human body, and model the relationship between force and the human skeleton using a GCN network;

[0131] The two sets of features are fused using a multilayer perceptron to generate the final force feedback feature vector.

[0132] This application proposes a Contact-Pose Coupling Representation (CPCR), which enhances the model's understanding of physical interactions by explicitly modeling the pose distribution under different contact scenarios. This representation comprises four core components: spatial distribution features of contact points, magnitude and direction features of contact forces, force mapping features of human joints, and features of the influence of object properties on pose, collectively forming a comprehensive force feedback representation.

[0133] In metaverse sports game scenarios, when a user's digital human catches a ball or pushes a virtual object, the network can accurately encode the contact mechanics features, enabling the digital human to exhibit contact responses that conform to physical laws, such as the arm sinking when catching a heavy ball and the fingertips lightly touching when catching a light ball, which greatly enhances the realism and immersion of the interaction.

[0134] For applications that require precise expression of strength, such as virtual fitness, this application can integrate a muscle activation model to map contact force feedback to the activation level of muscle groups, so that the digital human can present visual effects such as muscle tension and prominent blood vessels when exerting force, thereby enhancing the expression of strength.

[0135] In other special interactive scenarios, such as music performances, this application can also associate force feedback features with sound performance. For example, in virtual piano performances, the volume of the sound can be adjusted according to the force of the key press, and the digital human's fingers will also show the corresponding force state, achieving an immersive experience that unifies the three senses of sight, hearing and touch.

[0136] Step 4: Using reinforcement learning algorithms, the initial pose is optimized based on the force feedback feature vector to meet the physical constraints and generate the final full-body pose.

[0137] Specifically, it includes the following sub-steps:

[0138] Step 4.1, Strengthen the construction of the learning environment;

[0139] Build a reinforcement learning environment and adopt an initial stance As the starting point of the state space, the posture adjustment action is defined as the action space, and physical rationality and user feature similarity are used as reward functions to form a complete reinforcement learning framework.

[0140] In certain training scenarios, this application can construct a multi-stage progressive training environment. First, it learns basic posture adjustment strategies in a simplified physical scenario, and then gradually increases the environmental complexity and physical constraints, enabling the model to learn complex interaction skills more stably.

[0141] In other highly specialized application scenarios, such as virtual industrial operation training, this application can also introduce an imitation learning environment based on expert demonstrations, using the operation data of professionals as a reference to guide the model to learn operation postures and force control that conform to industry standards.

[0142] Step 4.2, Action Policy Network Training;

[0143] This application utilizes the Advantage Actor-Critic (A2C) algorithm to train the action policy network. The network uses sparse sensor data Behavioral patterns Contact force and torque Given the given input conditions, the output is the optimal attitude adjustment strategy.

[0144] The expression for the policy network is:

[0145] ;

[0146] in, This represents the policy function, which is used to generate the optimal attitude adjustment policy based on the input conditions. This represents the output posture parameters, including the position and rotation information of each joint of the human body; This represents sparse sensor data, which contains limited key point location information; It indicates the behavioral pattern, describing the intent and type of the current interaction; This represents the set of contact forces, which includes a three-dimensional force vector acting at each contact point; This represents the set of torques, which includes a three-dimensional torque vector acting at each contact point.

[0147] In terms of specific implementation, the action policy network of this application adopts a policy gradient reinforcement learning framework based on neural networks, which includes two parts: a policy network and a value network.

[0148] The policy network adopts a multilayer perceptron structure, containing 4 hidden layers with 256 neurons in each layer, and the output layer uses the Tanh activation function to limit the range of motion.

[0149] Value networks employ a similar structure to evaluate state value and guide policy optimization.

[0150] To improve the performance and stability of the policy network, this application proposes a hierarchical multi-task policy learning method, which decomposes the pose generation task into three sub-tasks: whole-body balance control, limb coordination control, and fine motor control, and trains them separately before integrating them. Furthermore, a memory-enhanced experience replay mechanism is introduced, paying particular attention to rare but important interaction scene samples to improve the model's performance in complex interaction scenarios.

[0151] This application also improves the A2C algorithm by proposing a physically constrained policy optimization that directly integrates physical constraints into the policy gradient calculation, as shown in the following formula:

[0152] ;

[0153] in, The parameters of the policy network are a set of trainable variables that control the behavior of the policy network. Indicates about parameters The gradient operator; This represents the objective function for policy performance. Represents the mathematical expectation; Indicates the state Select action The probability of is given by the parameter . The strategy network determines; The logarithm of the policy probability; This represents the reward value obtained after the action is performed; Representing state The estimated value function; This represents the advantage function, used to evaluate the degree of advantage of a selected action relative to the average level. Indicates the state Next action The degree to which physical constraints are violated; These are constraint weights used to balance the importance of policy optimization and physical constraints.

[0154] In the metaverse fitness coach application scenario, this policy network can generate a complete and biomechanically compliant movement sequence based on a small amount of key point location information of the user and the physical requirements of the exercise movement (such as maintaining balance and maintaining the correct joint angle). This allows the digital human to present a professional and natural posture when demonstrating fitness movements, providing users with accurate movement demonstrations.

[0155] For different types of physical interactions, this application can train a series of specialized action strategy sub-networks, such as balance strategy networks, fine manipulation strategy networks, and force control strategy networks, and then dynamically combine these sub-strategies according to the current interaction scenario to achieve more precise posture adjustment.

[0156] Furthermore, in multi-person collaborative scenarios, this application can also extend the policy network to take into account social interaction factors. For example, when virtually moving heavy objects, two digital humans can coordinate the direction and timing of their force application, exhibiting a natural collaborative posture.

[0157] Step 4.3, attitude optimization under physical constraints;

[0158] The attitude adjustment policy based on the output of the policy network is used to adjust the initial attitude. Optimize it to meet the following physical constraints:

[0159] The ground reaction force is in equilibrium with gravity;

[0160] Hand posture adapts to the shape and weight of the object being held;

[0161] There should be no unnatural interweaving of limbs;

[0162] The joint angle is within the physiologically feasible range for the human body;

[0163] The optimization process employs gradient descent, adjusting the attitude parameters by minimizing the degree of violation of physical constraints. The objective function is:

[0164] ;

[0165] in, This represents the total loss function for physical constraints, used to evaluate the physical rationality of the attitude. This represents the balance constraint loss, used to evaluate the center-of-gravity balance and stability of a digital human's pose. This represents the grasping adaptation loss, used to assess the degree of adaptation between hand posture and the shape and weight of the held object; It indicates collision avoidance and is used to assess and prevent unnatural entanglement between limbs; This indicates the joint range constraint loss, used to ensure that all joint angles are within the physiologically feasible range for the human body. , , , These represent the weighting coefficients for the balance constraint loss, grasping adaptation loss, collision avoidance loss, and joint range constraint loss, respectively. They are used to adjust the importance ratio of each loss item in the total loss and can be dynamically adjusted according to the needs of different interaction scenarios.

[0166] Step 4.4, final pose generation;

[0167] Taking into account the initial attitude Based on the optimization results of physical constraints, the final full-body pose is generated. During the generation process, the system prioritizes physical plausibility while preserving the user's personalized posture characteristics to the greatest extent possible, ensuring that the generated result conforms to physical laws and also possesses personalized features.

[0168] In applications that emphasize visual expressiveness, such as virtual performances, this application can introduce pose emphasis technology to appropriately exaggerate certain key action features, making the action performance more vivid and lifelike, while still maintaining the basis of physical rationality.

[0169] In specific professional applications, such as sports analysis, this application can also generate diverse posture variations, demonstrating different possible action schemes under the same physical conditions, assisting users in analyzing and selecting the optimal action strategy.

[0170] Step 5: Perform multi-objective loss function optimization and quality assessment on the generated final full-body pose to ensure the naturalness and stability of the pose;

[0171] Specifically, it includes the following sub-steps:

[0172] Step 5.1, Define the multi-objective loss function;

[0173] Define a multi-objective loss function for comprehensively evaluating attitude quality, which includes the following components:

[0174] Sparse sensor data matching loss : Evaluate the degree of matching between the generated posture and the actual sensor data;

[0175] NeRF Prior Consistency Loss : Evaluate the consistency between the generated pose and the user's personalized characteristics;

[0176] Physical constraint loss : Assess the physical plausibility of the posture;

[0177] Action timing smoothing loss : Evaluate the smoothness of transitions between consecutive postures.

[0178] The expression for the comprehensive loss function is:

[0179] ;

[0180] in, This represents the overall loss function, used to evaluate the overall quality of the generated pose; This represents the sparse sensor data matching loss and evaluates the degree of matching between the generated pose and the actual sensor data. This represents the consistency loss with NeRF priors, evaluating the degree of consistency between the generated pose and the user's personalized features; It represents the physical constraint loss and assesses the physical rationality of the attitude. It represents the action timing smoothness loss and evaluates the degree of smooth transition between consecutive poses; , , , These represent the weighting coefficients for sparse sensor data matching loss, NeRF prior consistency loss, physical constraint loss, and action timing smoothing loss, respectively. They are used to adjust the importance ratio of each loss item in the total loss and can be dynamically adjusted according to specific application scenarios.

[0181] In specific application scenarios, this application can introduce additional loss terms to meet special needs. For example, in virtual dance training, a rhythm consistency loss can be added. This ensures that the digital human's movements are synchronized with the music beat; in virtual speeches, expressive loss can be added. This enhances the expressiveness of body language.

[0182] Furthermore, in application scenarios with different cultural backgrounds, this application can also introduce cultural adaptation loss. This allows digital humans to conform to the behavioral norms and habits of specific cultural environments, such as differences in greeting gestures and social distances in different cultures.

[0183] Step 5.2, Real-time optimization and updates;

[0184] Based on a defined multi-objective loss function, the Adam optimization algorithm is used to optimize the generated poses in real time, maximizing pose quality while ensuring real-time performance. The optimization process employs a sliding window method, optimizing only the poses of the most recent few frames at a time to reduce computational complexity.

[0185] In environments with ample computing resources, this application can adopt a hierarchical optimization strategy. First, a rapid coarse optimization is performed to meet real-time requirements, and then a more refined optimization is performed during idle computing time to continuously improve pose quality, similar to the application of LOD (Level of Detail) technology in games to animation optimization.

[0186] In highly dynamic scenarios, such as virtual sports competitions, this application can also employ predictive optimization methods to pre-calculate possible posture changes based on current movement trends, perform optimization calculations in advance, and reduce real-time response latency.

[0187] Step 5.3, Adaptive weight adjustment;

[0188] This application dynamically adjusts the weights of each term in the loss function based on different interaction scenarios and user preferences. For example, it increases the weight of physical constraint loss in fine-grained operation scenarios and increases the weight of smoothing loss in fast-moving scenarios, thereby achieving scene-adaptive pose generation.

[0189] This application can introduce a weight learning mechanism based on user feedback. By recording users' preferences for different generated results, it can automatically learn the weight configuration that best suits the aesthetics and needs of specific users. For example, some users may value the smoothness of movements, while others may value the degree to which personal characteristics are preserved.

[0190] In long-term use scenarios, this application can also achieve gradual weight adaptation. As the user's usage time increases, the system gradually adjusts the weight configuration to better match the user's unique habits and preferences, providing a personalized experience.

[0191] Step 5.4, Quality Assessment and Feedback;

[0192] A posture quality assessment model is established to comprehensively evaluate the generated postures from aspects such as naturalness, stability, and the degree of preservation of personalized features. The assessment results are then fed back to the aforementioned steps to form a closed-loop optimization mechanism, thereby continuously improving the quality of posture generation.

[0193] In terms of specific implementation, the posture quality assessment model of this application adopts a multi-channel convolutional neural network structure, which includes three parallel branches: posture naturalness assessment channel, physical rationality assessment channel, and personalized feature similarity assessment channel.

[0194] The posture naturalness assessment channel is based on pre-trained human motion priors to determine whether the generated posture conforms to the laws of human motion.

[0195] The physical rationality assessment channel is based on physical simulation calculations to determine attitude stability and balance;

[0196] The personalized feature similarity assessment channel evaluates the degree of personal feature retention by comparing it with the user's historical posture data.

[0197] To improve the accuracy and generalization ability of the evaluation model, this application proposes a quality evaluation framework based on adversarial learning, similar to the discriminator in generative adversarial networks. Through adversarial training with a large amount of real human motion data and synthetic data, the evaluation model can accurately identify unnatural or physically inconsistent postures.

[0198] The evaluation results are represented as a quality score vector. ,in, Represents the overall quality assessment vector. , , They represent the first , , A quality score, This represents the total number of quality scores. Each score value ranges from 0 to 1, where 1 represents the highest quality and 0 represents the lowest quality.

[0199] This application also innovatively designs an adaptive feedback mechanism based on the evaluation results, dynamically adjusting the parameters and weights in the aforementioned pose generation and optimization steps according to the scores of different quality dimensions. For example, when the physical rationality score is found to be low, the weight of the physical constraint loss is automatically increased; when the personalized feature similarity score is insufficient, the regularization effect of the NeRF prior model is enhanced.

[0200] This closed-loop feedback mechanism can be represented as:

[0201] ;

[0202] in, Indicates time step Time The weight of the loss term, i.e., in the optimization process... In the nth iteration cycle, corresponding to the nth Weighting coefficients of each component of the loss function; Indicates the next time step Time The updated weights of each loss term, i.e., the adjusted weight values ​​used in the next iteration; This indicates that the quality assessment model is for the first... The evaluation score for each quality dimension is used to quantify the performance of the current pose in that dimension. For quality score The adjustment function dynamically calculates the adjustment coefficient of the weights based on the quality assessment results. When the quality score is low, the adjustment range is increased to strengthen the impact of the corresponding loss item, while the weights remain relatively stable when the quality score is high.

[0203] In metaverse speech scenarios, this evaluation model can monitor the posture quality of the speaker's digital human in real time, ensuring that the generated gestures, postures, and body movements are both consistent with the user's personal habits and remain natural and fluid, avoiding stiff or uncoordinated movements, and enhancing the persuasiveness and naturalness of the speech. When unnatural postures are detected, the system will immediately make adjustments to ensure that the digital human's performance always maintains a high level of quality.

[0204] In scenarios that support multi-user interaction, this application can introduce a social adaptability assessment module to evaluate the appropriateness of digital human postures in social interactions, such as maintaining an appropriate social distance, facing the conversation partner, and making responsive postures that match the content of the conversation, thereby improving the naturalness of multi-user interaction.

[0205] In demanding professional applications, such as virtual anchors, this application can also integrate professional performance evaluation standards to ensure that the digital human's movements conform to professional standards, such as upright posture, appropriate gestures, and matching facial expressions and tone of voice, thus meeting the quality requirements of professional occasions.

[0206] Implementation process example:

[0207] During the implementation of this collaborative design platform, a system deployment was carried out for a product design team consisting of 10 designers. Each designer was equipped with a standard VR headset and hand controllers, and the following implementation steps were performed:

[0208] Data collection and personalized modeling phase: Initial data collection was conducted for each designer for 15 minutes, using 8 surround cameras to record the designer performing a series of standard and typical work actions.

[0209] The collected data includes approximately 9,000 frames of multi-view video at a resolution of 1920×1080 and a frame rate of 30fps. In addition, interaction logs of each designer in the design software over the past 3 months were collected, including approximately 20 hours of operational data.

[0210] Based on the collected multi-view video data, the system constructed a neural radiation field human body model for each designer. The model contains an 8-layer fully connected network with a total of approximately 1,200,000 neurons, which can accurately express the designer's body shape and appearance features.

[0211] Meanwhile, by analyzing interaction behavior logs, a bidirectional LSTM network (2 layers, 256 hidden units per layer) is used to extract the behavioral sequence patterns of each designer, with particular attention paid to unique habits in various design operations, such as grabbing methods, operation rhythm, and posture preferences.

[0212] Real-time pose generation phase: In actual collaborative design sessions, the system only collects 6DoF data (90 samples per second) from the designer's head-mounted display and controllers, including position coordinates and rotation angles.

[0213] Based on this sparse data and combined with pre-learned behavioral patterns, the designer's full-body posture is inferred in real time through a 3-layer Transformer network (8 attention heads in each layer).

[0214] When designers need to manipulate virtual product prototypes, the system's differentiable physics engine accurately calculates the contact points and interaction forces between the hand and the product.

[0215] For example, when a designer rotates a product screw, the system can not only calculate the resistance of the screw and the contact point of the fingers, but also generate a force feedback feature vector through a force feedback feature encoding network. This vector guides the motion strategy network to adjust the designer's full-body posture so that it presents a force-exerting posture that conforms to the laws of physics, such as changes in wrist angle and slight forward leaning of the body.

[0216] Optimization and evaluation phase: The system uses a multi-objective loss function to optimize the generated pose in real time, ensuring physical rationality while maintaining the designer's personal characteristics.

[0217] In scenarios with intensive interaction (such as product assembly), the weight of physical constraint loss will be automatically increased to ensure that the posture can adapt to the needs of fine operation; while in the demonstration and explanation stage, the weight of expressiveness and personality characteristics will be increased accordingly.

[0218] The multi-channel quality assessment network continuously monitors the posture generation quality, evaluating it 30 times per second. When an unnatural posture is detected (such as an arm crossing the table or abnormal body balance), the posture optimization process is immediately triggered, and the optimization results are fed back to the posture generation module to form a closed-loop optimization.

[0219] Technical effectiveness verification:

[0220] To verify the technical effectiveness of this implementation method in practical applications, a systematic evaluation of the collaborative design platform was conducted, focusing on verifying the two core technical effects of "preservation of personalized features" and "realism of environmental interaction".

[0221] Verification of Personalized Feature Retention: For 10 designers, digital human poses were generated using both this implementation method and the control group method (traditional method based on skeletal mapping). The degree of personalized feature retention was then evaluated through two types of tests: first, an objective indicator test, which calculated the similarity of key features between the generated pose and the designer's real pose; second, a subjective recognition test, where colleagues of the designers were invited to attempt to identify the real designer corresponding to the digital human. The comparison results of personalized feature retention are shown in Table 1.

[0222] Table 1: Comparison of Personalization Feature Retention Results;

[0223]

[0224] Test results show that this implementation method significantly outperforms traditional methods in preserving user personalization characteristics, improving by nearly 40 percentage points. Particularly in preserving action habits, such as designers' unique operating rhythms, gesture habits, and posture preferences, this method demonstrates a clear advantage.

[0225] Colleague recognition tests further validated the digital human's ability to retain individual characteristics, with most colleagues able to correctly identify the corresponding designer simply by observing the digital human's movements.

[0226] Verification of Realism in Environmental Interaction: To evaluate the realism of the interaction between the digital human and the virtual environment, a series of test scenarios with varying levels of difficulty were designed, including fine motor skills (such as tightening screws), strength-based skills (such as pushing heavy objects), and complex assembly tasks. Evaluation was conducted through both professional evaluators' ratings and the designer's own perceptual assessment. The comparison results of the realism in environmental interaction are shown in Table 2.

[0227] Table 2: Comparison Results of Realistic Environmental Interaction Effects;

[0228]

[0229] Test results show that this implementation method achieved significant results in terms of the realism of environmental interaction, with an average improvement of over 50 percentage points. Particularly in terms of force feedback, the digital human can naturally adjust its posture according to the different physical properties of objects, such as the body tilt angle when pushing heavy objects and the difference in force when lifting light and heavy objects. User immersion evaluations reflect the designer's level of identification with their digital human image, indicating that this method has achieved good results in realistically reproducing user intentions.

[0230] Comprehensive verification results demonstrate that this implementation method successfully achieved the expected technical effects in the application of the Metaverse Collaborative Design Platform. It not only preserved the user's personalized characteristics but also significantly enhanced the realism of environmental interaction, providing designers with a highly natural and immersive collaborative experience. These technical advantages directly translate into practical application value: collaborative design efficiency increased by 35%, solution communication costs decreased by 40%, and user satisfaction increased by 42%.

[0231] The embodiments of the present invention have been described above. However, the embodiments are not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make more equivalent embodiments under the guidance of the present embodiments, and all of them are within the protection scope of the present embodiments.

Claims

1. A deep learning-based meta-universe digital human generation method, characterized in that, The method comprises the following steps: A neural radiance field human model is constructed from multi-view video data, and user behavior sequence patterns are learned to establish a personalized pose prior model; Based on the established personalized pose prior model, sparse sensor data of the user's head and hands are collected, combined with behavior intention recognition, and a preliminary full-body skeletal pose is generated using a motion inference network based on Transformer; When the digital human interacts with objects in a virtual environment, the preliminary full-body skeletal pose is analyzed by a differentiable physics engine to calculate contact points, contact forces, and moments, generating a force feedback feature vector; Using a reinforcement learning algorithm, the preliminary pose is optimized according to the force feedback feature vector to meet the physical constraint conditions, generating the final full-body pose; An action policy network is trained using the advantage actor-critic algorithm, which takes sparse sensor data, behavior patterns, contact forces, and moments as conditional inputs and outputs the optimal pose adjustment strategy; The action policy network uses a neural network-based policy gradient reinforcement learning framework, which includes a policy network and a value network; The policy network uses a multi-layer perceptron structure, which includes four hidden layers with 256 neurons each, and the output layer uses a Tanh activation function to limit the action amplitude; The value network uses the same structure as the policy network to evaluate the state value and guide the optimization direction of the policy; The final full-body pose is optimized by a multi-objective loss function and quality evaluation to ensure the naturalness and stability of the pose.

2. The method of claim 1, wherein the method is based on deep learning. The constructed neural radiance field human model uses a multi-layer perceptron network structure, taking position coordinates and viewing angles in three-dimensional space as input, and outputting the color and density of the corresponding points to form a complete 3D human representation.

3. The method of claim 1, wherein the method further comprises: The learning of user behavior sequence patterns uses a long short-term memory network, which receives user pose sequence data at the input layer, captures the temporal dependence of the sequence through a bidirectional LSTM layer, and generates a behavior pattern feature representation at the output layer.

4. The method of claim 1, wherein the method is based on deep learning. The personalized pose prior model uses a hybrid structure of variational autoencoder and graph neural network to encode user features into latent space vectors and establish a mapping relationship between poses and user features.

5. The method of claim 1, wherein the method further comprises: The action inference network based on Transformer includes an encoder and a decoder. The encoder processes sparse sensor data, and the decoder generates full-body skeletal poses combined with behavior intention features and prior pose information, and introduces a body structure perception attention mechanism.

6. The method of claim 1, wherein the method is based on deep learning. The differentiable physics engine includes a differentiable collision detection module, a differentiable constraint solver, and a differentiable integrator, making the physical interaction process accessible and adjustable by optimization algorithms.

7. The method of claim 1, wherein the method is based on deep learning. The reinforcement learning algorithm uses the advantage actor-critic algorithm to train the action policy network, directly integrates physical constraints into policy gradient calculation, and optimizes the pose to meet balance, grasping adaptability, joint range, and other physical constraints.

8. The method of claim 1, wherein the method is based on deep learning. The multi-objective loss function includes sparse sensor data matching loss, consistency loss with the prior model, physical constraint loss, and action temporal smoothing loss, and the weights of each loss term can be dynamically adjusted according to the interaction scenario.

9. The method of claim 1, wherein the method is based on deep learning. The quality evaluation adopts a multi-channel convolutional neural network structure, and includes three parallel branches of posture naturalness evaluation, physical rationality evaluation, and individualized feature similarity evaluation.

10. A deep learning-based meta-universe digital human generation system for performing the deep learning-based meta-universe digital human generation method of any one of claims 1-9. The method comprises the following steps: An individualized posture data acquisition and pre-training module is used to construct a neural radiance field human model and an individualized posture prior model; A sparse sensor data acquisition and fusion module is used to acquire user head and hand data and generate a preliminary full-body posture; A physical environment interaction force perception and modeling module is used to calculate contact force feedback features through a differentiable physics engine; A force feedback posture dynamic optimization and generation module is used to realize posture optimization under physical constraints based on reinforcement learning; A posture optimization and quality evaluation module is used to improve posture generation quality through multi-objective optimization and quality evaluation.

Citation Information

Patent Citations

  • Human body model establishing method, system and equipment and storage medium

    CN114202629A

  • Interactive indirect reasoning twin posture detection method and system

    CN114821006A