Overall human body action generation method oriented to control perception
By constructing a conditional variational autoencoder and sequence diffusion model, combining the object manipulation characteristics, high-quality and coordinated overall human movements are generated, which solves the problem of missing information generated by the action under sparse tracking signals of VR equipment, and improves the immersion and interaction quality in virtual reality.
Patent Information
- Application Number
- CN202510868369.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-26
AI Technical Summary
Existing VR devices cannot fully track the user's whole body movements, resulting in missing information and increasing the complexity of action generation in the overall human body movement generation. The overall movement quality generated is insufficient, affecting the immersion and interaction quality.
By constructing a conditional variational autoencoder, a sparse tracking signal is potentially coded to represent the sparse tracking signal, combining the sequence diffusion model and the sequence control network, the object control features are used to generate the overall human movement of the fusion manipulation intention, and the action encoding and decoding process is optimized.
It significantly improves the overall action nature and coordination of virtual characters in the VR environment, improves the action quality and interaction consistency, and enhances the generation speed and scene semantic consistency.
Smart Images

Figure CN120374814A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of virtual reality, and particularly relates to a method for generating overall human body movements oriented to manipulation perception. Background Art
[0002] As an important bridge connecting the real world and the digital world, virtual reality (VR) technology is continuously driving the development of the human-computer interaction experience towards a more natural and immersive direction. Generating realistic overall human body movements, especially the generation technology that simultaneously includes body and hand movements, is the key to realizing the simulation of real user behaviors in a virtual environment. However, despite the rapid progress of VR technology in aspects such as visual display and spatial positioning, the motion control signals provided by current mainstream VR devices (such as Meta Quest Pro, Apple Vision Pro, and PICO 4 Pro) are still very limited, unable to comprehensively track the user's entire body, resulting in a serious problem of information loss in the generation of overall human body movements.
[0003] Meanwhile, the different behaviors shown by users when manipulating objects in a VR scenario will significantly affect the overall human body movements, further increasing the complexity of motion modeling, making it a challenging task to generate high-quality overall human body movements based on sparse tracking signals. Currently, researchers mainly rely on various data-driven models to solve this problem, ranging from regression methods to probabilistic generation models. Although these methods have proposed various feasible optimal motion estimation strategies, they ignore the constraints on the distribution of the potential motion space under the conditions of joint sparse motion control and manipulation content, resulting in an overly broad estimated range of the generated overall motion, and thus making it difficult to accurately and efficiently estimate the body and hand movements simultaneously during manipulation. Problems such as insufficient rationality of body postures and low frame rates of overall motion generation have, to a certain extent, restricted the immersion and interaction quality of users in a virtual environment. Summary of the Invention
[0004] To achieve the above object, the technical solution of the present invention is as follows: A method for generating overall human body movements oriented to manipulation perception, comprising the following steps:
[0005] Step 1, perform potential coding representation on the human body and hand movements encoded from sparse tracking signals by constructing a conditional variational autoencoder;
[0006] Step 2, gradually reconstruct the noisy motion encoding in combination with a sequence diffusion model to obtain the initial body and hand movement potential encoding;
[0007] Step 3: Use the sequence control network to extract object manipulation features, guide the DDPM to generate action latent encodings that fuse manipulation intentions, and finally generate the final overall human action sequence that fuses manipulation semantics through the decoding module of the CVAE.
[0008] Preferably, in Step 1, a conditional variational autoencoder is constructed, including: a body action encoding module, a hand action encoding module, and corresponding action decoding modules. The body and hand action encoding modules encode human body and hand actions through sparse tracking signals and output body and hand action encodings in the latent action space respectively.
[0009] Preferably, in Step 2, a sequence diffusion model is constructed, including: a body DDPM and a hand DDPM. Given the sparse tracking signal and the noisy body latent encoding as inputs, the body DDPM reconstructs the initial body action latent encoding. Given this initial body action latent encoding, the sparse tracking signal, and the noisy hand latent encoding, the hand DDPM reconstructs the initial hand action latent encoding.
[0010] Preferably, in Step 3, a sequence control network is constructed, including a body control module and a hand control module. Given the object manipulation representation information, the body control module extracts manipulation features and inputs them to the body DDPM; the output body latent encoding and the initial body action latent encoding are input into a linear layer together to obtain the body action latent encoding that fuses manipulation information; similarly, the hand control module extracts manipulation features, inputs the manipulation features and the initial hand action latent encoding into the hand control module together, and generates the hand action latent encoding that fuses manipulation information; finally, the optimized body and hand action encodings are input into the corresponding action decoding modules constructed in Step 1 to generate the final overall action that fuses manipulation intentions.
[0011] Preferably, Step 3 includes object manipulation features to guide the overall action encoding of manipulation perception, including:
[0012] The user actively performs manipulation behaviors through the controller buttons of the VR device. Once confirmed, the body and hand actions will be further refined;
[0013] Use the standard SMPL-X human skeleton model to represent the body and hands, which consists of 10 real values. For joint j in all joints A of SMPL-X, its local rotation is defined in the set function at the time stamp where A(j) represents the ordered set of the ancestor joints of joint j, and the global rotation of this joint is calculated by successively multiplying the local rotations of all its ancestor joints i from the current joint j to the root joint, and is represented by the following formula:
[0014] (1),
[0015] Sample the vertices of the object in the left - hand and right - hand coordinate systems respectively, rather than directly using all vertices or uniform sampling. For any vertex on the interaction object , its projection on the unit sphere is:
[0016] (2),
[0017] By projecting all vertices in the interaction object onto the surface of the unit sphere to achieve normalization processing, where is the set of all vertices of the interaction object , represents the \(i\) - th point in \(V\), represents the average value of the vertices in \(V\);
[0018] Then, uniformly sample vertices on the unit sphere and convert each sampled point to the left / right - hand coordinate system, which are represented as the polar coordinate forms in the left - hand system and the right - hand system respectively ;
[0019] Define the distance feature vector , which provides continuous information about the proximity between the hand and the surface of the interaction object at time stamp \(t\). At each time stamp \(t\), uniformly sample points on the left and right hands respectively, and calculate the minimum distance between each sampled point and the nearest point on the surface of the interaction object, and construct the distance feature vector .
[0020] Preferably, after constructing the effective manipulation features, use the manipulation features to generate the overall action with manipulation intention through progressive manipulation guidance. The optimization process includes:
[0021] 1) Potential motion learning optimization, the object of optimization is the conditional variational auto - encoder obtained in the first - stage training. Among them, the action encodings of the human body and the hand are embedded into the latent space of the CVAE and used as an important reference for generating the overall action encoding in steps two and three. In addition, the body and hand action decoders in the CVAE rely on the overall action encoding information to reconstruct the overall action sequence;
[0022] 2) Train the sequence diffusion model in step two so that it can learn complex body postures and hand details;
[0023] 3) Use the manipulation representation to train the sequence control network to optimize the overall action encoding.
[0024] Preferably, in (1), the object of optimization is the conditional variational autoencoder obtained by training in the first step. Among them, the action encodings of the human body and the hand are embedded into the latent space of the CVAE and used as an important reference for generating the overall action encoding in the second and third steps. In addition, the body and hand action decoders in the CVAE rely on these encoding information to reconstruct the overall action sequence.
[0025] The loss function of the CVAE is:
[0026] (3),
[0027] The first term is the KL divergence loss , which is used to constrain the latent space to approach the standard normal distribution, where represents the weight of the KL divergence loss, represents the latent distribution predicted by the encoder and the standard normal distribution between the KL divergence;
[0028] The second term is the joint rotation reconstruction loss , which is used to constrain the predicted joint rotation to approach the true joint rotation, where represents the weight of this loss, represents the smooth-L1 loss function, and are the predicted and true joint rotation values respectively;
[0029] The third term is the joint position loss calculated based on forward kinematics , which is used to constrain the predicted joint position to approach the true joint position, where represents the weight of this loss, and are and joint positions obtained by forward kinematics calculation.
[0030] Preferably, in (2), the sequence diffusion model in the second step is trained so that it can learn complex body postures and hand details. Under the condition that the input is the noisy body action latent encoding and the sparse tracking signal, the loss function of the body diffusion model is:
[0031] (4),
[0032] where and represent the predicted and true body action latent encodings respectively;
[0033] Next, fix the parameters in the Body DDPM and train the hand diffusion model. Its loss function is as follows:
[0034] (5),
[0035] where and represent the predicted and true latent encodings of hand actions respectively.
[0036] Preferably, in 3), the manipulation representation is used to train the sequence control network to optimize the overall action encoding. After fixing the parameters of the body diffusion model and the hand diffusion model, the body control module is optimized through the following formula:
[0037] (6),
[0038] This loss function includes the rotation loss of body joints and the position loss of body joints Two terms, where is the weight of this position loss, where and represent the latent encodings of body actions of predicted and true fused manipulation information respectively, and are the joint positions calculated based on and ;
[0039] The loss function of the hand control module is as follows:
[0040] (7),
[0041] where, and represent the latent encodings of hand actions of predicted and true fused manipulation information respectively, and are the hand joint positions calculated based on and . In addition to the hand joint rotation loss and the hand joint position loss , it also includes the displacement loss of hand key vertices, where represents the weight of this displacement loss, is the hand vertex position sampled in the manipulation representation, is the vertex position predicted by this method for the output hand action.
[0042] Preferably, by training a lightweight online network to fit the sequence diffusion model constructed in Steps 2 and 3 and guided by the control network, the model can perform inference in fewer steps and predict clear latent action representations.
[0043] In Steps 2 and 3, for each DDPM, a 24-layer DiT denoiser is trained as the teacher model and distilled into a target 6-layer denoiser model. Since the teacher model and the target model share the network architecture and feature dimensions, the target denoiser not only learns the final output of the teacher model but also the outputs of some of its intermediate layers. In the diffusion inference process of Steps 2 and 3, initially 5 denoising steps are adopted; subsequently, the denoising steps are reduced through the distillation process, and finally only one denoising step is required to complete the inference process.
[0044] Compared with the prior art, the beneficial effects of the present invention are as follows: By introducing a specific object operation representation, the present invention significantly improves the naturalness and coordination of the overall actions generated by virtual characters in the VR environment, demonstrates higher action quality and interaction consistency during object interaction, greatly improves the generation speed while ensuring action quality, and lays a solid foundation for generating more complex scene-semantic consistent interaction actions. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 is the execution flow of the overall human action generation system of the present invention;
[0046] Figure 2 is an instantaneous display diagram of the overall actions when virtual humans of different body types in the present invention manipulate the camera;
[0047] Figure 3 is the distillation flow chart of the present invention;
[0048] Figure 4 is the object manipulation feature display diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] The following further clarifies the present invention in conjunction with the drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.
[0050] Embodiment: This embodiment proposes a method for generating overall human actions oriented to manipulation perception. By constructing a conditional variational autoencoder (CVAE), latent coding representations of human body and hand actions encoded from sparse tracking signals are obtained, and then combined with a sequence diffusion model (DDPM) to gradually reconstruct the noisy action encoding, obtaining initial latent encodings of the body and hand actions;
[0051] Furthermore, the sequence control network is used to extract object manipulation features, guiding the DDPM to generate action latent codes that fuse manipulation intentions. Finally, the decoding module of the CVAE is used to generate the final overall human action sequence that fuses manipulation semantics. This method can achieve overall human action generation for manipulation perception under the input conditions of sparse tracking signals and object manipulation representation information, improving the physical rationality and temporal consistency of the generated actions in the virtual reality environment.
[0052] As Figure 1 shown, the processing flow of this method mainly includes three steps:
[0053] Step 1: Construct a conditional variational autoencoder (CVAE), including: a body action encoding module, a hand action encoding module, and corresponding action decoding modules; the body and hand action encoding modules are designed to accurately encode human and hand actions using sparse tracking signals and output the body and hand action codes in the latent action space respectively; the body and hand action decoding modules are responsible for generating accurate human and hand actions based on these latent codes.
[0054] Step 2: Construct a sequence diffusion model, including: a body denoising diffusion probability model (DDPM) and a hand DDPM; given the sparse tracking signal and the noisy body latent code as inputs, the body DDPM reconstructs the initial body action latent code; given this initial body action latent code, the sparse tracking signal, and the noisy hand latent code, the hand DDPM reconstructs the initial hand action latent code.
[0055] Step 3: Construct a sequence control network, including a body control module and a hand control module; given the object manipulation representation information, the body control module extracts manipulation features and inputs them to the body DDPM; the output body latent code and the initial body action latent code are input into a linear layer together to obtain the body action latent code that fuses manipulation information; similarly, the hand control module extracts manipulation features and inputs the manipulation features and the initial hand action latent code into the hand control module together to generate the hand action latent code that fuses manipulation information; finally, the optimized body and hand action codes are input into the corresponding action decoding modules constructed in Step 1 to generate the final overall action that fuses manipulation intentions.
[0056] According to the above three steps, the final overall action that fuses manipulation intentions can be generated through the sparse tracking signal of the VR device and the object manipulation representation information in the virtual scene. To better describe the process of this method, the following definitions used in the article are given:
[0057] Conditional Variational Autoencoder (CVAE): The Variational Autoencoder (VAE) is a generative model that combines an autoencoder with a probabilistic graphical model. It maps the input to a normal distribution in the latent space through an encoder, samples from it, and then reconstructs the data through a decoder, optimizing by maximizing the Evidence Lower Bound (ELBO). CVAE introduces conditional information on the basis of VAE, making the generation process controllable and suitable for tasks that require specific conditional guidance.
[0058] Sparse Tracking Signals: Refers to the limited tracking signals provided by the headset and controllers in VR devices. Compared with full-body joint sensors, it only captures the poses of the head and hands.
[0059] Encoder / Decoder: The encoder compresses high-dimensional input into a low-dimensional latent representation, extracting key features; the decoder reconstructs the latent representation into the original data space, and the two constitute the autoencoder framework.
[0060] Latent Code: In an autoencoder or generative model, it is the low-dimensional vector output by the encoder, which implicitly contains the key features and structure of the data and serves as an intermediate representation in the generation process. Its distribution characteristics affect the interpolation ability and generation quality of the model and are the core parameters for controlling the generated content.
[0061] Diffusion Probability Model (DDPM): A generative model based on Markov chains. It gradually adds noise to the data through the forward process and then trains the reverse process to learn to denoise step by step to recover the data distribution. DDPM generates high-quality samples by optimizing the noise prediction task and has become one of the mainstream generative methods due to its stable training and excellent results.
[0062] Model Distillation: A technique for optimizing model deployment through knowledge transfer. Its core is to transfer the knowledge of a complex Teacher Model to a lightweight Student Model. During the training process, the Student Model minimizes the output difference with the Teacher Model and the task loss function to achieve model compression and performance approximation while significantly reducing the computational complexity.
[0063] Object Manipulation Representation:
[0064] The goal of this method is to use the sparse signals provided by VR devices to generate overall human actions that conform to manipulation perception. To achieve this goal, first, an efficient method that can quickly extract and accurately represent manipulation content is needed. Therefore, this method proposes an effective object manipulation feature to guide the overall action encoding of manipulation perception, such as Figure 4 shown, consisting of the following:
[0065] Manipulation state label at timestamp t and the shape features of the avatar model SMPL-X , the characteristics of the manipulated object and the manipulation action features .
[0066] Manipulation state label:
[0067] is a binary parameter, and the user can only be in the manipulation or non-manipulation state at timestamp t. This method allows the user to actively perform manipulation behaviors through the controller buttons of the VR device. Once confirmed, the above three other features will be used to further refine the body and hand movements, making the generated movements more natural and realistic during the manipulation process.
[0068] Shape features:
[0069] This method uses the standard SMPL-X human skeleton model to represent the body and hands, where is the body shape parameter of SMPL-X, consisting of 10 real values. For joint j in all joints A of SMPL-X, its local rotation is defined in the set function at timestamp , and A(j) represents the ordered set of the ancestor joints of joint j. The global rotation of this joint is calculated by successively multiplying the local rotations of all its ancestor joints i from the current joint j to the root joint, and is represented by the following formula:
[0070] (1),
[0071] In physically reasonable overall movements, virtual humans with different body shapes will have different local rotations under the same manipulation behaviors and sparse tracking signals. Figure 2 shows the instantaneous overall movements of virtual humans with different body shapes when manipulating the camera. Compared with the body shape in part (a) of Figure 2 , the thinner virtual human in part (b) of Figure 2 has problems with hand and foot penetration.
[0072] Characteristics of the manipulated object:
[0073] To generate accurate movements during the interaction, this method introduces object sampling points to represent the geometric shape and spatial position of the manipulated object. To reduce the feature complexity and enhance the expression effect, this method samples vertices of the object in the left and right hand coordinate systems respectively, rather than directly using all vertices or uniform sampling. For any vertex on the interaction object , its projection on the unit sphere is:
[0074] (2),
[0075] This step aims to project all vertices in the interactive object onto the surface of the unit sphere to achieve normalization, where is the set of all vertices of the interactive object , represents the i-th point in V, represents the average value of the vertices in V, that is, the geometric center.
[0076] Then, uniformly sample vertices on the unit sphere and convert each sampling point to the coordinate system of the left / right hand, which are represented as the polar coordinate forms of the left hand system and the right hand system respectively .
[0077] Manipulation action feature:
[0078] To generate a responsive and stable overall action and promote smooth transition of actions during manipulation, this method defines a distance feature vector , which provides continuous information about the proximity between the hand and the surface of the interactive object at timestamp t. At each timestamp t, uniformly sample points on the left and right hands respectively, and calculate the minimum distance between each sampling point and the nearest point on the surface of the interactive object, and construct a distance feature vector with this.
[0079] Progressive manipulation guidance optimization:
[0080] After constructing effective manipulation features, this method uses progressive manipulation guidance to generate an overall action with manipulation intent using the manipulation features.
[0081] This optimization process includes three steps:
[0082] Latent Motion Learning Optimization; Initial Holistic Motion Code Optimization; Manipulation-Aware Holistic Motion Code Optimization.
[0083] Latent motion learning optimization step:
[0084] In this step, the object of optimization is the conditional variational autoencoder (CVAE) obtained from the first-stage training. Among them, the action encodings of the human body and hand are embedded into the latent space of the CVAE and serve as important references for generating the overall action encoding in the second and third stages. In addition, the body and hand action decoders in the CVAE rely on this encoding information to reconstruct the overall action sequence. To ensure the accuracy and naturalness of the actions, the optimization goal of this stage is to minimize the reconstruction error caused by inaccurate encoding.
[0085] The loss function of the CVAE is:
[0086] (3),
[0087] The first term is the KL divergence loss , which is used to constrain the latent space to approach the standard normal distribution. Among them represents the weight of the KL divergence loss, represents the latent distribution predicted by the encoder and the standard normal distribution between the KL divergences.
[0088] The second term is the joint rotation reconstruction loss , which is used to constrain the predicted joint rotation to approach the true joint rotation. Among them represents the weight of this loss, represents the smooth-L1 loss function, and are the predicted and true joint rotation values respectively.
[0089] The third term is the joint position loss calculated based on forward kinematics , which is used to constrain the predicted joint position to approach the true joint position. Among them represents the weight of this loss, and are and The joint positions obtained by forward kinematics calculation.
[0090] Latent motion learning optimization step:
[0091] In this step, the sequence diffusion model in the second stage is trained to enable it to learn complex body postures and hand details. Under the condition that the input is the noisy body action latent encoding and the sparse tracking signal, the loss function of the body diffusion model (BodyDDPM) is:
[0092] (4),
[0093] Among them and respectively represent the predicted and true latent encodings of body movements.
[0094] Next, fix the parameters in Body DDPM and train the Hand Diffusion Model (Hand DDPM), whose loss function is:
[0095] (5),
[0096] where and respectively represent the predicted and true latent encodings of hand movements.
[0097] Overall motion optimization steps for manipulation perception:
[0098] In this step, the manipulation representation is used to train the sequence control network to optimize the overall motion encoding. After fixing the parameters of the body diffusion model and the hand diffusion model (DDPM), the body control module is optimized by the following formula:
[0099] (6),
[0100] This loss function includes the rotation loss of body joints and the position loss two terms. Where and respectively represent the predicted and true latent encodings of body movements with fused manipulation information, and are the joint positions calculated based on and .
[0101] Similarly, the loss function of the hand control module is as follows:
[0102] (7),
[0103] where, and respectively represent the predicted and true latent encodings of hand movements with fused manipulation information, and are the hand joint positions calculated based on and . In addition to the hand joint rotation loss and the hand joint position loss , it also includes the displacement loss of the key vertices of the hand. Where represents the weight of this loss, is the hand vertex position sampled in the manipulation representation, is the vertex position predicted by the output of this method for hand movements.
[0104] Model Distillation Step: In this step, a lightweight online network is trained to fit the sequence diffusion model constructed in Phases II and III and guided by the control network, enabling the model to perform inference with fewer steps and predict clear latent action representations.
[0105] Specifically, in the second and third phases, for each DDPM, we train a 24-layer DiT denoiser as the Teacher Model and distill it into a target 6-layer denoiser model. Since the teacher model and the target model share the network architecture and feature dimensions, the target denoiser learns not only the final output of the teacher model but also the outputs of some of its intermediate layers.
[0106] During the diffusion inference process in the second and third phases, initially 5 denoising steps are adopted; subsequently, a distillation process as shown in Figure 3 is used to reduce the denoising steps, and finally only one denoising step is required to complete the inference process.
[0107] It should be noted that the above content only illustrates the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. For those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements all fall within the protection scope of the claims of the present invention.
Claims
1. An overall human motion generation method for manipulation perception, characterized in that, It includes the following steps: Step 1: Perform latent encoding representation on human and hand actions encoded from sparse tracking signals by constructing a conditional variational autoencoder; Step 2: Combine a sequence diffusion model to gradually reconstruct the noisy action encoding to obtain the initial body and hand action latent encodings; Step 3: Use a sequence control network to extract object manipulation features, guide the DDPM to generate action latent encodings integrating manipulation intentions, and finally generate the final overall human action sequence integrating manipulation semantics through the decoding module of the CVAE.
2. The overall human motion generation method for manipulation perception according to claim 1, wherein In Step 1, constructing the conditional variational autoencoder includes: a body action encoding module, a hand action encoding module, and corresponding action decoding modules. The body action encoding module and the hand action encoding module encode human and hand actions through sparse tracking signals and output body and hand action encodings in the latent action space respectively.
3. The method for generating an overall human body motion oriented to manipulation perception according to claim 1, wherein, In Step 2, constructing the sequence diffusion model includes: a body DDPM and a hand DDPM. Given the sparse tracking signal and the noisy body latent encoding as inputs, the body DDPM reconstructs the initial body action latent encoding. Given this initial body action latent encoding, the sparse tracking signal, and the noisy hand latent encoding, the hand DDPM reconstructs the initial hand action latent encoding.
4. The overall human motion generation method for manipulation perception according to claim 1, wherein In Step 3, constructing the sequence control network includes a body control module and a hand control module. Given the object manipulation representation information, the body control module extracts manipulation features and inputs them to the body DDPM; the output body latent encoding and the initial body action latent encoding are input into a linear layer together to obtain the body action latent encoding integrating manipulation information; Similarly, the hand control module extracts manipulation features, inputs the manipulation features and the initial hand action latent encoding into the hand control module together to generate the hand action latent encoding integrating manipulation information; finally, the optimized body action encoding and hand action encoding are input into the corresponding action decoding modules constructed in Step 1 to generate the final overall action integrating manipulation intentions.
5. The overall human motion generation method for manipulation perception according to claim 1, characterized in that Step 3 includes object manipulation features to guide the overall action encoding of manipulation perception, including: The user actively performs manipulation behaviors through the controller buttons of the VR device. Once confirmed, the body and hand actions will be further refined; The body and hand are represented using the standard SMPL-X human body model, which consists of 10 real-valued numbers. For joint j in all joints A of SMPL-X, its local rotation is defined as a set function at timestamp where A(j) represents the ordered set of the ancestor joints of joint j. The global rotation of joint j is calculated by successively multiplying the local rotations of all its ancestor joints i from the current joint j to the root joint and is represented by the following formula: (1), Sample the vertices of the object in the left- and right-hand coordinate systems respectively, rather than directly using all vertices or uniform sampling. For any vertex on the interactive object , its projection on the unit sphere is: (2), By projecting all vertices in the interactive object onto the surface of the unit sphere to perform normalization processing, where is the set of all vertices of the interactive object , represents the i-th point in V, represents the average value of the vertices in V; Next, uniformly sample vertices on the unit sphere and transform each sampled point into the left / right hand coordinate system, represented as the polar coordinate forms in the left hand system and the right hand system respectively ; ; Define the distance feature vector , which provides continuous information about the proximity between the hand and the surface of the interactive object at timestamp t. At each timestamp t, points are uniformly sampled on the left and right hands respectively, and the minimum distance between each sampled point and the nearest point on the surface of the interactive object is calculated, and the distance feature vector is constructed based on this .
6. The method for generating an overall human motion oriented to manipulation perception according to claim 1, wherein After constructing effective manipulation features, use the manipulation features to generate the overall action with manipulation intentions through progressive manipulation guidance. The optimization process includes: 1) Latent motion learning optimization, where the object of optimization is the conditional variational autoencoder obtained in the first stage of training. Among them, the action encodings of the human body and hands are embedded into the latent space of the CVAE and serve as an important reference for generating the overall action encoding in Steps 2 and 3. In addition, the body and hand action decoders in the CVAE rely on the overall action encoding information to reconstruct the overall action sequence; 2) Train the sequence diffusion model in Step 2 so that it can learn complex body postures and hand details; 3) Use the manipulation representation to train the sequence control network to optimize the overall action encoding.
7. The method for generating an overall human motion oriented to manipulation perception according to claim 6, wherein 1), the object of optimization is the conditional variational autoencoder trained in Step 1. Among them, the action encodings of the human body and the hand are embedded into the latent space of the CVAE and serve as an important reference for generating the overall action encoding in Steps 2 and 3. In addition, the body and hand action decoders in the CVAE rely on the action encodings of the human body and the hand to reconstruct the overall action sequence. The loss function of the CVAE is: (3), The first term is the KL divergence loss , which is used to constrain the latent space to approach the standard normal distribution, where represents the weight of the KL divergence loss, represents the latent distribution predicted by the encoder and the standard normal distribution is the KL divergence between them; The second term is the joint rotation reconstruction loss , which is used to constrain the predicted joint rotation to approach the true joint rotation, where represents the weight of this loss, represents the smooth-L1 loss function, and are the predicted and true joint rotation values respectively; The third term is the joint position loss calculated based on forward kinematics , which is used to constrain the predicted joint positions to approach the true joint positions, where represents the weight of this loss, and are the joint positions obtained by and through forward kinematic calculation.
8. The overall human motion generation method for manipulation perception according to claim 6, wherein 2), the sequence diffusion model in Step 2 is trained to enable it to learn complex body postures and hand details. Under the condition that the input is the noisy body action latent encoding and the sparse tracking signal, the loss function of the body diffusion model is: (4), Among them and respectively represent the predicted and true latent encodings of body movements; Next, the parameters in Body DDPM are fixed, and the hand diffusion model is trained. Its loss function is: (5), Among them and respectively represent the predicted and true latent encodings of hand actions.
9. The method for generating an overall human body motion oriented to manipulation perception according to claim 6, wherein, 3), the manipulation representation is used to train the sequence control network to optimize the overall action encoding. After fixing the parameters of the body diffusion model and the hand diffusion model, the body control module is optimized through the following formula: (6), The loss function includes the rotation loss of body joints and the position loss of body joints in two terms, where is the weight of the position loss, where and respectively represent the latent encodings of the body actions of the predicted and true fused manipulation information and are the joint positions calculated based on and ; The loss function of the hand control module is as follows: (7), Among them, and respectively represent the potential encodings of hand actions for predicting and true fusion manipulation information. and are the hand joint positions calculated based on and In addition to the hand joint rotation loss and the hand joint position loss it also includes the displacement loss of the key vertices of the hand, where represents the weight of this displacement loss, is the hand vertex position sampled in the manipulation representation, is the vertex position predicted for the output hand action.
10. The overall human motion generation method for manipulation perception according to claim 1, wherein By training a lightweight online network to fit the sequence diffusion model constructed in Steps 2 and 3 and guided by the control network, the model can perform inference with fewer steps and predict a clear latent action representation. In Steps 2 and 3, for each DDPM, a 24-layer DiT denoiser is trained as the teacher model and distilled into the target 6-layer denoiser model. Since the teacher model and the target model share the network architecture and feature dimensions, the target denoiser not only learns the final output of the teacher model but also learns the output of some of its intermediate layers. In the diffusion inference process of Steps 2 and 3, initially 5 denoising steps are adopted; subsequently, the denoising steps are reduced through the distillation process, and finally only one denoising step is required to complete the inference process.
Citation Information
Patent Citations
Skeleton sequence identification method and system based on mask pattern auto-encoder
CN116434347A
Human body shape sensing sparse IMU motion capture method and system
CN118097775A
Robot multi-modal data generation method and system based on potential diffusion model
CN119903471A
Learning energy-based model with variational auto-encoder as amortized sampler
US20230169325A1
Method and apparatus for generating common-sense after-class exercise in low-resource scenario
WO2024197740A1