Free-form hand-object interaction generation method in field environment based on data-driven training
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-22
- Publication Date
- 2026-08-11
AI Technical Summary
[0003]然而,现有技术在数据集构建和生成方法两方面存在显著不足
Smart Images

Figure CN122551429A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision, human-computer interaction, and generative artificial intelligence, specifically to a method for generating free-form hand-object interaction in a field environment based on data-driven training. Background Technology
[0002] Hand-object interaction (HOI) is the foundation for humans to express intentions and perform tasks, and generating controllable interactions is crucial for augmented reality / virtual reality (AR / VR), robotics, and embodied artificial intelligence.
[0003] However, existing technologies have significant shortcomings in both dataset construction and generation methods. Regarding datasets, existing hand-object interaction datasets are mainly collected in laboratory environments, lacking diversity and realism in field scenes. Interaction types are primarily limited to grasping actions, with insufficient coverage of non-grasping interactions such as pushing, pressing, poking, and rotating. Occlusion of objects by hands in field videos makes 3D reconstruction difficult, and existing diffusion restoration methods are prone to geometric inconsistencies. Traditional 3D data acquisition requires specialized equipment and extensive manual annotation, resulting in high costs and limited scalability.
[0004] In terms of generation methods, existing methods are mainly limited to the grasping paradigm, with overly coarse control signals and strong inductive biases. They primarily aim to generate stable grasping postures, sacrificing interaction diversity. Even the latest methods using large language models for complex text control still have underlying model designs geared towards grasping interactions, lacking the ability to capture fine-grained control over diverse non-grasping interactions. The generated interactions often exhibit physical inconsistencies such as penetration and inaccurate contact, especially the severe global pose drift problem in free-form interaction generation. Therefore, there is an urgent need for a method that can automatically construct large-scale, diverse 3D interaction datasets from field videos and train on them to overcome the limitations of the grasping center and achieve fine-grained, controllable free-form hand-object interaction generation. Summary of the Invention
[0005] The purpose of this invention is to address the shortcomings of the existing technology and propose a data-driven training method for generating free-form hand-object interactions in the field. This method aims to automatically construct large-scale and diverse 3D interactive datasets from field videos and break through the limitations of traditional grasping centers to achieve diversified, controllable, and physically reasonable free-form hand-object interaction generation.
[0006] To achieve the above objectives, the present invention provides the following technical solution: The present invention provides a data-driven training-based method for generating free-form hand-object interactions in a field environment, characterized by the following steps: Step 1: Construct a set of hand-object interaction samples ;in, Let i be the i-th hand-object interaction sample, and ;in, Represents the mesh of the i-th object. This represents the state parameter of the i-th finely aligned hand. This represents the contact tag vector of the i-th object. This represents the vector of the i-th hand contact label. This represents the label of the i-th 17-dimensional hand component, covering the fingertips, fingertips, dorsal joints of the five fingers, as well as the palm and back of the hand. This represents the i-th short description text. This represents the i-th long description text; Step 2, according to ,right Sampling was performed to obtain object point clouds. ; Constructing and initializing the hand point cloud and will and The data are input into the contact map prediction module for processing, and the predicted contact map of the i-th object is obtained accordingly. and hand prediction contact map Thus, the i-th hand-object interaction sample is constructed. Contact loss , used to train the contact pattern prediction module; Step 3: Construct a multi-level conditional diffusion attitude generation module with multi-layer Transformer blocks as the backbone, and use... , , , , , As conditional inputs, the i-th hand-object interaction sample is constructed. diffusion loss By supervising the training of the multi-level conditional diffusion pose generation module, the i-th diffusion hand state parameters are obtained. ; Step 4: Construct a pose optimization network with a single-layer Transformer block as the backbone. and to , , and By performing single-step reasoning, we obtain attitude correction amount Further obtained rough results Thus constructing optimizer loss Optimize the network with a supervisory approach Through training, the i-th final hand state parameters are obtained. .
[0007] The method described in this invention is also characterized in that step 1 includes: Step 1.1: Obtain the i-th outdoor video clip. And extract the object-free frames from them. Hand-object interaction frame and prompt area ; Step 1.2, according to The i-th estimated object grid is obtained from the single-frame object 3D reconstruction network. ; Step 1.3, according to The i-th initial hand state parameters are obtained from the monocular hand parameter estimation network. The i-th two-dimensional hand key point and the i-th two-dimensional projection parameters ; Step 1.4, according to The set of points of the i-th object in a unified coordinate system is obtained from a large-scale pre-trained SAM3D model. and the i-th hand point set ; Step 1.5, The corresponding meshes are registered to the corresponding normalized reference point sets to obtain the i-th coarsely aligned hand parameters. and the i-th object mesh ; Step 1.6, according to right Iterative optimization is performed to obtain the i-th fine-aligned hand state parameters. ,in, Let be the i-th hand shape parameter. Let be the i-th hand posture parameter. Let i be the i-th global rotation parameter. Let i be the i-th global translation parameter; Step 1.7, according to Calculate the contact tag vector of the i-th object. The i-th hand contact tag vector and the label of the i-th hand component ;in, This indicates the number of point cloud samples of the object. This indicates the number of point cloud samples taken from the hand. Step 1.8: Based on the i-th short description text and contact labels Generate the i-th long description text. .
[0008] Furthermore, step 1.1 includes: Step 1.1.1: Based on the i-th outdoor video clip The segmentation model generates the mask sequence for the i-th object. and the i-th hand mask sequence ; Step 1.1.2, according to and The expanded crossover-union ratio determines the i-th interaction period. And select the interaction period The single frame of the object that is closest in time is taken as the i-th unoccluded frame of the object. ; Step 1.1.3, Extraction any candidate time Key points in the frame are matched using the RANSAC algorithm to obtain candidate time points. The corresponding affine transformation matrix and to Perform singular value decomposition to obtain Corresponding rotational components ; Step 1.1.4, according to Intersection and gradient comparison with mask, in Select from the corresponding frames to obtain the i-th hand-object interaction frame. ; Step 1.1.5: Use the affine transformation matrix Will object mask Transferred to Above, we obtain the i-th prompt area. .
[0009] Furthermore, step 1.5 includes: Depend on The i-th object point set is obtained by sampling. ;Depend on Synthesize the i-th initial hand point set Using the ICP point cloud registration method to ( , and normalized reference point set Preliminary alignment was performed to obtain and .
[0010] Furthermore, the iterative optimization in step 1.6 is from the camera center towards the cue area. The pixels within the array emit rays, and the set of hand contact points that are hit are... and the set of contact points of the object The area is identified as the 3D contact region; thus, it is constructed using equation (1). Hand refinement loss function And by minimizing Obtain the i-th fine-aligned hand state parameters : (1) In equation (1), This indicates the constraint that the 3D hand joint points projected onto the 2D key points are consistent, and ,in, Indicates by The first result obtained from forward computation A three-dimensional hand joint. This refers to the number of joints in the hand. Indicates the projection parameter of the i-th projection parameter A defined two-dimensional projection function, For the first Key points of a two-dimensional hand; This represents the constraint for aligning the hand point cloud and the object point cloud within the 3D contact area, and ,in, Indicates the chamfer distance; express The physical constraint loss, and we have: (2) In equation (2), express The contact distance constraint, and ; express The penalty penetrates the constraint of the hand vertex of the object, and ,in, Indicates by The obtained number Each hand-shaped mesh vertex This represents the number of vertices in the hand mesh. Let i be the signed distance function for the i-th object; express The hand's self-penetrating constraint, and ,in, Let be the signed self-penetrating function of the i-th hand; Indicates hand posture parameters Reasonable constraints, among which, This represents the hand anatomy constraint function calculated based on Manotorch.
[0011] Furthermore, in step 1.7 It is obtained through the following process: exist Upper fixed sampling There are 10 object points, and the scale factor is calculated based on the i-th object. Normalize all object points to obtain the i-th normalized object point set. ,in, Indicates the first A normalized object point; by Get hand point set ,in, Indicates the first A normalized hand point, Given the number of vertices in the hand mesh, the first vertex is calculated using equation (3). The nearest neighbor vote of the normalized object point on the hand point : (3) In equation (3), Indicates an indicator function; Indicates in Nearest neighbor mapping in; remember For voting distribution Quantiles are used to select the i-th high-frequency candidate point using equation (4). : (4) Use equation (5) to filter the set of the i-th core contact points. : (5) In equation (5), express arrive nearest neighbor distance, Indicates the absolute distance threshold. Indicates the relative distance threshold. Indicates in The nearest neighbor mapping; Using equation (6) to obtain Contact tags : (6) In equation (6), Normalized object point set The average nearest neighbor distance; This represents the scaling parameter. express Any core contact point in it.
[0012] Furthermore, step 1.8 includes: It is generated according to an action template, which is "applying a [degree of force] to the [object contact part] of the [object category] using [hand contact part], so that it [produces a result]"; where "[]" is used to identify replaceable slot variables in the action template. It is based on the object category, interaction intent, and the hand component being touched. And the area of contact object is generated.
[0013] Furthermore, step 2 includes: The contact image prediction module extracts the hand mask at the i-th component level. ,in, This represents a text parsing function; From the i-th object mesh The point cloud of the i-th object is obtained by sampling. and the i-th scaling factor ; Initialize the hand point cloud by initializing all hand parameters to zero. ; Calculate the features of the i-th object The i-th hand feature and the i-th coarse text feature The i-th fine text feature ;in, Indicates splicing, For point cloud feature extractors, It is a T5 text encoder; The contact map prediction module includes: object branches and hand branches; , and The input is processed in the contact graph prediction module of the conditional VQ-VAE architecture, and the object branch outputs the contact prediction of the i-th object. The hand branch outputs the prediction of the i-th hand contact. ; Using equation (7) to construct a contact diagram prediction module for... loss function : In equation (7), express Contact area overlap loss, express VQ-VAE codebook loss; Represents object branch pairs The contact area overlaps and the loss, and ; Indicates hand part support Contact area overlap loss, ;in, , They represent the first Predicted contact probability of an object point and a hand point. , They represent , The corresponding real contact label, It is a smoothing constant; Represents object branch pairs The codebook loss, and , Indicates hand part support The codebook loss, and ;in, and These represent the i-th hidden variable output by the object branch encoder and the hand branch encoder, respectively; They respectively represent the codebook and and The codebook vector that corresponds to the hidden variable; This indicates the gradient truncation operation; Weighting for commitment loss.
[0014] Furthermore, step 3 includes: Step 3.1, use Extracting body point clouds respectively global features Hand-shaped dot clouds global features ; from Local features of the extract from nearby bodies ;from Extracting local features of the hand from nearby areas ; Using a T5 encoder, respectively from Extracting coarse-grained text features and from Extracting fine-grained text features ; Step 3.2: Construct the first equation using equations (8) and (9). Global conditions of the Transformer block and local conditions : (8) (9) In equations (8) and (9), The number of Transformer layers for shallow fusion. This indicates that local conditions are not injected. Indicates a splicing operation; Step 3.3, let express The actual hand state parameters, express At the diffusion time step The noise-adding result is then processed by the multi-level conditional pose generation module using a denoising diffusion probability model based on multi-level Transformer blocks. And use equation (10) to predict the first Individual Hand-Object Interaction Sample Denoising internal hand state vector and order The corresponding external hand state parameters are denoted as ; (10) In equation (10), for The parameters, This represents the total number of layers in the Transformer block. Step 3.4: Use the FiLM mechanism to apply global conditions. and the Current features of the Transformer block Processing is performed to obtain the first result through equation (11). Features after modulation of layer Transformer blocks : (11) In equation (11), and These represent the scale and offset predicted by global conditions, respectively. Step 3.5, with As the first Query vector of layer ,by As the first Layer key vector and the layer value vector Thus, the cross-attention mechanism is used to obtain the first... Layer cross attention results ; Step 3.6, and After residual normalization, we obtain the first... Current features of the Transformer block Thus, the first The final feature of the layer Transformer block , and as Output ; Step 3.7: Construct a multi-level conditional pose generation module using equation (12). diffusion loss function : (12) In equation (12), This refers to Gaussian noise generated during the noise generation process. Indicates by The predicted denoised hand state vector; and From respectively The global rotation and translation of the hand obtained in the process; and for and Corresponding to the actual value; For the reason The calculated distance map from the hand joints to the object surface. This is a true distance map.
[0015] Furthermore, the pose optimization network in step 4 The two-stage optimization loss is constructed using equation (13). : (13) In equation (13), express Physical constraint loss, express Contact cycle consistency loss, express The weights are: (14) In equation (14), This represents the mapping from hand to object. This represents the mapping from object to hand. For the predicted hand contact area Corresponding points, Predicted object contact area The corresponding points, and we have: (15).
[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. In terms of data construction, this invention automatically constructs training samples through object individual frame extraction, reference point set estimation, similarity registration, and multi-level annotation. These samples can maintain consistency with the variable requirements of subsequent contact map prediction, text conditional modeling, and pose optimization, and support the construction of large-scale datasets.
[0017] 2. In terms of generation method, this invention breaks through the limitations of traditional grasping paradigms and can generate diverse daily interactive actions such as pushing, pressing, poking, and rotating; the coarse-to-fine hierarchical condition injection strategy effectively decouples global semantic guidance and local geometric refinement; and by combining text features extracted from a large language model with fine-grained hand part descriptions, it achieves precise control over interactive intent. 3. The cyclic consistency constraint of this invention effectively solves the global attitude drift problem through bidirectional mapping consistency; combined with single forward propagation and adaptive cyclic optimization, it ensures both inference efficiency and physical rationality. Attached Figure Description
[0018] Figure 1 This is an overview diagram of the technical solution of the present invention, which compares and shows the traditional laboratory grasping interaction (left) with the diverse daily interactions targeted by the present invention (right).
[0019] Figure 2 This is a flowchart of the overall hand-object interaction generation method; it includes schematic diagrams of contact graph prediction, multi-level conditional diffusion, and physical constraint optimization modules. Figure 3 This is an example of the generated result. Detailed Implementation
[0020] The technical solution of the present invention will now be further described in conjunction with specific embodiments and the accompanying drawings.
[0021] To address the limitations of existing hand-object interaction generation methods, which are mainly constrained by grasping paradigms, struggle to express free-form everyday interactions such as pushing, pressing, poking, and rotating, and lack fine-grained semantic control and physical plausibility, this invention proposes a data-driven training-based method for generating free-form hand-object interactions in a field environment. This method first constructs a set of 3D field training samples for training the generation method, and then completes network training and inference generation based on this training sample set. Figure 2 As shown, the specific steps include the following: Step 1: Construct a set of hand-object interaction samples ;in, Let i be the i-th hand-object interaction sample, and ;in, Represents the mesh of the i-th object. This represents the state parameter of the i-th finely aligned hand. This represents the contact tag vector of the i-th object. This represents the vector of the i-th hand contact label. This represents the label of the i-th 17-dimensional hand component, covering the fingertips, fingertips, dorsal joints of the five fingers, as well as the palm and back of the hand. This represents the i-th short description text. This represents the i-th long description text.
[0022] Step 1.1: Obtain the i-th outdoor video clip. And extract the object-free frames from them. Hand-object interaction frame and prompt area In practice, the Something-Something V2 dataset can be used as the original source, and the bounding boxes of the hand and object can be obtained using its associated Something-Else annotations. To improve the stability of 3D reconstruction, a multi-stage filtering pipeline can be adopted: first retain single-hand video clips, then retain single-object clips, and then further filter based on object size, visibility, and inter-frame stability. As a specific example, approximately 180,049 initial clips can be sequentially filtered into 137,578 single-hand clips, 82,728 single-object clips, 57,384 clips with satisfactory object size, 12,888 clips with good visibility, and 8,551 clips with stable inter-frame stability. The purpose of this is to reduce occlusion, multi-object clutter, and drastic viewpoint changes, ensuring that each video clip... Establish a more stable starting point for reconstruction. Step 1.1 can be further implemented according to the following process: Step 1.1.1: Based on the i-th outdoor video clip The SAM2 segmentation model generates the mask sequence for the i-th object. and the i-th hand mask sequence ; Step 1.1.2, according to and The expanded crossover-union ratio determines the i-th interaction period. And select the interaction period The single frame of the object that is closest in time is taken as the object's unoccluded frame. .
[0023] Step 1.1.3: Extract using RoMa any candidate time Key points in the frame are matched using the RANSAC algorithm to obtain candidate time points. The corresponding affine transformation matrix ,right Perform singular value decomposition: Separate Corresponding rotational components It is used to determine whether there is an excessive change in viewpoint in the candidate frame.
[0024] Step 1.1.4, according to Intersection and gradient comparison with mask, in Select from the corresponding frames to obtain the hand-object interaction frame. ; Step 1.1.5: Use the affine transformation matrix Will object mask Transferred to The prompt area is above. Compared to directly performing diffusion repair in the occluded area, it is easier to maintain the geometric consistency of the object and is also more suitable for large-scale automated processing.
[0025] Step 1.2, according to The i-th estimated object canonical mesh is obtained from the Trellis 3D object reconstruction network of a single frame. ; Step 1.3, according to The i-th initial hand state parameters are obtained from the monocular hand parameter estimation network HaMeR. The i-th two-dimensional hand key point and the i-th two-dimensional projection parameters ; Step 1.4, according to The set of points of the i-th object in a unified coordinate system is obtained from a large-scale pre-trained SAM3D model. and the i-th hand point set ; Step 1.5, The corresponding meshes are registered to the corresponding normalized reference point sets to obtain the i-th coarsely aligned hand parameters. and the i-th object mesh ; Specific implementation as the reason The i-th object point set is obtained by sampling. ;Depend on Synthesize the i-th initial hand point set Taking ICP registration of objects as an example, solve for the similarity transformation of the objects: Among them, the nearest neighbor mapping Used to establish the correspondence between the source point set and the target reference point set. Indicates the scale factor. Represents the rotation matrix. This represents the translation vector; after registration is complete, the object mesh can be updated as follows: The hand parameters are obtained by maintaining the local shape and pose, and only updating the global rotation, translation, and necessary intermediate scale values. Therefore, ( , and normalized reference point set Preliminary alignment was performed to obtain and .
[0026] Step 1.6, according to right Iterative optimization is performed to obtain the i-th fine-aligned hand state parameters. ,in, Let be the i-th hand shape parameter. Let be the i-th hand posture parameter. Let i be the i-th global rotation parameter. Let i be the i-th global translation parameter; The specific implementation of iterative optimization is to move from the camera center towards the cue area. The pixels within the array emit rays, and the set of hand contact points that are hit are... and the set of contact points of the object The area was identified as the 3D contact region. A hand refinement loss function was constructed using equation (1). The hand state parameters are obtained by minimizing this loss. : (1) In equation (1), This indicates the constraint that the 3D hand joint points projected onto the 2D key points are consistent, and ,in, Indicates by The first result obtained from forward computation A three-dimensional hand joint. This refers to the number of joints in the hand. Indicates the projection parameter of the i-th projection parameter A defined two-dimensional projection function, For the first Key points of a two-dimensional hand; Zoom in on the 3D contact area between the hand and the object, where... This indicates the chamfer distance. Let represent the physical constraint loss, and we have: (2) In equation (2), For contact distance constraints; This indicates a constraint that penalizes the hand's vertex that penetrates the object, and ,in Indicates by The obtained number Each hand-shaped mesh vertex This represents the number of vertices in the hand mesh. For the signed distance function of objects; This indicates a self-penetrating constraint on the hand, where For a signed self-penetrating function of the hand; Indicates hand posture parameters Reasonable constraints, among which This represents the hand anatomy constraint function calculated using Manotorch. In actual optimization, to maintain the stability of the individual hand shape, the shape parameters can be fixed, and only the pose, global rotation, translation, and scale can be optimized.
[0027] Step 1.7, according to Calculate the contact tag vector of the i-th object. The specific implementation is in Upper fixed sampling Each object point, and according to the object scale factor Normalizing these points yields the normalized object point set. ,in Indicates the first A normalized object point; by Get hand point set ,in Indicates the first The normalized hand point is calculated using equation (3). The nearest neighbor vote of the normalized object point on the hand point : (3) In equation (3), Indicates an indicator function; Represents the point set of objects Nearest neighbor mapping in; [denotes] For voting distribution Quantiles are used to select high-frequency candidate points using equation (4). : (4) Use formula (5) to screen core contact points : (5) In equation (5), express arrive nearest neighbor distance, Indicates the absolute distance threshold. Indicates the relative distance threshold. Indicates in The nearest neighbor mapping; then using equation (6) to obtain Contact label: (6) In equation (6), Normalized object point set The average nearest neighbor distance; This represents the scaling parameter. Represents a point set Any core contact point in it.
[0028] Hand side contact label This can be achieved using the same nearest neighbor voting, candidate point selection, core contact point filtering, and neighborhood expansion logic as the object side. Furthermore, the hand mesh is divided into 17 parts. .
[0029] Step 1.8, the i-th short description text Generated according to the action template, the action template is "Apply [force degree representation] to [object contact part] of [object category] using [hand contact part] to make it [produce result]", where "[]" is used to identify replaceable slot variables in the action template; the i-th long description text Based on object category, interaction intent, and the hand component being touched. And the area of contact with the object is generated. For example Figure 3 The "Press Calculator" feature can generate a short description text "Press Calculator" or a longer description text that further explains "Use your index finger to lightly press the calculator button for precise operation"; Figure 1 The phrase "push the water glass over" can be further described as using the pad of your index finger to apply a pushing force to the cap or body of the bottle, causing the water glass to tip over.
[0030] Once the sample set D is automatically constructed, quality control can be performed on four possible operations: "accept / fine-tune / regenerate / discard the result," and the hand-object interaction generation based on the training sample set can continue.
[0031] Step 2, from Take training samples During training, the contact map prediction module, attitude diffusion generation module, and attitude optimization module are optimized separately; during inference, they are executed in the order of "contact first, then attitude, then optimization".
[0032] according to ,right Sampling was performed to obtain object point clouds. Construct and initialize the hand point cloud. Before generating high-degree-of-freedom hand poses, the contact map prediction module first determines the local interaction target, starting by extracting the part-level hand mask. ,in, This represents a text parsing function that maps expressions such as "finger pad" and "palm" in natural language to 17-dimensional hand component annotations. ; by the i-th object mesh Sampling of object point clouds and scale factor Initialize the hand point cloud by initializing all hand parameters to zero. .
[0033] Calculate object features Hand features Bold text features Fine text features ,in Indicates splicing, For point cloud feature extractors, It is a T5 text encoder.
[0034] The contact map prediction module includes: object branches and hand branches; , and The input conditions are processed in the contact graph prediction module of the VQ-VAE architecture, and the object branch outputs the contact prediction of the i-th object. The hand branch outputs the prediction of the i-th hand contact. Therefore, the loss function of the contact map prediction module is constructed using equation (7). : (7) In equation (7), express Contact area overlap loss, express VQ-VAE codebook loss; Represents object branch pairs The contact area overlaps and the loss, and ; Indicates hand part support Contact area overlap loss, ;in, , They represent the first Predicted contact probability of an object point and a hand point. , They represent , The corresponding real contact label, It is a smoothing constant; Represents object branch pairs The codebook loss, and , Indicates hand part support The codebook loss, and ;in, and These represent the i-th hidden variable output by the object branch encoder and the hand branch encoder, respectively; They respectively represent the codebook and and The codebook vector that corresponds to the hidden variable; This indicates the gradient truncation operation; Weighting for commitment loss.
[0035] Step 3, according to ( , , , , , A multi-layered conditional diffusion pose generation module is constructed, with multi-layered Transformer blocks as the backbone, using diffusion loss. Supervised training yields the i-th diffusion hand state parameters. ; Step 3.1, use Extracting body point clouds respectively global features Hand-shaped dot clouds global features ; from Local features of the extract from nearby bodies ;from Extracting local features of the hand from nearby areas ; Using a T5 encoder, respectively from Extracting coarse-grained text features and from Extracting fine-grained text features ; Step 3.2: Construct the first equation using equations (8) and (9). Global conditions of the Transformer block and local conditions : (8) (9) In equations (8) and (9), The number of Transformer layers for shallow fusion. This indicates that local conditions are not injected. Indicates a splicing operation; Step 3.3, let express The actual hand state parameters, express At the diffusion time step The noise-adding result is then processed by the multi-level conditional pose generation module using a denoising diffusion probability model based on multi-level Transformer blocks. And use equation (10) to predict the first Individual Hand-Object Interaction Sample Denoising internal hand state vector and order The corresponding external hand state parameters are denoted as ; (10) In equation (10), for The parameters, The total number of layers in the Transformer block; in a specific embodiment, the multi-layer conditional pose generation module can adopt an 8-layer Transformer structure, including 4 attention heads, a latent dimension of 512, a feedforward network dimension of 1024, an activation function of GELU, and a dropout rate of 0.1.
[0036] Step 3.4: Use the FiLM mechanism to apply global conditions. and the Current features of the Transformer block Processing is performed to obtain the first result through equation (11). Features after modulation of layer Transformer blocks : (11) In equation (11), and These represent the scale and offset predicted by global conditions, respectively. Step 3.5, with As the first Query vector of layer ,by As the first Layer key vector and the layer value vector Thus, the cross-attention mechanism is used to obtain the first... Layer cross attention results ,in, For attention feature dimensions; Step 3.6, and After residual normalization, we obtain the first... Current features of the Transformer block Thus, the first The final feature of the layer Transformer block , and as Output The processing flow for each Transformer block is as follows: FiLM modulation → self-attention → cross-attention (if local conditions exist) → feedforward network → residual connections and layer normalization. During training, components of the global conditions are randomly discarded with a 10% probability to enhance robustness.
[0037] Step 3.7: Construct a multi-level conditional pose generation module using equation (12). diffusion loss function : (12) In equation (12), This refers to Gaussian noise generated during the noise generation process. Indicates by The predicted denoised hand state vector; and From respectively The global rotation and translation of the hand obtained in the process; and for and Corresponding to the actual value; For the reason The calculated distance map from the hand joints to the object surface. This is a true distance map.
[0038] The Adam optimizer can be used during training, with a learning rate of [missing information]. The batch size is 128, training lasts 1000 epochs, with a diffusion step T=1000, and data augmentation is performed on the objects with random rotation and scaling. To handle the long-tailed distribution of hand part contact frequencies, 17 parts are aggregated into 7 semantic categories, encoded as 7-bit binary labels, and balanced sampling is performed. During inference, Gaussian noise is used... Begin iterative noise reduction Each step involves prediction via Transformer and sampling based on the posterior distribution of DDPM. The final output is the hand posture parameters. .
[0039] Step 4, according to ( , , )and A pose optimization network is constructed using single-layer Transformer blocks (with the same structure as in step 3: 8 layers, 4 heads, 512 dimensions, but with independently trained parameters). First, obtain the attitude correction amount through single-step reasoning. Further rough results were obtained. The network uses optimizer loss. Supervised training; The attitude optimization network uses equation (13) to construct its optimization loss in both the single-step inference and iterative optimization stages. : (13) In equation (13), For the physical constraint loss in equation (2), express Contact cycle consistency loss, express The weights are: (14) In equation (14), This represents the mapping from hand to object. This represents the mapping from object to hand. For the predicted hand contact area Corresponding points, Predicted object contact area The corresponding point, and we have: (15) In one specific embodiment, it can be set , , , , A two-step optimization strategy can be adopted during testing: the first step is to quickly correct the global pose through a single forward propagation, and the second step is to... Execute on the basis The second iteration optimizes the learning rate. It is used to fine-tune local contact details.
[0040] Figure 3 The examples of tilting, scrolling, pinching, and pressing demonstrate that the generated results can vary with the text and the area of contact, rather than degenerating into a single grasping gesture.
[0041] This embodiment can be implemented by an electronic device, which includes a memory and a processor. The memory stores a program that supports the processor in executing the above-described method, and the processor is configured to execute the program stored in the memory. Accordingly, a computer program stored in a computer-readable storage medium, when run by the processor, can perform the above-described method steps.
Claims
1. A method for generating free-form hand-object interaction in a field environment based on data-driven training, characterized in that, comprising the steps of: Step 1: Construct a set of hand-object interaction samples ;in, Let i be the i-th hand-object interaction sample, and ;in, Represents the mesh of the i-th object. This represents the state parameter of the i-th finely aligned hand. This represents the contact tag vector of the i-th object. This represents the vector of the i-th hand contact label. This represents the label of the i-th 17-dimensional hand component, covering the fingertips, fingertips, dorsal joints of the five fingers, as well as the palm and back of the hand. This represents the i-th short description text. This represents the i-th long description text; Step 2, according to ,right Sampling was performed to obtain object point clouds. ; Constructing and initializing the hand point cloud and will and The data are input into the contact map prediction module for processing, and the predicted contact map of the i-th object is obtained accordingly. and hand prediction contact map Thus, the i-th hand-object interaction sample is constructed. Contact loss , used to train the contact pattern prediction module; Step 3: Construct a multi-level conditional diffusion attitude generation module with multi-layer Transformer blocks as the backbone, and use... , , , , , As conditional inputs, the i-th hand-object interaction sample is constructed. diffusion loss By supervising the training of the multi-level conditional diffusion pose generation module, the i-th diffusion hand state parameters are obtained. ; Step 4: Construct a pose optimization network with a single-layer Transformer block as the backbone. and to , , and By performing single-step reasoning, we obtain attitude correction amount Further obtained rough results Thus constructing optimizer loss Optimize the network with a supervisory approach Through training, the i-th final hand state parameters are obtained. .
2. The method of claim 1, wherein, said step 1 comprises: Step 1.1, obtaining the i-th outdoor video clip and extracting the object unoccluded frames therefrom , the hand-object interaction frames and the hint region ; Step 1.2, according to the i-th estimated object canonical mesh from the single-frame object 3D reconstruction network ; Step 1.3, obtaining the i-th initial hand state parameter from the monocular hand parameter estimation network , the i-th initial hand state parameter , the i-th two-dimensional hand key point , and the i-th two-dimensional projection parameter ; Step 1.4, according to The set of points of the i-th object in a unified coordinate system is obtained from a large-scale pre-trained SAM3D model. and the i-th hand point set ; Step 1.5, The corresponding meshes are registered to the corresponding normalized reference point sets to obtain the i-th coarsely aligned hand parameters. and the i-th object mesh ; Step 1.6, according to right Iterative optimization is performed to obtain the i-th fine-aligned hand state parameters. ,in, Let be the i-th hand shape parameter. Let be the i-th hand posture parameter. Let i be the i-th global rotation parameter. Let i be the i-th global translation parameter; Step 1.7, according to Calculate the contact tag vector of the i-th object. The i-th hand contact tag vector and the label of the i-th hand component ;in, This indicates the number of point cloud samples of the object. This indicates the number of point cloud samples taken from the hand. Step 1.8: Based on the i-th short description text and contact labels Generate the i-th long description text. .
3. The method according to claim 2, characterized in that, said step 1.1 comprises: Step 1.1.1: Based on the i-th outdoor video clip The segmentation model generates the mask sequence for the i-th object. and the i-th hand mask sequence ; Step 1.1.2, according to and The expanded crossover-union ratio determines the i-th interaction period. And select the interaction period The single frame of the object that is closest in time is taken as the i-th unoccluded frame of the object. ; Step 1.1.3, Extraction any candidate time Key points in the frame are matched using the RANSAC algorithm to obtain candidate time points. The corresponding affine transformation matrix and to Perform singular value decomposition to obtain Corresponding rotational components ; Step 1.1.4, according to and mask intersection ratio gradient, in selecting in the corresponding frame, obtaining the ith hand-object interaction frame ; Step 1.1.5, using the affine transformation matrix The object mask is transferred to the image to obtain the i-th hint region .
4. The method of claim 1, wherein, said step 1.5 comprises: Depend on The i-th object point set is obtained by sampling. ;Depend on Synthesize the i-th initial hand point set Using the ICP point cloud registration method to ( , and normalized reference point set Preliminary alignment was performed to obtain and .
5. The method of claim 2, wherein, The iterative optimization in step 1.6 is from the camera center towards the cue area. The pixels within the array emit rays, and the set of hand contact points that are hit are... and the set of contact points of the object The area is identified as the 3D contact region; thus, it is constructed using equation (1). Hand refinement loss function And by minimizing Obtain the i-th fine-aligned hand state parameters : (1) In equation (1), This indicates the constraint that the 3D hand joint points projected onto the 2D key points are consistent, and ,in, Indicates by The first result obtained from forward computation A three-dimensional hand joint. This refers to the number of joints in the hand. Indicates the projection parameter of the i-th projection parameter A defined two-dimensional projection function, For the first Key points of a two-dimensional hand; This represents the constraint for aligning the hand point cloud and the object point cloud within the 3D contact area, and ,in, Indicates the chamfer distance; express The physical constraint loss, and we have: (2) In equation (2), express The contact distance constraint, and ; express The penalty penetrates the constraint of the hand vertex of the object, and ,in, Indicates by The obtained number Each hand-shaped mesh vertex This represents the number of vertices in the hand mesh. Let i be the signed distance function for the i-th object; express The hand's self-penetrating constraint, and ,in, Let be the signed self-penetrating function of the i-th hand; Indicates hand posture parameters Reasonable constraints, among which, This represents the hand anatomy constraint function calculated based on Manotorch.
6. The method of claim 2, wherein, The step 1.7 in is obtained by the following procedure: exist Upper fixed sampling There are 10 object points, and the scale factor is calculated based on the i-th object. Normalize all object points to obtain the i-th normalized object point set. ,in, Indicates the first A normalized object point; by Get hand point set ,in, Indicates the first A normalized hand point, Given the number of vertices in the hand mesh, the first vertex is calculated using equation (3). The nearest neighbor vote of the normalized object point on the hand point : (3) In formula (3), denotes an indicator function; denotes a nearest neighbor mapping in Recall For the vote distribution The quantile, the i-th high-frequency candidate point is selected by formula (4) : (4) Screening the ith set of core contact points using formula (5) : (5) in formula (5), denotes to the nearest neighbor distance, denotes an absolute distance threshold, denotes a relative distance threshold, denotes the nearest neighbor mapping of the nearest neighbor mapping; The contact tag of formula (6) is obtained : (6) In equation (6), Normalized object point set The average nearest neighbor distance; This represents the scaling parameter. express Any core contact point in it.
7. The method of claim 2, wherein, Step 1.8 includes: It is generated according to an action template, which is "applying a [degree of force] to the [object contact part] of the [object category] using [hand contact part] to produce a [result]"; where "[]" is used to identify replaceable slot variables in the action template. It is based on the object category, interaction intent, and the hand component being touched. And the area of contact object is generated.
8. The method of claim 1, wherein, said step 2 comprises: The contact map prediction module extracts the i-th component-level hand mask wherein, denotes a text parsing function; From the i-th object mesh The point cloud of the i-th object is obtained by sampling. and the i-th scaling factor ; Generating an initialized hand point cloud from all-zero initialization of hand parameters ; computing the ith object feature , the ith hand feature , and the ith coarse text feature , the ith fine text feature ; wherein, denotes concatenation, is a point cloud feature extractor, is a T5 text encoder; The contact map prediction module includes: object branches and hand branches; , and The input is processed in the contact graph prediction module of the conditional VQ-VAE architecture, and the object branch outputs the contact prediction of the i-th object. The hand branch outputs the prediction of the i-th hand contact. ; Using equation (7) to construct a contact diagram prediction module for... loss function : In equation (7), express Contact area overlap loss, express VQ-VAE codebook loss; Represents object branch pairs The contact area overlaps and the loss, and ; Indicates hand part support Contact area overlap loss, ;in, , They represent the first Predicted contact probability of an object point and a hand point. , They represent , The corresponding real contact label, It is a smoothing constant; Represents object branch pairs The codebook loss, and , Indicates hand part support The codebook loss, and ;in, and These represent the i-th hidden variable output by the object branch encoder and the hand branch encoder, respectively; They respectively represent the codebook and and The codebook vector that corresponds to the hidden variable; This indicates the gradient truncation operation; Weighting for commitment loss.
9. The method according to claim 1, characterized in that, said step 3 comprises: Step 3.1, using extract global features of the object point cloud and global features of the hand point cloud respectively; From extracting local features of objects in the vicinity ; from extracting local features of hands in the vicinity ; coarse-grained text features are extracted from fine-grained text features are extracted from ; Step 3.2, constructing the first global condition for the layer Transformer block using formula (8) and formula (9) and the local condition : (8) (9) In equations (8) and (9), The number of Transformer layers for shallow fusion. This indicates that local conditions are not injected. Indicates a splicing operation; Step 3.3, let express The actual hand state parameters, express At the diffusion time step The noise-adding result is then processed by the multi-level conditional pose generation module using a denoising diffusion probability model based on multi-level Transformer blocks. And use equation (10) to predict the first Individual Hand-Object Interaction Sample Denoising internal hand state vector and order The corresponding external hand state parameters are denoted as ; ,(10) In equation (10), for The parameters, This represents the total number of layers in the Transformer block. Step 3.4: Use the FiLM mechanism to apply global conditions. and the Current features of the Transformer block Processing is performed to obtain the first result through equation (11). Features after modulation of layer Transformer blocks : (11) In formula (11), and respectively represent the scale and offset predicted by the global condition; Step 3.5, taking as the first layer query vector , taking as the first layer key vector and the first layer value vector respectively, so as to obtain the first layer cross-attention result by using the cross-attention mechanism; Step 3.6, and After residual normalization, we obtain the first... Current features of the Transformer block Thus, the first The final feature of the layer Transformer block , and as Output ; Step 3.7, constructing a multi-level conditional pose generation module with formula (12) to the diffusion loss function of : (12) In equation (12), This refers to Gaussian noise generated during the noise generation process. Indicates by The predicted denoised hand state vector; and From respectively The global rotation and translation of the hand obtained in the process; and for and Corresponding to the actual value; For the reason The calculated distance map from the hand joints to the object surface. This is a true distance map.
10. The method of claim 1, wherein, The pose optimization network in step 4 is constructed by using formula (13) to construct a two-stage optimization loss : (13) In equation (13), express Physical constraint loss, express Contact cycle consistency loss, express The weights are: (14) In equation (14), This represents the mapping from hand to object. This represents the mapping from object to hand. For the predicted hand contact area Corresponding points, Predicted object contact area The corresponding points, and we have: (15)。