Hand-object interaction data set expansion method and device based on Gaussian splashing
Through Gaussian splattering technology and super-resolution modules, the problem of single posture and insufficient realism in the existing technology is solved, and the effect of deep learning models is improved.
Patent Information
- Application Number
- CN202510242651.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-12-31
- Filing Date
- 2025-03-03
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art is difficult to generate hand-object interaction data with diverse postures and strong sense of reality, resulting in poor effectiveness of deep learning models in hand-object interaction scenarios.
Gaussian splashing technology is used to build a hand-to-object interaction model, combine rigid transformation and bone transformation to optimize posture, and use super-resolution modules and generative adversarial network to enhance the realism of data, and generate high-quality hand-to-object interaction data through the Encoder-Decoder network.
The end-to-end hand-to-object interaction model is realized to generate diverse and high-quality hand-to-object interaction data, which improves the generalization ability of deep learning models and the applicability of data sets.
Smart Images

Figure CN120339498A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer graphics, and particularly relates to a method and device for augmenting a hand-object interaction data set based on Gaussian splashing. Background Art
[0002] Hand-object interaction is the main way for people to interact with the outside world in daily life. Therefore, helping a computer understand the postures of the hand and the object in three-dimensional space and the relative position relationship between the hand and the object during hand-object interaction, and further analyzing the interaction intention of the user is crucial for related applications in the fields of mixed reality and robotics. In mixed reality, the computer can timely obtain the behavior state and mental state of the user, assist the user in completing the operation task, and improve the interaction efficiency; in the field of robotics, it can support the fine manipulator control function in the teleoperation scenario. The computer can perceive the operation habits of the user and remotely control the manipulator to complete relevant work in a dangerous scenario. At the same time, hand-object state perception can solve the collision between the hand and the object, form a good grasping operation plan, and improve the efficiency of teleoperation.
[0003] However, due to the high degree of freedom of the human hand and rich posture changes, and at the same time, there are serious mutual occlusions during the hand-object interaction process, resulting in difficulties in data collection and annotation. The quantity and annotation accuracy of the data set limit the effect of the deep learning model. Traditional three-dimensional data augmentation methods need to first scan the hand and the object with expensive scanning instruments to construct a mesh model, which is time-consuming, laborious and inefficient; the two-dimensional image generation strategy based on the generative model cannot perceive the accurate hand-object interaction relationship, and the generated interaction postures lack rationality. The existing methods based on neural rendering are difficult to end-to-end construct models of the hand (including single hand and both hands) and the object for the hand-object interaction scenario with a high degree of freedom, and the rendered images of the constructed human hand model have poor quality; there is a lack of diverse posture augmentation schemes, resulting in a lack of diversity in the data augmentation results and a large difference from the images of real data, and the effect of the deep learning model for hand-object interaction scenario perception cannot be effectively improved. To achieve a data augmentation method with diverse interaction postures and strong authenticity of rendering results, the following key problems need to be solved:
[0004] 1) It is necessary to realize automated and end-to-end construction of hand-object models based on the existing hand-object interaction data set. However, the posture of the hand-object interaction process has a high degree of freedom and there are serious mutual occlusions, resulting in difficulties in data annotation and a deviation between the annotation result and the true value; different from rigid objects, there will be a problem of inconsistent posture annotation when annotating the chain-like structure of the human hand, making the rendered images of the existing model construction methods have artifacts, blurring, etc. Since the synthesized images are quite different from the real images, the effect of the model deteriorates during model training.
[0005] 2) Existing data augmentation methods lack pose diversity. The two-dimensional image data generation method for hand-object interaction scenarios is limited by the shape of the object and the diversity of interaction postures, and cannot generate interaction images that meet reasonable grasping requirements. Existing three-dimensional data-based hand-object interaction pose generation schemes are limited to the interaction between a single hand and an object, and lack a pose optimization scheme for the interaction between two hands and an object. The singularity of poses in the augmented dataset causes the deep learning model to fail to learn diverse prior knowledge of interaction postures, and it cannot achieve accurate perception of hand-object interaction scenarios when facing new interaction scenarios.
[0006] In summary, a data augmentation method that satisfies pose diversity and strong image realism is a common core problem for many mainstream applications. Summary of the Invention
[0007] The object of the present invention is to provide a technical solution for augmenting a hand-object interaction dataset. Taking the color image of the hand-object interaction scene from a sparse perspective as the input, based on the Gaussian splash technology, it realizes the augmentation of a dataset with diverse hand-object interaction postures and strong authenticity of rendering results. The augmented data can be used to improve the effect of the deep learning model.
[0008] The technical solution adopted by the present invention to achieve the above object is as follows:
[0009] A method for augmenting a hand-object interaction dataset based on Gaussian splash, comprising the following steps:
[0010] 1) Data preprocessing: Collect the color image of the hand-object interaction from a sparse perspective and perform preprocessing;
[0011] 2) Gaussian splash modeling and rendering: According to the foreground image obtained by preprocessing, use the Gaussian splash method to construct a human hand model and an object model and perform projection rendering, and optimize the Gaussian kernel parameters;
[0012] 3) Pose optimization and rendering: Based on the initialized hand-object interaction pose, transform the human hand model and the object model to calculate the contact area in the observation space; Combine the calculated contact area with the predicted contact area to optimize the pose parameters; Render and generate a rough image according to the optimized hand-object interaction pose;
[0013] 4) Super-resolution reconstruction: Perform super-resolution reconstruction on the rough image to generate a clear image;
[0014] 5) Data fusion and augmentation: Fuse the generated clear image with the background image to obtain an augmented dataset of hand-object interaction containing background information.
[0015] Further, the steps of preprocessing in step 1) include:
[0016] Segment the color image of the hand-object interaction from a sparse perspective, and extract the foreground image of the hand-object interaction;
[0017] Scale and translate the foreground image to obtain an image with a unified resolution, and make the principal point of the internal parameters of the image located at the center point of the image.
[0018] Furthermore, the steps of constructing the human hand model and the object model by the Gaussian splashing method and performing projection rendering in step 2) include:
[0019] Define Gaussian kernels on the human hand model and the object model in the standard space, and calculate the center point and covariance matrix of the Gaussian kernels;
[0020] Based on the bone transformation and skinning weights, transform the Gaussian kernel of the human hand model in the standard space to the observation space;
[0021] Based on the rigid transformation of the object pose, transform the Gaussian kernel of the object model in the standard space to the observation space;
[0022] Project the 3D Gaussian kernel in the observation space into a 2D Gaussian kernel, and use the neural rendering method to render the image.
[0023] Furthermore, in step 2), based on the total loss function L SSIM composed of the L1 loss L1, the SSIM loss L R and the distance regularization loss L HO optimize the parameters of the Gaussian kernel.
[0024] Furthermore, the representation of the total loss function L HO is as follows:
[0025] L HO = (1 - λ SSIM )L1 + λ SSIM L SSIM + λ R L R
[0026] Among them, λ SSIM , λ R are hyperparameters, and L1 and L SSIM are used to constrain the difference between the rendered image and the real image to be as small as possible;
[0027] The loss function L R is used to constrain the position of the vertices of the optimized object model to be as close as possible to the surface of the initial object model, and its calculation formula is as follows:
[0028]
[0029] Among them, N v represents the number of vertices of the object model, v iDenote the vertices of the object model as \(v\), \(M\) represents the initial object model, and \(dist(\cdot)\) represents the distance from a vertex to the nearest face.
[0030] Further, the step of initializing the hand-object interaction pose in step 3) includes randomly rotating along each axis and translating the hand and the object.
[0031] Further, in step 3), based on the initialized hand-object interaction pose, the hand model and the object model are transformed to obtain the vertex positions in the observation space, and then input into ContactNet to predict the contact area.
[0032] Further, in step 3), based on the consistency loss \(L_{cons}\) C , the contact loss \(L_{cont}\) H , and the penetration loss \(L_{pen}\) P , the total loss \(L\) Poss is used to optimize the pose parameters.
[0033] Further, the total loss \(L\) Poss is expressed as follows:
[0034]
[0035] where \(\lambda_1\) C , \(\lambda_2\) H , \(\lambda_3\) P are hyperparameters, \(\Omega\) represents the calculated contact area, and \(\hat{\Omega}\) represents the contact area predicted by ContactNet.
[0036] Further, in step 4), the Encoder-Decoder deep learning model is used to perform super-resolution reconstruction on the rough image, and the steps include:
[0037] Construct a super-resolution training set using the rough images and the corresponding ground truth images;
[0038] Use the Encoder-Decoder model to encode and decode the features of the rough images in the super-resolution training set to generate super-resolution images;
[0039] Use the generative adversarial network to train the Encoder-Decoder model, and use the discriminator to judge the authenticity of the generated images;
[0040] Use the trained Encoder-Decoder model to perform super-resolution reconstruction on the rough images.
[0041] Further, when training the Encoder-Decoder model, the loss function used is composed of the L1 loss, the VGG loss, and the generative adversarial loss.
[0042] A hand-object interaction dataset augmentation device based on Gaussian splashing, which is used to implement the above method, includes:
[0043] A data preprocessing module, which is used to collect color images of hand-object interaction from sparse viewpoints and perform preprocessing;
[0044] A hand-object model construction module, which is used to construct a human hand model and an object model by using the Gaussian splashing method according to the foreground image obtained by preprocessing, perform projection rendering, and optimize the Gaussian kernel parameters;
[0045] A pose optimization module, which is used to transform the human hand model and the object model based on the initialized hand-object interaction pose to calculate the contact area in the observation space; combine the calculated contact area with the predicted contact area to optimize the pose parameters; and render and generate a rough image according to the optimized hand-object interaction pose;
[0046] A super-resolution module, which is used to perform super-resolution reconstruction on the rough image to generate a clear image;
[0047] An augmented data processing module, which is used to fuse the generated clear image with the background image to obtain an augmented dataset of hand-object interaction containing background information.
[0048] The beneficial effects achieved by the present invention are as follows:
[0049] 1. Based on the Gaussian splashing technology, the present invention realizes the construction of an end-to-end hand-object interaction model, and can generate high-quality hand-object interaction data from sparse-viewpoint color images.
[0050] 2. The present invention uses the Gaussian kernel modeling method to construct a hand-object model in the standard space, and converts it to the observation space through a rigid transformation and a bone transformation matrix to realize more accurate modeling of hand-object interaction relationships.
[0051] 3. The present invention proposes a pose optimization module, which extends single-handed interaction to two-handed interaction, optimizes the hand-object contact area, enhances the pose diversity of the dataset, and improves the generalization ability of the deep learning model.
[0052] 4. The present invention uses the super-resolution module to improve the detail quality of the rendered image through an Encoder-Decoder network, reduce the blur and artifacts caused by occlusion and annotation inconsistencies, and enhance the realism of the data.
[0053] 5. The present invention uses a generative adversarial network to optimize the rendered image, and uses L1 loss, VGG loss, and generative adversarial loss to constrain the generation process, making the synthetic data closer to real images.
[0054] 6. The present invention adopts background fusion technology to combine the generated hand-object interaction images with the COCO2017 background data, making the background noise distribution of the augmented data closer to that of real captured data and improving the applicability of the dataset.
[0055] 7. The data augmentation method of the present invention has a high degree of automation, can efficiently generate large-scale and diverse hand-object interaction data, reduce the costs of data collection and annotation, and improve the quality of the dataset.
[0056] 8. The present invention enhances the realism and diversity of the augmented data, can provide rich pose prior information for deep learning models, and improves the learning effect of hand-object interaction tasks.
[0057] 9. The method of the present invention is applicable to professional research and popular applications, meeting the needs for high-quality hand-object interaction data in different scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 is the overall flowchart of a hand-object interaction dataset augmentation method based on Gaussian splash in an embodiment.
[0059] Figure 2 is an example diagram of the augmented dataset generated in an embodiment.
[0060] Figure 3 is a schematic diagram of the application of the augmented dataset generated in an embodiment.
[0061] Figure 4 is the effect diagram of the improvement of the deep learning model by the augmented data generated in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0062] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to specific embodiments and the accompanying drawings.
[0063] An embodiment of the present invention provides a hand-object interaction dataset augmentation method based on Gaussian splash, and its overall process is as Figure 1 shown, specifically including the following processing steps:
[0064] Step 1: Data preprocessing
[0065] The data used is hand-object interaction color images under sparse viewpoints.
[0066] First, segment the input color image to obtain the foreground pixels of hand-object interaction, and then scale and translate it to obtain an image with a resolution of W×H, and at the same time, the principal point of the internal parameters is at the center point of the picture on.
[0067] Step 2: Gaussian splash modeling and rendering
[0068] Based on the input sparse-view images, use the Gaussian splashing method to end-to-end construct human hand models (including single-hand and two-hand models) and object models. The constructed models can use the input hand-object interaction poses and camera poses to render new images.
[0069] First, define Gaussian kernels on the human hand model (i.e., MANO-HD) and object model (i.e., object mesh model) in the standard space respectively. The representation of the Gaussian kernel is as follows:
[0070]
[0071] where x c represents the center point of the Gaussian kernel, Σ represents the covariance matrix. At the same time, the parameters of the Gaussian kernel also include: volume density σ and line-of-sight related color c.
[0072] Define the center point x of the Gaussian kernel c on the surface of the object model, and its calculation formula is as follows:
[0073] x c = βV = β1v1 + β2v2 + β3v3
[0074] where β represents the trainable weight parameter and satisfies β1 + β2 + β3 = 1, and V represents the three vertices of the patch on the object model.
[0075] The covariance matrix Σ = RSS T R T , where the rotation matrix where r1 points to the normal of the patch of the object model, r2 represents the direction from the center point m of the patch to the vertex v1, and r3 is obtained by calculating the orthonormalization of r1 and r2. The size matrix where s1 = s2 = ‖m - v1‖, s3 = <v2, R3>, and ε represents a very small value.
[0076] For the Gaussian kernel of the human hand model, additionally define the skinning weight W = (w1, w2,..., w n ), which represents the degree to which each Gaussian kernel is affected by the 3D key points of the human hand. It is obtained by weighted summation of the skinning weights of each vertex on the MANO-HD patch and the weight parameter β. Then, the Gaussian kernel of the human hand model in the standard space can be transformed to the observation space using the bone transformation matrix B and the skinning weight. The specific calculation formula is as follows:
[0077]
[0078] where, represents the center point of the Gaussian kernel in the standard space of the human hand, Represents the center point of the Gaussian kernel in the observation space.
[0079] For the Gaussian kernel of the object model, based on the input object pose \(T\in SE(3)\), the Gaussian kernel of the object model in the standard space is transformed to the observation space by a rigid transformation. The specific calculation formula is as follows:
[0080]
[0081] Where, Represents the center point of the Gaussian kernel in the object standard space, Represents the center point of the Gaussian kernel in the observation space.
[0082] Then, project the 3D Gaussian kernel into a 2D Gaussian kernel and render the image using the neural rendering method. The specific rendering formula is as follows:
[0083]
[0084] Where, \(\alpha\) i Represents the opacity of the \(i\)-th Gaussian kernel on the ray, and \(c\) i Represents the color of the \(i\)-th Gaussian kernel.
[0085] Finally, calculate the error between the input real image and the rendered image based on the loss function to optimize the parameters of the Gaussian kernel. The loss functions used include the L1 loss \(L1\), the SSIM loss \(L\) SSIM and the distance regularization loss \(L\) R , and the complete loss function is expressed as:
[0086] \(L\) HO =(1 - \(\lambda\) SSIM )\(L1+\lambda\) SSIM \(L\) SSIM +\(\lambda\) R \(L\) R
[0087] Where, \(\lambda\) SSIM , \(\lambda\) R are hyperparameters, which are set to 0.2 and 0.5 respectively. \(L1\) and \(\lambda\) SSIM are used to constrain the difference between the rendered image and the real image to be as small as possible.
[0088] The loss function \(L\) R is used to constrain the position of the vertices of the optimized object model to be as close as possible to the surface of the initial object model. Its calculation formula is as follows:
[0089]
[0090] Where, \(N\) v represents the number of vertices of the object model, and \(v\) iDenote the vertices of the object model as \(v\), \(M\) represents the initial object model, and \(dist(·)\) represents the distance from the vertex to the nearest patch.
[0091] Step 3: Pose Optimization and Rendering
[0092] To generate diverse hand-object interaction poses and enable the augmented dataset to provide rich pose priors, the GraspTTA method is improved. The optimization method for single-handed hand-object interaction poses is extended to scenarios involving single-handed, two-handed, and object interactions. This method uses the pre-trained ContactNet model in GraspTTA to predict the contact area during interaction.
[0093] First, initialize the input hand-object interaction pose, including: randomly rotate the input hand-object interaction pose by [0, 20°] along the x-axis, y-axis, and z-axis respectively. Then, translate the root joint of the human hand away from the center of the object by a distance of 5% of the distance from the hand to the object. Finally, add a random translation of [0, 6 cm].
[0094] Then, based on the initialized hand-object interaction pose, transform the human hand model and the object model to obtain the vertex positions in the observation space, and then input them into ContactNet to predict the contact area. Meanwhile, based on the human hand model and the object model in the observation space, the current contact area \(\Omega\) can be calculated, and then the self-supervised consistency loss \(L\) is used. C To constrain the actually calculated contact area \(\Omega\) to be as consistent as possible with the contact area predicted by the network, its calculation formula is as follows:
[0095]
[0096] Meanwhile, the contact loss \(L_c\) is used. H To constrain the hand vertices that are too close to the object to contact the object, the penetration loss \(L_p\) is used. P To constrain the object vertices inside the human hand model to contact the human hand. The total pose optimization loss is expressed as:
[0097]
[0098] During the optimization process, only the pose parameters of the left and right hands are optimized. \(l\) and \(r\) represent the left and right hands respectively, and \(\lambda_c\), \(\lambda_p\), \(\lambda_s\) C , \(\lambda_c\) H , \(\lambda_p\) P are hyperparameters, which are set to 1, 1, 17 respectively. This iterative optimization process is carried out 200 times to obtain the optimized two-handed hand-object interaction poses.
[0099] Finally, combine the optimized hand-object interaction poses with the human hand model and the object model constructed in the previous step for rendering, and a rough image \(I\) can be obtained.C for enhancing the photo-realism of images for subsequent super-resolution modules.
[0100] Step 4: Super-resolution reconstruction
[0101] To address the issue of blurred and artifact-ridden rendered images caused by inconsistent pose annotations in the training of the Gaussian splashed human hand and object models, an Encoder-Decoder deep learning model is now used to aggregate local region features of the image to pixel points to infer and generate an image with sharper texture, enhancing the photo-realism of the image.
[0102] First, construct a dataset for training the super-resolution module. Use the poses in the test set as input to the constructed human hand model and object model for rendering to obtain a rough image I C , and then pair it with the corresponding real image I GT to form an image pair as the training set.
[0103] Use the StyleUNet network architecture in the StyleAvatar method and the Encoder-Decoder network as the generator. For the input rough image I C , calculate the bounding box based on its foreground pixels, then crop and resize the image to a resolution of 1024×1024. Input the rough image and a random feature vector into the generator for feature encoding and decoding to generate a super-resolution image I R , and then use the training model of the generative adversarial network to set up a discriminator to judge the authenticity of the generated image.
[0104] The loss functions used include the L1 loss L1, the VGG loss L VGG and the generative adversarial loss L GAN , and the total loss L S can be expressed as:
[0105] L S = λ1L1 + λ VGG L VGG + L GAN
[0106] where the L1 loss L1 and the VGG loss L VGG are used to constrain the generated image I R to be as similar as possible to the real image I GT , and L GAN is used to constrain the generator and discriminator in the generative adversarial network. λ1 and λ VGG are hyperparameters, set to 5 and 0.03 respectively.
[0107] Step 5: Data fusion and augmentation
[0108] The goal of the augmented data processing stage is to generate Image I R Add background information to ensure that the background noise distribution of the augmented image data is as consistent as possible with that of the real collected data. During the specific operation process, the training set of COCO2017 is used as the background image data I B For the generated Image I R Perform segmentation to obtain the foreground hand-object pixels to get the mask image M, and then adjust the background image I B The resolution is consistent with the generated Image I R Then, splice the images according to the following formula:
[0109] I=(1 - M)I B +I R
[0110] Where, I represents the finally generated augmented data, which has the characteristics of diverse interactive postures and strong authenticity of rendering results.
[0111] The embodiment of the present invention provides a hand-object interaction data set augmentation device based on Gaussian splash for implementing the above method, which includes:
[0112] The data preprocessing module is responsible for preprocessing the data of the input hand-object interaction color image, including foreground pixel acquisition of the image, translation and scaling of the image.
[0113] The hand-object model construction module is responsible for constructing the human hand and object models based on the Gaussian splash method for the input color image under sparse viewpoints. At the same time, the constructed models support pose editing operations, and new hand-object interaction images can be rendered using the input hand-object interaction poses and camera poses.
[0114] The pose optimization module is responsible for generating diverse hand-object interaction poses. Based on the input hand and object poses, iterative optimization methods are used to generate reasonable hand and object interaction poses. Inputting the poses generated by this module into the human hand and object models can generate rough hand-object interaction images.
[0115] The super-resolution module is responsible for improving the quality of the rough images. Based on the input rough hand-object interaction images, an Encoder-Decoder deep learning model is used to generate super-resolution images with stronger realism.
[0116] The augmented data processing module is responsible for adding background information to the generated super-resolution images, using the training set of COCO2017 as the background image to splice with the super-resolution images, and ensuring that the background of the augmented image data is as consistent as possible with the background noise distribution of the real collected data.
[0117] Figure 2The figure shows an example diagram of the augmented dataset generated in this embodiment, which can automatically generate diverse real hand-object interaction images for multi-category objects, featuring diverse interaction postures and high authenticity of rendering results.
[0118] Figure 3 The figure shows the application of the augmented dataset generated in this embodiment. In the prior art, the original dataset is used to train a deep learning model. Due to the limitation of the diversity of the dataset postures, the performance of the model on the test set is poor. After the data augmentation of the present invention, the original dataset and the augmented dataset are combined. After training the deep learning model, the performance of the model on the test set is significantly improved.
[0119] Figure 4 The figure shows the improvement of the augmented data generated in this embodiment on the deep learning model. The test is mainly carried out on two types of models, namely the hand-object interaction reconstruction model and the contact area prediction model. The first row is the input test image, the second row is the training of the deep learning model without data augmentation, and the third row is the training of the deep learning model with data augmentation. It can be found that after training the deep learning model with data augmentation, the prediction accuracy of the model can be significantly improved.
[0120] In another embodiment, the application object of the present invention - the human hand can be extended to objects similar to the human hand, such as the whole or part of the human body, robots, animals, plants, robotic arms, and robotic hands, etc. The application object of the present invention - objects can be extended to other rigid or non-rigid objects, etc.
[0121] In another embodiment, an electronic device (such as a computer, a server, etc.) is provided, which includes a memory and a processor. The memory stores a computer program, and the computer program is configured to be executed by the processor. The computer program includes instructions for executing the steps in the above-mentioned method.
[0122] In another embodiment, a computer-readable storage medium (such as ROM / RAM, disk, optical disc) is provided. When the computer program stored in the computer-readable storage medium is executed by a computer, the steps of the above-mentioned method are implemented.
[0123] The above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Those of ordinary skill in the art can modify or equivalently replace the technical solutions of the present invention without departing from the principles and scope of the present invention. The protection scope of the present invention shall be subject to what is described in the claims.
Claims
1. A method for augmenting a hand-object interaction dataset based on Gaussian splashing, characterized in that It includes the following steps: 1) Collect the hand-object interaction color images from sparse viewpoints and perform preprocessing; 2) According to the foreground images obtained by preprocessing, use the Gaussian splashing method to construct the human hand model and object model, perform projection rendering, and optimize the Gaussian kernel parameters; 3) Based on the initialized hand-object interaction pose, transform the human hand model and object model to calculate the contact area in the observation space; combine the calculated contact area with the predicted contact area to optimize the pose parameters; Render and generate a rough image according to the optimized hand-object interaction pose; 4) Perform super-resolution reconstruction on the rough image to generate a clear image; 5) Fuse the generated clear image with the background image to obtain an augmented dataset of hand-object interaction containing background information.
2. The method according to claim 1, characterized in that, The steps of preprocessing in step 1) include: Segment the hand-object interaction color images from sparse viewpoints to extract the foreground images of hand-object interaction; Scale and translate the foreground images to obtain images with a unified resolution, and make the principal point of the internal parameters of the images located at the center point of the images.
3. The method according to claim 1, characterized in that, The steps of using the Gaussian splashing method to construct the human hand model and object model and perform projection rendering in step 2) include: Define Gaussian kernels on the human hand model and object model in the standard space, and calculate the center points and covariance matrices of the Gaussian kernels; Based on bone transformation and skinning weights, transform the Gaussian kernel of the human hand model in the standard space to the observation space; Based on the rigid transformation of the object pose, transform the Gaussian kernel of the object model in the standard space to the observation space; Project the 3D Gaussian kernel in the observation space into a 2D Gaussian kernel, and use the neural rendering method to render the images.
4. The method according to claim 1 or 2, characterized in that, In step 2), based on the total loss function $L$ composed of the L1 loss $L1$, the SSIM loss $L$ SSIM , and the distance regularization loss $L$ R , the parameters of the Gaussian kernel are optimized. The expression of the total loss function $L$ HO is as follows: HO L HO = (1 - λ SSIM )L1 + λ SSIM L SSIM + λ R L R Among them, λ SSIM , λ R are hyperparameters, and L1 and L SSIM are used to constrain the difference between the rendered image and the real image to be as small as possible; Loss function L R It is used to constrain the position of the vertices of the optimized object model to be as close as possible to the surface of the initial object model, and its calculation formula is as follows: Among them, N v represents the number of vertices of the object model, v i represents the vertex of the object model, M represents the initial object model, and dist(·) represents the distance from the vertex to the nearest patch.
5. The method according to claim 1, wherein The steps of initializing the hand-object interaction pose in step 3) include random rotation along each axis and hand-object translation.
6. The method according to claim 1, wherein In step 3), based on the initialized hand-object interaction pose, transform the human hand model and object model to obtain the vertex positions in the observation space, and then input them into ContactNet to predict the contact area.
7. The method according to claim 1, wherein In step 3), the total loss L C composed of the consistency loss L H , the contact loss L P , and the penetration loss L Poss is used to optimize the pose parameters and is expressed as follows: Poss :[[]]END]] Among them, λ C , λ H , λ P are hyperparameters, Ω represents the calculated contact area, represents the contact area predicted by ContactNet.
8. The method according to claim 1, wherein The steps of using the Encoder-Decoder deep learning model to perform super-resolution reconstruction on the rough image in step 4) include: Construct a super-resolution training set using the rough images and corresponding real images; Use the Encoder-Decoder model to perform feature encoding and decoding on the rough images in the super-resolution training set to generate super-resolution images; Use the generative adversarial network to train the Encoder-Decoder model, and use the discriminator to judge the authenticity of the generated images; Use the trained Encoder-Decoder model to perform super-resolution reconstruction on the rough images.
9. The method according to claim 8, characterized in that, When training the Encoder-Decoder model, the loss function used is composed of L1 loss, VGG loss, and generative adversarial loss.
10. A hand-object interaction dataset augmentation device based on Gaussian splashing, which is used to implement the method described in any one of claims 1-9, and is characterized in that, It includes: A data preprocessing module for collecting hand-object interaction color images from sparse viewpoints and performing preprocessing; A hand-object model construction module for constructing the human hand model and object model using the Gaussian splashing method according to the foreground images obtained by preprocessing, performing projection rendering, and optimizing the Gaussian kernel parameters; A pose optimization module, which is used to transform the human hand model and the object model based on the initialized hand-object interaction pose to calculate the contact area in the observation space; combine the calculated contact area with the predicted contact area to optimize the pose parameters; Render and generate a rough image according to the optimized hand-object interaction pose; A super-resolution module, which is used to perform super-resolution reconstruction on the rough image to generate a clear image; An augmented data processing module, which is used to fuse the generated clear image with the background image to obtain an augmented data set of hand-object interaction containing background information.
Citation Information
Cited By
Gaussian hand-object interaction rendering denoising method
CN120931520A