Virtual gesture generation method and system based on contact modeling and contact area alignment

By constructing a joint contact representation and contact area alignment mechanism, using a conditional variational autoencoder for joint modeling and contact area alignment constraints, the shortcomings of virtual gesture generation methods in the prior art in capturing contact details and contact area consistency are solved, and the accuracy and diversity of generated postures are improved.

CN119916946AActive Publication Date: 2025-05-02NANCHANG UNIV
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510397125.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-05-02
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

The existing virtual gesture generation methods have shortcomings in capturing detailed information during the contact process, resulting in the generated grab poses lacking fine modeling of contact mechanics and ergonomics, and ignore the consistency constraints of the contact area, resulting in the possibility of local contact incoherence or global pose unreasonable problems.

Method used

By constructing a mechanism for alignment between joint contact representation and contact area, joint modeling is performed using the contact position, contact part and contact direction in the opponent-object point cloud feature of the conditional variational autoencoder, a joint contact representation is generated, and the contact area is aligned and constrained at the global and local levels to optimize the hand posture parameters.

Benefits of technology

It improves the accuracy and reality of virtual gestures in complex interactive scenarios, enhances the fineness of modeling of contact mechanics and ergonomics, and ensures that the generated posture has higher rationality and adaptability in terms of physical laws and diversity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119916946A_ABST
    Figure CN119916946A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer vision, and discloses a virtual gesture generation method and system based on contact modeling and contact area alignment, and the method comprises the steps: obtaining and preprocessing a hand-object contact data sample, and carrying out the point cloud feature extraction of the hand-object contact data sample, and obtaining a hand-object point cloud initial feature; performing joint contact modeling on an object contact position, a hand contact position, a hand contact part and a hand contact direction in the hand-object point cloud initial feature by using a conditional variation auto-encoder to obtain a joint contact representation, and obtaining a hand-object predicted contact diagram by using a contact prediction model based on the joint contact representation; according to the hand-object prediction contact graph, alignment constraint is carried out on the contact areas of the hand and the object in the global level and the local level, the global posture parameters and the local joint posture parameters of the hand are optimized, and the final grabbing gesture is obtained. According to the method, the accuracy and the reality of the virtual gesture in a complex interaction scene can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a method and system for generating virtual gestures based on contact modeling and contact area alignment. Background Art

[0002] Virtual gesture generation is to infer different ways of human hand interaction with a given object, and is widely used in fields such as virtual reality, game development, and human-computer interaction. There are two main types of existing methods for static virtual gesture generation. One type is limited to modeling only the contact area on the surface of the object, and using the contact graph representation applied to the point cloud of the target object. For example, ContactOpt estimates the hand posture using an image-based method, infers the contact on the surface of the object by training the contact model based on real contact data, and gradually optimizes the hand posture; existing methods also propose to predict the contact area of ​​the object based on the point cloud features of the object, so as to constrain the contact area of ​​the human hand as close to the contact area of ​​the object as possible by reconstruction loss and penetration loss. The other type is that with the emergence of large-scale datasets, recent studies use generative models supervised by large-scale datasets to generate human hand grasping postures, and use contact analysis functions to encourage the hand contact points to be close to the object but not to penetrate each other.

[0003] Although the existing technology has made some progress in virtual gesture generation, it still has significant limitations. First, the existing contact modeling method is too single and only focuses on the contact area on the surface of the object, which cannot fully capture the detailed information of the contact process, such as the contact part of the hand (such as fingertips, palm) and its contact direction. This single contact representation results in the lack of detailed modeling of contact mechanics and ergonomics in the generated grasping posture, which may affect the authenticity and rationality of the posture. Secondly, although the method based on the generative model can generate realistic postures, it ignores the consistency constraints of the contact area. Specifically, these methods do not coordinately optimize the global posture generation (such as the overall posture of the hand) and the local contact area (such as the contact position between the finger and the object), resulting in the generated posture may have problems of local contact incoherence or unreasonable global posture. In addition, the existing methods are insufficient in modeling the diversity of contact areas, making it difficult to generate a variety of grasping methods that conform to daily usage habits. Summary of the invention

[0004] The purpose of the embodiments of the present application is to provide a method and system for generating virtual gestures based on contact modeling and contact area alignment. By constructing a joint contact representation and contact area alignment mechanism, the problems of insufficient authenticity and limited diversity of grasping postures caused by incomplete contact information representation and lack of local and global contact constraints in the prior art are solved, thereby improving the accuracy and realism of virtual gestures in complex interactive scenarios.

[0005] In order to solve the above technical problems, this application is implemented as follows: In a first aspect, an embodiment of the present application provides a method for generating a virtual gesture based on contact modeling and contact area alignment, the method comprising: Acquire and preprocess hand-object contact data samples, perform point cloud feature extraction on the hand-object contact data samples, and obtain initial features of the hand-object point cloud; Using a conditional variational autoencoder to perform joint contact modeling on the object contact position, hand contact position, hand contact part, and hand contact direction in the initial features of the hand-object point cloud to obtain a joint contact representation, and decoding the joint contact representation to obtain a hand-object predicted contact map; According to the hand-object predicted contact graph, the contact areas of the hand and the object are aligned and constrained at global and local levels, and the global posture parameters and local joint posture parameters of the hand are optimized to obtain the final grasping gesture.

[0006] As an optional implementation of the first aspect of the present application, the steps of acquiring and preprocessing hand-object contact data samples, performing point cloud feature extraction on the hand-object contact data samples, and obtaining initial features of the hand-object point cloud include: using the point cloud network model PointNet++ to extract the initial features of the point cloud of the hand-object contact data samples.

[0007] As an optional implementation of the first aspect of the present application, a conditional variational autoencoder is used to perform joint contact modeling on the object contact position, hand contact position, hand contact part and hand contact direction in the initial features of the hand-object point cloud, and in the step of obtaining a joint contact representation, the joint contact representation includes an object contact position map, a hand contact position map, a hand contact part map and a hand contact direction map, which are represented by the product of conditional distribution probabilities: ,in, Represents the contact position of the object Object Model is the conditional distribution probability of the conditional input, Diagram showing hand contact position Initialized hand and object models The conditional distribution probability of the conditional input is Diagram showing hand contact areas The object contact position diagram The conditional distribution probability as an additional conditional input, Hand contact direction diagram Additional hand contact area diagram The conditional distribution probability for the conditional input.

[0008] As an optional implementation of the first aspect of the present application, the method further includes: when the object model is used as a conditional input, a latent normalization process is performed through a KL minimization constraint to obtain a normal distribution conforming to the normal distribution. The latent distribution of .

[0009] As an optional implementation of the first aspect of the present application, in the step of using a conditional variational autoencoder to perform joint contact modeling on the object contact position, hand contact position, hand contact part and hand contact direction in the initial features of the hand-object point cloud, the training loss function corresponding to the conditional variational autoencoder includes: a contact map reconstruction loss, which is used to constrain the difference between the predicted contact map and the true contact map; a cross entropy KL loss, which is used to constrain the latent layer distribution to be a standard normal distribution; and a cosine loss in the direction dimension, which is used to constrain the consistency between the hand contact direction and the true direction.

[0010] As an optional implementation manner of the first aspect of the present application, according to the hand-object predicted contact map, the contact areas of the hand and the object are aligned and constrained at the global and local levels respectively, and the global posture parameters and local joint posture parameters of the hand are optimized to obtain the final grasping gesture. The steps include: in the global posture optimization stage, the global posture of the hand is obtained by minimizing the alignment loss of the difference between the hand contact position map and the object contact position map; in the local posture optimization stage, the local joint posture of the hand is obtained by minimizing the difference from the predicted direction and the difference from the part position using the hand contact direction map and the hand contact part map; the global posture and local joint posture are combined to align the contact areas of the hand and the object and reduce mutual penetration, so as to obtain a reasonable grasping gesture.

[0011] As an optional implementation of the first aspect of the present application, it also includes: based on the penetration loss of the signed distance field SDF, constraining the directed distance from the hand to the surface of the object so that the directed distance is greater than or equal to zero.

[0012] In a second aspect, an embodiment of the present application provides a virtual gesture generation system based on contact modeling and contact area alignment, the system comprising: A data preparation module is used to obtain and preprocess hand-object contact data samples, perform point cloud feature extraction on the hand-object contact data samples, and obtain initial features of the hand-object point cloud; a joint contact modeling module, for performing joint contact modeling on the object contact position, hand contact position, hand contact part and hand contact direction in the initial features of the hand-object point cloud using a conditional variational autoencoder to obtain a joint contact representation, and decoding the joint contact representation to obtain a hand-object predicted contact map; The contact area alignment module is used to align the contact areas of the hand and the object at global and local levels according to the hand-object predicted contact map, optimize the global posture parameters and local joint posture parameters of the hand, and obtain the final grasping gesture.

[0013] In a third aspect, an embodiment of the present application provides an electronic device, comprising a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the method described in the first aspect.

[0014] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0015] Compared with the prior art, the present invention proposes a virtual gesture generation method based on contact modeling and contact area alignment: by extracting point cloud features from hand-object contact data samples, the geometric structure and spatial distribution information of the hand and the object can be effectively captured, while removing noise and non-critical details, forming high-dimensional and low-redundancy initial features, and providing an accurate input basis for subsequent modeling; by using conditional variational autoencoders to jointly model the object contact position, hand contact position, contact part and contact direction, multi-dimensional contact information can be coupled in the latent variable space, and the physical unreasonable problem caused by the single-factor independent modeling of the traditional method is solved by generative modeling, and a hand-object predicted contact map that conforms to the physical law is generated during the decoding process, taking into account both generation diversity and physical consistency; based on the predicted contact map, contact area alignment constraints are imposed at the global level (such as matching the palm center of mass with the object center of mass) and the local level (such as aligning the fingertips with the surface curvature of the object), and the mechanical stability of the overall hand posture and the refined adaptation of the local joints are achieved by hierarchically optimizing the global posture parameters and the local joint parameters, which not only reduces the optimization difficulty of the high-degree-of-freedom hand model, but also improves the grasping adaptability to special-shaped objects. As a result, the complete process from feature extraction, contact modeling to hierarchical optimization forms a closed loop: the point cloud features provide structured representation for joint modeling, the generated contact map provides optimization targets for alignment constraints, and the hierarchical optimization results are fed back to the contact modeling process to enhance physical rationality, ultimately achieving efficient generation of virtual grasping gestures that conform to physical laws, are stable, and have strong adaptability. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a flow chart of a method for generating a virtual gesture based on contact modeling and contact area alignment provided by a first embodiment of the present invention; Figure 2 It is a structural schematic diagram of a virtual gesture generation system based on contact modeling and contact area alignment provided by the second embodiment of the present invention. DETAILED DESCRIPTION

[0017] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0018] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here. In addition, the "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally represents that the objects associated with each other are in an "or" relationship.

[0019] In order to illustrate the technical solution described in this application, a specific embodiment is provided below for illustration.

[0020] Example 1 See also Figure 1 , which is a flowchart of a virtual gesture generation method based on contact modeling and contact area alignment proposed in the first embodiment of the present application. The steps of the proposed method are as follows.

[0021] Step S01: Acquire and preprocess hand-object contact data samples, perform point cloud feature extraction on the hand-object contact data samples, and obtain initial hand-object point cloud features.

[0022] In the data preparation and preprocessing stage, first, any existing public or collected large-scale hand-object interaction dataset (such as the GRAB dataset, HO3D v2 dataset, etc.) can be used, which provides interaction samples with hand annotations and object models. Then, the point cloud data of the object and hand are converted separately, and the data is normalized to ensure that all samples have similar scale distributions. Finally, in the registered hand-object samples, the actual contact points between the object and the hand are determined by the proximity algorithm, so as to obtain the real contact map annotations, which are used as supervisory signals for subsequent model training.

[0023] It should be noted that the GRAB dataset is a dataset of whole-body grasping objects, which includes 10 different objects grasping the complete 3D shape and posture sequence of 51 objects. The present invention uses the official training set and test set segmentation protocol to complete the experiment of this embodiment. HO3D v2 collects video sequences annotated with hand-object interaction postures. Since it only contains a dozen objects, this dataset is only used for auxiliary testing to verify the generalization ability of the model.

[0024] In the feature extraction stage, PointNet++ or other point cloud feature extraction networks are used on the hand point cloud and the object point cloud to obtain the hand feature vector and the object feature vector. The feature vector retains the geometric shape and local detail information of the object, providing conditional input for subsequent contact prediction.

[0025] Step S02: Use a conditional variational autoencoder to perform joint contact modeling on the object contact position, hand contact position, hand contact part and hand contact direction in the initial features of the hand-object point cloud to obtain a joint contact representation, decode the joint contact representation, and obtain a hand-object predicted contact map.

[0026] Specifically, the joint contact representation proposed in this embodiment is , including 4 contact representations, namely object contact position diagram , hand contact position diagram , hand contact area diagram And hand contact direction map . The hand contact position diagram Use hand vertices Indicates that M The number of vertices of the hand is 778. The other contacts are represented by a series of points , sampled on the surface of the object, and N represents 2048 sampling points on the surface of the object.

[0027] Object contact position diagram , where each , which indicates the possibility of being touched at each point on the surface of the object, and its value range is , 0 means the object is not in contact with the hand at all, and 1 means it is in full contact with the hand. The object contact position map intuitively shows which positions on the object surface may contact the hand, but due to the high degree of freedom of the hand, it is far from enough to rely solely on the object contact position map to infer the hand contact position and the way the hand grasps the object.

[0028] Hand contact position diagram , where each , represents the possibility of being touched at each point on the surface of the hand, and the value range is , 0 means no contact with the object, and 1 means full contact with the object. The hand contact position diagram intuitively shows which positions on the hand surface will contact the object and the contact possibility.

[0029] Hand contact area diagram , where each , indicating the hand label that is closest to the object, B To divide the hand into B The number of parts is recorded as a set , 1 means the existing hand tag is in contact with the object, and 0 means no contact.

[0030] Hand contact direction map , where each represents a direction vector, indicating the direction of the center of the hand relative to the contact point on the surface of the object. In the method of the present invention, each hand part is regarded as a unit sphere, and the contact point position can be represented by the direction vector Determine, that is, starting from the center of the hand, along the ray The position of the contact point between the hand and the object surface can be determined by searching until the signed distance field SDF = 0.

[0031] In some embodiments, the method of the present invention uses a conditional variational autoencoder structure to model the contact map with multimodal uncertainty. Therefore, the joint contact representation includes an object contact position map, a hand contact position map, a hand contact part map, and a hand contact direction map, which are represented by the product of conditional distribution probabilities: in, Represents the contact position of the object Object Model is the conditional distribution probability of the conditional input, Diagram showing hand contact position Initialized hand and object models The conditional distribution probability of the conditional input is Diagram showing hand contact areas The object contact position diagram The conditional distribution probability as an additional conditional input, Hand contact direction diagram Additional hand contact area diagram The conditional distribution probability for the conditional input.

[0032] Specifically, for the contact encoder, the object model is used as the original conditional input to obtain the latent distribution that conforms to the normal distribution , the probability of each contact graph distribution is , , and .in, represents the latent distribution of the contact position of the object, is the latent distribution of hand contact positions, represents the latent distribution of hand contact locations, Represents the latent distribution of hand contact direction. , represents the mean and standard deviation of the latent distribution of the object contact position, , represents the mean and standard deviation of the latent distribution of hand contact positions, , represents the mean and standard deviation of the latent distribution of hand contact locations, , Represents the mean and standard deviation of the latent distribution of hand contact directions.

[0033] For the latent normalization processing, the contact latent distribution obtained by the contact encoder is constrained by Kulback-Leibler divergence to make it as close to the standard normal distribution as possible. , and its calculation formula is as follows: in, represents the KL divergence constraint.

[0034] For the contact decoder, the output predicted contact map is given by the following formula: in, , , and They represent the relevant contact graph decoder, , , and This is the contact map corresponding to the prediction.

[0035] In some embodiments, the training loss function corresponding to the conditional variational autoencoder includes: a contact map reconstruction loss, which is used to constrain the difference between the predicted contact map and the true contact map; a cross entropy KL loss, which is used to constrain the latent layer distribution to be a standard normal distribution; and a cosine loss in the direction dimension, which is used to constrain the consistency between the hand contact direction and the true direction.

[0036] Specifically, the entire joint contact representation is based on the encoder-decoder architecture of the conditional variational autoencoder (i.e., the predictive contact prediction model), taking the object model as a condition, and the complete loss function is shown in the following formula: in, To calculate the loss between the predicted contact map and the true contact map, represents the weight value of the KL divergence constraint, Represents the KL divergence constraint; the specific definition is shown in the following formula: in, , and is the weight value, is the cross entropy loss, It represents the cosine loss in the direction dimension.

[0037] Step S03: According to the hand-object predicted contact map, the contact areas of the hand and the object are aligned and constrained at the global and local levels, and the global posture parameters and local joint posture parameters of the hand are optimized to obtain the final grasping gesture.

[0038] In some embodiments, in the global posture optimization stage, the global posture of the hand is obtained by minimizing the alignment loss of the difference between the hand contact position map and the object contact position map; in the local posture optimization stage, the local joint posture of the hand is obtained by minimizing the difference from the predicted direction and the difference from the part position using the hand contact direction map and the hand contact part map; the global posture and the local joint posture are combined to align the contact area of ​​the hand and the object and reduce mutual penetration, so as to obtain a reasonable grasping gesture.

[0039] Specifically, the above hand and object contact position map is predicted , , and The present invention proposes to use the hand-object contact area alignment method to respectively segment the hand joint model parameter and ,Right now Optimize. Specifically, the whole process is divided into the optimization of global posture parameters and local posture, shape and rotation parameters. The grasping generation model first initializes various parameters, obtains the initialized hand according to the MANO model and calculates the initial hand-to-object signed distance SDF according to the segmented hand joint model, and gradually iterates and optimizes the hand parameters through the contact graph through the hand-object contact area alignment method proposed below.

[0040] Optimize the alignment of object contact positions. For a given object model, for the object sampling points, calculate the signed distance SDF from the object to each hand part, and encourage the SDF of each point to approach 0, so that the hand is as close to the object contact position as possible. The calculation is as follows: in, Indicates the optimization of the contact position graph of the object, To predict the contact position map of the object, Represents the predicted contact map of the bth part of the hand.

[0041] For a given object model, in order to generate an accurate grasping posture for the object contact position map, the segmented hand joint model and the object model Object contact position map calculated by proximity algorithm And the object contact position map obtained by the contact prediction model The alignment should be maintained as shown in the following formula: in, The contact position map between the hand and the object is calculated by the proximity algorithm. Represents contact map alignment optimization.

[0042] Optimize the alignment of the hand contact points. Calculate the SDF of the signed distance from the hand to the object, as shown in the following formula: in, Indicates that the hand contact area is optimized. This is a diagram of the hand contact areas.

[0043] Hand contact direction alignment optimization: Encourage each hand contact direction to be aligned with the predicted contact direction as much as possible, calculated as shown in the following formula: in, This means that the hand contact direction is optimized. Represents the assignment weight.

[0044] Hand-object penetration optimization: To optimize the distance between the generated grasping gesture and the object, and prevent the generated hand gesture from penetrating the object, the specific definition is shown in the following formula: in, Represents the signed distance of each hand part relative to the object.

[0045] Regularization optimization: Finally, in order to reduce the complexity of the model, the hand parameters and Regularized optimization is performed, and the calculation formula is shown as follows: in, is the regularization optimization term.

[0046] The overall optimization function is shown in the following formula: The optimization function is applied to the optimization process of the global posture parameters and local posture, shape and rotation parameters of the gesture to obtain the final diversified and rationalized grasping gestures.

[0047] The evaluation indicators used in the present invention are physical rationality, stability and diversity: In order to evaluate the rationality of the generated grasping posture, two indicators are used: the mutual penetration volume and contact ratio of the hand and object model. Specifically, the mesh model of the hand and the object is converted into voxels composed of multiple 1 cubic millimeter cubes, and the mutual penetration volume is calculated by accurately measuring the overlapping voxels between the two. In addition, the proportion of grasping actions in which the hand contacts the object to all generated actions is taken as the contact ratio. Among them, the smaller the penetration volume and the larger the contact ratio, the more reasonable the generated grasping posture is.

[0048] To evaluate the stability of the generated gesture, place the generated grasping gesture and the object into a gravity simulator at the same time, and measure the average displacement of the object under the influence of gravity in the simulator. The smaller the average displacement, the more stable the generated grasping gesture.

[0049] In order to evaluate the diversity of the generated grasping postures, the generated grasping gestures are divided into 20 clusters using the K-means clustering method, and the type entropy and cluster size in each cluster are measured to evaluate the diversity of the grasping postures. Among them, the larger the type entropy value and the larger the cluster size, the better the diversity of the generated grasping postures.

[0050] As for experimental details and parameter settings, all experiments of the present invention were carried out on the same computer, and the detailed hardware and software configuration information of the computer is shown in Table 1.

[0051] Table 1 Hardware and software configuration information The parameters are set as follows: (1) Contact prediction model: During the training phase, random rotations are applied to the input objects for data augmentation to simulate possible natural changes. The batch size is set to 32, the standard Adam optimizer is used, and the initial learning rate is set to 1.6×10 -3, a total of 3000 rounds of training, where the number of sampling points of the input object model ,Will Decay from 0 to 5×10 -2 .

[0052] (2) Grasping Generative Model: For the grasping generative model, the optimization process is divided into two stages. In the first global pose translation and rotation optimization stage, the learning rate is set to 5×10 -2 , optimization iterations 200 times. In the second local pose optimization phase, the learning rate is 5×10 -3 , iterative optimization 1000 times. Among them, the alignment coefficient =0.01, object contact position diagram coefficient Contact position diagram with hands Set to 0.01, the penetration coefficient and direction contact map coefficient .

[0053] Example 2 See also Figure 2 , which is a structural schematic diagram of a virtual gesture generation system based on contact modeling and contact area alignment proposed in the second embodiment of the present application, wherein the system comprises: The data preparation module 100 is used to obtain and pre-process the hand-object contact data samples, perform point cloud feature extraction on the hand-object contact data samples, and obtain the initial features of the hand-object point cloud; The joint contact modeling module 200 is used to perform joint contact modeling on the object contact position, hand contact position, hand contact part and hand contact direction in the initial features of the hand-object point cloud using a conditional variational autoencoder to obtain a joint contact representation, and based on the joint contact representation, obtain a hand-object predicted contact map using a contact prediction model; The contact area alignment module 300 is used to align the contact areas of the hand and the object at both global and local levels according to the hand-object predicted contact map, optimize the global posture parameters and local joint posture parameters of the hand, and obtain the final grasping gesture.

[0054] A virtual gesture generation system based on contact modeling and contact area alignment in an embodiment of the present application may be a device, or a component, integrated circuit, or chip in a terminal. The device may be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device may be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device may be a server, a network attached storage (NAS), a personal computer (PC), etc., which is not specifically limited in the embodiment of the present application.

[0055] In the embodiment of the present application, a device for a virtual gesture generation system based on contact modeling and contact area alignment is provided. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.

[0056] The virtual gesture generation system based on contact modeling and contact area alignment provided in the embodiment of the present application can achieve Figure 1 In the method embodiment, various processes of a virtual gesture generation method based on contact modeling and contact area alignment are implemented, and to avoid repetition, they are not described here.

[0057] Optionally, an embodiment of the present application also provides an electronic device, including a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, each process of an embodiment of a virtual gesture generation method based on contact modeling and contact area alignment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.

[0058] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the embodiment of the above-mentioned virtual gesture generation method based on contact modeling and contact area alignment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0059] The processor is a processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0060] It should be noted that, in this article, the term "comprises", "includes" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "including one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the method and device in the embodiment of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0061] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0062] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

Claims

1. A virtual gesture generation method based on contact modeling and contact area alignment, characterized in that: include: Acquire and preprocess hand-object contact data samples, perform point cloud feature extraction on the hand-object contact data samples, and obtain initial features of the hand-object point cloud; Using a conditional variational autoencoder to perform joint contact modeling on the object contact position, hand contact position, hand contact part, and hand contact direction in the initial features of the hand-object point cloud to obtain a joint contact representation, and decoding the joint contact representation to obtain a hand-object predicted contact map; According to the hand-object predicted contact graph, the contact areas of the hand and the object are aligned and constrained at global and local levels, and the global posture parameters and local joint posture parameters of the hand are optimized to obtain the final grasping gesture.

2. A virtual gesture generation method based on contact modeling and contact area alignment according to claim 1, characterized in that: The steps of acquiring and preprocessing the hand-object contact data samples, extracting point cloud features from the hand-object contact data samples, and obtaining the initial features of the hand-object point cloud include: The point cloud network model PointNet++ is used to extract the initial point cloud features of hand-object contact data samples.

3. The method for generating virtual gestures based on contact modeling and contact area alignment according to claim 1, characterized in that: In the step of using a conditional variational autoencoder to perform joint contact modeling on the object contact position, hand contact position, hand contact part and hand contact direction in the initial features of the hand-object point cloud to obtain a joint contact representation, the joint contact representation includes an object contact position map, a hand contact position map, a hand contact part map and a hand contact direction map, which are represented by the product of conditional distribution probabilities: in, Represents the contact position of the object Object Model is the conditional distribution probability of the conditional input, Diagram showing hand contact position Initialized hand and object models The conditional distribution probability of the conditional input is Diagram showing hand contact areas The object contact position diagram The conditional distribution probability as an additional conditional input, Hand contact direction diagram Additional hand contact area diagram The conditional distribution probability for the conditional input.

4. The method for generating virtual gestures based on contact modeling and contact area alignment according to claim 3, characterized in that: Also includes: When the object model is used as a conditional input, the latent layer is normalized through the KL minimization constraint to obtain a normal distribution. The latent distribution of .

5. The method for generating virtual gestures based on contact modeling and contact area alignment according to claim 4, characterized in that: In the step of performing joint contact modeling on the object contact position, hand contact position, hand contact part and hand contact direction in the initial features of the hand-object point cloud using a conditional variational autoencoder, the training loss function corresponding to the conditional variational autoencoder includes: The contact map reconstruction loss is used to constrain the difference between the predicted contact map and the true contact map; Cross entropy KL loss, used to constrain the latent layer distribution to be a standard normal distribution; The cosine loss of the direction dimension is used to constrain the consistency between the hand contact direction and the true direction.

6. A virtual gesture generation method based on contact modeling and contact area alignment according to any one of claims 1 to 5, characterized in that: According to the hand-object predicted contact graph, the contact areas of the hand and the object are aligned and constrained at global and local levels, and the global posture parameters and local joint posture parameters of the hand are optimized to obtain the final grasping gesture, including the following steps: In the global pose optimization stage, the global pose of the hand is obtained by minimizing the alignment loss of the difference between the hand contact position map and the object contact position map; In the local posture optimization stage, the hand contact direction map and the hand contact part map are used to obtain the local joint posture of the hand by minimizing the difference with the predicted direction and the difference with the part position; The global pose and local joint pose are combined to align the contact area between the hand and the object and reduce mutual penetration, thus obtaining a reasonable grasping gesture.

7. The method for generating virtual gestures based on contact modeling and contact area alignment according to claim 6, characterized in that: Also includes: The penetration loss based on the signed distance field SDF is used to constrain the directed distance from the hand part to the object surface so that the directed distance is greater than or equal to zero.

8. A virtual gesture generation system based on contact modeling and contact area alignment, characterized in that: The system comprises: A data preparation module is used to obtain and preprocess hand-object contact data samples, perform point cloud feature extraction on the hand-object contact data samples, and obtain initial features of the hand-object point cloud; a joint contact modeling module, for performing joint contact modeling on the object contact position, hand contact position, hand contact part and hand contact direction in the initial features of the hand-object point cloud using a conditional variational autoencoder to obtain a joint contact representation, and decoding the joint contact representation to obtain a hand-object predicted contact map; The contact area alignment module is used to align the contact areas of the hand and the object at global and local levels according to the hand-object predicted contact map, optimize the global posture parameters and local joint posture parameters of the hand, and obtain the final grasping gesture.

9. An electronic device, characterized in that: The method comprises a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of a method for generating a virtual gesture based on contact modeling and contact area alignment as described in any one of claims 1 to 7 are implemented.

10. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of a method for generating a virtual gesture based on contact modeling and contact area alignment as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Hand posture estimation and tracking method based on depth data

    CN110286749A

  • Multi-fingered dexterous hand grabbing gesture planning method based on deep neural network

    CN114643586A

  • Method and system for generating grabbing posture of dexterous hand guided by limited grabable area

    CN116704160A

  • Track optimization method for grabbing operation of dexterous manipulator

    CN117773922A

  • Intra-class multi-specification tool oriented dexterous hand grabbing posture generation method

    CN119089989A