Virtual Gesture Generation Method and System Based on Contact Modeling and Contact Area Alignment

By constructing a joint contact representation and contact area alignment mechanism in virtual gesture generation technology, using conditional variational autoencoder for joint modeling and aligning constraints on the contact area, the problems of incomplete contact information representation and missing local and global contact constraints in the prior art are solved, and more efficient virtual gesture generation is achieved, improving the authenticity and diversity of the posture.

CN119916946BActive Publication Date: 2025-06-10NANCHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510397125.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-06-10
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

The existing virtual gesture generation technology has shortcomings in capturing detailed information during the contact process and collaborating on optimizing global and local contact areas, resulting in limited authenticity and diversity of generated grab poses.

Method used

By constructing a mechanism for alignment between joint contact representation and contact area, joint modeling is performed using the conditional variational autoencoder opponent-object point cloud features to generate joint contact representations, and alignment constraints are performed on the contact area at the global and local levels to optimize hand posture parameters.

Benefits of technology

Improve the accuracy and reality of virtual gestures in complex interactive scenarios, enhance the fine modeling ability of contact mechanics and ergonomics, and improve the diversity and physical consistency of generated grasping postures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119916946B_ABST
    Figure CN119916946B_ABST
Patent Text Reader

Abstract

This application belongs to the field of computer vision technology and discloses a virtual gesture generation method and system based on contact modeling and contact area alignment. The method includes: acquiring and preprocessing hand-object contact data samples, extracting point cloud features from the hand-object contact data samples to obtain initial hand-object point cloud features; using a conditional variational autoencoder to perform joint contact modeling on the object contact position, hand contact position, hand contact part, and hand contact direction in the initial hand-object point cloud features to obtain a joint contact representation, and based on the joint contact representation, using a contact prediction model to obtain a hand-object predicted contact map; according to the hand-object predicted contact map, performing alignment constraints on the contact areas of the hand and the object at both the global and local levels, optimizing the global pose parameters and local joint pose parameters of the hand to obtain the final grasping gesture. This method can improve the accuracy and realism of virtual gestures in complex interaction scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly relates to a virtual gesture generation method and system based on contact modeling and contact area alignment. Background Art

[0002] Virtual gesture generation infers different ways of hand interaction with a given object and is widely applied in fields such as virtual reality, game development, and human-computer interaction. For static virtual gesture generation, there are mainly two categories of existing methods. One category is limited to only modeling the contact area on the surface of the object for use in the contact map representation of the target object point cloud. For example, ContactOpt uses an image-based method to estimate the hand pose, infers the contact on the object surface by training a contact model with real contact data, and gradually optimizes the hand pose. There are also existing methods that propose predicting the contact area of the object based on the object point cloud features, using constraints such as reconstruction loss and penetration loss to make the contact area of the human hand as close as possible to the object contact area. The other category is with the emergence of large-scale datasets, recent research uses generative models supervised by large-scale datasets to generate human hand grasping postures, and at the same time uses a contact analysis function to encourage the hand contact points to be close to the object but not penetrate each other.

[0003] Although the existing technologies have made certain progress in virtual gesture generation, there are still significant limitations. First, the existing contact modeling methods are too single, only focusing on the contact area on the surface of the object, and unable to fully capture the detailed information in the contact process, such as the hand contact parts (such as fingertips, palm) and their contact directions. This single contact representation results in the generated grasping postures lacking fine modeling of contact mechanics and ergonomics, which may affect the authenticity and rationality of the postures. Second, although the methods based on generative models can generate realistic postures, they ignore the consistency constraints of the contact area. Specifically, these methods do not co-optimize the global posture generation (such as the overall hand posture) and the local contact area (such as the contact position of the finger with the object), resulting in problems such as local contact incoherence or unreasonable global postures in the generated postures. In addition, the existing methods have insufficient modeling ability for the diversity of the contact area and are difficult to generate various grasping methods that conform to daily usage habits. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to provide a virtual gesture generation method and system based on contact modeling and contact area alignment. By constructing a joint contact representation and a contact area alignment mechanism, the problems of insufficient authenticity of grasping postures and limited diversity caused by incomplete representation of contact information and lack of local and global contact constraints in the existing technologies are solved, thereby improving the accuracy and realism of virtual gestures in complex interaction scenarios.

[0005] To solve the above technical problems, the present application is implemented as follows:

[0006] In a first aspect, an embodiment of the present application provides a virtual gesture generation method based on contact modeling and contact area alignment. The method includes:

[0007] Obtain and preprocess hand-object contact data samples, extract point cloud features from the hand-object contact data samples, and obtain initial hand-object point cloud features;

[0008] Use a conditional variational autoencoder to perform joint contact modeling on the object contact position, hand contact position, hand contact part, and hand contact direction in the initial hand-object point cloud features to obtain a joint contact representation, and decode the joint contact representation to obtain a predicted hand-object contact map;

[0009] According to the predicted hand-object contact map, align and constrain the contact areas of the hand and the object at both the global and local levels, optimize the global pose parameters and local joint pose parameters of the hand, and obtain the final grasping gesture.

[0010] As an alternative implementation of the first aspect of the present application, the steps of obtaining and preprocessing hand-object contact data samples and extracting point cloud features from the hand-object contact data samples to obtain initial hand-object point cloud features include: using the point cloud network model PointNet++ to extract the initial hand-object point cloud features of the hand-object contact data samples.

[0011] As an alternative implementation of the first aspect of the present application, in the step of using a conditional variational autoencoder to perform joint contact modeling on the object contact position, hand contact position, hand contact part, and hand contact direction in the initial hand-object point cloud features to obtain a joint contact representation, the joint contact representation includes an object contact position map, a hand contact position map, a hand contact part map, and a hand contact direction map, and is represented by the product of conditional distribution probabilities: , where represents the object contact position map is the conditional distribution probability with the object model as the conditional input, represents the hand contact position map is the conditional distribution probability with the initialized hand and the object model jointly as the conditional input, represents the hand contact part map is the conditional distribution probability with the object contact position map as an additional conditional input, represents the hand contact direction map is the conditional distribution probability with the additional hand contact part map as the conditional input.

[0012] As an alternative implementation of the first aspect of the present application, it further includes: when an object model is input as a condition, latent normalization processing is performed through KL minimization constraint to obtain a latent distribution that conforms to a normal distribution of the latent layer.

[0013] As an alternative implementation of the first aspect of the present application, in the step of jointly modeling the object contact position, hand contact position, hand contact part, and hand contact direction in the hand-object point cloud initial features by using a conditional variational autoencoder, the training loss function corresponding to the conditional variational autoencoder includes: a contact map reconstruction loss for constraining the difference between the predicted contact map and the true contact map; a cross-entropy KL loss for constraining the latent distribution to be a standard normal distribution; and a cosine loss in the direction dimension for constraining the consistency between the hand contact direction and the true direction.

[0014] As an alternative implementation of the first aspect of the present application, according to the hand-object predicted contact map, aligning constraints are respectively performed on the contact areas of the hand and the object at the global and local levels, and the global pose parameters and local joint pose parameters of the hand are optimized to obtain the final grasping gesture. The steps include: in the global pose optimization stage, an alignment loss that minimizes the difference between the hand contact position map and the object contact position map is used to obtain the global pose of the hand; in the local pose optimization stage, the hand contact direction map and the hand contact part map are used to obtain the local joint pose of the hand by minimizing the difference from the predicted direction and the difference in part position; the global pose and the local joint pose are combined to align the contact areas of the hand and the object and reduce mutual penetration to obtain a reasonable grasping gesture.

[0015] As an alternative implementation of the first aspect of the present application, it further includes: a penetration loss based on the signed distance field SDF for constraining the directed distance from the hand part to the object surface to make the directed distance greater than or equal to zero.

[0016] In a second aspect, an embodiment of the present application provides a virtual gesture generation system based on contact modeling and contact area alignment. The system includes:

[0017] A data preparation module for acquiring and preprocessing hand-object contact data samples, extracting point cloud features from the hand-object contact data samples to obtain hand-object point cloud initial features;

[0018] A joint contact modeling module for jointly modeling the object contact position, hand contact position, hand contact part, and hand contact direction in the hand-object point cloud initial features by using a conditional variational autoencoder to obtain a joint contact representation, and decoding the joint contact representation to obtain a hand-object predicted contact map;

[0019] A contact area alignment module, which is used to perform alignment constraints on the contact areas of the hand and the object at both the global and local levels according to the hand-object predicted contact map, optimize the global pose parameters and local joint pose parameters of the hand, and obtain the final grasping gesture.

[0020] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0021] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0022] Compared with the prior art, the present invention proposes a virtual gesture generation method based on contact modeling and contact area alignment: by extracting point cloud features from hand-object contact data samples, it can effectively capture the geometric structures and spatial distribution information of the hand and the object, while removing noise and non-critical details, forming initial features with high dimension and low redundancy, providing an accurate input basis for subsequent modeling; using a conditional variational autoencoder to jointly model the object contact position, hand contact position, contact part, and contact direction, it can couple multi-dimensional contact information in the latent variable space, solve the physically unreasonable problems caused by single-factor independent modeling in traditional methods through generative modeling, and generate a hand-object predicted contact map that conforms to physical laws during the decoding process, taking into account both generation diversity and physical consistency; based on the predicted contact map, contact area alignment constraints are imposed at the global level (such as matching the palm centroid with the object centroid) and the local level (such as aligning the fingertip with the object surface curvature), and by hierarchically optimizing the global pose parameters and local joint parameters, the mechanical stability of the overall hand posture and the refined adaptation of local joints are achieved, which not only reduces the optimization difficulty of the high-degree-of-freedom hand model but also improves the grasping adaptability to irregular objects. Thus, a complete process from feature extraction, contact modeling to hierarchical optimization forms a closed loop: the point cloud features provide a structured representation for joint modeling, the generated contact map provides an optimization target for alignment constraints, and the hierarchical optimization results are fed back to the contact modeling process to enhance physical rationality, ultimately achieving efficient generation of virtual grasping gestures that conform to physical laws, are stable, and have strong adaptability. Description of the Drawings

[0023] Figure 1 is a flowchart of a virtual gesture generation method based on contact modeling and contact area alignment provided by the first embodiment of the present invention;

[0024] Figure 2 is a schematic structural diagram of a virtual gesture generation system based on contact modeling and contact area alignment provided by the second embodiment of the present invention. Detailed implementation manners

[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0026] The terms "first", "second", etc. in the description and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein. In addition, "and / or" in the description and claims means at least one of the connected objects. The character " / " generally indicates that the related objects before and after are in an "or" relationship.

[0027] In order to illustrate the technical solutions described in the present application, the following will be described through specific embodiments.

[0028] Embodiment 1

[0029] Please refer to Figure 1 , which is a flowchart of a virtual gesture generation method based on contact modeling and contact area alignment proposed in the first embodiment of the present application. The steps of the proposed method are as follows.

[0030] Step S01: Obtain and preprocess the hand-object contact data sample, extract the point cloud features of the hand-object contact data sample, and obtain the initial hand-object point cloud features.

[0031] In the data preparation and preprocessing stage, first, any existing publicly available or collected large-scale hand-object interaction datasets (such as the GRAB dataset, HO3D v2 dataset, etc.) can be used. This dataset provides interaction samples with hand annotations and object models. Then, the point cloud data of the object and the hand are respectively converted, and the data is normalized to ensure that all samples have a similar scale distribution. Finally, in the registered hand-object samples, the actual contact points between the object and the hand are determined through the proximity algorithm, so as to obtain the real contact icon annotation, which is used as a supervision signal for subsequent model training.

[0032] It should be noted that the GRAB dataset is a dataset for whole-body object grasping, which includes the complete 3D shapes and pose sequences of 51 objects to be grasped by 10 different objects. The experiments of this embodiment are completed using the official training set and test set splitting protocol. HO3D v2 collects video sequences with hand-object interaction pose annotations. Since it only contains a dozen object objects, this dataset is only used for auxiliary testing to verify the generalization ability of the model.

[0033] In the feature extraction stage, the PointNet++ or other point cloud feature extraction networks are respectively used for the hand point cloud and the object point cloud to obtain the hand feature vector and the object feature vector. This feature vector retains the geometric shape and local detail information of the object, providing conditional input for subsequent contact prediction.

[0034] Step S02: Use the conditional variational autoencoder to perform joint contact modeling on the object contact position, hand contact position, hand contact part, and hand contact direction in the initial features of the hand-object point cloud to obtain the joint contact representation, and decode the joint contact representation to obtain the hand-object predicted contact map.

[0035] Specifically, the joint contact representation proposed in this embodiment , includes 4 contact representations, namely the object contact position map , the hand contact position map , the hand contact part map and the hand contact direction map . Among them, the hand contact position map is represented by the hand vertices , where M is the number of hand vertices 778. The other contact representations are a series of points , sampled on the object surface, and N represents 2048 object surface sampling points.

[0036] The object contact position map , where each , represents the possibility of each point on the object surface being contacted, and the value range is , 0 means the object is not in contact with the hand at all, and 1 means it is in complete contact with the hand. The object contact position map intuitively shows which positions on the object surface may be in contact with the hand. However, due to the high degree of freedom of the hand, it is far from sufficient to infer the hand contact position and the way the hand grasps the object only relying on the object contact position map.

[0037] The hand contact position map , where each , represents the possibility of each point on the hand surface being contacted, and the value range is , 0 indicates no contact with the object, and 1 indicates full contact with the object. The hand contact position map intuitively shows which positions on the surface of the hand will come into contact with the object and the contact possibility.

[0038] Hand contact part map , where each , represents the hand label closest to the contact with the object, B is to divide the hand into B number of parts, denoted as the set , 1 indicates that the existing hand label is in contact with the object, and 0 indicates no contact.

[0039] Hand contact direction map , where each represents a direction vector, indicating the direction of the center of the hand part relative to the contact point on the object surface. In the method of the present invention, each hand part is regarded as a unit sphere, and the contact point position can be determined by the direction vector , that is, starting from the center of the hand part, searching along the ray until the signed distance field SDF = 0, then the contact point position between the hand part and the object surface can be determined.

[0040] In some embodiments, the method of the present invention adopts a conditional variational autoencoder structure to perform multi-modal uncertainty modeling on the contact map. Therefore, the joint contact representation includes the object contact position map, the hand contact position map, the hand contact part map, and the hand contact direction map, which is represented by the product of conditional distribution probabilities:

[0041]

[0042] Among them, represents the conditional distribution probability of the object contact position map conditioned on the object model , represents the conditional distribution probability of the hand contact position map conditioned on the initialized hand and the object model jointly as the conditional input, represents the conditional distribution probability of the hand contact part map which is the conditional distribution probability with the object contact position map as an additional conditional input, represents the conditional distribution probability of the hand contact direction map conditioned on the additional hand contact part map .

[0043] Specifically, for the contact encoder, taking the object model as the original conditional input to obtain a latent distribution , each contact map distribution probability is , , and . Among them, represents the latent distribution of the object contact position, is the latent distribution of the hand contact position, represents the latent distribution of the hand contact part, represents the latent distribution of the hand contact direction. , represent the mean and standard deviation of the latent distribution of the object contact position, , represent the mean and standard deviation of the latent distribution of the hand contact position, , represent the mean and standard deviation of the latent distribution of the hand contact part, , represent the mean and standard deviation of the latent distribution of the hand contact direction.

[0044] For the latent normalization process, the contact latent distributions obtained by the contact encoder are respectively constrained by the Kulback-Leibler divergence to make them as close as possible to the standard normal distribution , and its calculation formula is as follows:

[0045]

[0046] Among them, represents the KL divergence constraint.

[0047] For the contact decoder, the output predicted contact map is as shown in the following formula:

[0048]

[0049]

[0050]

[0051]

[0052] Among them, , , and respectively represent the relevant contact map decoders, , , and are the predicted contact maps obtained correspondingly.

[0053] In some embodiments, the training loss function corresponding to the conditional variational autoencoder includes: a contact map reconstruction loss for constraining the difference between the predicted contact map and the true contact map; a cross-entropy KL loss for constraining the latent distribution to be a standard normal distribution; and a cosine loss in the direction dimension for constraining the consistency between the hand contact direction and the true direction.

[0054] Specifically, the entire joint contact representation is based on the encoder-decoder architecture of the conditional variational autoencoder (i.e., the predicted contact prediction model), taking the object model as a condition. The complete loss function is shown in the following formula:

[0055]

[0056] where is the loss between the predicted contact map and the true contact map, represents the weight value of the KL divergence constraint, represents the KL divergence constraint; the specific definition is shown in the following formula:

[0057]

[0058] where , and are weight values, is the cross-entropy loss, then represents the cosine loss in the direction dimension.

[0059] Step S03: According to the hand-object predicted contact map, perform alignment constraints on the contact areas of the hand and the object at both the global and local levels, optimize the global pose parameters and local joint pose parameters of the hand, and obtain the final grasping gesture.

[0060] In some embodiments, in the global pose optimization stage, the global pose of the hand is obtained by minimizing the alignment loss between the hand contact position map and the object contact position map; in the local pose optimization stage, using the hand contact direction map and the hand contact part map, the local joint pose of the hand is obtained by minimizing the difference from the predicted direction and the difference in part positions; combining the global pose and the local joint pose, align the contact area between the hand and the object and reduce mutual penetration to obtain a reasonable grasping gesture.

[0061] Specifically, the above-mentioned hand and object contact position maps , , and are predicted. The present invention proposes to use the hand-object contact area alignment method to separately optimize the parameters of the segmented hand joint model and , that is, Optimize. Specifically, the whole process is divided into the optimization of global pose parameters and local pose, shape and rotation parameters. The grasping generation model first initializes various parameters, calculates the initial directed distance SDF from the hand to the object based on the MANO model and the initialized hand and segmented hand joint models, and gradually iteratively optimizes the hand parameters through the hand-object contact area alignment method proposed below by the contact map.

[0062] Object contact position alignment optimization. For a given object model, for the object sampling points, calculate the directed distance SDF from the object to each hand part, and encourage the SDF of each point to tend to 0, so that the hand gets as close as possible to the object contact position. The calculation is as follows:

[0063]

[0064] Among them, represents the optimization of the object contact position map, is the predicted object contact position map, represents the predicted contact map of the b-th part of the hand.

[0065] For a given object model, for the object contact position map, to generate an accurate grasping pose, the object contact position map calculated by the segmented hand joint model and the object model through the nearest neighbor algorithm and the object contact position map obtained by the contact prediction model should be kept aligned, as shown in the following formula:

[0066]

[0067]

[0068] Among them, is the nearest neighbor algorithm, calculating the contact position map between the generated hand and the object, represents the optimization of the contact map alignment.

[0069] Hand contact part alignment optimization. Calculate the directed distance SDF from the hand to the object, and the calculation is as follows:

[0070]

[0071] Among them, represents the optimization of the hand contact part, is the hand contact part map.

[0072] Hand contact direction alignment optimization: Encourage each hand contact direction to be as aligned as possible with the predicted contact direction, and the calculation is as follows:

[0073]

[0074] Among them, indicates the optimization of the hand contact direction, indicating the assignment weight.

[0075] Hand-object penetration optimization: It is the optimization of the distance between the generated grasping gesture and the object, preventing the generated hand posture from penetrating the object. The specific definition is shown in the following formula:

[0076]

[0077] Among them, represents the directed distance of each hand part relative to the object.

[0078] Regularization optimization: Finally, in order to reduce the complexity of the model, the hand parameters and are optimized by regularization. The calculation formula is shown in the following formula:

[0079]

[0080] Among them, is the regularization optimization term.

[0081] The overall optimization function is shown in the following formula:

[0082]

[0083] Applying the optimization function to the optimization process of the global pose parameters and local pose, shape and rotation parameters of the gesture, the final diverse and reasonable grasping gestures are obtained.

[0084] For the evaluation metrics used in the present invention, they are physical rationality, stability, and diversity respectively:

[0085] In order to evaluate the rationality of the generated grasping postures, two metrics, the mutual penetration volume and the contact ratio of the hand and object models, are adopted. Specifically, the mesh models of the hand and the object are converted into voxels composed of multiple 1 cubic millimeter cubes, and the mutually penetrating volume is calculated by accurately measuring the overlapping voxels between the two. In addition, the proportion of the grasping actions where the hand contacts the object among all the generated actions is used as the contact ratio. Among them, the smaller the penetration volume and the larger the contact ratio, the more reasonable the generated grasping posture.

[0086] To evaluate the stability of the generated postures, the generated grasping gestures and the object are simultaneously placed in a mock gravity simulator, and the magnitude of the average displacement of the object under the influence of gravity in the simulator is measured. The smaller the average displacement, the more stable the generated grasping posture.

[0087] To evaluate the diversity of generated grasping postures, the K-means clustering method is used to divide the generated grasping gestures into 20 clusters, and the class entropy and cluster size in each cluster are measured to evaluate the diversity of grasping postures. Among them, the larger the class entropy value and the larger the cluster size, the better the diversity of the generated grasping postures.

[0088] For the experimental details and parameter settings, all experiments of the present invention are carried out on the same computer, and the detailed hardware and software configuration information of the computer is shown in Table 1.

[0089] Table 1 Hardware and software configuration information

[0090]

[0091] The parameter settings are as follows:

[0092] (1) Contact prediction model: In the training stage, random rotation is used for data augmentation of the input object to simulate possible natural changes. The batch size is set to 32, and the standard Adam optimizer is used, with the initial learning rate set to 1.6×10 -3 , and a total of 3000 rounds of training are carried out, where the number of sampling points of the input object model , and is decayed from 0 to 5×10 -2 .

[0093] (2) Grasp generation model: For the grasp generation model, the optimization process is divided into two stages. In the first global pose translation and rotation optimization stage, the learning rate is set to 5×10 -2 , and the optimization is iterated 200 times. In the second local pose optimization stage, the learning rate is 5×10 -3 , and the iterative optimization is 1000 times. Among them, the alignment coefficient = 0.01, the object contact position map coefficient and the hand contact position map are set to 0.01, the penetration coefficient and the direction contact map coefficient .

[0094] Embodiment 2

[0095] Please refer to Figure 2 , which shows a schematic structural diagram of a virtual gesture generation system based on contact modeling and contact area alignment proposed in the second embodiment of the present application. The system includes:

[0096] A data preparation module 100, configured to obtain and preprocess hand-object contact data samples, extract point cloud features from the hand-object contact data samples, and obtain initial hand-object point cloud features;

[0097] The joint contact modeling module 200 is configured to perform joint contact modeling on the object contact position, hand contact position, hand contact part, and hand contact direction in the initial hand-object point cloud features by using a conditional variational autoencoder, obtain a joint contact representation, and based on the joint contact representation, obtain a hand-object predicted contact map by using a contact prediction model;

[0098] The contact area alignment module 300 is configured to, according to the hand-object predicted contact map, perform alignment constraints on the contact areas of the hand and the object at both the global and local levels, optimize the global pose parameters and local joint pose parameters of the hand, and obtain a final grasping gesture.

[0099] A virtual gesture generation system based on contact modeling and contact area alignment in an embodiment of the present application may be a device, or a component, an integrated circuit, or a chip in a terminal. The device may be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device may be a server, a network attached storage (NAS), a personal computer (PC), etc., and the embodiments of the present application do not make specific limitations.

[0100] A device of a virtual gesture generation system based on contact modeling and contact area alignment in an embodiment of the present application. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, and the embodiments of the present application do not make specific limitations.

[0101] A virtual gesture generation system provided in an embodiment of the present application can implement Figure 1 each process implemented in a method embodiment of a virtual gesture generation method based on contact modeling and contact area alignment. To avoid repetition, it will not be elaborated here.

[0102] Optionally, an embodiment of the present application further provides an electronic device, including a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements each process of the above method embodiment of a virtual gesture generation method based on contact modeling and contact area alignment, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0103] An embodiment of the present application further provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above embodiment of a virtual gesture generation method based on contact modeling and contact area alignment, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0104] Wherein, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0105] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without more limitations, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may also be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.

[0106] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal (which may be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in various embodiments of the present application.

[0107] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.

Claims

1. A virtual gesture generation method based on contact modeling and contact area alignment, characterized in that: include: Acquire and preprocess hand-object contact data samples, perform point cloud feature extraction on the hand-object contact data samples, and obtain initial features of the hand-object point cloud; The conditional variational autoencoder is used to perform joint contact modeling on the object contact position, hand contact position, hand contact part and hand contact direction in the initial features of the hand-object point cloud to obtain a joint contact representation, and the joint contact representation is decoded to obtain a hand-object predicted contact map, which is represented by the product of conditional distribution probabilities: in, Represents the contact position of the object Object Model is the conditional distribution probability of the conditional input, Diagram showing hand contact position Initialized hand and object models The conditional distribution probability of the conditional input is Diagram showing hand contact areas The object contact position diagram The conditional distribution probability as an additional conditional input, Hand contact direction diagram Additional hand contact area diagram The conditional distribution probability of the conditional input; According to the hand-object predicted contact map, the contact areas of the hand and the object are respectively aligned and constrained at global and local levels, and the global posture parameters and local joint posture parameters of the hand are optimized to obtain the final grasping gesture, which specifically includes: in the global posture optimization stage, the global posture of the hand is obtained by minimizing the alignment loss of the difference between the hand contact position map and the object contact position map; in the local posture optimization stage, the local joint posture of the hand is obtained by minimizing the difference from the predicted direction and the difference from the part position using the hand contact direction map and the hand contact part map; the global posture and local joint posture are combined to align the contact areas of the hand and the object and reduce mutual penetration, so as to obtain a reasonable grasping gesture.

2. A virtual gesture generation method based on contact modeling and contact area alignment according to claim 1, characterized in that: The steps of acquiring and preprocessing the hand-object contact data samples, extracting point cloud features from the hand-object contact data samples, and obtaining the initial features of the hand-object point cloud include: The point cloud network model PointNet++ is used to extract the initial point cloud features of hand-object contact data samples.

3. The method for generating virtual gestures based on contact modeling and contact area alignment according to claim 1, characterized in that: Also includes: When the object model is used as a conditional input, the latent layer is normalized through the KL minimization constraint to obtain a normal distribution. The latent distribution of .

4. The method for generating virtual gestures based on contact modeling and contact area alignment according to claim 3, characterized in that: In the step of performing joint contact modeling on the object contact position, hand contact position, hand contact part and hand contact direction in the initial features of the hand-object point cloud using a conditional variational autoencoder, the training loss function corresponding to the conditional variational autoencoder includes: The contact map reconstruction loss is used to constrain the difference between the predicted contact map and the true contact map; Cross entropy KL loss, used to constrain the latent layer distribution to be a standard normal distribution; The cosine loss of the direction dimension is used to constrain the consistency between the hand contact direction and the true direction.

5. The method for generating virtual gestures based on contact modeling and contact area alignment according to claim 1, characterized in that: Also includes: The penetration loss based on the signed distance field SDF is used to constrain the directed distance from the hand part to the object surface so that the directed distance is greater than or equal to zero.

6. A virtual gesture generation system based on contact modeling and contact area alignment, characterized in that: The system is used to implement the steps of a virtual gesture generation method based on contact modeling and contact area alignment as claimed in claim 1, including: A data preparation module is used to obtain and preprocess hand-object contact data samples, perform point cloud feature extraction on the hand-object contact data samples, and obtain initial features of the hand-object point cloud; a joint contact modeling module, for performing joint contact modeling on the object contact position, hand contact position, hand contact part and hand contact direction in the initial features of the hand-object point cloud using a conditional variational autoencoder to obtain a joint contact representation, and decoding the joint contact representation to obtain a hand-object predicted contact map; The contact area alignment module is used to align the contact areas of the hand and the object at global and local levels according to the hand-object predicted contact map, optimize the global posture parameters and local joint posture parameters of the hand, and obtain the final grasping gesture.

7. An electronic device, characterized in that: The method comprises a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of a method for generating a virtual gesture based on contact modeling and contact area alignment as described in any one of claims 1 to 5 are implemented.

8. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of a virtual gesture generation method based on contact modeling and contact area alignment as described in any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Hand posture estimation and tracking method based on depth data

    CN110286749A

  • Multi-fingered dexterous hand grabbing gesture planning method based on deep neural network

    CN114643586A