Hand-object interaction three-dimensional reconstruction method and device based on language guidance

By using a language-guided 3D reconstruction method for hand-object interaction, we can acquire and optimize the point cloud and commands of hand-object interaction, construct a gravitational field, solve the problems of insufficient accuracy and fit in the reconstruction of 3D models of hand-object interaction in existing technologies, and realize high-precision hand-object interaction modeling in complex scenes.

CN121190656APending Publication Date: 2025-12-23Artificial Intelligence and Robotics Innovation Center of Hong Kong Institute of Innovation, Chinese Academy of Sciences
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511112455.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-12-23

AI Technical Summary

Technical Problem

In hand-object interaction scenarios, existing technologies have low accuracy and fit of 3D model reconstruction, especially when there is occlusion and irregular contact between complex interactive elements such as hands and instruments. The prediction results are unstable and the reconstructed model does not fit the actual situation well.

Method used

The language-guided hand-object interaction 3D reconstruction method obtains the initial interaction point cloud and interaction command of the hand-object interaction action, extracts the point cloud features of the hand and the instrument and the semantic features of the command, predicts the contact probability and area of ​​the instrument, constructs a gravitational field, optimizes the point position of the hand, and comprehensively considers the information of the hand, the instrument and the interaction command to perform 3D reconstruction.

Benefits of technology

It achieves accurate and highly fitting 3D model reconstruction of hand-object interaction in complex interactive scenarios, improving the prediction accuracy and fitting of hand-instrument interaction, and is applicable to tasks such as intraoperative navigation, virtual training and robot interaction in intelligent medical assistance systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121190656A_ABST
    Figure CN121190656A_ABST
Patent Text Reader

Abstract

The invention provides a hand-object interaction three-dimensional reconstruction method and device based on language guidance, and the method comprises the steps: extracting hand point cloud features and instrument point cloud features in an initial interaction point cloud, and the instruction semantic features of an interaction instruction; based on the hand point cloud feature, the instrument point cloud feature and the instruction semantic feature, obtaining an instrument contact area; constructing a gravitational field of the hand part point based on the distance between the hand part point and an instrument contact site in the instrument contact area; and based on the gravitational field, optimizing the three-dimensional space position of the hand part point to obtain a three-dimensional reconstruction model. According to the method provided by the invention, the interaction instruction reflecting the interaction intention is introduced, the interaction action of the hand and the instrument is accurately reflected, position solving is carried out on the basis of the constructed gravitational field, and three-dimensional model reconstruction with accurate prediction and high fitting degree of the hand and the instrument is realized; and a more reliable and more accurate hand and instrument interaction model is provided for tasks such as intraoperative navigation in an intelligent medical auxiliary system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of three-dimensional reconstruction, in particular to a language-guided hand-device interaction three-dimensional reconstruction method and device. BACKGROUND

[0002] With the development of intelligent medical auxiliary systems, high-precision modeling between hands and devices is of great significance in intraoperative navigation, virtual training, and robot interaction tasks. Existing technologies mostly rely on image recognition, point cloud registration, or data-driven pose estimation to achieve three-dimensional point cloud reconstruction during hand-device interaction.

[0003] However, the above methods are acceptable in handling simple three-dimensional point cloud reconstruction of interactive elements, but when dealing with fine and complex dynamic relationships between interactive elements, such as occlusion and irregular contact between hands and devices, they expose the shortcomings of unstable prediction results and low fitting degree of reconstructed three-dimensional models. SUMMARY

[0004] The present application provides a language-guided hand-device interaction three-dimensional reconstruction method and device to solve the problem of low accuracy and fitting degree of three-dimensional model reconstruction in hand-device interaction scenarios in existing technologies.

[0005] The present application provides a language-guided hand-device interaction three-dimensional reconstruction method, comprising: Obtaining initial interaction point clouds of hand-device interaction actions and interaction instructions; Extracting hand point cloud features and device point cloud features in the initial interaction point clouds, and instruction semantic features of the interaction instructions; Based on the hand point cloud features, the device point cloud features, and the instruction semantic features, obtaining the contact probability of each device site in the device element and the device contact area; Based on the distance between the hand sites in the hand element and the device contact sites in the device contact area, constructing the gravitational field of the hand sites; Based on the gravitational field, optimizing the three-dimensional spatial position of the hand sites to obtain a three-dimensional reconstruction model of the hand-device interaction action.

[0006] According to the language-guided hand-device interaction three-dimensional reconstruction method provided by the present application, the contact probability of each device site in the device element and the device contact area are obtained based on the hand point cloud features, the device point cloud features, and the instruction semantic features, comprising: Fusing the instruction semantic features with the site point cloud features of each hand site to obtain hand fusion features; Fusing the instruction semantic features with the site point cloud features of each device site to obtain device fusion features; predicting a contact probability of each instrument site based on the hand fusion feature and the instrument fusion feature; selecting the instrument contact region from the instrument elements based on the contact probability of each instrument site.

[0007] According to the application, a language-guided hand-object interaction three-dimensional reconstruction method is provided, which constructs a gravitational field based on the distance between the hand sites in the hand elements and the instrument contact sites in the instrument contact region, including: constructing a potential energy function based on the contact probability of the instrument contact sites and the distance between the hand sites and the instrument contact sites; adding a gradient smoothing regularization term to the potential energy function to obtain the gravitational field; The gradient smoothing regularization term is constructed based on the Laplacian norm of the first derivative of the potential energy function in space.

[0008] According to the application, a language-guided hand-object interaction three-dimensional reconstruction method is provided, which optimizes the three-dimensional spatial position of the hand sites based on the gravitational field to obtain a three-dimensional reconstruction model of the hand-object interaction action, including: calculating the current shape index and the initial shape index between each hand contact site in the hand contact region; constructing a contact constraint based on the shape deviation loss between the current shape index and the initial shape index; constructing an optimization constraint function based on the contact constraint and the hand structure constraint; optimizing the three-dimensional spatial position of the hand sites based on the optimization constraint function and the gravitational field to obtain the three-dimensional reconstruction model.

[0009] According to the application, a language-guided hand-object interaction three-dimensional reconstruction method is provided, which includes a Laplace constraint in the hand structure constraint; The construction process of the Laplace constraint includes: determining the adjacent regions of each hand site; constructing a Laplace constraint based on the three-dimensional spatial position of each hand site and the adjacent hand sites in the adjacent regions.

[0010] According to the application, a language-guided hand-object interaction three-dimensional reconstruction method is provided, which also includes a joint angle constraint in the hand structure constraint: The construction process of the joint angle constraint includes: calculating the joint angle of the hand joint based on the three-dimensional spatial position of the hand site corresponding to the hand joint; Based on the joint angle, a joint angle constraint is constructed.

[0011] The application further provides a language-guided hand-object interaction three-dimensional reconstruction device, comprising: An acquisition unit acquires an initial interaction point cloud of a hand-object interaction action and an interaction instruction; A feature extraction unit extracts hand point cloud features and instrument point cloud features in the initial interaction point cloud and instruction semantic features of the interaction instruction; A prediction unit obtains contact probabilities of each instrument site in an instrument element and an instrument contact region based on the hand point cloud features, the instrument point cloud features and the instruction semantic features; A gravitational field construction unit constructs a gravitational field of a hand site in a hand element based on a distance between the hand site and an instrument contact site in the instrument contact region; A reconstruction unit optimizes a three-dimensional space position of the hand site based on the gravitational field to obtain a three-dimensional reconstruction model of the hand-object interaction action.

[0012] The application further provides an electronic device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the language-guided hand-object interaction three-dimensional reconstruction method according to any one of the above when executing the program.

[0013] The application further provides a non-transitory computer readable storage medium having a computer program stored thereon, wherein the computer program is executable on a processor to implement the language-guided hand-object interaction three-dimensional reconstruction method according to any one of the above.

[0014] The application further provides a computer program product comprising a computer program, wherein the computer program is executable on a processor to implement the language-guided hand-object interaction three-dimensional reconstruction method according to any one of the above.

[0015] The language-guided hand-object interaction three-dimensional reconstruction method and device provided by the application can obtain an initial interaction point cloud and an interaction instruction, extract relevant features, predict an instrument contact region and construct a gravitational field, and finally optimize a hand site position to obtain a three-dimensional reconstruction model of a hand-object interaction action, which comprehensively considers information of a hand, an instrument and an interaction instruction, can accurately reflect an interaction state of the hand and the instrument, and based on the constructed gravitational field, can solve a position, thereby realizing three-dimensional model reconstruction with high prediction accuracy and high hand-instrument adhesion in a complex interaction scene, and providing a more reliable and more accurate hand-instrument interaction model for intraoperative navigation, virtual training and robot interaction in an intelligent medical auxiliary system. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort on the basis of these drawings.

[0017] Figure 1 is a flowchart of a language-guided hand-object interaction three-dimensional reconstruction method provided by the present application; Figure 2 is a flowchart of a method for solving three-dimensional space positions based on a gravitational field provided by the present application; Figure 3 is a structural schematic diagram of a language-guided hand-object interaction three-dimensional reconstruction device provided by the present application; Figure 4 is a structural schematic diagram of an electronic device provided by the present application. DETAILED DESCRIPTION

[0018] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below in combination with the drawings in the present application. Obviously, the described embodiments are some embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort fall within the scope of protection of the present application.

[0019] It should be noted that the prior art relies on methods based on image recognition, point cloud registration or data-driven pose estimation to reconstruct three-dimensional models of interactive elements. Among them, the method based on image recognition extracts the features of hands and instruments by analyzing the image information in the surgical scene, and then constructs a three-dimensional model, but this method is easily affected by image quality, lighting conditions and occlusion factors; the point cloud registration technology matches and fuses the point cloud data of hands and instruments obtained from different perspectives, but in complex interactive scenes, due to the complex and variable motion of hands and instruments, the point cloud data may have serious noise and missing, resulting in a significant decrease in registration accuracy; data-driven pose estimation mainly uses a large amount of labeled data to train the model to predict the pose of hands and instruments, but such methods have very high requirements for the quality and quantity of data, and in actual application, complex and variable surgical scenes often cannot provide sufficient comprehensive and accurate data to support the training and optimization of the model.

[0020] Therefore, in a medical scene, the accuracy and fitting of the three-dimensional model reconstruction when the hand and the instrument interact based on the above three-dimensional model reconstruction method are not high. In view of this problem, the application provides a language-guided hand-object interaction three-dimensional reconstruction method to realize three-dimensional model reconstruction of a hand-object interaction action with high prediction accuracy and high fitting. Figure 1 is a flowchart of the language-guided hand-object interaction three-dimensional reconstruction method provided by the application, as Figure 1 indicated, the method comprises: Step 110, obtaining an initial interaction point cloud of a hand-object interaction action and an interaction instruction.

[0021] Here, the hand-object interaction action refers to an object that needs to be three-dimensional model reconstructed, including a hand and an instrument that interacts with the hand. Here, the hand-object interaction action may be, for example, a hand-object interaction action such as a hand holding a forceps. In addition, the interaction instruction here refers to a text description that can reflect the hand-object interaction intention, which may be, for example, "hold the hemostat".

[0022] Specifically, the depth image corresponding to the hand-object interaction action can be obtained by a depth camera, and the initial interaction point cloud can be constructed by recognizing the depth image. Alternatively, the initial interaction point cloud can be obtained by scanning the hand-object interaction action based on a three-dimensional scanning device. In addition, the action instruction corresponding to the hand-object interaction action can be input by a person based on the hand-object interaction action. Alternatively, the action semantic information in the image can be recognized by image recognition on the image corresponding to the hand-object interaction action, and the recognized action semantic information is taken as the interaction instruction.

[0023] Step 120, extracting hand point cloud features and instrument point cloud features in the initial interaction point cloud, and instruction semantic features of the interaction instruction.

[0024] Specifically, for hand point cloud feature extraction, first, the hand point cloud can be segmented from the initial interaction point cloud by using a point cloud segmentation and feature extraction algorithm based on deep learning, such as the PointNet++ network. Then, the key points, pose vectors and other features of the hand point cloud are extracted by a feature encoding network, such as the PointNet++ network or Point-BERT (Bidirectional Encoder Representations from Transformers, Transformer-based bidirectional encoder representation). Similarly, for instrument point cloud feature extraction, the instrument point cloud can be segmented from the initial interaction point cloud based on a point cloud segmentation and feature extraction algorithm based on deep learning. Then, the geometric shape, size, motion trajectory and other features of the instrument point cloud are extracted by a feature encoding network to obtain the instrument point cloud features.

[0025] In addition, for the instruction semantic features, the interactive instruction can be converted into a vector representation by natural language processing techniques such as word embedding or recurrent neural networks or large language models, and the semantic features thereof can be extracted.

[0026] It should be noted that by extracting the instruction semantic features of the interactive instruction, important information for subsequent analysis of the interaction relationship between the hand and the instrument is provided, which helps to accurately understand the interaction intention and state, and thus improves the accuracy of three-dimensional model reconstruction.

[0027] In step 130, based on the hand point cloud features, the instrument point cloud features, and the instruction semantic features, the contact probability of each instrument site in the instrument element and the instrument contact area are obtained.

[0028] Specifically, the instruction semantic features can be fused with the hand point cloud features and the instrument point cloud features respectively to obtain the fused hand point cloud features and the fused instrument point cloud features. Then, the fused features can be analyzed by using a machine learning algorithm or a deep learning model to predict the contact probability of each instrument site in the instrument element. The machine learning algorithm can be, for example, a random forest, a support vector machine, etc. The deep learning model can be a model constructed based on a neural network. The neural network model can adopt a double-branch architecture and can be based on an interactive attention network of a Transformer to predict the contact probability of each hand site in the hand element and the contact probability of each instrument site in the instrument element, respectively.

[0029] Further, according to a set probability threshold, the instrument site set with a contact probability greater than the preset probability threshold can be taken as the instrument contact area. Here, the instrument contact area refers to the instrument part area that is likely to be in contact with the hand, which is determined according to the contact probability of each instrument site, and provides a reference for subsequent construction of the gravitational field and optimization of the hand site position.

[0030] It should be noted that by comprehensively considering the information of the hand, the instrument, and the interactive instruction, the contact condition between the instrument and the hand is accurately predicted, which provides a reasonable basis for subsequent construction of the gravitational field, and helps to improve the accuracy of optimization of the hand site.

[0031] In step 140, based on the distance between the hand site in the hand element and the instrument contact site in the instrument contact area, a gravitational field of the hand site is constructed.

[0032] Specifically, for each hand point, the distance between it and all instrument contact points in the instrument contact area is calculated. According to the reciprocal of the distance or other preset function relationship, the gravitational field of the hand point is constructed. It should be noted that the size and direction of the gravitational field can be used to reflect the "attraction" of the hand point from different instrument contact points. Thus, by constructing the gravitational field, the spatial relationship between the hand point and the instrument contact area can be converted into a calculable physical quantity, providing a clear guiding direction for the optimization of the hand point, which helps the hand point to move to a reasonable interaction position.

[0033] In step 150, the three-dimensional spatial position of the hand point is optimized based on the gravitational field to obtain a three-dimensional reconstruction model.

[0034] Specifically, an optimization algorithm can be used to adjust the three-dimensional spatial position of the hand point according to the gravitational field of the hand point. The optimization algorithm can use gradient descent or Newton's method. It should be noted that in the optimization process, the motion constraints and physiological structure limitations of the hand itself are considered to ensure that the optimized hand point position conforms to the actual motion law of the hand. The optimization process is repeated until the preset convergence condition is met, and a three-dimensional reconstruction model of the hand-object interaction action is obtained. After obtaining the three-dimensional reconstruction model, a visual output of the three-dimensional reconstruction model can be obtained by mapping the three-dimensional reconstruction model.

[0035] It should be noted that by optimizing the position of the hand point, the interaction between the hand and the instrument is more natural and reasonable, and the accuracy and fit of the hand-object interaction reconstruction point cloud are improved. Thus, by realizing accurate and realistic hand-object interaction three-dimensional reconstruction, it can be widely used in many fields. For example, in the medical field, it can provide high-quality data support for intraoperative navigation and virtual training, and help precision medicine. In industrial manufacturing, product assembly processes and robot operations can be optimized; in virtual reality games, player hand operation immersion can be enhanced; in rehabilitation therapy, it helps to design personalized training programs and improve patient hand function recovery.

[0036] The method provided by the embodiments of the present application obtains the initial interaction point cloud and the interaction instruction, extracts relevant features, predicts the instrument contact area and constructs the gravitational field, and finally optimizes the position of the hand point to obtain a three-dimensional reconstruction model of the hand-object interaction action. This process considers the information of the hand, the instrument and the interaction instruction, and can accurately reflect the interaction state of the hand and the instrument. Based on the constructed gravitational field, the position is solved, the three-dimensional model reconstruction is realized with high accuracy and high hand-instrument fit in complex interaction scenarios, and a more reliable and accurate hand-instrument interaction model is provided for intraoperative navigation, virtual training and robot interaction in intelligent medical auxiliary systems.

[0037] Based on any of the above embodiments, step 130 comprises: The instruction semantic features are fused with the point cloud features of each hand part point respectively to obtain hand fusion features; The instruction semantic features are fused with the point cloud features of each instrument point respectively to obtain instrument fusion features; The contact probabilities of the instrument points are predicted based on the hand fusion features and the instrument fusion features; The instrument contact region is selected from the instrument elements based on the contact probabilities of the instrument points.

[0038] Here, the hand fusion features refer to the feature set obtained by fusing the instruction semantic features with the point cloud features of each hand part point. It can be understood that the hand fusion features comprehensively include the point cloud features of the hand itself such as geometry and motion, and the operation intention information contained in the interaction instruction, which can more comprehensively and accurately describe the state and possible interaction behavior of the hand in the current interaction scene, and provide a key basis for subsequent prediction of the contact probability of the instrument point.

[0039] In addition, the instrument fusion features refer to the feature set formed by fusing the instruction semantic features with the point cloud features of each instrument point respectively. It can be understood that the instrument fusion features not only include the point cloud features of the instrument itself such as shape, size and motion, but also incorporate the operation target information embodied in the interaction instruction, which is helpful to more accurately understand the role of the instrument in the interaction process and the potential interaction relationship with the hand, and thus assist in predicting the contact probability of the instrument point.

[0040] Specifically, first, the encoded instruction semantic feature vector can be fused with the point cloud feature vector of each hand part point, and the fused features of all hand part points are taken as the hand fusion features. It should be noted that the fusion operation here can be realized based on direct splicing, and attention mechanism can also be used to weight and fuse the importance of different hand part point cloud features according to the instruction semantic features, so as to enhance the contribution of the point cloud features of key hand part points.

[0041] Here, the fused features of any hand part point can be represented by the following formula, as shown in the following formula: ; In the formula, represents the fused features of any hand part point; represents cross-modal attention fusion, or feature splicing and linear projection operation; represents the point cloud features of the hand part point; represents the instruction semantic features.

[0042] Similarly, the instruction semantic feature can be fused with the site point cloud feature of each instrument site respectively, and the fused features of all instrument sites are taken as instrument fusion features. Similarly, the fusion operation here can be realized based on the direct splicing method, and the attention mechanism can also be used to weight and fuse the importance of different instrument site point cloud features according to the instruction semantic feature, so as to enhance the contribution of the site point cloud features of the key instrument sites.

[0043] Here, the fused feature of any instrument site can be represented by the following formula, as shown in the following formula: ; In the formula, represents the fused feature of any instrument site; represents cross-modal attention fusion, or feature splicing and linear projection operation; represents the site point cloud feature of the instrument site; represents the instruction semantic feature.

[0044] Further, a multi-modal fusion prediction model can be constructed, such as a neural network-based regression model or classification model. The model takes the hand fusion feature and the instrument fusion feature as input, and performs feature transformation and interaction through network structures such as fully connected layers and convolutional layers, to output the contact probability of each instrument site. Here, the contact probability of each instrument site is represented as , where represents the total number of instrument sites in the instrument element.

[0045] Finally, a plurality of instrument sites can be selected from the instrument element by the contact probability of each instrument site to obtain the instrument contact region. For example, the contact probabilities of each instrument site can be arranged in descending order, and the top M instrument sites can be taken as the instrument sites. Alternatively, the instrument sites with a contact probability greater than a probability threshold can be taken as the instrument sites. Here, the number of top-ranked and the probability threshold can be set according to actual needs and application scenarios. For example, for a scene requiring high-precision interaction, a higher threshold can be set to ensure that only instrument sites with a larger contact probability are selected as the contact region; for a general scene, the threshold can be appropriately reduced.

[0046] The method provided in this invention fuses the semantic features of instructions with the point cloud features of hand and instrument sites, respectively, to obtain hand fusion features and instrument fusion features. This allows for the prediction of the contact probability of each instrument site and the determination of the instrument contact area. This process fully utilizes interactive instruction information and comprehensively considers the physical characteristics of the hand and instrument, enabling a more accurate and comprehensive understanding of the interaction between the hand and the instrument. This provides richer information for subsequent accurate prediction of instrument site contact probabilities, improving prediction accuracy. Compared to traditional methods, this approach improves prediction accuracy, clarifies key interaction areas, and provides a more reliable foundation for subsequently constructing the gravitational field of hand sites and optimizing hand site positions. It is applicable to various precision medical operations, such as clamping, pushing, and rotation, and exhibits strong generalization and interaction stability.

[0047] Based on any of the above embodiments, step 140 includes: A potential energy function is constructed based on the contact probability of the instrument contact point and the distance between the hand part and the instrument contact point; By adding a gradient smoothing regularization term to the potential energy function, the gravitational field is obtained. The gradient smoothing regularization term is constructed based on the Laplace norm of the first derivative of the potential function in space.

[0048] Specifically, the potential energy function can be constructed using the contact probability of the instrument contact point and the distance between the hand position and the instrument contact point. Here, the potential energy function can be constructed using the following formula, as shown below: In the formula, Let x represent the potential energy function at the hand position. Indicates the contact area of ​​the instrument Any instrument contact point in the process; This represents the contact probability at the contact point p of the device; This represents the distance between the hand part point x and the instrument contact point p; This indicates the range of influence of the gravitational field.

[0049] It should be noted that the constructed potential energy function takes into account both the contact probability of the instrument contact point and the distance between the hand part and the instrument contact point. It can accurately reflect the magnitude and direction of the "attraction" effect from the instrument contact point on the hand part, providing a basis for the subsequent construction of the gravitational field.

[0050] Further, considering that the gravitational potential field in the gravity guiding mechanism has a strong influence on the deformation path, in order to make the three-dimensional spatial position of the hand point more smooth, a gradient smoothing regularization term can be introduced into the potential energy function to construct the gravitational field of the hand point. Here, the gradient smoothing regularization term can be calculated by the following steps: first, the first derivative of the potential energy function can be calculated. Then, the Laplacian norm of the first derivative in space is calculated.

[0051] It should be noted that the introduction of the gradient smoothing regularization term can make the gravitational field smoother, avoid the occurrence of unstable or unreasonable situations in the hand point optimization process due to the local sharp change of the potential energy function, and improve the stability and accuracy of the hand point position optimization.

[0052] The method provided by the embodiment of the application constructs a potential energy function based on the contact probability of the instrument contact point and the distance between the hand point and the instrument contact point, and introduces a gradient smoothing regularization term based on the Laplacian norm of the first derivative of the potential energy function to obtain the gravitational field. This process comprehensively considers the physical relationship between the hand and the instrument and the spatial smoothness requirement, and can effectively guide the hand point to move to a reasonable position. Compared with the traditional method, the scheme improves the stability and accuracy of the hand point position optimization, makes the hand-object interaction reconstructed point cloud more consistent with the actual interaction, and provides a more reliable and more accurate solution for hand-instrument interaction modeling in an intelligent medical auxiliary system, which helps to improve the effect of intraoperative navigation, virtual training and other applications.

[0053] It should be noted that, in order to avoid the problems of excessive deformation and structure collapse of the hand contact area under semantic driving, a conservative shape preserving constraint can be introduced to the predicted hand contact area. Based on any of the above embodiments, step 150 comprises: calculating the current shape indicator and the initial shape indicator between each hand contact point in the hand contact area; constructing a contact constraint based on the shape deviation loss between the current shape indicator and the initial shape indicator; constructing an optimization constraint function based on the contact constraint and the hand structure constraint; optimizing the three-dimensional spatial position of the hand point based on the optimization constraint function and the gravitational field to obtain the three-dimensional reconstruction model.

[0054] Here, the hand contact area can be calculated based on the calculation method of the instrument contact area. In detail, a multi-modal fusion prediction model can be constructed, such as a neural network-based regression model or a classification model. The model takes hand fusion features and instrument fusion features as input, and performs feature transformation and interaction through network structures such as fully connected layers and convolutional layers to output the contact probability of each hand point. Here, the contact probability of each instrument point is represented as wherein represents the total number of hand points in the hand element. Then, a plurality of hand points can be selected from the hand element by the contact probability of each hand point, to obtain a hand contact region.

[0055] Specifically, first, the shape features can be expressed by defining a plurality of shape indicators, such as including relative distance, local normal, curvature, etc. Then, the shape indicators under the initial state of the hand and the current state are calculated respectively to obtain the current shape indicators and the initial shape indicators. Among them, the shape indicators defined above are calculated for each hand contact point in the hand contact region under the initial state of the hand, and these indicator values are stored as initial shape indicators. In addition, during the iteration process of the three-dimensional spatial position of the hand points, the same shape indicators can be calculated for each hand contact point in the hand contact region after iteration in real time to obtain the current shape indicators.

[0056] Further, the contact constraint can be constructed by calculating the shape offset loss between the current shape indicators and the initial shape indicators. In detail, for each hand contact point, the difference between its current shape indicator and the initial shape indicator is calculated, and then the weighted sum of these differences is obtained to obtain the shape offset loss.

[0057] It should be noted that by calculating the current shape indicators and the initial shape indicators, the shape change of the hand contact region during the interaction process can be quantified, which provides a basis for subsequent construction of the contact constraint. Thus, the contact constraint here helps to maintain the geometric continuity and contact stability of the contact region.

[0058] Then, the hand structure constraint can be constructed by the hand bone length constraint, the joint angle constraint, and the hand internal topology consistency constraint. Next, the optimization constraint function is constructed by combining the contact constraint and the hand structure constraint. Further, the total optimization objective function can be constructed by combining the optimization constraint function and the gravitational field. The three-dimensional spatial position of the hand points can be optimized by selecting a suitable optimization algorithm. Taking the gradient descent method as an example, first, the gradient of the total optimization objective function with respect to the three-dimensional spatial position variable of the hand points is calculated, and then the iterative update is performed according to the iteration rule, wherein the iteration rule includes the learning rate and the number of iterations, until the convergence condition is met, such as the gradient being less than a certain threshold or the number of iterations reaching a maximum value.

[0059] During the optimization process, the three-dimensional spatial position of the hand points after each iteration is recorded, and when the optimization converges, the final hand point position information is combined with the instrument point position information to construct a three-dimensional reconstruction model.

[0060] The method provided by the embodiment of the application can make the hand points move reasonably to the instrument contact area under the premise of meeting the interaction requirements and the hand structure constraints, so as to obtain an accurate three-dimensional reconstruction model, and provide more accurate results for hand-instrument interaction modeling in an intelligent medical auxiliary system.

[0061] It should be noted that, in order to prevent the hand from being damaged in the deformation process, such as point cloud overlap, palm structure breakage, and abnormal connection of knuckles, based on any of the above embodiments, the hand structure constraint comprises a Laplace constraint; The construction process of the Laplace constraint comprises: determining the adjacent area of each hand point; constructing the Laplace constraint based on the three-dimensional space positions of the hand points and the adjacent hand points in the adjacent area.

[0062] Specifically, the hand mesh model can be constructed based on the hand point cloud, for each vertex in the mesh model, that is, the hand point, all edges connected to the vertex can be traversed, and other vertices connected through the edges can be found, and these vertices and the local area around them constitute the adjacent area of the hand point. Then, for each hand point, the distance difference between the average position of the adjacent hand points in the adjacent area and the hand point itself can be calculated, and the Laplace coordinates are calculated. It can be understood that the goal of the Laplace constraint is to keep the local geometry of the hand point unchanged during the optimization process, that is, the Laplace coordinates remain as consistent as possible before and after optimization.

[0063] Then, the Laplace constraint can be constructed by keeping the Laplace coordinates of the current hand point and each hand point in the initial MANO hand model skeleton unchanged as the optimization goal through the Laplace coordinates of each hand point in the initial MANO hand model skeleton.

[0064] It should be noted that the Laplace constraint can effectively maintain the local geometry of the hand point, and ensure that the local geometric structure remains relatively rigid and consistent during the optimization process, avoid unreasonable deformation of the hand shape due to excessive optimization, and make the reconstructed hand model more consistent with the topology of the real hand.

[0065] Based on any of the above embodiments, the hand structure constraint further comprises a joint angle constraint: The construction process of the joint angle constraint comprises: calculating the joint angle of the hand joint based on the three-dimensional space positions of the hand points corresponding to the hand joint; constructing the joint angle constraint based on the joint angle.

[0066] Specifically, first, for each joint of the hand, its corresponding key hand point combination is determined, for example, the metacarpophalangeal joint of the index finger corresponds to the index finger root key point, the palm connection point, etc. Then, in the process of calculating the joint angle, for the joint angle of any joint, the three-dimensional spatial position of the hand point corresponding to the joint can be used to calculate the joint angle of the joint. In some complex hand pose estimation and constraint construction, the joint angle can also be calculated by constructing the local coordinate system of the joint. Taking the metacarpophalangeal joint as an example, first, the initial pose of the joint is determined, and then the rotation matrix of the joint is calculated according to the change of the three-dimensional spatial position of the hand point. The Euler angle or quaternion representing the rotation angle parameter of the joint is extracted from the rotation matrix, so as to obtain the joint angle.

[0067] Further, a reasonable angle range can be set for each joint according to the physiological structure and normal activity range of the human hand to construct the joint angle constraint. For example, the metacarpophalangeal joint of the finger can usually move between 0° and 90°, and the metacarpophalangeal joint of the thumb can have a larger range of motion.

[0068] It should be noted that by setting the joint angle range constraint, it is ensured that the pose of the hand model conforms to the physiological structure of the human hand, and unreasonable joint bending or stretching is avoided, so that the reconstructed or optimized hand model is more realistic and credible.

[0069] It should be further noted that the above optimization constraint function can include contact constraints and / or hand structure constraints. The hand structure constraints can include at least one of Laplace constraints and joint angle constraints.

[0070] Based on any of the above embodiments, Figure 2 is a flowchart of the method for solving the three-dimensional spatial position based on the gravitational field provided by the application, as Figure 2 shown, the method comprises: Step 210, constructing a gravitational field according to the interaction area. The interaction area here refers to the instrument contact area. In detail, the gravitational field is constructed based on the distance between the hand points in the hand elements and the instrument contact points in the instrument contact area.

[0071] Step 220, constructing a diffusion bridge based on the gravitational field.

[0072] Step 230, constructing an iterative optimization and structure constraint condition. The structure constraint condition includes at least one of the shape offset loss, the Laplace constraint, and the joint angle constraint.

[0073] Step 240, iteratively solving the three-dimensional spatial position of the hand point.

[0074] Based on any of the above embodiments, the interaction instruction can be obtained through image recognition by an image recognition model. Here, the image recognition model can be trained based on a multi-modal large model, and the trained image recognition model can automatically recognize the current interaction state from the image containing the hand-object interaction action and generate the interaction instruction as input. For example, the image recognition model observes that the doctor holds a scalpel in the current picture and automatically generates "holding a scalpel"; or recognizes that the hemostat is lifted in the clamping action and generates "release the hemostat". Thus, the automatic language generation process is realized through the image-language large model, without the need for manual input, and is suitable for general deployment driven by unlabeled data.

[0075] Based on any of the above embodiments, Figure 3 is a structural schematic diagram of a language-guided hand-object interaction three-dimensional reconstruction device provided by the present application, as Figure 3 shown, the device comprises: An acquisition unit 310 acquires an initial interaction point cloud of a hand-object interaction action and an interaction instruction. A feature extraction unit 320 extracts hand point cloud features and instrument point cloud features in the initial interaction point cloud, and instruction semantic features of the interaction instruction. A prediction unit 330 obtains the contact probability of each instrument site in the instrument element and the instrument contact area based on the hand point cloud features, the instrument point cloud features, and the instruction semantic features. A gravitational field construction unit 340 constructs a gravitational field of a hand site based on the distance between the hand site in the hand element and the instrument contact site in the instrument contact area. A reconstruction unit 350 optimizes the three-dimensional spatial position of the hand site based on the gravitational field to obtain a three-dimensional reconstruction model of the hand-object interaction action.

[0076] The device provided by the embodiment of the present application acquires an initial interaction point cloud and an interaction instruction, extracts relevant features, predicts an instrument contact area and constructs a gravitational field, and finally optimizes the position of a hand site to obtain a three-dimensional reconstruction model of a hand-object interaction action. This process comprehensively considers the information of the hand, the instrument, and the interaction instruction, and can accurately reflect the interaction state of the hand and the instrument. Based on the constructed gravitational field, the position is solved, the three-dimensional model reconstruction with high hand-instrument adhesion is predicted accurately in a complex interaction scene, and a more reliable and more accurate hand-instrument interaction model is provided for intraoperative navigation, virtual training, and robot interaction in an intelligent medical auxiliary system. task.

[0077] Based on any of the above embodiments, the prediction unit is specifically configured to: fuse the instruction semantic features with the site point cloud features of each hand site to obtain hand fusion features; Fusing the instruction semantic features with the site point cloud features of the respective instrument sites respectively to obtain instrument fusion features; Based on the hand fusion features and the instrument fusion features, the contact probability of the respective instrument sites is predicted; Based on the contact probability of the respective instrument sites, the instrument contact region is selected from the instrument elements.

[0078] Based on any of the above embodiments, the gravitational field construction unit is specifically configured to: Based on the contact probability of the instrument contact sites and the distance between the hand sites and the instrument contact sites, a potential energy function is constructed; A gradient smoothing regularization term is added to the potential energy function to obtain the gravitational field; The gradient smoothing regularization term is constructed based on the Laplacian norm of the first derivative of the potential energy function in space.

[0079] Based on any of the above embodiments, the reconstruction unit is specifically configured to: Calculate the current shape index and the initial shape index between the respective hand contact sites in the hand contact region; Based on the shape offset loss between the current shape index and the initial shape index, a contact constraint is constructed; Based on the contact constraint and the hand structure constraint, an optimization constraint function is constructed; Based on the optimization constraint function and the gravitational field, the three-dimensional spatial position of the hand sites is optimized to obtain the three-dimensional reconstruction model.

[0080] Based on any of the above embodiments, the hand structure constraint includes a Laplace constraint; The reconstruction unit is further specifically configured to: Determine the adjacent region of each hand site; Based on the three-dimensional spatial position of the respective hand sites and the adjacent hand sites in the adjacent region, a Laplace constraint is constructed.

[0081] Based on any of the above embodiments, the hand structure constraint further includes a joint angle constraint: The reconstruction unit is further specifically configured to: Based on the three-dimensional spatial position of the hand sites corresponding to the hand joints, the joint angle of the hand joints is calculated; Based on the joint angle, a joint angle constraint is constructed.

[0082] Figure 4 An example of an entity structure diagram of an electronic device is shown as Figure 4As shown, the electronic device can include a processor 410, a communications interface 420, a memory 430, and a communications bus 440, wherein the processor 410, the communications interface 420, and the memory 430 complete mutual communication through the communications bus 440. The processor 410 can invoke a logic instruction in the memory 430 to execute a language guidance-based hand-object interaction three-dimensional reconstruction method, which includes: acquiring an initial interaction point cloud of a hand-object interaction action and an interaction instruction; extracting hand point cloud features and instrument point cloud features in the initial interaction point cloud, and instruction semantic features of the interaction instruction; obtaining contact probabilities of each instrument site in an instrument element and an instrument contact area based on the hand point cloud features, the instrument point cloud features, and the instruction semantic features; constructing an attractive field of a hand site based on a distance between the hand site in a hand element and an instrument contact site in the instrument contact area; and optimizing a three-dimensional space position of the hand site based on the attractive field to obtain a three-dimensional reconstruction model of the hand-object interaction action.

[0083] In addition, the logic instruction in the memory 430 described above can be implemented in the form of a software function unit and sold or used as an independent product, and can be stored in a computer-readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.

[0084] In another aspect, the present application also provides a computer program product comprising a computer program, which can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor to enable a computer to perform the language-guided hand-object interaction three-dimensional reconstruction method provided by the above method, which comprises: obtaining an initial interaction point cloud of a hand-object interaction action and an interaction instruction; extracting hand point cloud features and instrument point cloud features in the initial interaction point cloud, and instruction semantic features of the interaction instruction; obtaining contact probabilities of each instrument site in an instrument element and an instrument contact area based on the hand point cloud features, the instrument point cloud features and the instruction semantic features; constructing a gravitational field of a hand site based on a distance between the hand site in a hand element and an instrument contact site in the instrument contact area; and optimizing a three-dimensional spatial position of the hand site based on the gravitational field to obtain a three-dimensional reconstruction model of the hand-object interaction action.

[0085] In yet another aspect, the present application also provides a non-transitory computer readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a language-guided hand-object interaction three-dimensional reconstruction method provided by the above method, which comprises: obtaining an initial interaction point cloud of a hand-object interaction action and an interaction instruction; extracting hand point cloud features and instrument point cloud features in the initial interaction point cloud, and instruction semantic features of the interaction instruction; obtaining contact probabilities of each instrument site in an instrument element and an instrument contact area based on the hand point cloud features, the instrument point cloud features and the instruction semantic features; constructing a gravitational field of a hand site based on a distance between the hand site in a hand element and an instrument contact site in the instrument contact area; and optimizing a three-dimensional spatial position of the hand site based on the gravitational field to obtain a three-dimensional reconstruction model of the hand-object interaction action.

[0086] The device embodiments described above are only illustrative, wherein the units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., they can be located in one place, or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the present embodiment scheme according to actual needs. Those skilled in the art can understand and implement it without creative labor.

[0087] Those skilled in the art can clearly understand the technical solutions of the various embodiments from the above description of the embodiments, and the various embodiments can be implemented by means of software with the necessary general hardware platforms, and of course, can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part of the prior art that makes a contribution, can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, and the like, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0088] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for some technical features therein; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A language-guided 3D reconstruction method based on hand-object interaction, characterized in that, include: Obtain the initial point cloud and interaction commands for hand-object interaction actions; Extract the hand point cloud features and instrument point cloud features from the initial interaction point cloud, as well as the instruction semantic features of the interaction command; Based on the hand point cloud features, the instrument point cloud features, and the instruction semantic features, the contact probability and instrument contact area of ​​each instrument point in the instrument element are obtained. Based on the distance between the hand location in the hand element and the instrument contact point in the instrument contact area, the gravitational field of the hand location is constructed. Based on the gravitational field, the three-dimensional spatial position of the hand part is optimized to obtain a three-dimensional reconstruction model of the hand-object interaction action.

2. The language-guided hand-object interaction-based 3D reconstruction method according to claim 1, characterized in that, The process of obtaining the contact probability and contact area of ​​each instrument point in the instrument element based on the hand point cloud features, the instrument point cloud features, and the instruction semantic features includes: The semantic features of the instructions are fused with the point cloud features of each hand location to obtain the hand fusion features; The semantic features of the instructions are fused with the point cloud features of each instrument site to obtain the instrument fusion features; Based on the hand fusion features and the instrument fusion features, the contact probability of each instrument site is predicted; Based on the contact probability of each instrument site, the instrument contact area is selected from the instrument elements.

3. The language-guided hand-object interaction-based 3D reconstruction method according to claim 1, characterized in that, The construction of the gravitational field based on the distance between the hand location in the hand element and the device contact point in the device contact area includes: A potential energy function is constructed based on the contact probability of the instrument contact point and the distance between the hand part and the instrument contact point; By adding a gradient smoothing regularization term to the potential energy function, the gravitational field is obtained. The gradient smoothing regularization term is constructed based on the Laplace norm of the first derivative of the potential function in space.

4. The language-guided hand-object interaction-based 3D reconstruction method according to any one of claims 1 to 3, characterized in that, The optimization of the three-dimensional spatial position of the hand location based on the gravitational field to obtain a three-dimensional reconstruction model of the hand-object interaction action includes: Calculate the current shape index and initial shape index between each hand contact point in the hand contact area; Based on the shape offset loss between the current shape index and the initial shape index, a contact constraint is constructed; Based on the aforementioned contact constraints and hand structure constraints, an optimization constraint function is constructed; Based on the optimization constraint function and the gravitational field, the three-dimensional spatial position of the hand part is optimized to obtain the three-dimensional reconstruction model.

5. The language-guided hand-object interaction-based 3D reconstruction method according to claim 4, characterized in that, The hand structure constraints include Laplace constraints; The process of constructing the Laplace constraint includes: Determine the adjacent regions of each hand part; Based on the three-dimensional spatial positions of each hand point and adjacent hand points in the adjacent regions, Laplace constraints are constructed.

6. The language-guided hand-object interaction-based 3D reconstruction method according to claim 4, characterized in that, The hand structure constraints also include joint angle constraints: The process of constructing the joint angle constraints includes: The joint angle of the hand joint is calculated based on the three-dimensional spatial position of the hand joint corresponding to the hand joint. Based on the joint angles, joint angle constraints are constructed.

7. A language-guided, hand-object interactive 3D reconstruction device, characterized in that, include: The acquisition unit acquires the initial interaction point cloud and interaction commands of the hand-object interaction action; The feature extraction unit extracts hand point cloud features and instrument point cloud features from the initial interaction point cloud, as well as instruction semantic features of the interaction command. The prediction unit, based on the hand point cloud features, the instrument point cloud features, and the instruction semantic features, obtains the contact probability and instrument contact area of ​​each instrument point in the instrument element. The gravitational field construction unit constructs the gravitational field of the hand part based on the distance between the hand part point in the hand element and the instrument contact point in the instrument contact area; The reconstruction unit optimizes the three-dimensional spatial position of the hand part based on the gravitational field to obtain a three-dimensional reconstruction model of the hand-object interaction action.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the language-guided hand-object interaction 3D reconstruction method as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the language-guided hand-object interaction 3D reconstruction method as described in any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the language-guided hand-object interaction 3D reconstruction method as described in any one of claims 1 to 6.