Hierarchical visual-haptic fusion-based hand pose estimation method and system
Patent Information
- Application Number
- CN202611186373.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-06
- Publication Date
- 2026-09-15
AI Technical Summary
[0005]本发明的目的是提供一种分层视觉触觉融合的在手位姿估计方法及系统,旨在解决现有机器人在手位姿估计中视觉易受遮挡产生漂移、触觉局部信息易产生位姿歧义,且难以根据视觉和触觉可靠性动态选择位姿细化起点的问题
[0042]This application generates multiple visual pose hypotheses and determines the visual anchor pose through visual branching, enabling the system to obtain the global position and orientation information of the target object. By combining tactile availability and visual uncertainty, it dynamically selects the tactile refinement starting point between the current visual anchor pose and the pose at the previous moment, avoiding tactile noise introduced in the non-contact stage and reducing pose drift caused by visual occlusion, reflection, or segmentation errors. Furthermore, it generates local candidate poses around the selected anchor point, constructs corresponding simulated tactile observations, and performs cross-modal fusion of real tactile observations, simulated tactile observations, and visual features. This utilizes global visual constraints to eliminate pose ambiguities caused by local tactile contact, and determines the current six-degree-of-freedom pose through pose residual correction and confidence optimization. Therefore, this application can improve the robot's on-hand pose estimation accuracy, continuous tracking stability, and task success rate under grasping, insertion, assembly, and external disturbance conditions.
Smart Images

Figure CN122746986A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robot perception and intelligent operation technology, specifically to a hierarchical visual-tactile fusion method and system for estimating hand pose. Background Technology
[0002] In tasks such as robot dexterity operation, precision assembly, insertion, grasping and adjustment, and tool use, the robot needs to continuously acquire the six-degree-of-freedom pose of the object being grasped relative to the robot's end effector, vision sensor, or world coordinate system in order to provide accurate spatial state information for subsequent trajectory planning and operation control.
[0003] Existing handheld pose estimation methods mainly include visual estimation methods, tactile estimation methods, and visual-tactile fusion methods. Visual estimation methods can obtain the overall contour and global geometric information of the target object using color images, depth images, or point clouds, making them suitable for pose initialization before grasping or when occlusion is minor. However, when there is robot finger occlusion, target object self-occlusion, surface reflection, missing depth, or inaccurate target segmentation, the quality of visual observation deteriorates, easily leading to pose jumps or cumulative drift. Tactile estimation methods can continue to acquire local contact deformation, pressure distribution, or contact geometry information after the target object is occluded. However, the effective sensing area of tactile sensors is usually small. For objects with symmetrical structures, similar local curved surfaces, or repetitive geometric features, different global poses may produce similar tactile observations, thus easily leading to multiple pose solutions and mismatches.
[0004] Existing visual-tactile fusion methods typically employ feature stitching, result weighting, or end-to-end joint estimation, failing to adequately distinguish the reliability differences between vision and touch at different operational stages, such as before grasping, effective contact, and severe occlusion. This can introduce invalid tactile noise in the non-contact stage or continue to use erroneous visual results as the basis for pose correction when vision is unreliable. Furthermore, some methods directly search or regress the target pose in the complete six-degree-of-freedom space, resulting in a large search range and a lack of an effective comparison mechanism between real tactile observations and simulated tactile observations corresponding to candidate poses, making it difficult to simultaneously consider global localization capability, local refinement accuracy, and continuous tracking stability. Therefore, it is necessary to propose an on-the-fly pose estimation technique that can dynamically adjust the estimation strategy based on the availability and reliability of vision and touch. Summary of the Invention
[0005] The purpose of this invention is to provide a hierarchical visual-tactile fusion method and system for on-hand pose estimation, which aims to solve the problems in existing robot on-hand pose estimation, such as visual drift caused by occlusion, pose ambiguity caused by local tactile information, and difficulty in dynamically selecting the starting point for pose refinement based on visual and tactile reliability.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] This invention provides a method for estimating hand pose through hierarchical visual-tactile fusion, comprising the following steps:
[0008] Acquire visual and tactile observations during the robot's manipulation of the target object;
[0009] Multiple visual pose hypotheses and their scores are generated based on the visual observations, and the visual pose hypothesis with the highest score is determined as the visual anchor pose.
[0010] Determine whether the tactile observation meets the tactile availability conditions; if not, determine the visual anchor pose as the current six-degree-of-freedom pose.
[0011] When the condition is met, the visual uncertainty is calculated based on the degree of dispersion of the top k visual poses with the highest scores in translation and rotation space. When the visual uncertainty does not exceed a preset threshold, the visual anchor pose is selected, and when it exceeds the preset threshold, the pose of the previous moment is selected as the tactile refinement anchor pose.
[0012] Multiple local candidate poses are generated around the tactile refined anchor pose, and simulated tactile observations corresponding to each local candidate pose are generated based on the geometric model of the target object and the contact model of the tactile sensor.
[0013] Extract tactile difference features between real tactile observation and simulated tactile observation, use the tactile difference features as queries and visual features as keys and values to perform cross-attention fusion, and output the pose residuals and confidence of each local candidate pose;
[0014] Update each local candidate pose according to the pose residual, and determine the updated pose with the highest confidence as the current six-degree-of-freedom pose, and save it as the previous pose for the next time step.
[0015] As a preferred embodiment of the first aspect of the present invention, the step of generating multiple visual pose hypotheses and their scores based on visual observation includes: segmenting the target region of an RGB-D image to obtain a target object mask; inputting the RGB-D image, the target object mask, and the target object geometric model into a visual pose estimation network to generate multiple visual pose hypotheses including three-dimensional translation and three-dimensional rotation; and using a pose scoring network to determine the score of each visual pose hypothesis.
[0016] As a preferred embodiment of the first aspect of the present invention, the tactile availability condition is determined based on at least one of the following information: the amplitude of change of the tactile image relative to the non-contact background image, the effective contact area in the tactile image, the pressure value or contact force value output by the tactile sensor, the closing state of the robot gripper, and the contact feedback output by the robot controller; when the corresponding information reaches a preset contact threshold, it is determined that the tactile observation meets the tactile availability condition.
[0017] As a preferred embodiment of the first aspect of the present invention, the visual uncertainty satisfies:
[0018] ;
[0019] in, Indicates visual uncertainty; Indicates the number of visual pose assumptions involved in the calculation; and They represent the first Translation vectors and rotation matrices for each visual pose hypothesis; and They represent the preceding Average translation and average rotation of a visual pose hypothesis; This indicates the rotational geodesic distance.
[0020] As a preferred embodiment of the first aspect of the present invention, the generation of multiple local candidate poses around the tactile refinement anchor pose includes: sampling multiple six-dimensional perturbation vectors, each including three-dimensional translational perturbation and three-dimensional rotational perturbation, in the tangent space of the Lie algebra SE(3), and generating local candidate poses according to the following formula:
[0021] ;
[0022] in, Indicates the first One local candidate pose; This indicates the anchor positioning posture as 204. This represents a six-dimensional perturbation vector that includes three-dimensional translational perturbations and three-dimensional rotational perturbations; The matrix form represents the six-dimensional perturbation vector.
[0023] As a preferred technical solution of the first aspect of the present invention, the simulated tactile observation is generated based on the contact position, contact depth, contact area and deformation relationship between the surface of the target object and the surface of the tactile sensor under local candidate poses, and has the same or spatially aligned data form as the real tactile observation; the simulated tactile observation is a tactile depth map or a tactile deformation map, and is generated by a finite element tactile simulator, a geometric contact model, a differentiable tactile renderer or a learning tactile generation model.
[0024] As a preferred embodiment of the first aspect of the present invention, the extraction of tactile difference features includes: extracting real tactile observation features and simulated tactile observation features respectively using a tactile encoder with shared weights, and obtaining tactile difference features based on the feature differences between the two; the cross-attention fusion satisfies:
[0025] ;
[0026] in, This indicates the tactile characteristics after fusion; Indicates tactile characteristics; Indicates visual features; , and These represent the projection matrices corresponding to the query, key, and value, respectively. By injecting the visual global context into the tactile local features, candidate poses that are similar in local contact shape but have unreasonable global poses can be suppressed.
[0027] As a preferred embodiment of the first aspect of the present invention, the pose residual includes translation residual and rotation residual; the training loss of the haptic refinement network includes pose residual regression loss and confidence regression loss, and the confidence labels satisfy:
[0028] ;
[0029] ;
[0030] in, For confidence level labels, For combined pose error, and These represent the translation and rotation errors of the updated pose relative to the true pose, respectively. and These are the translation error weights and rotation error weights, respectively. This refers to the temperature parameter.
[0031] The second invention provides a hierarchical visual-tactile fusion in-hand pose estimation system for performing the first aspect, comprising:
[0032] The visual acquisition module is used to acquire visual observations of the target object;
[0033] The tactile sensing module is used to acquire tactile observations between the robot's end effector and the target object;
[0034] The visual anchor point generation module is used to generate multiple visual pose hypotheses and their scores based on the visual observations, and to determine the visual anchor pose and visual features.
[0035] An uncertainty-aware routing module is used to determine whether the tactile observation is available. When the tactile observation is unavailable, it outputs the visual anchor pose. When the tactile observation is available, it calculates the visual uncertainty based on the translational and rotational discreteness of multiple visual pose assumptions, and selects the tactile refined anchor pose between the visual anchor pose and the pose at the previous moment based on the visual uncertainty.
[0036] A local candidate generation module is used to generate multiple local candidate poses around the tactile refined anchor pose.
[0037] The tactile simulation module is used to generate corresponding simulated tactile observations based on the geometric model of the target object, the contact model of the tactile sensor, and the candidate poses of each local area.
[0038] The cross-modal tactile refinement module is used to extract tactile difference features between real tactile observations and simulated tactile observations, and to perform cross-attention fusion of the tactile difference features with visual features to output the pose residuals and confidence scores corresponding to each local candidate pose.
[0039] The pose output module is used to update each local candidate pose according to the pose residual, determine the updated pose with the highest confidence as the current six-degree-of-freedom pose, and feed the current six-degree-of-freedom pose back to the uncertainty-aware routing module as the previous pose for the next moment.
[0040] As a preferred embodiment of the second aspect of the present invention, the cross-modal tactile refinement module includes a shared-weight tactile encoder, a visual feature extraction unit, a cross-attention unit, a residual regression unit, and a confidence prediction unit. The shared-weight tactile encoder is used to encode real tactile observations and simulated tactile observations respectively. The cross-attention unit uses the tactile difference features formed by the two as queries and uses visual features as keys and values for fusion. The residual regression unit is used to output translation residuals and rotation residuals. The confidence prediction unit is used to output the confidence of each local candidate pose.
[0041] The technical effects and advantages provided by the present invention in the above technical solution are as follows:
[0042] This application generates multiple visual pose hypotheses and determines the visual anchor pose through visual branching, enabling the system to obtain the global position and orientation information of the target object. By combining tactile availability and visual uncertainty, it dynamically selects the tactile refinement starting point between the current visual anchor pose and the pose at the previous moment, avoiding tactile noise introduced in the non-contact stage and reducing pose drift caused by visual occlusion, reflection, or segmentation errors. Furthermore, it generates local candidate poses around the selected anchor point, constructs corresponding simulated tactile observations, and performs cross-modal fusion of real tactile observations, simulated tactile observations, and visual features. This utilizes global visual constraints to eliminate pose ambiguities caused by local tactile contact, and determines the current six-degree-of-freedom pose through pose residual correction and confidence optimization. Therefore, this application can improve the robot's on-hand pose estimation accuracy, continuous tracking stability, and task success rate under grasping, insertion, assembly, and external disturbance conditions. Attached Figure Description
[0043] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0044] Figure 1 This is a schematic diagram of the hand pose estimation system of the present invention;
[0045] Figure 2 This is a flowchart illustrating the hand pose estimation method of the present invention;
[0046] Figure 3 This is a schematic diagram illustrating the local candidate pose generation, tactile simulation, and cross-modal tactile refinement under visual anchor point constraints of the present invention.
[0047] Figure 4 This is a schematic diagram illustrating the anchor positioning posture selection based on tactile availability and visual uncertainty according to the present invention.
[0048] In the diagram: 101, Visual Acquisition Module; 102, Tactile Acquisition Module; 103, Visual Anchor Point Generation Module; 104, Uncertainty-Aware Routing Module; 105, Local Candidate Generation Module; 106, Tactile Simulation Module; 107, Cross-Modal Tactile Refinement Module; 108, Pose Output Module; 201, Visual Observation; 202, Tactile Observation; 203, Visual Pose Hypothesis; 204, Anchor Pose; 205, Local Candidate Pose; 206, Simulated Tactile Observation; 207, Pose Residual and Confidence; 208, Current Six-DOF Pose. Detailed Implementation
[0049] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that the description of this disclosure will be more complete and fully convey the concept of the exemplary embodiments to those skilled in the art. The drawings are merely illustrative of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted.
[0050] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more exemplary embodiments. Numerous specific details are provided in the following description to give a full understanding of exemplary embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure may be practiced with one or more specific details omitted, or methods, components, steps, etc. In other instances, well-known structures, methods, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this disclosure.
[0051] Example 1
[0052] like Figure 1 As shown, this embodiment provides a hierarchical visual-tactile fusion on-hand pose estimation system, including a visual acquisition module 101, a tactile acquisition module 102, a visual anchor point generation module 103, an uncertainty-aware routing module 104, a local candidate generation module 105, a tactile simulation module 106, a cross-modal tactile refinement module 107, and a pose output module 108.
[0053] The visual acquisition module 101 is used to acquire visual observations 201 during the process of the robot gripping or manipulating a target object. The visual observations 201 are preferably RGB-D images, including color images and depth images aligned with the color image space. In other embodiments, the visual observations 201 can also be RGB images, depth maps, point clouds, multi-camera images, or combinations of the above data.
[0054] The tactile acquisition module 102 is used to acquire tactile observations 202 generated when the robot gripper, dexterous hand, or other end-effector comes into contact with a target object. The tactile observations 202 can be tactile images, tactile depth maps, deformation maps, pressure distributions, contact force signals, strain signals, or torque signals.
[0055] The visual anchor point generation module 103 receives visual observations 201, identifies or segments target objects in the visual observations 201, generates multiple visual pose hypotheses 203, and determines the visual anchor pose based on the scores of each visual pose hypothesis 203. The multiple visual pose hypotheses 203 are used to characterize multiple possible six-degree-of-freedom poses corresponding to the current visual observation.
[0056] The uncertainty-aware routing module 104 receives visual pose hypothesis 203, visual anchor pose and tactile observation 202, and determines the pose estimation path to be used based on whether the tactile observation 202 is available and the degree of dispersion of the multiple visual pose hypotheses 203.
[0057] When tactile observation 202 is unavailable, uncertainty perception routing module 104 sends the visual anchor pose to pose output module 108, which outputs the visual anchor pose as the current six-degree-of-freedom pose 208.
[0058] When tactile observation 202 is available, the uncertainty-aware routing module 104 further calculates the visual uncertainty and selects the anchor pose 204 between the current visual anchor pose and the pose at the previous moment based on the visual uncertainty. When the vision is reliable, the anchor pose 204 is the current visual anchor pose; when the vision is unreliable, the anchor pose 204 is the pose at the previous moment.
[0059] The local candidate generation module 105 receives the anchor pose 204 determined by the uncertainty-aware routing module 104 and generates multiple local candidate poses 205 around the anchor pose 204. The local candidate poses 205 are located in the local six-degree-of-freedom pose space near the anchor pose 204.
[0060] The tactile simulation module 106 generates corresponding simulated tactile observations 206 based on the geometric model of the target object, the contact model of the tactile sensor, and the candidate poses of each locality 205.
[0061] The cross-modal tactile refinement module 107 receives real tactile observations 202, simulated tactile observations 206, and visual features extracted by the visual anchor point generation module 103. It compares the features of real tactile observations 202 and simulated tactile observations 206, and outputs the pose residuals and confidence scores 207 corresponding to each local candidate pose 205 through cross-modal feature fusion with visual anchoring.
[0062] The pose output module 108 updates the corresponding local candidate pose 205 according to the pose residual, and selects the current six-DOF pose 208 from the updated local candidate poses based on the confidence level. The current six-DOF pose 208 is also saved as the pose of the previous time step for the next time step, and fed back to the uncertainty-aware routing module 104.
[0063] Through the above module relationships, the system can rely on visual global positioning when tactile feedback is unavailable, use tactile feedback to refine local pose when tactile feedback is available, and use the pose from the previous moment to maintain tracking continuity when the current vision is occluded or interfered with.
[0064] Example 2
[0065] like Figure 2 As shown, based on Example 1, this example provides a hierarchical visual-tactile fusion method for hand pose estimation, including the following steps:
[0066] S1, acquire visual and tactile observations. The visual acquisition module 101 acquires visual observations 201 of the target object, and the tactile acquisition module 102 acquires tactile observations 202 generated by the robot's end effector contacting the target object.
[0067] S2, generate visual pose hypotheses and visual anchor points, perform target recognition or target region segmentation on visual observation 201 to obtain target object mask; generate multiple visual pose hypotheses 203 based on visual observation 201 and target object mask, and determine the score of each visual pose hypothesis 203.
[0068] The visual pose hypothesis with the highest score is determined as the visual anchor pose. The visual anchor pose is used to provide the global position and orientation of the target object and to define the local search area for subsequent haptic refinement.
[0069] S3. Determine whether tactile observation is available. Based on the signal amplitude, contact area, pressure value, gripper status, or robot controller feedback of tactile observation 202, determine whether effective contact is formed between the robot's end effector and the target object. If tactile observation 202 is unavailable, execute step S4a; if tactile observation 202 is available, execute step S4b.
[0070] S4a outputs visual pose. When tactile observation 202 is unavailable, the tactile refinement process is not started. Instead, the visual anchor pose is directly output as the current six-degree-of-freedom pose 208. This processing method is suitable for the pre-grab, non-contact stage, or stage where the tactile sensor only outputs background noise. It can avoid invalid tactile signals from interfering with the visual estimation results.
[0071] S4b, Calculate visual uncertainty. When tactile observation 202 is available, select the top k visual pose hypotheses with the highest scores from multiple visual pose hypotheses 203, and calculate the visual uncertainty based on the degree of dispersion of the top k visual pose hypotheses in translation space and rotation space:
[0072] ;
[0073] in, Indicates visual uncertainty; Indicates the number of visual pose assumptions involved in the calculation; and They represent the first Translation vectors and rotation matrices for each visual pose hypothesis; and They represent the preceding Average translation and average rotation of a visual pose hypothesis; This indicates the rotational geodesic distance.
[0074] Visual uncertainty is used to evaluate the reliability of the current visual anchor point: the more concentrated the multiple visual pose assumptions are, the more reliable the visual results are; the more dispersed the multiple visual pose assumptions are, the greater the possibility that the current visual observation is affected by occlusion, reflection or segmentation error.
[0075] S5, selects the anchoring posture based on visual reliability, and addresses visual uncertainty. Compared with the preset visual uncertainty threshold Compare them.
[0076] When the following conditions are met: If the current visual result is deemed reliable, the current visual anchor pose is selected as the anchor pose 204.
[0077] When the following conditions are met: If the current visual result is determined to be unreliable, the pose of the previous moment is selected as the anchor pose 204.
[0078] S6, Generate local candidate poses, and generate multiple local candidate poses 205 around the selected anchor pose 204.
[0079] In one implementation, multiple six-dimensional perturbations are sampled in the Lie algebraic tangent space corresponding to SE(3), and local candidate poses are obtained through exponential mapping:
[0080] ;
[0081] in, Indicates the first One local candidate pose; This indicates the anchor positioning posture as 204. This represents a six-dimensional perturbation vector that includes three-dimensional translational perturbations and three-dimensional rotational perturbations; The matrix form represents the six-dimensional perturbation vector.
[0082] S7, Generate simulated tactile observations. Based on the geometric model of the target object, the contact model of the tactile sensor, and the candidate poses of each locality 205, generate corresponding simulated tactile observations 206 respectively:
[0083] ;
[0084] in, Indicates the first Simulated tactile observations corresponding to local candidate poses; Represents a tactile simulation function; Represents the geometric model of the target object.
[0085] Simulated tactile observations 206 can be obtained through a finite element tactile simulator, a geometric contact model, a differentiable tactile renderer, or a learning-based tactile generation model.
[0086] S8, perform cross-modal tactile refinement, encode features for real tactile observation 202 and simulated tactile observation 206 respectively, and obtain tactile features to characterize the contact geometry differences between the two; at the same time, extract visual features from visual observation 201 or visual anchor point generation module 103.
[0087] Perform cross-attention using tactile features as queries and visual features as keys and values:
[0088] ;
[0089] in, This indicates the tactile characteristics after fusion; Indicates tactile characteristics; Indicates visual features; , and These represent the projection matrices corresponding to the query, key, and value, respectively. By injecting the visual global context into the tactile local features, candidate poses that are similar in local contact shape but have unreasonable global poses can be suppressed.
[0090] S9 outputs the current six-DOF pose. The cross-modal haptic refinement module 107 outputs the translation residual, rotation residual, and confidence level corresponding to each local candidate pose 205.
[0091] ;
[0092] in, Indicates the translation residual; Represents rotational residuals; Indicates confidence level; This represents a tactile detailing network.
[0093] Applying the translation and rotation residuals to the corresponding local candidate pose 205 yields the updated pose. The updated pose with the highest confidence is then used as the current six-DOF pose 208. After the current six-DOF pose 208 is output, it is saved as the previous pose for the next time step, so that it can be used when the vision is unreliable.
[0094] Example 3
[0095] like Figure 3 As shown, the visual anchor point generation module 103 generates multiple visual pose hypotheses 203 based on visual observations 201, and determines the anchor pose 204 from the multiple visual pose hypotheses 203.
[0096] The local candidate generation module 105 performs local sampling around the anchor pose 204 to generate multiple local candidate poses 205. The local candidate poses 205 are used to characterize multiple possible pose states near the anchor pose 204.
[0097] The tactile simulation module 106 generates corresponding simulated tactile observations 206 for each local candidate pose 205. The tactile simulation process is used to predict the contact deformation or contact distribution that the tactile sensor should theoretically form when the target object is in the corresponding local candidate pose.
[0098] Real tactile observation 202 and simulated tactile observation 206 are input into a tactile encoder with shared weights to extract real tactile features and simulated tactile features, respectively. Based on the differences between real tactile features and simulated tactile features, tactile difference features of the corresponding local candidate pose 205 are obtained.
[0099] Visual features are derived from visual observation 201 or visual anchor point generation module 103. The cross-attention unit uses tactile difference features as queries and visual features as keys and values to inject global visual geometric information into local tactile contact features.
[0100] The residual regression unit outputs the translation and rotation residuals corresponding to each local candidate pose 205 based on the fused features; the confidence prediction unit outputs the confidence level corresponding to each local candidate pose 205.
[0101] The pose residual and confidence level together constitute the output information 207. Based on this, the pose output module 108 updates and optimizes the local candidate poses 205 to determine the current six-degree-of-freedom pose 208.
[0102] Example 4
[0103] like Figure 4 As shown, the uncertainty-aware routing module 104 first determines whether touch is available based on the touch observation 202.
[0104] When tactile observation 202 is unavailable, it indicates that there has been no effective contact between the robot's end effector and the target object. The system directly outputs the visual anchor pose to complete the pose estimation in the non-contact stage.
[0105] When tactile observation 202 is available, the top k visual pose hypotheses with the highest scores are selected from the visual pose hypotheses 203, and the visual uncertainty is calculated based on the degree of translational and rotational discretization. .
[0106] when When the current visual pose assumption is relatively concentrated and the current visual observation is reliable, the visual anchor pose is selected as the anchor pose 204, and the tactile refinement process begins.
[0107] when When the current visual pose assumption is relatively scattered, the current visual observation may be affected by occlusion, reflection, lack of depth or target segmentation error. The pose of the previous moment is selected as the anchor pose 204, and the tactile refinement process begins.
[0108] The routing method described above does not simply weight the visual and tactile results, but rather determines the generation center of the subsequent local candidate pose 205 by combining tactile availability and visual uncertainty. After the anchor pose 204 changes, the corresponding local candidate pose 205, simulated tactile observation 206, pose residual, and confidence level 207 all change accordingly, enabling the system to select an estimation path that is compatible with the current sensor reliability at different operational stages.
[0109] The foregoing has only described certain exemplary embodiments of the present invention by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.
Claims
1. A hierarchical visual-tactile fusion method for hand pose estimation, characterized in that, Includes the following steps: Acquire visual and tactile observations during the robot's manipulation of the target object; Multiple visual pose hypotheses and their scores are generated based on the visual observations, and the visual pose hypothesis with the highest score is determined as the visual anchor pose. Determine whether the tactile observation meets the tactile availability conditions; if not, determine the visual anchor pose as the current six-degree-of-freedom pose. When the condition is met, the visual uncertainty is calculated based on the degree of dispersion of the top k visual poses with the highest scores in translation and rotation space. When the visual uncertainty does not exceed a preset threshold, the visual anchor pose is selected, and when it exceeds the preset threshold, the pose of the previous moment is selected as the tactile refinement anchor pose. Multiple local candidate poses are generated around the tactile refined anchor pose, and simulated tactile observations corresponding to each local candidate pose are generated based on the geometric model of the target object and the contact model of the tactile sensor. Extract tactile difference features between real tactile observation and simulated tactile observation, use the tactile difference features as queries and visual features as keys and values to perform cross-attention fusion, and output the pose residuals and confidence of each local candidate pose; Update each local candidate pose according to the pose residual, and determine the updated pose with the highest confidence as the current six-degree-of-freedom pose, and save it as the previous pose for the next time step.
2. The on-hand pose estimation method for hierarchical visual-tactile fusion according to claim 1, characterized in that, The process of generating multiple visual pose hypotheses and their scores based on visual observation includes: segmenting the target region of the RGB-D image to obtain a target object mask; inputting the RGB-D image, the target object mask, and the target object geometric model into a visual pose estimation network to generate multiple visual pose hypotheses including three-dimensional translation and three-dimensional rotation; and using a pose scoring network to determine the score of each visual pose hypothesis.
3. The on-hand pose estimation method for hierarchical visual-tactile fusion according to claim 1, characterized in that, The tactile availability condition is determined based on at least one of the following information: the amplitude of change of the tactile image relative to the non-contact background image, the effective contact area in the tactile image, the pressure value or contact force value output by the tactile sensor, the closing state of the robot gripper, and the contact feedback output by the robot controller; when the corresponding information reaches a preset contact threshold, it is determined that the tactile observation meets the tactile availability condition.
4. The on-hand pose estimation method for hierarchical visual-tactile fusion according to claim 1, characterized in that, The visual uncertainty satisfies: ; in, Indicates visual uncertainty; Indicates the number of visual pose assumptions involved in the calculation; and They represent the first Translation vectors and rotation matrices for each visual pose hypothesis; and They represent the preceding Average translation and average rotation of a visual pose hypothesis; This indicates the rotational geodesic distance.
5. The on-hand pose estimation method for hierarchical visual-tactile fusion according to claim 1, characterized in that, The generation of multiple local candidate poses around the haptic refinement anchor pose includes: sampling multiple six-dimensional perturbation vectors, each including three-dimensional translational perturbation and three-dimensional rotational perturbation, in the tangent space of the Lie algebra SE(3), and generating local candidate poses according to the following formula: ; in, Indicates the first One local candidate pose; This indicates the anchor positioning posture as 204. This represents a six-dimensional perturbation vector that includes three-dimensional translational perturbations and three-dimensional rotational perturbations; The matrix form represents the six-dimensional perturbation vector.
6. The on-hand pose estimation method for hierarchical visual-tactile fusion according to claim 1, characterized in that, The simulated tactile observation is generated based on the contact position, contact depth, contact area, and deformation relationship between the surface of the target object and the surface of the tactile sensor under local candidate poses, and has the same or spatially aligned data form as the real tactile observation; the simulated tactile observation is a tactile depth map or a tactile deformation map, and is generated by a finite element tactile simulator, a geometric contact model, a differentiable tactile renderer, or a learning tactile generation model.
7. The on-hand pose estimation method for hierarchical visual-tactile fusion according to claim 1, characterized in that, The extraction of tactile difference features includes: extracting real tactile observation features and simulated tactile observation features using a tactile encoder with shared weights, and obtaining tactile difference features based on the feature differences between the two; the cross-attention fusion satisfies: ; in, This indicates the tactile characteristics after fusion; Indicates tactile characteristics; Indicates visual features; , and These represent the projection matrices corresponding to the query, key, and value, respectively. By injecting the visual global context into the tactile local features, candidate poses that are similar in local contact shape but have unreasonable global poses can be suppressed.
8. The on-hand pose estimation method for hierarchical visual-tactile fusion according to claim 1, characterized in that, The pose residuals include translation residuals and rotation residuals; the training loss of the haptic refinement network includes pose residual regression loss and confidence regression loss, and the confidence labels satisfy: ; ; in, For confidence level labels, For combined pose error, and These represent the translation and rotation errors of the updated pose relative to the true pose, respectively. and These are the translation error weights and rotation error weights, respectively. This refers to the temperature parameter.
9. A hierarchical visual-tactile fusion hand position estimation system, used to execute the hierarchical visual-tactile fusion hand position estimation method according to any one of claims 1-8, characterized in that, include: The visual acquisition module is used to acquire visual observations of the target object; The tactile sensing module is used to acquire tactile observations between the robot's end effector and the target object; The visual anchor point generation module is used to generate multiple visual pose hypotheses and their scores based on the visual observations, and to determine the visual anchor pose and visual features. An uncertainty-aware routing module is used to determine whether the tactile observation is available. When the tactile observation is unavailable, it outputs the visual anchor pose. When the tactile observation is available, it calculates the visual uncertainty based on the translational and rotational discreteness of multiple visual pose assumptions, and selects the tactile refined anchor pose between the visual anchor pose and the pose at the previous moment based on the visual uncertainty. A local candidate generation module is used to generate multiple local candidate poses around the tactile refined anchor pose. The tactile simulation module is used to generate corresponding simulated tactile observations based on the geometric model of the target object, the contact model of the tactile sensor, and the candidate poses of each local area. The cross-modal tactile refinement module is used to extract tactile difference features between real tactile observations and simulated tactile observations, and to perform cross-attention fusion of the tactile difference features with visual features to output the pose residuals and confidence scores corresponding to each local candidate pose. The pose output module is used to update each local candidate pose according to the pose residual, determine the updated pose with the highest confidence as the current six-degree-of-freedom pose, and feed the current six-degree-of-freedom pose back to the uncertainty-aware routing module as the previous pose for the next moment.
10. The hierarchical visual-tactile fusion in-hand pose estimation system according to claim 9, characterized in that, The cross-modal tactile refinement module includes a shared-weight tactile encoder, a visual feature extraction unit, a cross-attention unit, a residual regression unit, and a confidence prediction unit. The shared-weight tactile encoder is used to encode real tactile observations and simulated tactile observations respectively. The cross-attention unit uses the tactile difference features formed by the two as queries and visual features as keys and values for fusion. The residual regression unit is used to output translation residuals and rotation residuals. The confidence prediction unit is used to output the confidence of each local candidate pose.