Object pose estimation method

By fusing object structure, hand posture, and contact area features, the problem of pose estimation error caused by visual sensor occlusion during robot hand grasping is solved, achieving higher accuracy.

CN121962238APending Publication Date: 2026-05-01PAXINI TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PAXINI TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2025-12-16
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In existing technologies, when a robot's hand grasps an object, the visual sensor suffers from a loss of visual features due to occlusion, leading to accumulated pose estimation errors and tracking failure, resulting in low accuracy.

Method used

By acquiring the structural features of the target object, the hand pose features of the robot hand, and the contact area features between the hand and the target object, and fusing these features in a multi-dimensional manner to form a global fusion feature, pose estimation can be performed.

Benefits of technology

It effectively compensates for the information gaps caused by the lack of visual features, reduces error accumulation and tracking failure, and improves the accuracy of pose estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121962238A_ABST
    Figure CN121962238A_ABST
Patent Text Reader

Abstract

The invention provides an object pose estimation method. The object pose estimation method comprises the steps of obtaining object structure features of a target object, hand pose features of a robot hand and contact area features of the hand in contact with the target object; fusing the object structure feature and the contact area feature to obtain a first fusion feature; fusing the hand posture feature with the contact area feature to obtain a second fused feature; fusing the first fusion feature and the second fusion feature to obtain a global fusion feature; and performing pose estimation based on the global fusion features to obtain a target pose of the target object. According to the embodiment of the invention, the accuracy of object pose estimation can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robotics, and in particular to a method for estimating the pose of an object. Background Technology

[0002] Six-dimensional (6D) pose estimation is a core prerequisite for robots to complete hand operations (such as grasping and object repositioning). The accuracy of pose estimation directly determines the success rate of the operation task. Only by accurately obtaining the spatial position and rotational attitude of the object can the robot stably control the hand movements and avoid the object falling off or the operation deviation.

[0003] Currently, visual modality is the mainstream technique for pose estimation: these methods rely on red, green, blue-depth (RGB-D) cameras to acquire object images, extract visual features through deep learning models, and combine them with labeled datasets to complete pose estimation.

[0004] However, when a robot grasps an object with its hands, its fingers will cover most of the object's surface, causing the visual sensor to be unable to acquire complete object features. The image information it originally relied on is missing, which in turn causes the accumulation of pose estimation errors or even tracking failure, resulting in low accuracy of pose estimation. Summary of the Invention

[0005] This application provides a method for estimating the pose of an object, which can improve the accuracy of object pose estimation.

[0006] In a first aspect, embodiments of this application provide a method for estimating the pose of an object, the method comprising: Acquire the structural features of the target object, the hand posture features of the robot hand, and the contact area features of the hand contacting the target object; The first fused feature is obtained by fusing the object's structural features with the contact area features; the second fused feature is obtained by fusing the hand posture features with the contact area features. The first and second fusion features are fused to obtain the global fusion feature; Pose estimation is performed based on global fusion features to obtain the target pose of the target object.

[0007] Secondly, this application provides a pose estimation device for an object, the device comprising: The acquisition module is used to acquire the object structure features of the target object, the hand posture features of the robot hand, and the contact area features of the hand contacting the target object; The first fusion module is used to fuse the object's structural features with the contact area features to obtain a first fused feature; and to fuse the hand posture features with the contact area features to obtain a second fused feature. The second fusion module is used to fuse the first fusion feature and the second fusion feature to obtain a global fusion feature; The estimation module is used to perform pose estimation based on the global fusion features to obtain the target pose of the target object.

[0008] Thirdly, embodiments of this application provide an electronic device, which includes: a processor and a memory storing computer program instructions; When the processor executes computer program instructions, it implements a pose estimation method for an object as described in any of the embodiments of the first aspect.

[0009] Fourthly, embodiments of this application provide a computer storage medium storing computer program instructions, which, when executed by a processor, implement the pose estimation method for an object as described in any of the embodiments of the first aspect.

[0010] Fifthly, embodiments of this application provide a computer program product in which instructions, when executed by a processor of an electronic device, cause the electronic device to perform an object pose estimation method as described in any of the embodiments of the first aspect above.

[0011] In the object pose estimation method provided in this application embodiment, the object structure features, robot hand posture features, and contact area features between the hand and the object are acquired. Therefore, in scenarios where visual features are missing due to the robot hand grasping the object, the contact area features can reflect the actual contact state between the hand and the object, supplementing the lack of visual information and avoiding or reducing estimation bias when a single visual feature is missing. The object structure features and contact area features are fused to obtain a first fused feature, realizing the association between the object structure features and the contact state; the hand posture features and contact area features are fused to obtain a second fused feature, establishing a mapping relationship between the hand posture and the contact state; finally, the first and second fused features are further fused to obtain a global fused feature. This global fused feature integrates the association information of the object, hand, and contact, forming a comprehensive feature representation that can completely characterize the object pose. Pose estimation based on this global fused feature can fully utilize the complementary advantages of multi-dimensional features, effectively compensate for the information gaps caused by missing visual features, avoid or reduce the error accumulation and tracking failure problems caused by occlusion in traditional visual pose prediction, thereby improving the accuracy of object pose estimation. Attached Figure Description

[0012] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This is one of the flowcharts illustrating the object pose estimation method provided in the embodiments of this application; Figure 2 This is a second schematic flowchart of the object pose estimation method provided in the embodiments of this application; Figure 3 This is a schematic diagram of the structure of an object pose estimation device provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0014] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0015] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.

[0016] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the term "comprising" or any other variations thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.

[0017] 6D pose estimation of an object is a core prerequisite for a robot to complete hand operations. The accuracy of pose estimation directly determines the success rate of the operation task. Only by accurately obtaining the spatial position and rotational attitude of the object can the robot stably control the hand movements and avoid the object falling off or the operation deviating.

[0018] Currently, visual modality is the mainstream technique for pose estimation: these methods rely on RGB-D cameras to acquire object images, extract visual features through deep learning models, and combine them with labeled datasets to complete pose estimation. However, when a robot grasps an object with its hands, its fingers may obscure most of the object's surface, preventing the visual sensor from acquiring complete object features. This results in missing image information, leading to accumulated pose estimation errors or even tracking failure, ultimately resulting in low accuracy in pose estimation.

[0019] Based on this, although related technologies have attempted to extend to single tactile modalities or multimodal fusion, there are still many limitations.

[0020] While tactile modalities can compensate for visual occlusion, they only provide local information about the contact point and lack global spatial context of the object. They are prone to pose drift and their performance is affected by factors such as sensor type and aging.

[0021] The visual, tactile, and proprioceptive fusion methods in related technologies also share significant common shortcomings: most methods only utilize fingertip tactile sensation and do not fully utilize contact information from larger areas such as the fingertips, resulting in the inability to recover complete object features even in highly occluded scenarios; visual, tactile, and joint pose do not achieve effective spatiotemporal alignment or deep coupling, leading to low fusion efficiency and robustness; related technologies often rely on specific hand structures or fixed contact patterns, making it difficult to adapt to different hand shapes, multi-finger grasping, or complex occlusion situations; the simulation pipelines of related technologies are relatively idealized in terms of visual rendering and tactile simulation, making it difficult to directly transfer to real robots; moreover, the hand-object segmentation of related technologies still relies on a single visual model or manual prompts, failing to fully utilize tactile points as auxiliary information, and the multimodal collaborative perception system is still incomplete.

[0022] To address the problems existing in related technologies, embodiments of this application provide a method for estimating the pose of an object.

[0023] The object pose estimation method provided in the embodiments of this application will be introduced first below. For example... Figure 1 As shown, the method specifically includes the following steps: S100: Obtain the object structure features of the target object, the hand posture features of the robot hand, and the contact area features of the hand contacting the target object.

[0024] Optionally, in the embodiments of this application, the target object refers to the core object of pose estimation in the embodiments of this application, that is, the specific object that the robot hand is grasping, about to operate, or in an interactive state.

[0025] Object structural features are characteristic information that represents the inherent geometric properties of a target object. They are abstract features extracted from the target object's three-dimensional data (such as visual point clouds). Object structural features can reflect the inherent properties of the target object, such as shape, size, surface contour, and key structural parts, and do not change with the object's position, posture, or external contact state.

[0026] Hand pose features are characteristic information representing the spatial posture and structural morphology of a robot's hand, which can be extracted based on robot hand model data (such as hand model point clouds). Hand pose features can reflect the overall spatial position of the robot's hand, the distribution of finger joint angles, and the relative positional relationship between the hand and the target object, clarifying the specific posture of the hand during the interaction process.

[0027] Contact area features are characteristic information representing the state of the contact area between the robot hand and the target object. They can be extracted based on sensor data (such as tactile point clouds) when the hand contacts the target object. Contact area features can reflect contact state information such as the distribution of contact positions between the hand and the target object, the spatial arrangement of contact points, and the correlation of contact intensity.

[0028] Optionally, in one feasible implementation of this application, the target object can be photographed in real time using an RGB-D camera to collect three-dimensional point cloud data of the object's surface. The point cloud data is then encoded and abstracted using a feature extraction network to extract object structural features that can characterize the object's shape, contour, key structures, and other inherent attributes. Data such as hand joint angles and motion states are collected using the robot's built-in body sensors. Combined with a predefined hand model, a three-dimensional point cloud of the hand is generated. After feature extraction, hand posture features reflecting the hand's spatial posture and finger arrangement are obtained. The tactile sensor array mounted on the hand is used to perceive the contact point position and distribution when the hand contacts the target object in real time, forming a tactile point cloud. After feature encoding, contact area features characterizing the contact state between the two are obtained. The three types of data acquisition processes strictly follow the time synchronization principle to ensure that the scene state corresponding to the features is consistent.

[0029] Optionally, in another feasible implementation of this application, multimodal data of different target objects, different hand postures, and different contact scenarios can be collected in advance. After data cleaning, a standardized dataset is constructed. The dataset is processed by an offline trained feature extraction model to establish a feature library and store object structure feature templates, hand posture feature templates, and typical contact area feature templates. During real-time pose estimation, simplified visual point clouds of target objects, key hand joint data, and real-time tactile point clouds of the current scene are collected by sensors. The collected data is matched with the templates in the feature library to obtain object structure features, hand posture features, and contact area features that are suitable for the current scene.

[0030] S200, the object structural features and contact area features are fused to obtain a first fused feature; the hand posture features and contact area features are fused to obtain a second fused feature.

[0031] Optionally, in this embodiment, the first fusion feature is the feature information obtained by fusing the object structure features of the target object with the contact area features of the hand contacting the target object. A correlation is established between the inherent geometric properties of the object and the contact state. Through the fusion process, the inherent properties represented by the object structure features, such as shape, contour, and key structures, are made complementary and correlated with the interactive information reflected in the contact area features, such as the distribution of contact positions and the spatial arrangement of contact points.

[0032] The second fusion feature is the feature information obtained by fusing the robot hand's posture features with the contact area features of the hand contacting the target object. A mapping relationship between the hand's spatial posture and the contact state is constructed. During the fusion process, information such as the overall spatial position of the hand, finger joint angles, and hand shape contained in the hand posture features are deeply coupled with interactive information such as the contact distribution and contact point association represented by the contact area features. This feature not only clearly defines the robot hand's own spatial posture state but also clearly associates it with the spatial logic of "how the hand contacts the object," accurately reflecting the adaptation relationship between hand posture and contact state.

[0033] Optionally, in one feasible implementation of this application, the object structure features, hand posture features, and contact area features can be uniformly processed to ensure that the feature vectors of the three are consistent in dimension. Then, the object structure features and contact area features are directly concatenated to form a joint feature vector containing the object's inherent attributes and contact state, which is the first fusion feature. Similarly, the hand posture features and contact area features are concatenated in the same way to obtain the second fusion feature that integrates hand posture and contact information.

[0034] In another feasible implementation of this application, the weight coefficients of object structure features and contact area features can be learned, and weights can be assigned according to the contribution of features to subsequent pose estimation. The two types of features are weighted and summed to obtain the first fusion feature. The same logic is used to learn the weights of hand pose features and contact area features, and the second fusion feature is obtained by weighted calculation.

[0035] S300, the first fusion feature and the second fusion feature are fused to obtain a global fusion feature.

[0036] Optionally, in this embodiment, the global fusion feature is a comprehensive feature information that fully characterizes the target object, the robot hand, and their contact relationship by fusing the first fusion feature (the fusion result of object structural features and contact area features) and the second fusion feature (the fusion result of hand posture features and contact area features). It integrates two types of related information, "object-contact" and "hand-contact," and achieves complementarity and synergy of multi-dimensional information through feature interaction. This preserves the inherent structural attributes of the target object and the spatial posture state of the robot hand, while also clarifying the spatial logic and relationship of their contact, forming a complete feature representation covering the core information of the object, hand, and contact.

[0037] Optionally, in one feasible implementation of this application, the weight coefficients of the two fusion features can be learned, and the weights can be assigned according to the difference in contribution of the two types of fusion features to pose estimation. The weight coefficients will be dynamically adjusted according to different scenarios. Then, the first fusion feature and the second fusion feature are weighted and summed according to their corresponding weights to obtain the global fusion feature. This can highlight the role of key related information and suppress redundant interference.

[0038] S400, based on the global fusion features, pose estimation is performed to obtain the target pose of the target object.

[0039] Optionally, in this embodiment of the application, the target pose is based on global fusion features and is the six-dimensional pose information that can completely characterize the state of the target object in three-dimensional space, which is finally output by the pose estimation process. Specifically, it includes two core contents: the three-dimensional spatial position and the three-dimensional spatial rotation of the target object.

[0040] Optionally, in one feasible implementation of this application, a convolutional neural network is first used as a pose decoding network. The global fusion features obtained by S300 are input into the network. The decoding network gradually maps the feature dimensions through fully connected layers and outputs the three-dimensional position prediction value (x, y, z) and three-dimensional rotation related parameters (such as rotation matrix, quaternion or Euler angle) of the target object. At the same time, multiple pose candidate solutions and corresponding confidence scores are generated to form a preliminary pose prediction result.

[0041] Subsequently, a candidate selection mechanism is introduced to sort multiple pose candidate solutions based on confidence scores, select the Top-K candidate solutions with the highest confidence scores, and exclude abnormal candidates that deviate significantly from the reasonable range. Then, by calculating the matching degree between the candidate solutions and the object's structural features and contact area features, the candidate selection is further optimized, and the candidate solutions with the highest consistency with multi-dimensional related information are retained.

[0042] In the object pose estimation method provided in this application embodiment, the object structure features, robot hand posture features, and contact area features between the hand and the object are acquired. Therefore, in scenarios where visual features are missing due to the robot hand grasping the object, the contact area features can reflect the actual contact state between the hand and the object, supplementing the lack of visual information and avoiding or reducing estimation bias when a single visual feature is missing. The object structure features and contact area features are fused to obtain a first fused feature, realizing the association between the object structure features and the contact state; the hand posture features and contact area features are fused to obtain a second fused feature, establishing a mapping relationship between the hand posture and the contact state; finally, the first and second fused features are further fused to obtain a global fused feature. This global fused feature integrates the association information of the object, hand, and contact, forming a comprehensive feature representation that can completely characterize the object pose. Pose estimation based on this global fused feature can fully utilize the complementary advantages of multi-dimensional features, effectively compensate for the information gaps caused by missing visual features, avoid or reduce the error accumulation and tracking failure problems caused by occlusion in traditional visual pose prediction, thereby improving the accuracy of object pose estimation.

[0043] In one embodiment, acquiring the object structure features of the target object, the hand posture features of the robot hand, and the contact area features of the hand contacting the target object includes: The visual point cloud of the target object, the hand model point cloud of the hand, and the tactile point cloud of the hand in contact with the target object are obtained. The visual point cloud of the object, the hand model point cloud, and the tactile point cloud of the object are respectively encoded to obtain initial visual point cloud features, initial hand point cloud features, and initial tactile point cloud features; A self-attention mechanism is applied to the initial visual point cloud features, the initial hand point cloud features, and the initial tactile point cloud features respectively to obtain the object structure features, the hand posture features, and the contact area features.

[0044] Optionally, in this embodiment, the object visual point cloud is a set of three-dimensional data generated after the target object is acquired by a visual sensor (such as an RGB-D camera), consisting of a large number of discrete three-dimensional coordinate points (x, y, z). These coordinate points together delineate the surface contour, spatial shape, and geometric structure of the target object.

[0045] The hand model point cloud is a set of three-dimensional points generated based on a preset three-dimensional model of the robot hand or real-time sensing data. It includes the three-dimensional coordinate information of the entire hand and each finger and joint.

[0046] A tactile point cloud is a collection of three-dimensional data related to contact, collected by the tactile sensor array mounted on the robot's hand when it comes into contact with a target object. Each point cloud data includes information such as the three-dimensional spatial coordinates of the contact point, the contact pressure value, and the activation status of the sensing unit.

[0047] The initial visual point cloud features are the basic feature representations obtained after preliminary feature encoding of the object's visual point cloud. Compared with the final object structural features, they lack targeted information enhancement and key feature focus, and can only initially reflect the object's surface geometric attributes; while the object structural features suppress redundant information and can more accurately represent the object's core inherent geometric attributes.

[0048] The initial hand point cloud features are the basic feature representations obtained after preliminary feature encoding of the hand model point cloud. The core difference between them and hand pose features is that the initial hand point cloud features are only a basic abstraction of the hand structure and do not highlight the key pose information related to pose estimation; while hand pose features can more accurately reflect the real-time spatial state of the hand and provide effective support for establishing the association between the hand and the object.

[0049] Initial tactile point cloud features are the basic feature representations obtained after preliminary feature encoding of the object's tactile point cloud. Compared with contact area features, initial tactile point cloud features can only initially reflect the basic state of contact and cannot highlight information such as contact position association and core contact point distribution that are important for pose estimation; while contact area features can more accurately characterize the contact interaction state between the hand and the object, providing a data foundation to compensate for the lack of visual information.

[0050] Optionally, in one specific implementation of this application, multimodal raw point cloud acquisition and preprocessing are first performed. The object visual point cloud is directly acquired by an RGB-D camera, which acquires 2048 three-dimensional coordinate points on the surface of the target object. The hand model point cloud is extracted from a predefined robot hand model, which acquires 2048 three-dimensional coordinate points. The object tactile point cloud is acquired by acquiring the three-dimensional coordinates of 256 contact points through a hand tactile sensor array.

[0051] Subsequently, a unified data augmentation operation was performed on the three types of original point clouds. After the point cloud was completed using a Point Fractal Network (PF Net), multi-scale downsampling was performed. The object and hand point clouds were compressed in a three-level process of 2048 points → 512 points → 256 points, and the tactile point cloud was sampled in a process of 256 points → 128 points → 64 points. At the same time, Gaussian noise with a standard deviation of 0.005 was added to all point clouds to simulate the error of real sensors. During the training phase, the real device and simulation data sources were dynamically switched with a 50% probability to improve generalization. If rotation data was involved, the quaternion was converted into a rotation matrix and the first two columns were extracted as a 6-dimensional rotation representation.

[0052] For example, if the quaternion is (w,x,y,z), it can be converted into a rotation matrix as follows: R = [ [1-2y²-2z², 2xy-2zw, 2xz+2yw], [2xy+2zw, 1-2x²-2z², 2yz-2xw], [2xz-2yw, 2yz+2xw, 1-2x²-2y²] ].

[0053] Extract the first two columns (6 dimensions) of the rotation matrix as the rotation representation. That is, the elements of the first column are: [1-2y²-2z², 2xy+2zw, 2xz-2yw] (corresponding to the 3 elements of the first column of the matrix); the elements of the second column are: [2xy-2zw, 1-2x²-2z², 2yz+2xw] (corresponding to the 3 elements of the second column of the matrix).

[0054] Next, the point cloud feature encoding stage is adopted. PointNet is used to process the three types of preprocessed point clouds independently. Through three layers of feature abstraction (object visual point cloud and hand model point cloud are 2048→512→256, and object tactile point cloud is adapted to its sampling dimension), the object visual point cloud, hand model point cloud and object tactile point cloud are converted into 64-dimensional initial visual point cloud features, initial hand point cloud features and initial tactile point cloud features respectively, realizing the initial abstraction of point cloud data.

[0055] Finally, a self-attention mechanism enhancement step is executed, in which the enhanced object structure features, hand pose features, and contact area features are output sequentially through the encoder layer of the Transformer for each type of initial point cloud features, so that the features focus on the key local structures related to pose estimation.

[0056] In this application embodiment, addressing the issues of spatiotemporal misalignment and low fusion efficiency of visual, tactile, and proprioceptive data in related methods, this application aligns multimodal data through a preprocessing stage. Temporal alignment employs synchronous acquisition of visual image frames and tactile force data under the same timestamp. Spatial alignment utilizes dexterous hand proprioception (joint angles) to calculate the spatial arrangement of the object's tactile point cloud, ensuring that the object's tactile point cloud, hand model point cloud, and object's visual point cloud coordinate systems are consistent. This design avoids information bias during cross-modal feature fusion, enabling the Transformer module to efficiently learn intermodal relationships, achieving multimodal spatiotemporal alignment, and improving fusion efficiency.

[0057] In these alternative embodiments, multimodal point cloud data acquisition, combined with feature encoding and self-attention mechanisms, comprehensively acquires core information about objects, hands, and contacts. This provides high-quality feature support for subsequent fusion and pose estimation, effectively improving the accuracy and reliability of pose estimation in occluded scenarios.

[0058] In one embodiment, applying a self-attention mechanism to the initial visual point cloud features, the initial hand point cloud features, and the initial tactile point cloud features to obtain the object structure features, the hand posture features, and the contact area features includes: Generate a first query vector, a first key vector, and a first value vector corresponding to feature F; wherein feature F is any one of the initial visual point cloud feature, the initial hand point cloud feature, and the initial tactile point cloud feature; An attention matrix is ​​determined based on the first key vector and the first value vector; the attention matrix includes the attention value between any two point clouds in the feature F, and the attention value is used to characterize the degree of association importance between any two point clouds; Based on the attention matrix and the first query vector, an attention enhancement feature corresponding to the feature F is generated; the attention enhancement feature is any one of the object structure feature, the hand posture feature, and the contact area feature.

[0059] Optionally, in this embodiment, the first query vector is used to retrieve key content related to itself from the internal information of feature F. The first key vector is used to calculate the similarity with the first query vector, thereby determining the degree of association and matching between different point cloud features. The first value vector contains the specific feature information of each point cloud in feature F, and subsequently, based on the weight allocation of the attention matrix, this information is weighted and integrated to ultimately form an enhanced feature that focuses on key information.

[0060] The attention matrix is ​​a matrix calculated based on the first key vector and the first value vector, and its dimension corresponds to the number of point clouds in the feature F.

[0061] The attention value is a single element in the attention matrix, used to quantitatively characterize the degree of importance of the association between any two point clouds in feature F. Its value directly reflects the influence weight of one point cloud feature on another; the higher the value, the stronger the association between the two and the more critical their contribution to the final feature enhancement.

[0062] Optionally, in one specific implementation of this application, it can be implemented through a Transformer encoder layer, with the feature F (initial visual / hand / tactile point cloud features, all of which are N×D feature matrices composed of N D-dimensional points) as the processing object, as follows: First, linear projection is performed, and the feature F is input into three independent fully connected layers respectively. The first query vector Q, the first key vector K, and the first value vector V with dimension matching are generated through linear transformation, which provides the basis for attention association calculation.

[0063] Next, the attention matrix is ​​determined, and the attention weight matrix is ​​calculated. (D) k (where Q and K are the dimensions), element A in the matrix ij The attention value is a quantitative feature F that quantifies the importance of the relationship between point i and point j. A higher value indicates that the synergistic effect of the two on pose estimation is more critical.

[0064] Subsequently, attention-enhanced features are generated. First, the attention matrix A is used to weight and sum the first value vector V to obtain the preliminary enhanced features F. attn =AV, then perform residual join to F attn The feature is added to the original feature F and then normalized by layers to eliminate dimensionality differences. The feature is then further optimized by a feedforward network with two fully connected layers. Residual connections and layer normalization are performed again. The final output feature is the attention-enhanced feature, which corresponds to object structure features, hand pose features, or contact area features, thus achieving dynamic focusing on key local structures.

[0065] In these alternative embodiments, the importance of correlations between point cloud features is precisely quantified by generating an attention matrix. The self-attention mechanism allows the model to focus on feature correlations that are key to pose estimation, strengthening core information, suppressing redundancy, and improving the relevance and accuracy of feature representation.

[0066] In one embodiment, fusing the object structural features with the contact area features to obtain a first fused feature; and fusing the hand posture features with the contact area features to obtain a second fused feature, includes: The object structure features, hand posture features, and contact area features are fused with their respective point cloud spatial coordinates to generate object enhancement features, hand enhancement features, and tactile enhancement features. Based on the object enhancement features, the hand enhancement features, and the tactile enhancement features, they are fused through a cross-attention mechanism to obtain the first fused feature and the second fused feature.

[0067] Optionally, in this embodiment, object enhancement features make object features more closely match the actual spatial scene. Hand enhancement features clarify the hand's posture state and lock its spatial position, making the hand features more scene-related. Haptic enhancement features clearly locate the contact position and spatial distribution between the hand and the object.

[0068] Cross-attention is an attention mechanism used for cross-feature association fusion, and its core is to focus on the key associations between different features.

[0069] Optionally, in one specific implementation of this application, the feature and point cloud spatial coordinate fusion operation is first performed to match the original point cloud corresponding three-dimensional spatial coordinates for the object structure features, hand posture features, and contact area features (object structure features correspond to object visual point cloud coordinates, hand posture features correspond to hand model point cloud coordinates, and contact area features correspond to tactile point cloud coordinates). The feature vectors are combined with the corresponding coordinate vectors (x, y, z) using a feature concatenation method. For example, 64-dimensional object structure features are concatenated with 3-dimensional coordinates to form a 67-dimensional vector. The dimensions are unified and information integration is completed through a fully connected layer to generate object enhancement features, hand enhancement features, and tactile enhancement features, respectively.

[0070] Subsequently, a cross-attention mechanism is initiated for fusion. For the generation of the first fused feature, the object enhancement feature is used as the query vector (Q), and the tactile enhancement features are used as the key vector (K) and value vector (V). The similarity between Q and K is calculated to obtain an attention matrix, quantifying the importance of the association between them. Then, this matrix is ​​used to weight and sum V, and combined with residual connections and layer normalization, the output focuses on the first fused feature related to "object-contact". For the second fused feature, the hand enhancement feature is used as Q, and the tactile enhancement features are used as K and V. The above cross-attention calculation process is repeated to generate the second fused feature focusing on the "hand-contact" association.

[0071] In these alternative embodiments, features are fused with spatial coordinates, allowing the enhanced features to possess both attribute and location information, making them more relevant to the scene. The cross-attention mechanism accurately captures the correlation between object, hand, and contact features, and selectively fuses and generates fused features that focus on "object-contact" and "hand-contact," enhancing interaction information and improving pose estimation accuracy.

[0072] In one embodiment, fusing the object structural features, the hand pose features, and the contact area features with their respective corresponding point cloud spatial coordinates to generate object enhancement features, hand enhancement features, and tactile enhancement features includes: The object structural features, hand posture features, and contact area features are respectively stitched together with their corresponding point cloud spatial coordinates to obtain object stitching features, hand stitching features, and tactile stitching features; A self-attention mechanism is applied to the object stitching feature, the hand stitching feature, and the tactile stitching feature respectively to obtain the object enhancement feature, the hand enhancement feature, and the tactile enhancement feature.

[0073] Optionally, in one specific implementation of this application, the feature and point cloud spatial coordinates are first concatenated. For object structure features, hand pose features, and contact area features, their corresponding original point cloud three-dimensional spatial coordinates are extracted respectively. The D-dimensional feature vector of each point is directly concatenated with the 3-dimensional coordinate vector (x, y, z) to form (D+3)-dimensional object concatenation features, hand concatenation features, and tactile concatenation features. Subsequently, the three types of concatenation features are input into the standard Transformer encoder layer, and query, key, and value vectors are generated sequentially through linear projection. The attention matrix is ​​calculated to quantify the importance of the association between points. Then, after weighted summation, residual connection, layer normalization, and feedforward network processing, self-attention feature extraction is completed, and finally, object enhancement features, hand enhancement features, and tactile enhancement features are output respectively.

[0074] In these alternative embodiments, features are first concatenated with spatial coordinates to allow the features to carry location information; then, key associations are strengthened through a self-attention mechanism, and the resulting enhanced features combine attributes and spatial characteristics to improve pose estimation performance.

[0075] In one embodiment, based on the object enhancement features, the hand enhancement features, and the haptic enhancement features, cross-modal interaction fusion is performed through a cross-attention mechanism to obtain the first fused feature and the second fused feature, including: Using the object enhancement feature as the second query vector and the tactile enhancement feature as the second key vector and the second value vector, cross-attention calculation is performed to obtain the first fused feature; Using the hand enhancement features as the third query vector, the second key vector, and the second value vector, cross-attention calculation is performed to obtain the second fused feature.

[0076] Optionally, in one specific implementation of this application, the vector correspondence is first clarified, with the object enhancement feature as the second query vector Q1 and the hand enhancement feature as the third query vector Q2. The haptic enhancement feature is uniformly used as both the second key vector K1 and the second value vector V1 to ensure consistent interaction benchmarks. For the generation of the first fusion feature, Q1, K1, and V1 are linearly projected, and the similarity matrix between Q1 and K1 is calculated. Attention weights are obtained through softmax normalization, quantifying the importance of the association between each point of the object enhancement feature and the haptic feature. Then, V1 is weighted and summed, and the first fusion feature is output after residual connection and layer normalization. When generating the second fusion feature, Q2 is used as the query vector, and the same process is repeated for K1 and V1. Similarity calculation focuses on the matching relationship between hand parts and haptic information. After weight allocation, feature weighting, and normalization, the second fusion feature relating hand posture and haptic feedback is obtained.

[0077] In this application embodiment, in order to address the problems of haptic methods lacking global context, being prone to pose drift, and not being deeply coupled with proprioception, this application converts hand joint pose (proprioception) into hand model point cloud as an independent input modality. In cross-modal fusion, hand posture and contact force information are explicitly combined through "hand-haptic fusion features", which deeply couples proprioception and optimizes global pose estimation.

[0078] In these alternative embodiments, the haptic enhancement features are fixed as a second key vector and a second value vector, making them a “unified reference” for the interaction between the object and the hand features, since the haptic features themselves carry direct contact information between the object and the hand.

[0079] Meanwhile, object enhancement features and hand enhancement features are used as independent query vectors: object enhancement features are used as queries to actively retrieve contact information that matches the object's geometric attributes and spatial state from the tactile reference, focusing on "what kind of tactile feedback the object needs"; hand enhancement features are used as queries to accurately match the corresponding action adaptation logic in the tactile reference, focusing on "how the hand should respond to tactile feedback".

[0080] Compared to arbitrarily replacing query vectors (such as using tactile features as queries, or objects / hands as keys / values), this design makes the fused features more aligned with the core requirements of pose estimation, thereby improving the accuracy of pose estimation.

[0081] Using haptic enhancement features as the core hub, a cross-attention mechanism is used to establish directional interactions between objects, hands, and touch. This captures key associations between "object-touch" and "hand-touch," strengthening cross-modal interaction information and improving estimation accuracy.

[0082] In one embodiment, the first fusion feature and the second fusion feature are fused to obtain a global fusion feature, including: Using the first fusion feature as the fourth query vector and the second fusion feature as the third key vector and third value vector, cross-attention calculation is performed to obtain the global fusion feature.

[0083] Optionally, in one specific implementation of this application, the vector roles are first defined, with the first fused feature serving as the fourth query vector Q3, and the second fused feature serving as the third key vector K2 and the third value vector V2. Then, linear projection is performed on Q3, K2, and V2 through independent fully connected layers to ensure dimensionality matching. The similarity matrix between Q3 and K2 is calculated, and attention weights are obtained through softmax normalization to quantify the importance of the association between the "object-touch" and "hand-touch" features. V2 is weighted and summed using these weights, and then cross-attention calculation is completed through residual connections and layer normalization. Finally, global max pooling is performed on the output features to obtain a fixed-length global fused feature, integrating all preceding interaction information.

[0084] In these alternative embodiments, the first fusion feature is an "object-haptic" association feature, which corely carries the matching relationship between the geometric attributes of the object itself and haptic feedback. This is the "target core" of pose estimation, which ultimately revolves around the object's state. Using the first fusion feature as the query can proactively anchor the core question of "what kind of hand interaction does the object need?" If the second fusion feature (hand-haptic association) is used as the query, it is easy to fall into a local perspective of "what kind of interaction can the hand provide." Therefore, setting the first fusion feature as the query makes the fusion guided by the object's needs, ensuring that the global features are closely related to the target object, and improving the targeting of pose estimation.

[0085] In this application embodiment, addressing the problem of difficulty in feature extraction and soaring errors caused by hand occlusion in traditional vision methods during hand operations, this application utilizes multimodal fusion of visual point clouds, hand model point clouds, and tactile point clouds. By leveraging tactile sensors to directly perceive contact information, it compensates for missing data in visually occluded areas. Specifically, it employs point cloud completion, multi-scale feature extraction, and three-level Transformer cross-modal fusion to dynamically enhance object structural features with tactile features, significantly reducing pose errors caused by occlusion, effectively solving complex occlusion problems, and improving pose estimation accuracy.

[0086] In one embodiment, the step of estimating the pose based on the global fusion features to obtain the target pose of the target object includes: The target estimated position of the target object is obtained by using the position prediction head of the pose estimation model based on the position association features in the global fusion features; The rotation prediction head of the pose estimation model estimates multiple candidate rotation matrices of the target object and the confidence level of each candidate rotation matrix based on the rotation association features in the global fusion features. The candidate rotation matrix with the highest confidence among the multiple candidate rotations is determined as the target candidate rotation matrix; The target pose is determined based on the estimated target position and the target candidate rotation matrix.

[0087] Optionally, in the embodiments of this application, the pose estimation model is a deep learning model used to output the pose of the target object, and the core is to perform prediction based on global fusion features.

[0088] The position prediction head is a submodule of the pose estimation model, specifically responsible for extracting and predicting position information, and outputting the specific position coordinates of the target object in three-dimensional space.

[0089] Location association features are sub-features in the global fusion features that specifically characterize the spatial location of an object, integrating the spatial coordinates of the object and the hand, as well as their interactive positional relationships.

[0090] The estimated position of a target refers to the specific coordinates of the target object in three-dimensional space as predicted by the model.

[0091] The rotation prediction head is another sub-module of the pose estimation model, which focuses on predicting the rotation state of an object and outputs multiple candidate rotation matrices and their corresponding confidence scores.

[0092] Rotation-related features are posture-related information during the interaction between an object and a hand, integrating rotation-related logic of object geometric properties, hand posture, and tactile feedback.

[0093] Candidate rotation matrices are multiple matrices output by the rotation prediction head that describe the rotational state of an object. Each matrix corresponds to a possible object rotational pose, mathematically quantifying the object's rotation angle and direction in three-dimensional space.

[0094] The confidence score is a probability value assigned to each candidate rotation matrix by the rotation prediction head, used to characterize the reliability of the corresponding rotation matrix. The higher the score, the better the candidate rotation matrix matches the actual rotation state of the object, and it is the core criterion for selecting the optimal rotation state from multiple candidates.

[0095] Optionally, in one specific implementation of this application, position estimation is first performed. The first 16 channels are extracted from the global fusion features of dimension [batch_size, 32] as position association features (input dimension [batch_size, 16]), and input into the position prediction head. The position prediction head is a single-layer fully connected network. The 16-dimensional features are mapped to a 3-dimensional output through linear transformation, directly obtaining the three-dimensional translation vector (x, y, z) representing the spatial coordinates of the target object, i.e., the estimated position of the target. The output dimension is [batch_size, 3].

[0096] Next, rotation estimation is performed, and the last 16 channels of the global fusion features are extracted as rotation-related features (input dimension [batch_size, 16]). The input is the rotation prediction head, which contains parallel branches. One branch generates 16 candidate rotation matrices in 6D form through a multi-layer fully connected network to quantify the rotation state of the object. The other branch is processed by the last linear layer (input [batch_size, 16], output [batch_size, 16]) and outputs a 16-dimensional confidence score vector through the softmax activation function. Each element corresponds to the reliability of a candidate matrix.

[0097] The optimal rotation matrix is ​​then selected, and the confidence vectors are numerically compared to locate the index of the maximum value. The corresponding candidate rotation matrix is ​​then determined as the target candidate rotation matrix. Finally, the results are integrated, combining the estimated 3D coordinates of the target position with the target candidate rotation matrix to form the complete target pose. The predicted position, 6D rotation matrix, confidence scores of each candidate, and ground truth information are stored in JSON format to support subsequent analysis.

[0098] like Figure 2 As shown, in a complete embodiment, object point cloud, tactile point cloud, and hand point cloud are used as inputs. The three are sequentially processed by point cloud completion, multi-scale convolution, feature compression, self-attention enhancement, and data normalization to obtain object features, tactile features, and hand features, respectively. Then, the object features and tactile features are fused into object-tactile fusion features (i.e., the first fusion feature), and the hand features and tactile features are fused into hand-tactile fusion features (i.e., the second fusion feature). Finally, these two types of fusion features are integrated into global fusion features, and the global fusion features output position prediction and rotation prediction results, respectively, to achieve hierarchical collaboration and accurate utilization of multimodal information.

[0099] In these alternative embodiments, estimation accuracy is improved by decoupling predicted position and rotation. The rotation prediction head generates multiple candidate matrices and assigns confidence scores; errors are reduced by selecting the result with the highest confidence score. Finally, the position and optimal rotation information are integrated to ensure both accuracy and reliability of the target pose.

[0100] In one embodiment, the pose estimation model includes a feature extraction network for extracting the global fusion features, and the feature extraction network includes first network parameters; the position prediction head includes second network parameters, and the rotation prediction head includes third network parameters; The pose estimation model is trained in the following manner: Obtain a training sample set, which includes multiple training samples. Each training sample includes a sample point cloud and a label pose corresponding to the sample point cloud. The label pose includes a label position and a label rotation matrix. Based on the sample point cloud, the predicted position is estimated by the feature extraction network of the base model and the position prediction head. Based on the first loss function between the predicted position and the label position, the parameters of the second network are updated, but the parameters of the first network are not updated, until the first training stopping condition is met, and the first model is obtained. Based on the sample point cloud, the predicted pose is estimated by the feature extraction network, the position prediction head and the rotation prediction head of the first model. Based on the second loss function between the predicted pose and the label pose, the parameters of the first network, the second network and the third network are updated until the second training stopping condition is met, and the second model is obtained. Based on the sample point cloud, the predicted rotation matrix is ​​estimated by the feature extraction network of the second model and the rotation prediction head. The third network parameters are updated based on the third loss function between the predicted rotation matrix and the label rotation matrix, but the first network parameters are not updated, until the third training stopping condition is met, thus obtaining the pose estimation model.

[0101] Optionally, in the embodiments of this application, it should be noted that the specific definitions and explanations of the relevant terms involved in this embodiment can be referred to the relevant descriptions in the foregoing embodiments, and will not be repeated here. The corresponding basic concepts in the foregoing embodiments are modified in this embodiment by adding the qualifiers "sample" or "prediction" to form expressions such as "sample point cloud", "predicted position", "predicted rotation matrix", and "predicted pose" in accordance with the needs of the model training scenario. Their core technical meanings are consistent with the corresponding concepts in the foregoing embodiments, and are only used to clarify the data orientation and output result attributes during the training process.

[0102] Furthermore, the acquisition of the aforementioned global fusion features can be achieved either through the feature extraction network integration of the pose estimation model described in this embodiment, or by following the feature fusion steps disclosed in the aforementioned embodiments. The core technical principles and operational logic of the two implementation methods remain consistent, and both can provide effective feature support for subsequent pose estimation.

[0103] Optionally, in this embodiment, the network parameters are learnable parameters of the corresponding modules (feature extraction network, position prediction head, rotation prediction head) of the pose estimation model, including the weights and biases of each layer. Their core function is to enable the corresponding modules to extract effective information from the input features and complete position or rotation predictions through updates during training.

[0104] The first loss function is a metric function that quantifies the difference between the predicted location and the label location (such as L1 or L2 loss). It is used to provide gradient signals for updating the parameters of the location prediction head, guide the optimization of the second network parameters, and make the location prediction results gradually approach the true label location.

[0105] The second loss function is a comprehensive function that measures the overall difference between the predicted pose and the labeled pose (usually composed of a weighted sum of position loss and rotation loss). It provides a unified optimization objective for the collaborative updating of multiple parameters (first, second, and third network parameters) during the overall training phase of the model, ensuring an overall improvement in pose estimation performance.

[0106] The third loss function is a metric function that focuses on the deviation between the predicted rotation matrix and the label rotation matrix. It is specifically designed to provide gradients for fine-tuning the parameters of the rotation prediction head and optimize the rotation prediction accuracy.

[0107] Training termination conditions are the criteria for determining whether to terminate a corresponding training phase. These typically include the loss function value converging to a preset threshold, the error not decreasing significantly over multiple iterations, or the preset maximum number of iterations being reached. Meeting any one of these conditions will stop the current training phase, ensuring a balance between model training efficiency and effectiveness.

[0108] Optionally, in one specific implementation of this application, a training sample set is first prepared, including sample point clouds and corresponding label poses (label positions, label rotation matrices). The Adam optimizer is used, with an initial learning rate of 0.001. The learning rate decreases by 50% every two epochs. In terms of data allocation, the ratio of real machine data to simulation data is 50%:50% in the initial training (first epoch). The proportion of real sensor data is increased by 5% every two epochs thereafter, and adjusted to 55%:45% in the third epoch. The transition process from simulation to real machine is learned through a gradual guided approach.

[0109] In this application embodiment, addressing the issue that most related methods are designed for two-finger grippers and cannot handle the complex occlusion and configuration differences of multi-finger hands, this application dynamically switches between real device and simulation data and adds Gaussian noise during data augmentation to simulate multi-sensor errors. The hand model point cloud covers multiple fingers (including fingertips and other parts), and the tactile point cloud contains 256 contact points, fully capturing multi-finger contact information. At the same time, through a multi-stage training strategy (first fixing the feature extraction network, then fine-tuning it from the back end), the adaptability of the model to different hand configurations is enhanced. After testing on various dexterous hands, the generalization performance is significantly better than existing methods, enhancing the generalization ability and adapting to complex scenarios of multi-finger dexterous hands.

[0110] Furthermore, to address the issues of unrealistic rendering of simulated pipelines and domain adaptation failure caused by idealized haptic simulations, this application dynamically mixes real machine data and simulated data during training (initially in a 50%:50% ratio, with 5% real machine data added every 2 epochs) to gradually adapt to real sensor noise. At the same time, PF Net is used for point cloud completion and multi-scale downsampling to enhance robustness to the idealization defects of simulated data. When deployed in real-world scenarios, no complex domain adaptation is required, bridging the simulation-reality gap and improving deployment feasibility.

[0111] In the first stage of training, the feature extraction network of the fixed base model (without updating the parameters of the first network) is used to input the sample point cloud into the feature extraction network to obtain global fusion features, which are then output as predicted positions by the position prediction head. The first loss function (such as L2 loss) between the predicted position and the label position is calculated, and the parameters of the second network are updated through backpropagation until the first loss function converges to a preset threshold (such as translation error ≤ 2mm) or reaches the maximum number of iterations (such as 20 epochs), thus satisfying the first training stopping condition and obtaining the first model.

[0112] In the second stage, the first model is called. After inputting the sample point cloud, the predicted pose is obtained through the feature extraction network, the position prediction head, and the rotation prediction head. The second loss function (the weighted sum of position loss and rotation loss) is calculated between the predicted pose and the label pose. The parameters of the first, second, and third networks are updated synchronously. When the Symmetric Average Distance Error (ADD-S) index of the symmetrical object stabilizes within the preset range (e.g., the average point distance of the symmetrical object is ≤1.5mm), the second training stopping condition is met, and the second model is obtained.

[0113] In the third stage, the position prediction head is frozen. The sample point cloud is input into the feature extraction network of the second model and the rotation prediction head to obtain the predicted rotation matrix. The parameters of the third network are updated based on the third loss function (weighted L1 loss + ranking loss), while the parameters of the first network are not updated. When the rotation error is ≤1 degree, the third training stopping condition is met, and the pose estimation model is finally obtained.

[0114] In these alternative embodiments, the training process employs a three-stage strategy: first, optimizing the position prediction head individually; then, collaboratively updating the parameters of the entire network; and finally, focusing on fine-tuning the rotation prediction. This avoids parameter conflicts caused by single-stage training and, through dynamic adjustments of fixing and updating, allows each module to accurately adapt to task requirements, improving the accuracy and generalization ability of the model's pose estimation.

[0115] In one embodiment, the step of estimating the predicted rotation matrix using the feature extraction network of the second model and the rotation prediction head includes: Based on the sample point cloud, the global fusion features of the samples are obtained through the feature extraction network. The rotation prediction head estimates multiple candidate rotation matrices based on the global fusion features of the samples, and the prediction confidence of each candidate rotation matrix is ​​obtained. The prediction candidate rotation matrix with the highest prediction confidence is determined as the prediction rotation matrix; The third loss function includes a weighted error term and a candidate ranking loss; The weighted error term is determined based on the pose error and the target weight. The pose error is determined based on each of the predicted candidate rotation matrices and the label rotation matrix. The target weight is determined based on the prediction confidence. The candidate ranking loss is determined based on the error between the confidence score of the positive sample and the confidence scores of all predicted candidate rotation matrices, wherein the positive sample is the predicted candidate rotation matrix with the smallest error to the label rotation matrix among all the predicted candidate rotation matrices.

[0116] Optionally, in this embodiment, the weighted error term is calculated by multiplying the pose error by the target weight, and is used to quantify the weighted rotation prediction bias, which allows the model to focus on error optimization of high-quality candidates and improve the targeting of loss calculation.

[0117] The target weights are derived from the prediction confidence of the predicted candidate rotation matrix, which makes candidates with small errors and high confidence account for a higher proportion in the loss calculation, guiding the model to learn the matching relationship between candidate quality and confidence.

[0118] The candidate ranking loss is based on the confidence scores of positive samples and all candidates. By calculating the error, the score difference between positive samples and other candidates is widened, ensuring that the confidence ranking is consistent with the candidate quality.

[0119] Positive samples are selected from multiple predicted candidate rotation matrices as "high-quality benchmarks," specifically referring to the candidates with the smallest error relative to the label rotation matrix, serving as the benchmark for distinguishing other candidates in the candidate ranking loss calculation.

[0120] Optionally, in one specific implementation of this application, the sample point cloud is first processed by a feature extraction network to generate global fusion features of the samples; these features are then input into a rotation prediction head, which outputs N=16 prediction candidate rotation matrices (6-dimensional vectors flattened with the first two columns of the rotation matrix). (This indicates that the prediction confidence score is also output.) , This represents the set of candidate rotation vectors output by the rotation prediction head; The data type of this vector set is real number; B refers to the batch size, which represents the number of samples processed simultaneously during a single training or inference session (e.g., B=32 represents processing 32 sample point clouds in a single session); N: refers to the number of candidate rotation matrices corresponding to each sample, that is, the number of rotation pose candidates output by the rotation prediction head for a single sample (combined with the previous technical solution, N=16, which means that 16 candidate rotation matrices are generated for each sample).

[0121] Next, the third loss function is calculated: The first step is to calculate the L1 error (i.e., pose error) of the candidate pose. For each sample i and candidate j, the formula is:

[0122] The error between each predicted candidate rotation matrix and the label rotation matrix is ​​obtained. ,in This is the 6-dimensional true vector corresponding to the label rotation matrix. Here, i is the sample index, ranging from [1, B], corresponding to the i-th sample in the batch data (for example, i=5 represents the point cloud of the 5th sample in the batch and its label rotation matrix); j is the candidate index, ranging from [1, N], corresponding to the j-th candidate rotation vector of the i-th sample (for example, j=3 represents the 3rd candidate rotation vector of the i-th sample). This represents the L1 error (i.e., pose error) corresponding to the j-th candidate rotation vector of the i-th sample.

[0123] The second step is to calculate the weighted error term: First, use softmax to convert the confidence score into the target weight.

[0124] in, This represents the target weight corresponding to the j-th candidate rotation matrix of the i-th sample; This represents the original prediction confidence score of the j-th candidate rotation matrix for the i-th sample; The summation index of the candidate rotation matrix is ​​used only to traverse all candidates of a single sample to avoid confusion with the current candidate index j, and its value range is [1, N].

[0125] Then follow the formula:

[0126] The weighted error term is obtained. Among them, This is the weighted error term.

[0127] The third step is to calculate the candidate ranking loss: First, select the candidate rotation matrix with the smallest error for each sample i as the positive sample.

[0128] in, This represents the candidate index of the positive sample corresponding to the i-th sample, that is, from all the candidate rotation matrices of this sample, the candidate rotation matrix with the smallest L1 pose error is selected as the positive sample.

[0129] Then, through cross-entropy loss:

[0130] Increase the confidence gap between positive samples and other candidates. Among these, For candidate ranking loss, It is the positive sample of the i-th sample (index) The original confidence score of ).

[0131] Finally, the total loss function is used:

[0132] The loss calculation is completed, and the third network parameters of the rotating prediction head are updated accordingly. Among these, For the third loss function, This is a loss balancing hyperparameter used to adjust the ranking loss. The contribution percentage of the total loss should be considered to avoid a single loss dominating the training process. In these alternative embodiments, the weighted error term allows the model to focus on optimizing high-quality candidates, the candidate ranking loss strengthens the matching between confidence and candidate quality, and the dual loss synergistically improves the accuracy of rotation estimation, thereby enhancing the accuracy and stability of rotation prediction.

[0133] Optionally, in the embodiments of this application, based on the prior art deficiencies outlined in the background section and the technical details of the specific implementation methods, the present invention has the following outstanding advantages: To address the issues of rotation estimation being susceptible to noise interference and the lack of fault tolerance mechanisms in single output, this application outputs 16 candidate rotation matrices and confidence scores in parallel during pose decoupling prediction. The loss function combines weighted L1 loss (implicitly learning candidate quality) and ranking loss (explicitly widening the gap between positive and negative samples) to ensure that high-precision candidates are selected first. The stability of rotation estimation is improved through the multi-candidate ranking loss function.

[0134] To address the issues of segmentation relying on manual prompts and the lack of end-to-end optimization for multimodal interactions, this application uses tactile point clouds as "implicit prompts" to replace manual annotation (traditional methods use manual mouse clicks on objects to be segmented in images, while this method uses the tactile sense of a dexterous hand to achieve automated segmentation). The entire pipeline is trained end-to-end from data acquisition to pose prediction. Multimodal features are collaboratively inferred through Transformer to avoid module isolation, thus optimizing segmentation and perception end-to-end and improving system integration.

[0135] It should be noted that the various optional implementation methods described in the embodiments of this application can be combined with each other or implemented individually without conflict, and the embodiments of this application do not limit this.

[0136] Figure 3 A schematic diagram of the structure of an object pose estimation device provided in another embodiment of this application is shown. For ease of explanation, only the parts related to the embodiments of this application are shown.

[0137] Reference Figure 3 The object pose estimation device may include: The acquisition module 301 is used to acquire the object structure features of the target object, the hand posture features of the robot hand, and the contact area features of the hand contacting the target object. The first fusion module 302 is used to fuse the object structural features with the contact area features to obtain a first fusion feature; and to fuse the hand posture features with the contact area features to obtain a second fusion feature. The second fusion module 303 is used to fuse the first fusion feature and the second fusion feature to obtain a global fusion feature; The estimation module 304 is used to perform pose estimation based on the global fusion features to obtain the target pose of the target object.

[0138] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application, and are devices corresponding to the above-mentioned methods. All implementation methods in the above-mentioned method embodiments are applicable to the embodiments of this device. For details on its specific functions and the technical effects it brings, please refer to the method embodiment section, which will not be repeated here.

[0139] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0140] Figure 4 A schematic diagram of the hardware structure of the electronic device provided in an embodiment of this application is shown.

[0141] The device may include a processor 401 and a memory 402 storing program instructions.

[0142] When processor 401 executes the program, it implements the steps in any of the above method embodiments.

[0143] For example, the program can be divided into one or more modules / units, one or more of which are stored in memory 402 and executed by processor 401 to complete this application. The one or more modules / units can be a series of program instruction segments capable of performing a specific function, which describe the execution process of the program in the device.

[0144] Specifically, the processor 401 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.

[0145] Memory 402 may include mass storage for data or instructions. For example, and not limitingly, memory 402 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 402 may include removable or non-removable (or fixed) media. Where appropriate, memory 402 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 402 is non-volatile solid-state memory.

[0146] Memory may include read-only memory (ROM), random access memory (RAM), disk storage media devices, optical storage media devices, flash memory devices, and electrical, optical, or other physical / tangible memory storage devices. Therefore, typically, memory includes one or more tangible (non-transitory) readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the methods according to one aspect of this disclosure.

[0147] The processor 401 implements any of the methods described above by reading and executing program instructions stored in the memory 402.

[0148] In one example, the electronic device may also include a communication interface 403 and a bus 410. The processor 401, memory 402, and communication interface 403 are connected via the bus 410 and communicate with each other.

[0149] The communication interface 403 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.

[0150] Bus 410 includes hardware, software, or both, that couples components of an online data traffic metering device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 410 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, this application contemplates any suitable bus or interconnect.

[0151] Furthermore, in conjunction with the methods in the above embodiments, this application embodiment can provide a storage medium for implementation. This storage medium stores program instructions; when these program instructions are executed by a processor, they implement any of the methods in the above embodiments.

[0152] This application also provides a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled. The processor is used to run programs or instructions to implement the various processes of the above method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0153] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0154] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above method embodiments and achieve the same technical effects. To avoid repetition, it will not be described again here.

[0155] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.

[0156] The functional modules shown in the above block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on machine-readable media or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable media" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer grids such as the Internet, intranets, etc.

[0157] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.

[0158] The aspects of this disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and program products according to embodiments of this disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to create a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by special-purpose hardware performing the specified functions or actions, or can be implemented by a combination of special-purpose hardware and computer instructions.

[0159] The above are merely specific embodiments of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.

Claims

1. A method for estimating the pose of an object, characterized in that, The method includes: The structural features of the target object, the hand posture features of the robot hand, and the contact area features of the hand contacting the target object are obtained. The object's structural features are fused with the contact area features to obtain a first fused feature; the hand posture features are fused with the contact area features to obtain a second fused feature. The first fusion feature and the second fusion feature are fused together to obtain the global fusion feature; Based on the global fusion features, pose estimation is performed to obtain the target pose of the target object.

2. The method according to claim 1, characterized in that, The acquisition of the object structure features of the target object, the hand posture features of the robot hand, and the contact area features of the hand contacting the target object includes: The visual point cloud of the target object, the hand model point cloud of the hand, and the tactile point cloud of the hand in contact with the target object are obtained. The visual point cloud of the object, the hand model point cloud, and the tactile point cloud of the object are respectively encoded to obtain initial visual point cloud features, initial hand point cloud features, and initial tactile point cloud features; A self-attention mechanism is applied to the initial visual point cloud features, the initial hand point cloud features, and the initial tactile point cloud features respectively to obtain the object structure features, the hand posture features, and the contact area features.

3. The method according to claim 2, characterized in that, The step of applying a self-attention mechanism to the initial visual point cloud features, the initial hand point cloud features, and the initial tactile point cloud features respectively to obtain the object structure features, the hand posture features, and the contact area features includes: Generate a first query vector, a first key vector, and a first value vector corresponding to feature F; wherein feature F is any one of the initial visual point cloud feature, the initial hand point cloud feature, and the initial tactile point cloud feature; An attention matrix is ​​determined based on the first key vector and the first value vector; the attention matrix includes the attention value between any two point clouds in the feature F, and the attention value is used to characterize the degree of association importance between any two point clouds; Based on the attention matrix and the first query vector, an attention enhancement feature corresponding to the feature F is generated; the attention enhancement feature is any one of the object structure feature, the hand posture feature, and the contact area feature.

4. The method according to claim 1, characterized in that, The object's structural features and contact area features are fused to obtain a first fused feature; The hand posture features are fused with the contact area features to obtain a second fused feature, including: The object structure features, hand posture features, and contact area features are fused with their respective point cloud spatial coordinates to generate object enhancement features, hand enhancement features, and tactile enhancement features. Based on the object enhancement features, the hand enhancement features, and the tactile enhancement features, they are fused through a cross-attention mechanism to obtain the first fused feature and the second fused feature.

5. The method according to claim 4, characterized in that, The step of fusing the object structural features, the hand pose features, and the contact area features with their respective corresponding point cloud spatial coordinates to generate object enhancement features, hand enhancement features, and tactile enhancement features includes: The object structural features, hand posture features, and contact area features are respectively stitched together with their corresponding point cloud spatial coordinates to obtain object stitching features, hand stitching features, and tactile stitching features; A self-attention mechanism is applied to the object stitching feature, the hand stitching feature, and the tactile stitching feature respectively to obtain the object enhancement feature, the hand enhancement feature, and the tactile enhancement feature.

6. The method according to claim 4, characterized in that, Based on the object enhancement features, the hand enhancement features, and the tactile enhancement features, cross-modal interaction fusion is performed through a cross-attention mechanism to obtain the first fused feature and the second fused feature, including: Using the object enhancement feature as the second query vector and the tactile enhancement feature as the second key vector and the second value vector, cross-attention calculation is performed to obtain the first fused feature; Using the hand enhancement features as the third query vector, the second key vector, and the second value vector, cross-attention calculation is performed to obtain the second fused feature.

7. The method according to claim 1, characterized in that, The first fusion feature and the second fusion feature are fused to obtain a global fusion feature, including: Using the first fusion feature as the fourth query vector and the second fusion feature as the third key vector and third value vector, cross-attention calculation is performed to obtain the global fusion feature.

8. The method according to claim 1, characterized in that, The step of estimating the pose based on the global fusion features to obtain the target pose of the target object includes: The target estimated position of the target object is obtained by using the position prediction head of the pose estimation model based on the position association features in the global fusion features; The rotation prediction head of the pose estimation model estimates multiple candidate rotation matrices of the target object and the confidence level of each candidate rotation matrix based on the rotation association features in the global fusion features. The candidate rotation matrix with the highest confidence among the multiple candidate rotations is determined as the target candidate rotation matrix; The target pose is determined based on the estimated target position and the target candidate rotation matrix.

9. The method according to claim 8, characterized in that, The pose estimation model includes a feature extraction network for extracting the global fusion features, and the feature extraction network includes first network parameters; the position prediction head includes second network parameters, and the rotation prediction head includes third network parameters. The pose estimation model is trained in the following manner: Obtain a training sample set, which includes multiple training samples. Each training sample includes a sample point cloud and a label pose corresponding to the sample point cloud. The label pose includes a label position and a label rotation matrix. Based on the sample point cloud, the predicted position is estimated by the feature extraction network of the base model and the position prediction head. Based on the first loss function between the predicted position and the label position, the parameters of the second network are updated, but the parameters of the first network are not updated, until the first training stopping condition is met, and the first model is obtained. Based on the sample point cloud, the predicted pose is estimated by the feature extraction network, the position prediction head and the rotation prediction head of the first model. Based on the second loss function between the predicted pose and the label pose, the parameters of the first network, the second network and the third network are updated until the second training stopping condition is met, and the second model is obtained. Based on the sample point cloud, the predicted rotation matrix is ​​estimated by the feature extraction network of the second model and the rotation prediction head. The third network parameters are updated based on the third loss function between the predicted rotation matrix and the label rotation matrix, but the first network parameters are not updated, until the third training stopping condition is met, thus obtaining the pose estimation model.

10. The method according to claim 9, characterized in that, The step of estimating the predicted rotation matrix using the feature extraction network of the second model and the rotation prediction head includes: Based on the sample point cloud, the global fusion features of the samples are obtained through the feature extraction network. The rotation prediction head estimates multiple candidate rotation matrices based on the global fusion features of the samples, and the prediction confidence of each candidate rotation matrix is ​​obtained. The prediction candidate rotation matrix with the highest prediction confidence is determined as the prediction rotation matrix; The third loss function includes a weighted error term and a candidate ranking loss; The weighted error term is determined based on the pose error and the target weight. The pose error is determined based on each of the predicted candidate rotation matrices and the label rotation matrix. The target weight is determined based on the prediction confidence. The candidate ranking loss is determined based on the error between the confidence score of the positive sample and the confidence scores of all predicted candidate rotation matrices, wherein the positive sample is the predicted candidate rotation matrix with the smallest error between it and the label rotation matrix.