Pose optimization method and device of object
By constructing an electromagnetic force constraint model and using visual and tactile geometric information combined with physical constraints to optimize pose, the problems of occlusion robustness and physical consistency in existing technologies are solved, and high-accuracy pose estimation in complex scenes is achieved.
Patent Information
- Application Number
- CN202511914944.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-18
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2045-12-18
AI Technical Summary
Existing visual or visual-tactile pose estimation techniques have significant limitations in terms of occlusion robustness, physical consistency, multi-hardware generalization ability, real-time performance, and zero-shot adaptation ability, especially in robot hand occlusion scenarios where it is difficult to provide robust, accurate, and interpretable pose results.
By acquiring the visual point cloud features of the object, the geometric point cloud features of the robot's end effector, and the tactile point cloud features of the object, an electromagnetic force constraint model is constructed. The first and second constraint terms of the electromagnetic force constraint model are used to optimize the initial pose, ensuring that the pose conforms to the physical interaction logic, avoiding penetration and improving accuracy.
It does not rely on a large amount of haptic-visual pairing data for training, improves the accuracy of object pose estimation, reduces the impact of sensor type differences and wear on model estimation results, and achieves robust pose optimization in complex scenes.
Smart Images

Figure CN121340374B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robotics, and in particular to a method and apparatus for optimizing the pose of an object. Background Technology
[0002] Six-dimensional (6D) pose estimation is a key technology for robots to perform fine manipulation tasks, and its accuracy directly determines the success rate of robot operations. In complex hand-operated scenarios, continuous contact between the robot's fingers and objects can lead to severe visual occlusion, posing a serious challenge to pose estimation methods based on a single visual modality.
[0003] To overcome the limitations of visual information, related technologies have proposed visual-tactile fusion methods, which attempt to maintain the stability of pose estimation under occlusion conditions by combining complementary information from visual and tactile sensors. These methods typically rely on a large amount of acquired tactile-visual paired data for model training to learn the mapping relationship from multimodal signals to pose.
[0004] However, due to the diverse types and significant differences in physical characteristics of tactile sensors, and their susceptibility to wear, calibration errors, and other factors in practical deployments, the data distribution collected by different sensors varies considerably. This results in poor accuracy in estimating object poses using models trained on data from specific sensors. Summary of the Invention
[0005] This application provides a method and apparatus for optimizing the pose of an object, which can improve the accuracy of object pose estimation.
[0006] In a first aspect, embodiments of this application provide a method for optimizing the pose of an object, the method comprising:
[0007] Acquire feature data and the initial pose of the target object under the feature data; the feature data includes the object visual point cloud features of the target object, the geometric point cloud features of the robot's end effector, and the object tactile point cloud features of the robot contacting the target object.
[0008] Determine whether the feature data satisfies at least one physical feasibility constraint; wherein the physical feasibility constraint includes at least one of contact constraint, penetration constraint, and motion constraint;
[0009] If the feature data does not meet any of the physical feasibility constraints, an electromagnetic force constraint model is constructed based on the feature data. The electromagnetic force constraint model includes a first constraint term and a second constraint term. The first constraint term is determined based on the contact relationship between the tactile point cloud features and the visual point cloud features of the object, and the second constraint term is determined based on the relative positional relationship between the geometric point cloud features and the visual point cloud features of the object.
[0010] Based on the first and second constraints, the initial pose is optimized to obtain the target pose of the target object.
[0011] Secondly, this application provides a pose optimization device for an object, the device comprising:
[0012] An acquisition module is used to acquire feature data and the initial pose of the target object estimated by the robot under the feature data; the feature data includes the object visual point cloud features of the target object, the geometric point cloud features of the robot's end effector, and the object tactile point cloud features of the robot contacting the target object.
[0013] The judgment module is used to determine whether the feature data satisfies at least one physical feasibility constraint; wherein the physical feasibility constraint includes at least one of contact constraint, penetration constraint and motion constraint;
[0014] A construction module is used to construct an electromagnetic force constraint model based on the feature data when the feature data does not meet any physical feasibility constraint; the electromagnetic force constraint model includes a first constraint term and a second constraint term; the first constraint term is determined based on the contact relationship between the tactile point cloud features of the object and the visual point cloud features of the object, and the second constraint term is determined based on the relative positional relationship between the geometric point cloud features and the visual point cloud features of the object;
[0015] An optimization module is used to optimize the initial pose based on the first constraint and the second constraint to obtain the target pose of the target object.
[0016] Thirdly, embodiments of this application provide an electronic device, which includes: a processor and a memory storing computer program instructions;
[0017] When the processor executes computer program instructions, it implements a pose optimization method for an object as described in any of the embodiments of the first aspect.
[0018] Fourthly, embodiments of this application provide a computer storage medium storing computer program instructions, which, when executed by a processor, implement the pose optimization method for an object as described in any of the embodiments of the first aspect.
[0019] Fifthly, embodiments of this application provide a computer program product in which instructions, when executed by a processor of an electronic device, cause the electronic device to perform an object pose optimization method as described in any of the embodiments of the first aspect above.
[0020] In the object pose optimization method and apparatus provided in this application embodiment, firstly, feature data including the visual point cloud features of the target object, the geometric point cloud features of the robot's end effector, and the tactile point cloud features of the object are acquired, along with an initial pose estimated based on these feature data. Subsequently, it is determined whether the feature data satisfies at least one physical feasibility constraint; wherein, the physical feasibility constraint includes at least one of contact constraints, penetration constraints, and motion constraints. If the feature data does not satisfy any physical feasibility constraint, it indicates that the initial pose estimated based on the feature data has an accuracy defect. Therefore, an electromagnetic force constraint model can be constructed using a first constraint term determined based on the contact relationship between the object's tactile point cloud features and the object's visual point cloud features, and a second constraint term determined based on the relative positional relationship between the geometric point cloud features and the object's visual point cloud features. The first constraint term can accurately capture the actual contact state between the robot and the target object, while the second constraint term can effectively avoid the penetration problem between the object and the robot. By leveraging the synergistic effect of the first and second constraint terms of the electromagnetic force constraint model, the initial pose is optimized to obtain the target pose of the object. This eliminates the need for model training based on a large amount of tactile-visual paired data, freeing the model from dependence on specific tactile sensor data. It also avoids or reduces the impact of inconsistent data distribution caused by differences in sensor type, wear, or calibration errors on the model estimation results, thereby improving the accuracy of object pose estimation. Attached Figure Description
[0021] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is one of the flowcharts illustrating the object pose optimization method provided in the embodiments of this application;
[0023] Figure 2 This is a schematic diagram of the structure of an object pose optimization system provided in an embodiment of this application;
[0024] Figure 3 This is a second schematic flowchart of the object pose optimization method provided in the embodiments of this application;
[0025] Figure 4 This is the third flowchart illustrating the object pose optimization method provided in the embodiments of this application;
[0026] Figure 5 This is a schematic diagram of the structure of an object pose optimization device provided in an embodiment of this application;
[0027] Figure 6This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0028] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.
[0029] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.
[0030] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the term "comprising" or any other variations thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.
[0031] 6D pose estimation is a key fundamental capability in robot manipulation tasks. Its purpose is to determine the position and orientation of an object in three-dimensional space, thereby providing reliable input for state-based control, planning and manipulation.
[0032] In recent years, with the advancement of visual perception technology, vision-based pose estimation methods have made significant progress, including instance-level methods for specific known objects, category-level methods with class generalization capabilities, and novel object estimation methods for unknown objects. These methods rely on large amounts of image data and synthetic data augmentation strategies, and exhibit high accuracy in standard scenarios.
[0033] However, in real-world robot manipulation scenarios, especially tasks involving frequent contact and occlusion such as grasping and hand manipulation, pure vision methods face significant limitations. On one hand, vision methods heavily rely on clear visual observations. When an object is severely occluded by the robot's hand or the environment, the lack of visual cues can lead to acute errors or cumulative drift in pose estimation, or even complete tracking failure. On the other hand, due to the lack of modeling of the physical interaction process, pure vision estimation results may exhibit physically unreasonable situations such as objects penetrating the robot, object trajectories not conforming to dynamic laws, or inconsistencies with tactile contact states.
[0034] To alleviate the limitations of visual methods in occluded scenarios, researchers have begun exploring visual-tactile fusion methods, using tactile signals as supplementary cues for occluded areas. Related techniques improve the robustness of pose estimation by combining red, green, blue-depth (RGB-D) cameras with tactile signals, utilizing graph neural networks to process tactile data, and incorporating proprioceptive information or shape completion modules. Tactile sensors can provide crucial contact geometry information in occluded areas, demonstrating advantages in grasping and hand manipulation scenarios.
[0035] However, existing visual-haptic methods still face several technical bottlenecks: First, multimodal training relies on a large amount of paired visual-haptic data, but haptic data is difficult to collect, sensors wear out quickly, and there are significant differences between various sensor types, making it difficult to construct large-scale datasets. Second, existing models are often designed for specific types of haptic sensors, making it difficult to generalize to new sensor types or different robot platforms. Third, multimodal deep networks are computationally intensive, making it difficult to meet the real-time requirements of robot manipulation tasks. In addition, although some methods can use haptics to refine visual results, they still rely on specific sensor outputs or additional training of velocity estimation models, lacking adaptability to new objects and new haptic sensors.
[0036] From a system integration perspective, visual-haptic methods also face practical problems such as strong calibration dependence, incomplete physical constraint modeling, and difficulty in supporting dynamic and complex manipulation tasks. For example, hand-eye calibration and haptic calibration need to maintain accuracy over a long period of time, otherwise multimodal fusion is prone to failure; related technologies have failed to simultaneously integrate multiple physical constraints such as contact consistency, penetration avoidance, and kinematic constraints, which may lead to the estimated pose violating physical feasibility; in dynamic manipulation scenarios, the pose estimation methods of related technologies are difficult to guarantee continuity and stability.
[0037] In summary, existing visual or visual-haptic pose estimation techniques suffer from significant limitations in occlusion robustness, physical consistency, multi-hardware generalization ability, real-time performance, and zero-shot adaptation. Therefore, there is an urgent need for a pose estimation method that does not rely on haptic training data, can fuse visual and haptic geometric information, and incorporates physical constraints for online optimization. This method would provide more robust, accurate, and physically interpretable pose results during actual robot manipulation.
[0038] To address the problems existing in related technologies, embodiments of this application provide a method and apparatus for optimizing the pose of an object.
[0039] The object pose optimization method provided in the embodiments of this application will be described below. Figure 1 As shown, the method specifically includes the following steps:
[0040] S100, acquire feature data and the initial pose of the target object under the feature data; the feature data includes the object visual point cloud features of the target object, the geometric point cloud features of the robot's end effector, and the object tactile point cloud features of the robot contacting the target object.
[0041] Optionally, in the embodiments of this application, the target object refers to the direct object of pose optimization in the embodiments of this application, that is, the specific object that the robot needs to obtain a precise 6D pose in the manipulation task (such as grasping, in-hand adjustment, fine assembly, etc.).
[0042] A robot is a primary device that performs manipulation tasks, possessing end effectors for interacting with target objects, sensing modules, and data processing capabilities. Its core function is to precisely manipulate the target object by contacting it through the end effector.
[0043] The initial pose is a preliminary 6D pose of the target object estimated by the robot based on the feature data acquired by the S100, and serves as the basic reference benchmark for subsequent optimization processes. The initial pose is initially calculated by the robot's visual perception module, tactile sensing analysis module, or multimodal fusion, but may contain deviations due to factors such as visual occlusion, sensor noise, and lack of physical constraints.
[0044] Visual point cloud features of an object are digital representations of the geometric shape of a target object. They can reflect the shape outline and structural details of the target object and are three-dimensional geometric features of the object generated based on visual information.
[0045] Geometric point cloud features are digital representations of the geometry of a robot's end effector. End effectors include, but are not limited to, various contact components such as dexterous hands, mechanical grippers, vacuum suction cups, or combinations thereof. Geometric point cloud features primarily focus on areas where the end effector may come into contact with the target object (such as the fingertips of a dexterous hand, the gripping surface of a mechanical gripper, and the interactive surface at the end of a robotic arm), and can reflect the spatial contour of the end effector.
[0046] Object tactile point cloud features are three-dimensional digital features related to contact with a target object, collected by the robot through tactile sensors. The corresponding contact area is in three-dimensional space and can reflect the actual contact position and contact state between the robot's end effector and the target object.
[0047] Optionally, in one feasible implementation of this application, for acquiring the visual point cloud features of an object, the mesh model of the target object reconstructed by a visual sensor is used as input, and point cloud data is extracted from the model surface through an adaptive sampling method based on curvature. Curvature, as a geometric indicator measuring the degree of local bending of a surface, can effectively reflect the feature-rich structure of the model, such as edges, corners, and sharp transition areas. Therefore, the curvature of the object mesh is first calculated, and a non-uniform sampling probability distribution is constructed based on the curvature value, so that areas with higher curvature have a higher probability of being sampled, thereby ensuring the resolution of the point cloud in key geometric detail areas. In the actual sampling process, the Intrinsic Shape Signature (ISS) method can be used to preferentially extract local salient points, and then the sampling points are encoded by a lightweight multi-layer perceptron (MLP) to form the visual point cloud features of the object containing spatial coordinates and local geometric descriptions.
[0048] For geometric point cloud features, adaptive sampling is also performed based on the end effector's mesh model, but the sampling focus is concentrated on areas that may come into contact with the object, such as the fingertips, finger pads, and inner gripper surfaces, to improve the accuracy of subsequent contact relationship determination. Geometric point cloud features can be generated through similar curvature-weighted sampling and MLP encoding, ensuring consistency in geometric representation with the object's visual point cloud features. Furthermore, the end effector's position in the camera coordinate system can be calculated based on its current joint angles and forward kinematics, thus achieving correct spatial alignment.
[0049] The object's tactile point cloud features are derived from tactile sensors installed on the robot's hand. First, the raw readings from the tactile sensors are preprocessed, including noise reduction, touch point localization, and effective region selection. Then, a lightweight convolutional neural network is used to extract the object's tactile point cloud features. Each object's tactile point cloud feature includes its three-dimensional coordinates, tactile response intensity (used as a confidence score), and a feature vector generated based on local tactile texture. Subsequently, a pre-calibrated hand-eye matrix (or a coordinate transformation module based on online learning) is used to transform the object's tactile point cloud features from the tactile sensor coordinate system to the camera coordinate system, placing them in a unified coordinate space with the object's visual and geometric point cloud features.
[0050] After generating the visual point cloud features, geometric point cloud features, and tactile point cloud features of the object, the initial pose of the target object is further calculated based on the visual estimation module (e.g., a pose regression model based on deep learning or a point cloud registration algorithm). It should be noted that the initial pose is initially calculated from the collected feature data, but the feature data may be biased due to factors such as visual occlusion, sensor noise, and lack of physical constraints, which may lead to inaccurate initial pose estimation.
[0051] S200, determine whether the feature data satisfies at least one physical feasibility constraint; wherein the physical feasibility constraint includes at least one of contact constraint, penetration constraint and motion constraint.
[0052] Optionally, in this embodiment, physical feasibility constraints are a series of criteria used to determine whether feature data and corresponding initial poses conform to physical interaction rules. They are the core basis for triggering pose optimization and may specifically include at least one of contact constraints, penetration constraints, and motion constraints. Physical feasibility constraints are set based on the actual interaction logic between the robot and the target object, aiming to exclude pose estimation results that are physically infeasible.
[0053] Contact constraints are used to verify whether the contact relationship between the robot and the target object conforms to the logic of real interaction. For example, the shortest Euclidean distance from the tactile point cloud features of each object to the visual point cloud features of the object is calculated. If the distance exceeds a preset contact distance threshold, the tactile point cloud features of the object are determined not to meet the contact constraints.
[0054] Penetration constraints are used to prevent physically impractical spatial penetration between a target object and a robot. For example, the number of spatial overlap points between the visual point cloud features of an object and the geometric point cloud features of the robot's end effector is counted. If the percentage of overlap points exceeds a preset overlap ratio threshold, the penetration constraint is considered to be violated.
[0055] Motion constraints are used to ensure that the motion trajectory of a target object conforms to the requirements of dynamic smoothness. For example, the motion trajectory of the initial pose of multiple adjacent frames is fitted, and the trajectory curvature value is calculated. If the curvature value exceeds a preset curvature threshold, it is determined that the trajectory has abrupt changes and does not meet the motion constraints.
[0056] S300, if the feature data does not satisfy any physical feasibility constraint, construct an electromagnetic force constraint model based on the feature data; the electromagnetic force constraint model includes a first constraint term and a second constraint term; the first constraint term is determined based on the contact relationship between the tactile point cloud features of the object and the visual point cloud features of the object, and the second constraint term is determined based on the relative positional relationship between the geometric point cloud features and the visual point cloud features of the object.
[0057] Optionally, in this embodiment, the electromagnetic force constraint model is an abstract physical model used to model the interaction between the robot and the target object. Its core is to analogize the object's visual point cloud features, geometric point cloud features, and tactile point cloud features as charged particles, quantifying the spatial relationship and interaction requirements among them by simulating the laws of electromagnetic interaction. This model does not rely on large-scale training data; instead, it deduces the interaction relationship through physical laws, providing adjustment direction for the optimization of the initial pose and ensuring that the optimized pose conforms to both geometric constraints and physical interaction logic.
[0058] The first constraint term corresponds to the mechanical quantification index in the electromagnetic force constraint model that characterizes the "proximity requirement" between the tactile point cloud features and the visual point cloud features of an object. Its core logic is to simulate the law of attraction between opposite charges in electromagnetism (i.e., electromagnetic attraction). When there is a contact relationship or further contact is required between the tactile point cloud features (representing the robot's contact area) and the visual point cloud features (representing the target object's surface), the quantification calculation of the first constraint term guides the two to move closer in space, thereby capturing the actual contact state between the robot and the target object and correcting any contact deviations that may exist in the initial pose.
[0059] Contact relationship refers to the spatial association between contact points in the tactile point cloud features of an object and object points in the visual point cloud features of the object.
[0060] The second constraint term is a mechanical quantification index in the electromagnetic force constraint model that characterizes the "distance from the requirement" between geometric point cloud features and object visual point cloud features. Its specific value is determined based on the relative positional relationship between the two. Its core logic is the principle of mutual repulsion of like charges in electromagnetism (i.e., electromagnetic repulsion force), used to avoid the problem of penetration between the target object and the robot in space. When the relative position of the object point and the robot point meets the penetration risk judgment condition, the quantification calculation of the second constraint term will guide the two to move away from each other in space, ensuring that the optimized target object pose and the robot model do not overlap spatially, meeting the physical feasibility requirements.
[0061] Optionally, in one feasible implementation of this application, it is first determined whether the current feature data satisfies any of the physical feasibility constraints.
[0062] Specifically, this method checks whether the location of the tactile point in the object's tactile point cloud features can be found in a corresponding contact area on the object's point cloud surface. If there is an excessively large gap between the tactile point and the object's surface, the contact constraint is considered not satisfied. Furthermore, this method checks for significant geometric overlap between the object's visual point cloud features and geometric point cloud features. If negative distance or depth penetration exists, the penetration constraint is determined to be violated. If any of these physical feasibility constraints are mismatched, it indicates that the initial pose obtained from the feature data contains errors and requires further optimization.
[0063] After confirming that no physical feasibility constraint is met, the contact relationship is transformed into the first constraint term based on the spatial distance between the object's visual point cloud features and the object's tactile point cloud features using the electromagnetic force formula. The closer the distance, the larger the first constraint term. The second constraint term is constructed based on the relative positional relationship between the geometric point cloud features and the object's visual point cloud features, determining the penetration risk between the geometric point cloud features and the object's visual point cloud features. The higher the penetration risk, the larger the second constraint term. Finally, an electromagnetic force constraint model containing the synergistic effect of the two constraints is formed.
[0064] It should be noted that when the feature data satisfies all physical feasibility constraints, i.e., contact constraints, penetration constraints, and motion constraints, the initial pose is considered to have sufficient physical rationality. In this case, there is no need to construct an electromagnetic force constraint model or perform further optimization; instead, the initial pose is directly output as the target pose.
[0065] S400, optimize the initial pose according to the first constraint and the second constraint to obtain the target pose of the target object.
[0066] Optionally, in one feasible implementation of this application, the initial pose optimization is achieved through the synergistic effect of the first constraint and the second constraint. First, the quantized values and directions of action of the first and second constraints are extracted. The first constraint pulls the object's visual point cloud features closer to the object's tactile point cloud features, while the second constraint pushes the object's visual point cloud features away from the geometric point cloud features. Together, they constrain and form a dynamically balanced electromagnetic force field. Then, using the initial pose as a reference, the combined effect of the two constraints is used as the driving signal for pose adjustment, establishing a mapping relationship between the constraints and pose changes. A coordinate transformation algorithm is then used to calculate the object's position correction and posture rotation in three-dimensional space.
[0067] During the adjustment process, the relative positional changes of the object's visual point cloud features, tactile point cloud features, and geometric point cloud features are fed back in real time, and the magnitude and orientation of the first and second constraint terms are dynamically updated. Through multiple rounds of iterative optimization, until the object's visual point cloud features and tactile point cloud features meet the contact requirements and there is no risk of penetration with the geometric point cloud features, the output adjusted pose is the final target pose of the target object, ensuring that it conforms to the laws of physical interaction.
[0068] In a method for optimizing the pose of an object provided in this application embodiment, firstly, feature data including the visual point cloud features of the target object, the geometric point cloud features of the robot's end effector, and the tactile point cloud features of the object are acquired, along with an initial pose estimated based on these feature data. Then, it is determined whether the feature data satisfies at least one physical feasibility constraint; wherein, the physical feasibility constraint includes at least one of contact constraints, penetration constraints, and motion constraints. If the feature data does not satisfy any physical feasibility constraint, it indicates that the initial pose estimated based on the feature data has an accuracy defect. Therefore, an electromagnetic force constraint model can be constructed using a first constraint term determined based on the contact relationship between the object's tactile point cloud features and the object's visual point cloud features, and a second constraint term determined based on the relative positional relationship between the geometric point cloud features and the object's visual point cloud features. The first constraint term can accurately capture the actual contact state between the robot and the target object, while the second constraint term can effectively avoid the penetration problem between the object and the robot. By leveraging the synergistic effect of the first and second constraint terms of the electromagnetic force constraint model, the initial pose is optimized to obtain the target pose of the object. This eliminates the need for model training based on a large amount of tactile-visual paired data, freeing the model from dependence on specific tactile sensor data. It also avoids or reduces the impact of inconsistent data distribution caused by differences in sensor type, wear, or calibration errors on the model estimation results, thereby improving the accuracy of object pose estimation.
[0069] In one embodiment, the object visual point cloud features include multiple visual feature points; the object tactile point cloud features include multiple tactile feature points;
[0070] The determination of whether the feature data satisfies at least one physical feasibility constraint includes determining whether the contact constraint is satisfied:
[0071] For any visual feature point, obtain the shortest distance from the visual feature point to all the tactile feature points;
[0072] Based on the shortest distance, determine the contact probability that the visual feature point and the tactile feature point will come into contact;
[0073] If the average contact probability of all the visual feature points is less than a preset contact threshold, it is determined that the contact constraint is not satisfied.
[0074] Optionally, in this embodiment, the shortest distance refers to the minimum linear distance between a single visual feature point and all tactile feature points. The spatial distance between two points can be calculated using the Euclidean distance formula based on the three-dimensional coordinates of the visual feature point and each tactile feature point, and then the minimum distance value can be selected.
[0075] Contact probability is a quantitative indicator that represents the likelihood of actual contact between a visual feature point and all tactile feature points, calculated using a preset mapping rule based on the shortest distance from the visual feature point to all tactile feature points. The smaller the shortest distance, the closer the visual feature point is to the contact area detected by the tactile sensor, and the closer the contact probability is to 1; conversely, the larger the shortest distance, the closer the contact probability is to 0.
[0076] Optionally, in one specific implementation of this application, for each visual feature point, all tactile feature points are traversed first, and the spatial distance from the visual feature point to each tactile feature point is calculated using the Euclidean distance formula. The minimum value among these distances is then selected as the shortest distance. Subsequently, based on the probabilistic contact model (Gaussian process), using the formula:
[0077]
[0078] Calculate the contact probability that the visual feature point will make contact with the contact feature point corresponding to the shortest distance, where The sensor noise parameter (default value 0.001m) is used; the smaller the shortest distance, the closer the contact probability is to 1. After the contact probabilities of all visual feature points are calculated, these probability values are averaged to obtain the average contact probability, which is then compared with the preset adaptive contact threshold. (The initial value is 0.6, which can be adjusted online based on historical contact stability) comparison is performed; if the average contact probability is less than the threshold, the feature data is determined not to meet the contact constraints, and the subsequent corresponding processing flow is triggered; if the average contact probability is not less than the threshold, it means that the current feature data meets the requirements of the contact constraints.
[0079] In these alternative embodiments, the contact state of the "object-robot" is quantified by calculating the shortest distance, contact probability, and average contact probability between visual feature points and tactile feature points. This can effectively identify contact deviations in the initial pose and ensure the reliability of contact constraint checks.
[0080] In one embodiment, the object visual point cloud features include multiple visual feature points;
[0081] The determination of whether the feature data satisfies at least one physical feasibility constraint includes determining whether it satisfies a penetration constraint:
[0082] Based on the geometric point cloud features, a symbolic distance function field is constructed; the symbolic distance function field is used to characterize the shortest distance from any spatial point to the surface of the end effector.
[0083] Based on the symbolic distance function field, determine the target distance from each of the visual feature points to the surface of the end effector;
[0084] The penetration depth of each visual feature point through the surface of the end effector is determined based on the target distance;
[0085] If the average penetration depth of all the aforementioned visual feature points is greater than a preset depth threshold, it is determined that the penetration constraint is not satisfied.
[0086] Optionally, in the embodiments of this application, the symbolic distance function field is a spatial geometric field constructed based on geometric point cloud features. Its core is to assign a value to any point in space. The absolute value of the value represents the shortest Euclidean distance from the point to the surface of the end effector. The sign is used to distinguish the positional relationship: if the point is outside the surface of the end effector, the value is positive; if the point is inside the surface of the end effector (i.e., penetration occurs), the value is negative.
[0087] The target distance is the output value obtained by substituting the visual feature point into the signed distance function field. It directly corresponds to the shortest distance and positional relationship between the visual feature point and the end effector surface. When the target distance is positive, it indicates that the visual feature point is outside the end effector and no penetration has occurred; when the target distance is negative, it indicates that the visual feature point is inside the end effector and penetration has occurred.
[0088] Penetration depth is a quantitative indicator defined for visual feature points with negative target distances. Its value is the absolute value of the target distance, representing the depth to which the visual feature point penetrates into the surface of the end effector. If the target distance is positive, the penetration depth is 0 (no penetration); the smaller the negative value of the target distance, the greater the penetration depth, indicating a higher degree of spatial overlap between the object and the end effector.
[0089] Optionally, in one specific implementation of this application, a Signed Distance Function (SDF) field is first constructed in real time based on the geometric point cloud features; then, each visual feature point is substituted into the SDF field to obtain the target distance (i.e., SDF value) from the corresponding point to the end effector surface; for the target distance, if it is negative, the absolute value is taken as the penetration depth of the visual feature point (if it is positive, the penetration depth is 0), and the penetration degree is quantified by combining the penetration energy formula:
[0090]
[0091] in, =1000.0, The total number of visual feature points. Let i be the i-th visual feature point.
[0092] After the penetration depth of all visual feature points has been calculated, these depth values are averaged to obtain the average penetration depth, which is then compared with an adaptive preset depth threshold. (Initial value is 0.005m) Compare; if the average penetration depth is greater than the threshold, it is determined that the feature data does not meet the penetration constraint and the corresponding processing flow is triggered; if it is not greater, it means that the current feature data meets the requirements of the penetration constraint.
[0093] In these alternative embodiments, the penetration degree between the object and the end effector surface is quantified by constructing a signed distance function field, thus realizing a digital representation of the penetration constraint. When significant penetration is detected, the electromagnetic force constraint model is triggered to optimize the pose, effectively preventing illegal penetration between the object and the robot model and significantly improving the physical rationality of the pose estimation.
[0094] In one embodiment, determining whether the feature data satisfies at least one physical feasibility constraint includes determining whether a motion constraint is satisfied:
[0095] Extract target feature points from the visual point cloud features and the tactile point cloud features of the object;
[0096] Based on the feature data, determine the displacement of each target feature point between consecutive frames;
[0097] If the average displacement of all the target feature points is greater than a preset displacement threshold, or if the standard deviation of the displacement of all the target feature points is greater than a preset standard deviation threshold, then the motion constraint is determined not to be satisfied.
[0098] Optionally, in the embodiments of this application, the target feature points are representative key geometric points selected from the visual point cloud features and tactile point cloud features of the object. Typically, they are selected from regions rich in features such as edges and corners, or from the core points of the tactile contact area.
[0099] Displacement between consecutive frames refers to the change in the three-dimensional spatial coordinates of the same target feature point at two adjacent data acquisition times (i.e., two consecutive frames). It quantifies the range of motion of the target feature point over a short period of time.
[0100] Optionally, in one specific implementation of this application, the scale-invariant feature transform (SIFT) feature points of the object's visual point cloud features and tactile point cloud features are first extracted as target feature points (SIFT feature points have local feature representativeness and can stably reflect the key geometric information of the object and the contact area).
[0101] Subsequently, based on the feature data from two consecutive frames, the three-dimensional spatial coordinates of each target feature point in the preceding and following frames are obtained. The displacement of the feature point between consecutive frames is obtained by calculating the Euclidean distance of the coordinates. After the displacements of all target feature points are calculated, these displacement values are averaged to obtain the average displacement. At the same time, the standard deviation of the displacement is calculated to characterize the dispersion of the displacement. Finally, the average displacement is compared with an adaptive preset displacement threshold. (Initial value is 0.02m) Compare the displacement standard deviation with the preset standard deviation threshold. If the average displacement is greater than the displacement threshold (indicating that the object's motion amplitude is too large), or the displacement standard deviation is greater than the standard deviation threshold (indicating that the object's motion trajectory is irregular and there are abnormal fluctuations), then the feature data is determined not to meet the motion constraints. If neither of the two indicators exceeds the corresponding threshold, then the current feature data meets the requirements of the motion constraints.
[0102] In these alternative embodiments, by extracting target feature points, calculating continuous frame displacement, and combining the average displacement and standard deviation as dual indicators to check the motion state, it is possible to identify jitter problems with excessive object motion amplitude and capture abnormal fluctuations with irregular trajectories, thereby improving the reliability of determining whether the motion smoothness requirements are met.
[0103] In one embodiment, constructing the electromagnetic force constraint model based on the feature data includes:
[0104] The visual point cloud features and geometric point cloud features of the object are used as the first charged particles, and the tactile point cloud features of the object are used as the second charged particles; the first charged particles and the second charged particles are different terms of positively charged particles and negatively charged particles.
[0105] Based on Coulomb's law, the second constraint term is determined according to the charge relationship between the visual point cloud features and the geometric point cloud features of the object, and the first distance between the visual point cloud features and the geometric point cloud features of the object; the first distance is positively correlated with the second constraint term.
[0106] Based on Coulomb's law, the first constraint term is determined according to the charge relationship between the visual point cloud features and the tactile point cloud features of the object, and the second distance between the visual point cloud features and the tactile point cloud features of the object. The second distance is negatively correlated with the first constraint term.
[0107] The electromagnetic force constraint model is constructed based on the first constraint term and the second constraint term.
[0108] Optionally, in the embodiments of this application, the first charged particle is an abstract analogy between the visual point cloud features and the geometric point cloud features of the object, and is equivalent to particles with the same charge property (such as all being positively charged particles). This abstraction is to simulate the physical requirement of avoiding spatial penetration between the object and the robot by using the law of "like charges repel each other" in electromagnetism.
[0109] The second charged particle is an abstract analogy to the tactile point cloud features of an object, and its charge properties are opposite to those of the first charged particle (e.g., it can be set as a negatively charged particle). By setting this opposite charge, the electromagnetic law of "opposite charges attract each other" can be used to simulate the interaction requirement that the contact area between the object and the robot needs to be in close contact.
[0110] The charge relationship refers to the association of charge attributes (i.e., same or opposite) between a first charged particle and a second charged particle. In this application, the first charged particle corresponding to the object / robot and the second charged particle corresponding to touch have opposite charge relationships (used to generate the electromagnetic attraction corresponding to the first constraint term), while the first charged particles corresponding to the object and the robot have the same charge relationship (used to generate the electromagnetic repulsion corresponding to the second constraint term). The direction of the electromagnetic force is clearly defined by setting the charge attributes.
[0111] Optionally, in one specific implementation of this application, the feature data is first abstracted and mapped, and the visual point cloud features and geometric point cloud features of the object are uniformly equivalent to "first charged particles" (e.g., set as positive charged particles), and the tactile point cloud features of the object are equivalent to "second charged particles" (set as negative charged particles).
[0112] Subsequently, the second constraint term is calculated based on Coulomb's law: considering the visual point cloud features and geometric point cloud features of the object, the relationship between the like charges of the two is used, combined with the first distance from the visual feature point to the robot feature point (i.e., the spatial Euclidean distance), and the magnitude of the electromagnetic repulsion force is calculated by substituting it into the Coulomb force formula, thus obtaining the second constraint term.
[0113] Next, the first constraint term is calculated: considering the visual point cloud features and tactile point cloud features of the object, the opposite charge relationship between the two is used, combined with the second distance from the visual feature point to the tactile feature point, and the magnitude of the electromagnetic attraction is calculated by substituting into the Coulomb force formula to obtain the first constraint term.
[0114] Finally, the first and second constraints corresponding to all visual feature points are integrated to clarify the direction of action and quantification value of different constraints, forming an electromagnetic force constraint model that can simultaneously represent the "object-tactile fit requirement" and the "object-robot penetration avoidance requirement", providing a force-driven basis for subsequent pose optimization.
[0115] In these alternative embodiments, point cloud features are abstracted into heterogeneous / homogeneous charged particles, and an electromagnetic force constraint model is constructed using Coulomb's law. This model simulates the need for tactile contact between objects through the attraction of heterogeneous charges, while mitigating the risk of objects penetrating robots through the repulsive force of like charges, providing an accurate and physically logical force-driven basis for pose optimization.
[0116] In one embodiment, optimizing the initial pose based on the first constraint and the second constraint to obtain the target pose of the target object includes:
[0117] A loss function is constructed based on the first constraint term and the second constraint term; wherein the loss function is positively correlated with the second constraint term and negatively correlated with the first constraint term;
[0118] By optimizing the model, iterative optimization is performed based on the feature data and the loss function to obtain the adjusted pose.
[0119] The initial pose is adjusted according to the adjusted pose to obtain the target pose.
[0120] Optionally, in the embodiments of this application, the loss function is positively correlated with the second constraint term (the larger the second constraint term, the greater the risk of penetration between the object and the robot or the distance is too close, and the larger the loss value), and negatively correlated with the first constraint term (the larger the first constraint term, the higher the degree of contact between the object and the tactile sensor, and the smaller the loss value). By transforming physical interaction constraints into quantifiable mathematical indicators, a clear iterative direction is provided for pose optimization (i.e., minimizing the ideal pose state corresponding to the loss function).
[0121] The optimization model is an algorithmic framework for implementing pose iterative optimization. Its core function is to solve for the pose parameters that minimize the loss value based on feature data and the loss function mentioned above.
[0122] Optionally, in one specific implementation of this application, starting from the initial pose, the visual point cloud features, geometric point cloud features, and tactile point cloud features of the object are input. A gradient descent algorithm is used to calculate the gradient direction of the loss function, and the rotation and translation parameters of the pose are iteratively adjusted. In each iteration, the energy terms corresponding to the first and second constraint terms are updated based on the current pose, thereby updating the loss function value, and the pose parameters are then fine-tuned along the gradient direction. This iterative process continues until the loss function converges to a preset threshold (or reaches the maximum number of iterations), resulting in an adjusted pose that minimizes the loss function. Finally, the rotation and translation parameters corresponding to this adjusted pose are applied to the initial pose to complete the position and orientation correction, obtaining the target pose that meets the constraints of "object-tactile fit and object-robot non-penetration".
[0123] In these alternative embodiments, physical constraints are transformed into quantitative optimization objectives through a loss function, taking into account both the fit between the object and the tactile feedback and the robot's need to avoid penetration; the iterative optimization model precisely adjusts the pose parameters so that the final target pose meets the physical logic of actual interaction, effectively improving the accuracy and rationality of pose estimation and avoiding unreasonable deviations in the initial pose.
[0124] In one embodiment, the object visual point cloud features include multiple object point clouds of the target object at different times;
[0125] The step of constructing the loss function based on the first constraint term and the second constraint term includes:
[0126] The angular velocity of the target object is determined based on the point cloud of the multiple objects.
[0127] The magnetic field energy of the electromagnetic force constraint model is determined based on the angular velocity and the distance between the visual point cloud features and the tactile point cloud features of the object.
[0128] The loss function is determined based on the first constraint, the second constraint, and the magnetic field energy; wherein the magnetic field energy is positively correlated with the loss function.
[0129] Optionally, in the embodiments of this application, the angular velocity of the target object is a physical quantity calculated based on the pose changes of the object's point cloud at different times, used to characterize the rotation rate and direction of the target object per unit time.
[0130] Magnetic field energy is one of the components of the loss function. It is positively correlated with angular velocity (the greater the angular velocity, the higher the magnetic field energy) and also positively correlated with the loss function. Its core function is to constrain the rotation amplitude of the object's motion, avoid unstable situations such as rotational jitter during pose optimization, and improve the dynamic stability of pose adjustment.
[0131] Optionally, in one specific implementation of this application, the angular velocity of the target object is first estimated based on multiple object point clouds at different times by means of pose transformation (such as rotation matrix transformation) of consecutive frames. Subsequently, the object points in the object's visual point cloud features were combined. tactile points in the tactile point cloud features of objects Substituting into the magnetic field energy formula, we can calculate the magnetic field energy:
[0132]
[0133] in, For magnetic field energy, These are the weighting coefficients for the magnetic field energy term, for example... =0.1, the magnetic field energy is positively correlated with the angular velocity, and is used to constrain the smoothness of the object's rotation.
[0134] Simultaneously, the potential energy of the first constraint term and the potential energy of the second constraint term are calculated separately: the potential energy of the first constraint term is obtained through the following formula:
[0135]
[0136] in, Potential energy is the first constraint term, and the point charge of the object is... =-1, tactile point charge =+1 (for illustrative purposes only) For example, the weighting coefficients of the potential energy term in the first constraint term. =0.3, =1e-5, the first constraint term potential energy is used to characterize the degree of fit between the object and the touch.
[0137] The potential energy of the second constraint term is obtained through the following formula:
[0138]
[0139] in, The potential energy is the second constraint term, and the point charge of the end effector in the geometric point cloud features is the point charge. =-1 (same as the point charge of the object) For example, the weighting coefficients of the potential energy term in the second constraint term. =0.3, The end effector point is the feature of the geometric point cloud. The second constraint term, potential energy, is used to characterize the penetration constraint between the object and the robot.
[0140] Finally, an L2 regularization term can be introduced to avoid overfitting:
[0141]
[0142] in, For the regularization term value, The weight coefficients for the regularization term, for example =0.3, These are the pose transformation parameters (i.e., pose adjustment parameters).
[0143] Finally, the above four items are integrated into the total loss function:
[0144]
[0145] The loss function is constructed to provide a quantified target for subsequent pose optimization.
[0146] In this embodiment, an electromagnetic force constraint model is used to uniformly model hand-object interaction, transforming the object's tactile point cloud features, geometric point cloud features, and visual point cloud features into physical constraints, ensuring that the pose estimation results conform to physical feasibility. This compensates for the shortcomings of pure vision methods that are prone to failure in occluded scenes, and avoids the problems of penetration or unreasonable movement caused by the lack of physical constraints in visual-tactile methods.
[0147] In terms of methodology, the feasibility check module uses a probabilistic model (Gaussian process) and SDF-based penetration detection to evaluate contact, penetration, and motion constraints in real time, and adaptive thresholds to improve robustness to noise. The optimized loss function includes a first constraint term (tactile-object alignment), a second constraint term (avoiding robot-object penetration), and magnetic field energy (smoothing motion constraints), which together ensure the physical consistency of the pose.
[0148] The advantage of this design is that it can maintain stable pose tracking, reduce accumulated errors, and avoid acute failures in severely occluded in-hand manipulation tasks. For example, when vision is completely blocked, tactile signals can still "pull" the object's pose to the correct position through the first constraint.
[0149] In these alternative embodiments, the constructed loss function integrates the first constraint term, the second constraint term, and the magnetic field energy. The first constraint term ensures the fit between the object and the touch, the second constraint term avoids the risk of penetration by the robot, and the magnetic field energy constrains the rotational stability. Combined with the regularization term, it avoids overfitting and effectively improves the accuracy, stability, and physical rationality of the pose estimation.
[0150] In one embodiment, the step of iteratively optimizing the model based on the feature data and the loss function to obtain the adjusted pose includes:
[0151] Based on the feature data, the optimized model is used to determine the pose adjustment.
[0152] The object visual point cloud features are updated based on the adjusted pose to obtain the updated object visual point cloud features.
[0153] The loss value is determined based on the updated object visual point cloud features and the loss function;
[0154] If the loss value does not meet the preset stopping condition, the model parameters of the optimization model are adjusted, and the process returns to the step of determining the adjusted pose based on the feature data through the optimization model, until the loss value meets the preset stopping condition, and the adjusted pose after iterative optimization is obtained.
[0155] Optionally, in one specific implementation of this application, initialization is performed first. The translation parameters of the object pose are determined by the weighted average of the activated tactile points (with the weights being confidence levels). The rotation parameters of the object pose can be obtained from the initial visual pose or proprioception. Then, iterative optimization is initiated. The first two iterations use the Adaptive Moment Estimation (Adam) algorithm (with a learning rate of 1e-5) for warm-up. The last three iterations switch to the Limited-memory BFGS (L-BFGS) algorithm (maximum of 5 iterations) to accelerate convergence.
[0156] In each iteration, first construct and adjust the pose. The corresponding SE(3) matrix is used to describe the matrix representation of the joint transformation of the target object in three-dimensional space, which can be understood as the update amount of the pose adjustment on the rotation and translation components. Then, the adjusted pose is applied to the original object visual point cloud features to obtain the updated object visual point cloud features P_O_optimized; then, based on the updated object visual point cloud features, the potential energy of the first constraint term, the potential energy of the second constraint term, the magnetic field energy and the L2 regularization term are calculated and substituted into the total loss function to obtain the current loss value; the gradient of the loss function with respect to the optimization variables is calculated through backpropagation, the model parameters are adjusted according to the selected optimization algorithm (such as Adam's gradient update, L-BFGS's quasi-Newton iteration), and the step is returned to redetermine the adjusted pose.
[0157] If the change in loss value is less than 1e-5 during the iteration process, the process will terminate early. If the preset stopping condition is not met, the iteration will continue until the maximum number of iterations is reached or the loss value converges. Finally, the adjusted pose after the iteration optimization is completed will be output.
[0158] In other implementations, when the interaction intensity changes (e.g., a sudden change in object speed or a change in contact pattern), the magnitude of the charge can be adjusted (i.e., ...). This indirectly changes the strength of the potential energy of the first / second constraint term, and combined with the optimization of relative pose transformation, makes pose estimation more stable and converges faster in complex tasks (such as hand-to-hand hand exchange).
[0159] In this embodiment, the present application avoids the dependence of existing visual-haptic methods on large-scale haptic datasets. Through optimization algorithms and physical constraint modeling, it can be directly applied to unknown objects and novel haptic sensors without the need for pre-training with haptic data. This solves the core problems of difficult data collection, large gap between simulation and real-world application, and poor generalization ability in related technologies.
[0160] Specifically, the optimization algorithm is based on an electromagnetic force constraint model, whose parameters are general physical constants or adaptive settings, independent of training data specific to any object or sensor. Simultaneously, the tactile signal processing module uses a lightweight CNN and learnable coordinate transformations to extract general features from raw sensor data, eliminating the need for calibration or retraining for specific sensors. The advantage of this design is that, even with zero-shot setups, the method can quickly adapt to new objects and new tactile sensors, significantly improving its practicality and scalability.
[0161] In these alternative embodiments, an iterative optimization and dynamic update mechanism is adopted, combined with the Adam-L-BFGS hybrid strategy. This approach not only rapidly approaches the optimal solution region through preheating iterations, but also accurately adjusts the pose using an efficient convergence algorithm, effectively improving the accuracy and stability of pose optimization and avoiding local optima.
[0162] In one embodiment, the step of iteratively optimizing the model based on the feature data and the loss function to obtain the adjusted pose includes:
[0163] By optimizing the model, iterative optimization is performed based on the feature data and the loss function to obtain the adjusted pose and covariance matrix; the diagonal elements of the covariance matrix are used to characterize the variance of the adjusted pose in the position and rotation components, and the off-diagonal elements of the covariance matrix are used to characterize the covariance between different adjusted poses.
[0164] The method further includes:
[0165] Based on the covariance matrix, the model parameters of the optimization model are adjusted, and the optimized model with adjusted parameters is used to perform iterative optimization based on the feature data and the loss function to obtain the adjusted pose.
[0166] Optionally, in the embodiments of this application, the covariance matrix is a symmetric matrix that characterizes the uncertainty and correlation of each parameter (position component and rotation component) of the pose adjustment. Its dimension is consistent with the dimension of the pose parameters (e.g., a 6×6 covariance matrix corresponds to a 6-dimensional pose).
[0167] The covariance in the covariance matrix specifically refers to the statistics corresponding to the off-diagonal elements of the matrix. It is used to quantify the degree of linear correlation between different parameters in pose adjustment (such as x-axis translation and rotation around the y-axis, z-axis translation and rotation around the x-axis, etc.). The sign of this value reflects the direction of the correlation (positive or negative correlation), and the magnitude of the absolute value reflects the strength of the correlation: the larger the absolute value, the stronger the correlation between the changes of the two parameters; the closer the absolute value is to 0, the more independent the changes of the two parameters are.
[0168] Optionally, in one specific implementation of this application, the optimization model is first started, and iterative processes are carried out based on feature data and total loss function. After the first two warm-up iterations using the Adam algorithm, the algorithm is switched to L-BFGS. During the L-BFGS iteration process, the approximate Hessian matrix of the optimization process is stored. The Hessian matrix is used to describe the curvature of the function. After the adjusted pose is obtained, the 6×6 covariance matrix is calculated by the inverse of the Hessian matrix.
[0169] Then, the model parameters are adjusted and optimized based on the covariance matrix: if the variance of a parameter is large (high values of diagonal elements), the optimization weight of that parameter is increased; if the covariance of two parameters is high (large absolute values of off-diagonal elements), the constraint relationship between the parameters is adjusted. After adjustment, the optimization model is restarted, iterating based on feature data and loss function until the loss value meets the preset stopping condition, and finally outputting the adjusted pose; simultaneously, through... The process of updating the initial pose to the target pose is completed, achieving a combination of pose optimization and uncertainty constraints. For the target pose, To adjust posture, This is the initial pose.
[0170] In these alternative embodiments, a 6×6 covariance matrix of the adjusted pose is output to provide a confidence reference for robot manipulation strategies (such as motion planning or force control), quantifying the uncertainty of pose estimation. When the uncertainty is high (such as severe occlusion), the robot can switch to a conservative strategy (e.g., increasing the safety threshold of force control to avoid collisions caused by inaccurate pose estimation), improving mission safety.
[0171] Optionally, in the embodiments of this application, the data preprocessing stage adopts neural point cloud representation, which greatly reduces the storage and computational overhead of point cloud. At the same time, it focuses on the key interactive areas of the object through adaptive sampling, which further improves the processing efficiency. During testing, a hybrid optimization strategy of "Adam warm-up + L-BFGS acceleration" is adopted to ensure rapid convergence within 5 iterations, further compressing the time consumption.
[0172] The advantage of this design is that in real-time robot manipulation scenarios such as grasping and redirection, the pose update frequency is high, which can respond to the dynamic changes in the motion state of the object in a timely manner, providing accurate and efficient pose feedback for the control system and ensuring the smoothness and stability of the manipulation task.
[0173] On the other hand, the tactile signal processing module of this application adopts a lightweight CNN architecture, which outputs a general feature point cloud rather than a sensor-specific signal, and is compatible with various types of devices such as visual tactile sensors and force tactile sensors; the coordinate transformation network supports online learning of hand-eye relationship, which greatly reduces the dependence on precise calibration; at the same time, it outputs standard SE(3) pose and covariance matrix, which can be seamlessly connected with mainstream robot frameworks.
[0174] The core effect of this design is that users can flexibly replace tactile sensors or robot end effectors according to actual needs without redesigning the model architecture or adjusting the core algorithm, which significantly reduces the cost and complexity of technology deployment and improves the practicality and adaptability of the solution.
[0175] It should be noted that the various optional implementation methods described in the embodiments of this application can be combined with each other or implemented individually without conflict, and the embodiments of this application do not limit this.
[0176] Alternatively, in another embodiment of this application, such as Figure 2 As shown, a pose optimization system for an object is also provided. This system takes multi-source data as input, physical constraints as its core, and accurate pose output as its goal. Specifically, it includes the following modules:
[0177] The first layer is the data preprocessing layer. This layer receives input data from visual sensors, tactile sensors, and the robot's proprioception, and sequentially executes the processing flow of adaptive point cloud sampling, tactile feature extraction, and coordinate transformation network to transform multi-source data into standardized feature representations.
[0178] Next is the probabilistic feasibility check layer. Based on the features processed above, it sequentially performs contact checks, collision checks, SDF penetration checks, and motion checks. From the dimensions of contact rationality, spatial non-penetration, and motion stability, it evaluates the physical feasibility of the current pose, providing a basis for subsequent optimization.
[0179] The next step is the electromagnetic force system optimization layer. This module constructs an electromagnetic force constraint model and loss function based on the feasibility check results. It then performs iterative optimization through a hybrid optimizer (such as Adam+L-BFGS), while simultaneously achieving adaptive parameter adjustment and uncertainty estimation. This combines pose parameters with physical constraints to complete precise optimization.
[0180] Finally, the output layer outputs the optimized and accurate SE(3) pose, the corresponding uncertainty estimation results (i.e., the covariance matrix), and the visualized output content, providing clear and reliable pose feedback for the robot manipulation system.
[0181] Specifically, such as Figure 3 As shown, the data preprocessing layer starts with the 3D mesh model of the target object and the 3D mesh model of the robot's end effector as inputs. It first receives the raw data from the tactile sensor synchronously, and performs curvature calculation and feature detection on both types of 3D mesh models. Then, based on the curvature detection results, adaptive sampling is performed, and the features extracted from the tactile sensor data by a lightweight CNN are combined and input into the feature fusion and transformation module to complete the integration and spatial transformation of multi-source features, and output the neural point cloud representation and encoding results. The encoding results are processed by the SDF encoding network and the pose update network in sequence, and finally the output of the data preprocessing layer is completed through the feature adaptation module (i.e., the object's visual point cloud features, geometric point cloud features, and object's tactile point cloud features are obtained), providing standardized feature inputs for the subsequent pose optimization process.
[0182] like Figure 4 As shown, for the probabilistic feasibility check layer, the transformed point cloud (i.e., object visual point cloud features, geometric point cloud features, and object tactile point cloud features) is first input, and then the feasibility check stage is entered. This stage simultaneously performs contact constraint checks, penetration constraint checks, and motion constraint checks (see the aforementioned embodiments for details, which will not be repeated here). Then, the results of the three checks are integrated and evaluated through probabilistic fusion and decision logic, and then it is determined whether optimization is needed: if optimization is determined to be needed, the optimization process is triggered; if optimization is determined not to be needed, the current pose is retained, and the process ends.
[0183] Figure 5 A schematic diagram of the structure of an object pose optimization device provided in another embodiment of this application is shown. For ease of explanation, only the parts related to the embodiments of this application are shown.
[0184] Reference Figure 5 The object pose optimization device may include:
[0185] The acquisition module 501 is used to acquire feature data and the initial pose of the target object estimated by the robot under the feature data; the feature data includes the object visual point cloud features of the target object, the geometric point cloud features of the robot's end effector, and the object tactile point cloud features of the robot contacting the target object.
[0186] The judgment module 502 is used to determine whether the feature data satisfies at least one physical feasibility constraint; wherein the physical feasibility constraint includes at least one of contact constraint, penetration constraint and motion constraint;
[0187] Construction module 503 is used to construct an electromagnetic force constraint model based on the feature data when the feature data does not meet any physical feasibility constraint; the electromagnetic force constraint model includes a first constraint term and a second constraint term; the first constraint term is determined based on the contact relationship between the tactile point cloud features of the object and the visual point cloud features of the object, and the second constraint term is determined based on the relative positional relationship between the geometric point cloud features and the visual point cloud features of the object;
[0188] The optimization module 504 is used to optimize the initial pose according to the first constraint term and the second constraint term to obtain the target pose of the target object.
[0189] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application, and are devices corresponding to the above-mentioned methods. All implementation methods in the above-mentioned method embodiments are applicable to the embodiments of this device. For details on its specific functions and the technical effects it brings, please refer to the method embodiment section, which will not be repeated here.
[0190] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0191] Figure 6 A schematic diagram of the hardware structure of the electronic device provided in an embodiment of this application is shown.
[0192] The device may include a processor 601 and a memory 602 storing program instructions.
[0193] When the processor 601 executes the program, it implements the steps in any of the above method embodiments.
[0194] For example, the program can be divided into one or more modules / units, one or more of which are stored in memory 602 and executed by processor 601 to complete this application. The one or more modules / units can be a series of program instruction segments capable of performing a specific function, which describe the execution process of the program in the device.
[0195] Specifically, the processor 601 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0196] Memory 602 may include mass storage for data or instructions. For example, and not limitingly, memory 602 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 602 may include removable or non-removable (or fixed) media. Where appropriate, memory 602 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 602 is non-volatile solid-state memory.
[0197] Memory may include read-only memory (ROM), random access memory (RAM), disk storage media devices, optical storage media devices, flash memory devices, and electrical, optical, or other physical / tangible memory storage devices. Therefore, typically, memory includes one or more tangible (non-transitory) readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the methods according to one aspect of this disclosure.
[0198] The processor 601 implements any of the methods described in the above embodiments by reading and executing program instructions stored in the memory 602.
[0199] In one example, the electronic device may also include a communication interface 603 and a bus 610. The processor 601, memory 602, and communication interface 603 are connected via the bus 610 and communicate with each other.
[0200] The communication interface 603 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.
[0201] Bus 610 includes hardware, software, or both, that couples components of an online data traffic metering device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 610 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, any suitable bus or interconnect is contemplated herein.
[0202] Furthermore, in conjunction with the methods in the above embodiments, this application embodiment can provide a storage medium for implementation. This storage medium stores program instructions; when these program instructions are executed by a processor, they implement any of the methods in the above embodiments.
[0203] This application also provides a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled. The processor is used to run programs or instructions to implement the various processes of the above method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.
[0204] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0205] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above method embodiments and achieve the same technical effects. To avoid repetition, it will not be described again here.
[0206] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.
[0207] The functional modules shown in the above block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on machine-readable media or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable media" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer grids such as the Internet, intranets, etc.
[0208] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0209] The aspects of this disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and program products according to embodiments of this disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to create a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by special-purpose hardware performing the specified functions or actions, or can be implemented by a combination of special-purpose hardware and computer instructions.
[0210] The above are merely specific embodiments of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.
Claims
1. A method for optimizing the pose of an object, characterized in that, The method includes: Acquire feature data and the initial pose of the target object estimated by the robot based on the feature data; the feature data includes the visual point cloud features of the target object, the geometric point cloud features of the robot's end effector, and the tactile point cloud features of the target object when the robot contacts the robot. Determine whether the feature data satisfies at least one physical feasibility constraint; wherein the physical feasibility constraint includes at least one of contact constraint, penetration constraint, and motion constraint; If the feature data does not meet any of the physical feasibility constraints, an electromagnetic force constraint model is constructed based on the feature data; the electromagnetic force constraint model includes a first constraint term and a second constraint term; the first constraint term is determined based on the contact relationship between the tactile point cloud features and the visual point cloud features of the object, and the second constraint term is determined based on the relative positional relationship between the geometric point cloud features and the visual point cloud features of the object; Based on the first constraint and the second constraint, the initial pose is optimized to obtain the target pose of the target object; The step of constructing an electromagnetic force constraint model based on the feature data includes: The visual point cloud features and geometric point cloud features of the object are used as the first charged particles, and the tactile point cloud features of the object are used as the second charged particles; the first charged particles and the second charged particles are different terms of positively charged particles and negatively charged particles. Based on Coulomb's law, the second constraint term is determined according to the charge relationship between the visual point cloud features and the geometric point cloud features of the object, and the first distance between the visual point cloud features and the geometric point cloud features of the object; the first distance is positively correlated with the second constraint term. Based on Coulomb's law, the first constraint term is determined according to the charge relationship between the visual point cloud features and the tactile point cloud features of the object, and the second distance between the visual point cloud features and the tactile point cloud features of the object. The second distance is negatively correlated with the first constraint term. The electromagnetic force constraint model is constructed based on the first constraint term and the second constraint term.
2. The method according to claim 1, characterized in that, The object's visual point cloud features include multiple visual feature points; the object's tactile point cloud features include multiple tactile feature points; The determination of whether the feature data satisfies at least one physical feasibility constraint includes determining whether the contact constraint is satisfied: For any visual feature point, obtain the shortest distance from the visual feature point to all the tactile feature points; Based on the shortest distance, determine the contact probability that the visual feature point and the tactile feature point will come into contact; If the average contact probability of all the visual feature points is less than a preset contact threshold, it is determined that the contact constraint is not satisfied.
3. The method according to claim 1, characterized in that, The object's visual point cloud features include multiple visual feature points; The determination of whether the feature data satisfies at least one physical feasibility constraint includes determining whether it satisfies a penetration constraint: Based on the geometric point cloud features, a symbolic distance function field is constructed; the symbolic distance function field is used to characterize the shortest distance from any spatial point to the surface of the end effector. Based on the symbolic distance function field, determine the target distance from each of the visual feature points to the surface of the end effector; The penetration depth of each visual feature point through the surface of the end effector is determined based on the target distance; If the average penetration depth of all the aforementioned visual feature points is greater than a preset depth threshold, it is determined that the penetration constraint is not satisfied.
4. The method according to claim 1, characterized in that, The determination of whether the feature data satisfies at least one physical feasibility constraint includes determining whether it satisfies a motion constraint: Extract target feature points from the visual point cloud features and the tactile point cloud features of the object; Based on the feature data, determine the displacement of each target feature point between consecutive frames; If the average displacement of all the target feature points is greater than a preset displacement threshold, or if the standard deviation of the displacement of all the target feature points is greater than a preset standard deviation threshold, then the motion constraint is determined not to be satisfied.
5. The method according to claim 1, characterized in that, The step of optimizing the initial pose based on the first constraint and the second constraint to obtain the target pose of the target object includes: A loss function is constructed based on the first constraint term and the second constraint term; wherein the loss function is positively correlated with the second constraint term and negatively correlated with the first constraint term; By optimizing the model, iterative optimization is performed based on the feature data and the loss function to obtain the adjusted pose. The initial pose is adjusted according to the adjusted pose to obtain the target pose.
6. The method according to claim 5, characterized in that, The object visual point cloud features include multiple object point clouds of the target object at different times; The step of constructing the loss function based on the first constraint term and the second constraint term includes: The angular velocity of the target object is determined based on the point cloud of the multiple objects. The magnetic field energy of the electromagnetic force constraint model is determined based on the angular velocity and the distance between the visual point cloud features and the tactile point cloud features of the object. The loss function is determined based on the first constraint, the second constraint, and the magnetic field energy; wherein the magnetic field energy is positively correlated with the loss function.
7. The method according to claim 5, characterized in that, The step of iteratively optimizing the model based on the feature data and the loss function to obtain the adjusted pose includes: Based on the feature data, the optimized model is used to determine the pose adjustment. The object visual point cloud features are updated based on the adjusted pose to obtain the updated object visual point cloud features. The loss value is determined based on the updated object visual point cloud features and the loss function; If the loss value does not meet the preset stopping condition, the model parameters of the optimization model are adjusted, and the process returns to the step of determining the adjusted pose based on the feature data through the optimization model, until the loss value meets the preset stopping condition, and the adjusted pose after iterative optimization is obtained.
8. The method according to claim 5, characterized in that, The step of iteratively optimizing the model based on the feature data and the loss function to obtain the adjusted pose includes: By optimizing the model, iterative optimization is performed based on the feature data and the loss function to obtain the adjusted pose and covariance matrix; the diagonal elements of the covariance matrix are used to characterize the variance of the adjusted pose in the position and rotation components, and the off-diagonal elements of the covariance matrix are used to characterize the covariance between different adjusted poses. The method further includes: Based on the covariance matrix, the model parameters of the optimization model are adjusted, and the optimized model with adjusted parameters is used to perform iterative optimization based on the feature data and the loss function to obtain the adjusted pose.
9. A device for optimizing the pose of an object, characterized in that, The device includes: An acquisition module is used to acquire feature data and the initial pose of the target object estimated by the robot under the feature data; the feature data includes the object visual point cloud features of the target object, the geometric point cloud features of the robot's end effector, and the object tactile point cloud features of the robot contacting the target object. The judgment module is used to determine whether the feature data satisfies at least one physical feasibility constraint; wherein the physical feasibility constraint includes at least one of contact constraint, penetration constraint and motion constraint; A construction module is used to construct an electromagnetic force constraint model based on the feature data when the feature data does not meet any physical feasibility constraint; the electromagnetic force constraint model includes a first constraint term and a second constraint term; the first constraint term is determined based on the contact relationship between the tactile point cloud features of the object and the visual point cloud features of the object, and the second constraint term is determined based on the relative positional relationship between the geometric point cloud features and the visual point cloud features of the object; An optimization module is used to optimize the initial pose based on the first constraint and the second constraint to obtain the target pose of the target object. The step of constructing an electromagnetic force constraint model based on the feature data includes: The visual point cloud features and geometric point cloud features of the object are used as the first charged particles, and the tactile point cloud features of the object are used as the second charged particles; the first charged particles and the second charged particles are different terms of positively charged particles and negatively charged particles. Based on Coulomb's law, the second constraint term is determined according to the charge relationship between the visual point cloud features and the geometric point cloud features of the object, and the first distance between the visual point cloud features and the geometric point cloud features of the object; the first distance is positively correlated with the second constraint term. Based on Coulomb's law, the first constraint term is determined according to the charge relationship between the visual point cloud features and the tactile point cloud features of the object, and the second distance between the visual point cloud features and the tactile point cloud features of the object. The second distance is negatively correlated with the first constraint term. The electromagnetic force constraint model is constructed based on the first constraint term and the second constraint term.
Citation Information
Patent Citations
Robot control method, robot and storage medium
CN118254185A
Electromagnetic force compensation method and system based on virtual displacement
CN120595876A