A method and apparatus for evaluating grasp stability
Patent Information
- Application Number
- CN202610408348.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2026-03-30
- Filing Date
- 2026-03-31
- Publication Date
- 2026-08-28
AI Technical Summary
[0003]有鉴于此,本申请提供一种抓取稳定性评估方法和装置,用以解决现有技术中抓取稳定性评估准确性不足的问题
[0019]The grasping stability assessment method and apparatus provided in this application achieve accurate assessment of grasping stability by introducing multimodal visual and tactile perception in a unified three-dimensional space during robot grasping. Specifically, this application first calibrates the grasping system and establishes the transformation relationship between multiple coordinate systems in the grasping system, so that the visual perception data and tactile perception data acquired when grasping contact can be uniformly converted into a three-dimensional geometric representation in the same coordinate system, thereby explicitly establishing the spatial correspondence between visual and tactile information. On this basis, a multimodal three-dimensional representation of the target object is constructed based on the visual and tactile three-dimensional geometric representations, and a completed three-dimensional geometric representation of the target object is obtained based on the multimodal three-dimensional representation. By introducing the completed three-dimensional geometric representation of the target object, the grasping stability assessment can comprehensively consider the overall geometric structural features of the target object in addition to local contact information. Furthermore, the completed three-dimensional geometric representation is input into a pre-trained grasping stability assessment model to assess the grasping stability of the target object and output the assessment result. In summary, this application integrates visual and tactile multimodal perception information in a unified three-dimensional space and introduces approximate reasoning results of the complete shape of the target object as the evaluation basis. This enables the grasping stability assessment to combine local contact geometry and overall object shape characteristics simultaneously. Thus, even when visual observation is obstructed or tactile perception is partial and incomplete, a reliable assessment of grasping stability can still be achieved, effectively improving the accuracy and robustness of the assessment results.
Smart Images

Figure CN122656972A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of robot grasping and control, and in particular to a method and apparatus for assessing grasping stability. Background Technology
[0002] Grasping stability assessment is a crucial step in robot operation to determine the reliability of the current grasping state. Its goal is to determine, based on sensor information, whether the target object can be stably grasped after grasping has occurred or contact has been established. Existing grasping stability assessment methods mainly include tactile information-based assessment schemes and vision-tactile fusion assessment schemes. Tactile methods directly utilize signals such as contact pressure and shear force to determine the grasping state; while simple in structure, they lack an understanding of the overall shape of the target object. Vision-tactile fusion methods supplement global context by introducing visual information, improving assessment accuracy to some extent. However, existing vision-tactile fusion methods are mostly based on two-dimensional images or low-dimensional features for modeling, lacking a unified three-dimensional spatial representation between visual and tactile information. Their spatial correspondence usually relies on implicit learning by the network, making it difficult to form clear and interpretable geometric alignment relationships. Furthermore, these methods mainly rely on observable local perceptual information during grasping, failing to effectively incorporate the complete geometry of the target object. When vision is obstructed or sensor information is incomplete, the reliability of the assessment results is easily affected. Furthermore, grasping stability assessment models typically rely on data-driven training, but existing visual-tactile grasping datasets still have shortcomings in spatial alignment annotation, geometric integrity, and pose diversity, limiting the model's adaptability to complex scenes. Therefore, there is an urgent need for a method that integrates visual and tactile information in a unified 3D space and combines it with the overall geometry of the target object for grasping stability assessment. Summary of the Invention
[0003] In view of this, this application provides a grasping stability assessment method and apparatus to solve the problem of insufficient accuracy in grasping stability assessment in the prior art.
[0004] Specifically, this application is implemented through the following technical solution:
[0005] The first aspect of this application provides a method for evaluating crawling stability, the method comprising:
[0006] The grasping system is calibrated to obtain the transformation relationships between multiple coordinate systems of the grasping system;
[0007] When the robot grasps and makes contact with the target object, it acquires the visual perception data of the target object and the tactile perception data corresponding to the contact.
[0008] Based on the transformation relationship, the visual perception data and the tactile perception data are respectively converted into corresponding three-dimensional geometric representations in the same coordinate system. The three-dimensional geometric representations include the visual three-dimensional geometric representations obtained from the visual perception data and the tactile three-dimensional geometric representations obtained from the tactile perception data.
[0009] Based on the aforementioned three-dimensional geometric representation, a multimodal three-dimensional representation of the target object is constructed;
[0010] Based on the multimodal 3D representation, a complete 3D geometric representation of the target object is obtained;
[0011] Based on the completed 3D geometric representation, the tactile 3D geometric representation, and the pre-trained grasping stability evaluation model, the grasping stability of the target object is evaluated, and the evaluation results are output.
[0012] A second aspect of this application provides a crawling stability evaluation system, the system comprising an acquisition module, a construction module, an inference module, and an evaluation module, wherein:
[0013] The acquisition module is used to calibrate the grasping system and obtain the transformation relationship between multiple coordinate systems of the grasping system;
[0014] The acquisition module is used to acquire visual perception data of the target object and tactile perception data corresponding to the contact when the robot grasps and makes contact with the target object.
[0015] The construction module is used to convert the visual perception data and the tactile perception data into corresponding three-dimensional geometric representations in the same coordinate system based on the conversion relationship. The three-dimensional geometric representations include a visual three-dimensional geometric representation obtained from the visual perception data and a tactile three-dimensional geometric representation obtained from the tactile perception data.
[0016] The construction module is used to construct a multimodal three-dimensional representation of the target object based on the three-dimensional geometric representation;
[0017] The reasoning module is used to obtain a complete three-dimensional geometric representation of the target object based on the multimodal three-dimensional representation;
[0018] The evaluation module is used to evaluate the grasping stability of the target object based on the completed three-dimensional geometric representation, the tactile three-dimensional geometric representation, and the pre-trained grasping stability evaluation model, and output the evaluation results.
[0019] The grasping stability assessment method and apparatus provided in this application achieve accurate assessment of grasping stability by introducing multimodal visual and tactile perception in a unified three-dimensional space during robot grasping. Specifically, this application first calibrates the grasping system and establishes the transformation relationship between multiple coordinate systems in the grasping system, so that the visual perception data and tactile perception data acquired when grasping contact can be uniformly converted into a three-dimensional geometric representation in the same coordinate system, thereby explicitly establishing the spatial correspondence between visual and tactile information. On this basis, a multimodal three-dimensional representation of the target object is constructed based on the visual and tactile three-dimensional geometric representations, and a completed three-dimensional geometric representation of the target object is obtained based on the multimodal three-dimensional representation. By introducing the completed three-dimensional geometric representation of the target object, the grasping stability assessment can comprehensively consider the overall geometric structural features of the target object in addition to local contact information. Furthermore, the completed three-dimensional geometric representation is input into a pre-trained grasping stability assessment model to assess the grasping stability of the target object and output the assessment result. In summary, this application integrates visual and tactile multimodal perception information in a unified three-dimensional space and introduces approximate reasoning results of the complete shape of the target object as the evaluation basis. This enables the grasping stability assessment to combine local contact geometry and overall object shape characteristics simultaneously. Thus, even when visual observation is obstructed or tactile perception is partial and incomplete, a reliable assessment of grasping stability can still be achieved, effectively improving the accuracy and robustness of the assessment results. Attached Figure Description
[0020] Figure 1 A flowchart of Embodiment 1 of the crawling stability evaluation method provided in this application;
[0021] Figure 2 A schematic diagram illustrating a recipient application scenario as shown in an exemplary embodiment of this application;
[0022] Figure 3 A visualization of a portion of the 3DA-VTG dataset as shown in an exemplary embodiment of this application;
[0023] Figure 4 This is a schematic diagram of the structure of Embodiment 1 of the gripping stability evaluation device provided in this application. Detailed Implementation
[0024] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application.
[0025] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used herein are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0026] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0027] The following specific embodiments are given to illustrate the technical solution of this application in detail.
[0028] Figure 1 This is a flowchart of an embodiment of the crawling stability evaluation method provided in this application. Please refer to... Figure 1 The method provided in this embodiment may include:
[0029] S101. Calibrate the grasping system and obtain the transformation relationship between multiple coordinate systems of the grasping system.
[0030] Specifically, this application employs a simplified object handover scenario, which can be defined as a collaborative action process in which a "giver" transfers an object to a "receiver." To focus on the recipient's perception and learning, and to avoid the combined complexity of the complete handover process, this application adopts the following simplified setup: the object is first stably held in the air by the giver to provide constraint on the object; the recipient performs a grasping action based on the calculated grasping posture and collects corresponding visual and tactile data; subsequently, as the giver releases the constraint, the handover result is recorded as success or failure based on whether the object slips.
[0031] It should be noted that the method of this application is not only applicable to object handover scenarios, but can also be extended to other grasping tasks such as desktop grasping. In these scenarios, the unified three-dimensional representation of visual and tactile information, as well as the completion of shape approximation, can still be used to effectively evaluate the grasping stability.
[0032] Figure 2 This is a schematic diagram illustrating a recipient application scenario as shown in an exemplary embodiment of this application. Please refer to... Figure 2The receiving device in the grasping system provided in this embodiment includes a robotic arm, an end effector, and various sensors to acquire visual and tactile perception information of a target object in a real environment. The robotic arm is a multi-degree-of-freedom industrial or research robotic arm, with a two-finger gripper-type end effector installed at its end for grasping the target object. Each fingertip of the end effector is equipped with a tactile sensor to collect tactile perception data when in contact with the target object during grasping. The tactile sensor can be a high-resolution tactile sensor based on a flexible medium, with its sensing surface covered by a flexible gel layer to generate perceptible deformation information during contact.
[0033] It should be noted that the end effector can be a two-finger gripper, a multi-finger end effector, or other types of end effectors, as long as it can reliably grip and manipulate the target object during the grasping process. Regardless of its specific structure, the fingertips of the end effector must be equipped with tactile sensors to collect tactile perception data when in contact with the target object, thereby obtaining local contact information. This application does not impose any limitations; this embodiment uses a two-finger gripper end effector as an example for illustration.
[0034] Please continue to refer to Figure 2 The visual perception component of this grasping system includes a vision sensor, specifically an RGB-D camera, which is fixedly mounted outside the robotic arm or within the workspace in an "eye outside the hand" configuration. This camera collects visual perception data of the target object before or during grasping, preventing the end effector structure from obstructing visual observation. Through this system arrangement, the grasping system can simultaneously acquire global visual information and local tactile contact information of the target object during the grasping process.
[0035] It should be noted that the visual sensor can be a visual sensor capable of acquiring or recovering three-dimensional information of the target object's shape, including but not limited to RGB-D cameras, binocular cameras, structured light cameras, or RGB cameras combining multi-view imaging and three-dimensional reconstruction algorithms. No limitation is made in this application; this application uses an RGB-D camera as an example for illustration. For the specific implementation principles and methods of using an RGB camera, please refer to the descriptions in related technologies, which will not be repeated here.
[0036] Specifically, calibration tools or plates are placed on components such as the robotic arm, end effector, vision sensor, and tactile sensor to collect calibration data in each coordinate system. The pose relationships between these coordinate systems are then calculated using appropriate calibration methods to establish a unified coordinate reference frame for the grasping system. The coordinate transformation relationships for the vision sensor include its extrinsic and intrinsic parameter matrices; the coordinate transformation relationships for the tactile sensor include its extrinsic and projection matrices.
[0037] Optionally, in one possible implementation, the calibration of the grasping system involves obtaining the transformation relationships between multiple coordinate systems of the grasping system; including:
[0038] (1) Collect calibration data from different observation perspectives.
[0039] Specifically, the calibration data includes visual calibration data and tactile calibration data. The visual calibration data is used to establish the imaging model of the visual sensor and its spatial relationship with the target reference coordinate system, and the tactile calibration data is used to establish the imaging model of each tactile sensor and its spatial relationship with the target reference coordinate system.
[0040] Specifically, before calibration, the grasping center point corresponding to the current grasping task is first determined. The grasping center point is a spatial reference point corresponding to the current grasping pose, used to characterize the position of the grasping action on the calibration object. It is not limited to the geometric center or centroid of the calibration object, but is determined by the specific requirements of the grasping task. During the grasping task generation process, the system determines the grasping pose to be executed based on the geometric information of the calibration object, prior grasping strategies, or manually set grasping requirements. This grasping pose corresponds to a grasping center pixel position in the visual image. This pixel position represents the projection position of the grasping pose in the image coordinate system and serves as the two-dimensional image reference point for the current grasping task. After obtaining the grasping center pixel position, the system combines the depth image synchronously acquired by the RGB-D camera and uses the imaging model of the visual sensor to back-project the grasping center pixel position onto the target reference coordinate system, obtaining the corresponding three-dimensional spatial point. This three-dimensional spatial point is defined as the grasping center point and serves as a unified spatial reference for subsequent hemispherical view sampling, camera pose construction, and visual calibration calculations.
[0041] Specifically, taking the center point of the capture as the center of the unit sphere, select on the surface of the sphere... The upper hemisphere region is used as the effective sampling area, where z represents the component of the unit vector of the visual sensor's observation viewpoint direction on the z-axis of the target reference coordinate system, i.e., the projection component of the camera's optical axis direction on this coordinate axis. By setting a lower threshold for the z-component, the angle between the camera's line of sight and the normal of the target area can be constrained, eliminating low-angle observation directions that are close to horizontal, thereby reducing occlusion caused by the gripper or the object itself, while retaining a certain tilt angle to enhance the diversity of spatial observation. The value range of z is an empirical threshold range predetermined based on the grasping scene and occlusion characteristics, which can be determined according to the actual scene requirements, and is not limited in this embodiment.
[0042] Specifically, for the same candidate capture pose, within the aforementioned effective sampling area, the system performs uniform sampling of the camera's viewpoint direction. This uniform sampling refers to parameterizing the azimuth and pitch components of the viewpoint direction while satisfying the z-constant, ensuring that the generated multiple viewpoint directions are approximately uniformly distributed in the sense of spherical area, thereby avoiding the viewpoint from concentrating in a localized region.
[0043] Specifically, during calibration data acquisition, the system first controls the end effector of the grasping system to move to a predetermined calibration posture, ensuring stable contact between multiple tactile sensors mounted on the fingertip and the surface of the calibration object. While maintaining this contact state, the system, from the corresponding viewing angle, uses an RGB-D camera to acquire color and depth images of the calibration object, while simultaneously using each tactile sensor to acquire corresponding tactile image data. The color and depth images together constitute the visual calibration data, and the tactile image data constitutes the tactile calibration data.
[0044] (2) Based on the predetermined imaging model and the calibration data, calculate the transformation relationship between the image coordinate system and the visual sensor coordinate system, and the transformation relationship between the image coordinate system and each tactile sensor coordinate system, to obtain the visual intrinsic parameter matrix and multiple tactile intrinsic parameter matrices.
[0045] Specifically, both the visual sensor and each tactile sensor employ an image-based perception method, and their imaging process can be equivalent to a pinhole imaging model. Under the corresponding calibration posture, the system acquires the visual two-dimensional image collected by the visual sensor and the tactile two-dimensional image collected by each tactile sensor. Based on the calibration features extracted from the two-dimensional images, a geometric constraint relationship is established between the image coordinate system and the corresponding sensor coordinate system, thereby calculating the corresponding intrinsic parameter matrix.
[0046] Specifically, for a vision sensor, at each sampling viewpoint corresponding to the camera pose, the RGB-D camera acquires a two-dimensional image and a depth image containing the calibration object. The two-dimensional image is used to extract two-dimensional pixel feature points on the calibration object, and a mapping relationship between pixel coordinates and three-dimensional points in the vision sensor coordinate system is established based on a pinhole imaging model. Preferably, in one possible implementation, the depth image is used to back-project the two-dimensional pixel feature points to their three-dimensional spatial positions in the vision sensor coordinate system, and to verify the geometric consistency of the feature points in space, thereby eliminating invalid samples caused by depth anomalies, overhanging points, or reflection noise.
[0047] Furthermore, both the visual and tactile sensors employ a pinhole imaging model, where the pixel coordinates and the three-dimensional points in each sensor coordinate system satisfy the following relationship:
[0048]
[0049] in, The pixel coordinates of the feature points in the image coordinate system The coordinates of this feature point in each sensor coordinate system are the three-dimensional coordinates. Here are the intrinsic parameter matrices for each sensor. For a visual sensor, the scale factor is... Equal to the depth of a 3D point along the optical axis in the visual sensor coordinate system In this embodiment, the depth image is obtained from an RGB-D camera. For the tactile sensor, since its imaging object is located on a structurally fixed or constrained gel contact plane, its scale factor remains constant during calibration or can be uniquely determined by structural parameters. Therefore, it is equivalently incorporated into the intrinsic parameter matrix of the tactile sensor during the model building stage and is no longer solved separately during modeling. Intrinsic parameter matrices of each sensor. Represented as:
[0050]
[0051] in, and These are the equivalent focal lengths of the camera in the horizontal and vertical directions, respectively. and This represents the pixel position of the camera principal point in the image coordinate system. The intrinsic parameter matrix is used for spatial transformation between the various sensor coordinate systems and the image coordinate system.
[0052] (3) Based on the calibration data and the visual intrinsic parameter matrix, calculate the transformation relationship between the visual sensor coordinate system and the target reference coordinate system to obtain the first extrinsic parameter matrix.
[0053] Specifically, the target reference coordinate system has its origin at the grasping center point, and its coordinate axes are aligned with the robot's base coordinate system or the task-defined coordinate system, thus ensuring that the grasping center's position in space is aligned with the robot's motion reference coordinates. Furthermore, for each sampling viewpoint, the system extracts the three-dimensional coordinates of the calibration object's feature points in the world coordinate system based on the known three-dimensional geometry of the calibration object. And combined with the corresponding image pixel coordinates Establish the following spatial mapping relationship:
[0054]
[0055] in, Let be the rotation matrix of the vision sensor relative to the target reference coordinate system. Let be the translation vector. By solving the observation equations from multiple viewpoints, the system employs the PnP algorithm or a pose estimation method based on nonlinear minimization to calculate the optimal rotation matrix. With translation vector Furthermore, the transformation relationship between the visual sensor coordinate system and the target reference coordinate system is constructed, and the first extrinsic parameter matrix is represented as follows:
[0056]
[0057] in, This is the first extrinsic parameter matrix, used to realize the spatial transformation between the visual sensor coordinate system and the target reference coordinate system.
[0058] (4) Based on the installation pose information of each of the tactile sensors, calculate the transformation relationship between the coordinate system of each tactile sensor and the target reference coordinate system to obtain multiple second external parameter matrices.
[0059] Specifically, the tactile sensing component consists of image-based tactile sensors mounted on the robot's two fingertips. Each tactile sensor includes an elastic gel layer and an internal imaging unit. Two tactile sensors, numbered i∈{0,1}, are independently calibrated. During calibration, the system obtains the intrinsic extrinsic parameter relationship between the coordinate system of the internal imaging unit of each tactile sensor and its gel surface coordinate system through pre-calibrated or manufacturer-provided mounting geometry parameters. This extrinsic parameter relationship is expressed in homogeneous matrix form. The extrinsic parameter matrix describes the rigid body transformation relationship between the imaging unit coordinate system of the tactile sensor and the gel contact coordinate system.
[0060] Furthermore, by utilizing the robot's forward kinematics model or other methods capable of obtaining the end effector's pose in the target reference coordinate system, and combining joint angle information, end effector pose, and the fingertip fixation offset of the haptic sensor, the real-time pose matrix of each haptic sensor in the target reference coordinate system is calculated. By chaining matrix multiplication, a second extrinsic parameter matrix of the tactile sensor relative to the target reference coordinate system is constructed:
[0061]
[0062] in, Indicates the first The spatial pose of a tactile sensor coordinate system and a target reference coordinate system is used to describe the coordinate transformation relationship between the two.
[0063] (5) Based on the visual intrinsic parameter matrix and the first extrinsic parameter matrix, and each of the tactile intrinsic parameter matrices and the corresponding second extrinsic parameter matrices, establish the transformation relationship between multiple coordinate systems of the grasping system.
[0064] Specifically, a back-projection model is used to uniformly map the effective pixels in the two-dimensional perceptual image to a three-dimensional spatial representation in the target reference coordinate system. For any perceptual modality, the back-projection process can be expressed as:
[0065]
[0066] in, For pixels The corresponding depth value, This is the intrinsic parameter matrix of the vision sensor. Let be the first extrinsic parameter matrix of the vision sensor relative to the target reference coordinate system. This is a set of pixels with effective depth. Through the above back-projection process, two-dimensional perceptual information can be uniformly converted into a three-dimensional spatial point set in the target reference coordinate system.
[0067] Specifically, for the visual sensor, the system establishes a spatial mapping relationship between the visual sensor image coordinate system and the target reference coordinate system based on the intrinsic parameter matrix and the first extrinsic parameter matrix of the visual sensor. Based on this, and combined with the depth image acquired by the RGB-D camera, the system performs back-projection operations on the effective pixels on the visual image plane to generate a three-dimensional point cloud in the target reference coordinate system. The calculation method of the three-dimensional point cloud can be expressed as follows:
[0068]
[0069] in, This is the intrinsic parameter matrix of the vision sensor. Let be the first extrinsic parameter matrix of the vision sensor relative to the target reference coordinate system. These are the pixel coordinates in the visual image. For pixels Depth value at that location, This is a set of pixels that provide effective depth information. Through the above processing, the system converts the two-dimensional visual perception information into a three-dimensional visual point cloud representation aligned with the grasping center.
[0070] Specifically, for tactile sensors, the system targets each tactile sensor... Based on its corresponding tactile intrinsic parameter matrix and the second extrinsic parameter matrix A spatial mapping relationship is established between the tactile sensor image coordinate system and the target reference coordinate system. Based on this, the system performs a back-projection operation on the effective pixels on the tactile imaging plane, converting the two-dimensional tactile image information into three-dimensional contact points in the target reference coordinate system, thereby realizing the three-dimensional representation of tactile perception data in a unified reference coordinate system. The back-projection process can be implemented based on the following relationship:
[0071]
[0072] in, These are the pixel coordinates in the tactile image. This is the scale factor for tactile imaging.
[0073] In summary, the system completes the spatial alignment between the coordinate systems of each sensor and the target reference coordinate system, and realizes the fusion of visual and tactile perception data in a unified coordinate system, providing a foundation for subsequent grasping state estimation, multimodal spatial reconstruction and grasping strategy optimization.
[0074] S102. When the robot grasps and makes contact with the target object, it acquires the visual perception data of the target object and the tactile perception data corresponding to the contact.
[0075] Specifically, the robot approaches the target object according to a pre-planned grasping posture and performs the grasping action until the gripper makes contact with the target object. During the grasping process, visual perception data of the target object is collected in real time by vision sensors installed on the grasping system. This visual perception data includes two-dimensional images and depth images acquired by the vision sensors, used to characterize the spatial geometric information of the target object at the moment of grasping. Simultaneously, the system controls multiple tactile sensors installed on the gripper tips to collect tactile perception data when in contact with the surface of the target object. This tactile perception data includes tactile images acquired by tactile imaging units, used to characterize the spatial mapping relationship between the tactile sensor imaging units and the contact surface of the target object.
[0076] Specifically, both visual perception data and tactile perception data correspond to the same grasping action, ensuring that the collected data are consistent in time and space, thereby enabling accurate multimodal data alignment in the subsequent coordinate unification process based on transformation relationships.
[0077] S103. Based on the transformation relationship, the visual perception data and the tactile perception data are respectively converted into corresponding three-dimensional geometric representations in the same coordinate system, wherein the three-dimensional geometric representations include the visual three-dimensional geometric representations obtained from the visual perception data and the tactile three-dimensional geometric representations obtained from the tactile perception data.
[0078] Specifically, based on the transformation relationships between multiple coordinate systems obtained in the calibration step, the system transforms visual perception data and tactile perception data into a unified target reference coordinate system to construct a corresponding three-dimensional geometric representation. The three-dimensional geometric representation is used to characterize the geometric morphology information of the target object or tactile contact area in three-dimensional space. The three-dimensional geometric representation can be implemented in various data forms, including but not limited to three-dimensional point cloud representation, three-dimensional mesh model, voxel representation, implicit geometric function representation, or feature field representation. Specifically, the visual three-dimensional geometric representation is obtained from the visual perception data through coordinate transformation and three-dimensional reconstruction; the tactile three-dimensional geometric representation is obtained from the tactile perception data through coordinate transformation and geometric reconstruction. In this embodiment, the three-dimensional geometric representation preferably adopts a point cloud form, that is, it describes the geometric structure of the target object or contact area through a set of three-dimensional spatial coordinate points and their corresponding geometric or semantic attributes. The visual three-dimensional geometric representation, tactile three-dimensional geometric representation, multimodal three-dimensional representation, and completed three-dimensional geometric representation mentioned in subsequent steps are all specific forms of the above-mentioned three-dimensional geometric representation under different processing stages or different information sources.
[0079] Optionally, in one possible implementation, the conversion of visually perceived data into a corresponding three-dimensional geometric representation in the same coordinate system includes:
[0080] (1) Obtain a predetermined reference pixel point, and based on the reference pixel point and the visual perception data, obtain the two-dimensional image region of the target object.
[0081] Specifically, to construct a visual 3D geometric representation of the target object, it is necessary to first extract the 2D image region of the target object from the visual perception data to remove background information and ensure the accuracy of the spatial correspondence of the subsequent 3D point cloud. To this end, the system performs segmentation processing on the visual perception data to extract the 2D image region of the target object. The segmentation processing can employ one or more of the following methods: cue-based segmentation, object detection-based segmentation, etc., without limitation. Optionally, this embodiment employs a grasping center-guided visual segmentation method to address challenges posed by gripper occlusion or complex visual backgrounds that may occur during the grasping process.
[0082] Specifically, the system acquires a predetermined reference pixel. This reference pixel can be the center pixel or a manually specified positive cue pixel. In this embodiment, the center pixel is used as the positive cue pixel. For details on the specific implementation principle and method of the center pixel acquisition, please refer to the relevant descriptions, which will not be repeated here.
[0083] Specifically, the center pixel of the grasped object is used as a positive cue input to a visual segmentation algorithm, such as the SAM model, to indicate the approximate location of the target object in the image, thereby generating a corresponding segmentation mask. When the center pixel is visible, this strategy can directly generate a complete and reliable 2D image region of the target object; when the center pixel is occluded by the grasper or is not visible, additional positive cue pixels can be manually provided to ensure the integrity and accuracy of the segmentation mask. Through the above processing, the 2D image region obtained by the system not only removes background information but also retains the visible part of the target object, providing a reliable mask foundation for subsequent construction of a 3D geometric representation based on intrinsic and extrinsic parameter matrices.
[0084] (2) Based on the two-dimensional image region, the visual perception data and the visual intrinsic parameter matrix, the visual three-dimensional geometric representation in the visual sensor coordinate system is obtained.
[0085] Specifically, in this embodiment, the visual sensor is an RGB-D camera. The depth value of each pixel is directly obtained using depth image information, and its three-dimensional coordinates are calculated by combining the intrinsic parameter matrix. For example, any pixel in a two-dimensional image... For example, its corresponding depth value is denoted as By combining the focal length and principal point parameters contained in the intrinsic parameter matrix of the vision sensor, the three-dimensional coordinates of the pixel in the vision sensor coordinate system can be calculated based on the imaging geometry. Specifically:
[0086]
[0087] in, This is the intrinsic parameter matrix of the vision sensor. By performing the above back-projection operation on each pixel within the two-dimensional image region of the target object, a visual three-dimensional geometric representation in the vision sensor coordinate system can be obtained. This visual three-dimensional geometric representation is expressed in the form of a three-dimensional point cloud.
[0088] It should be noted that if depth information is missing, such as with an RGB camera as the visual sensor, meaning only a two-dimensional image can be obtained, then a depth estimation model or other 3D reconstruction methods can be used to generate the corresponding depth, thus completing the conversion from a 2D image to a 3D point cloud. For the specific implementation principles and directions of converting a 2D image into a visual 3D geometric representation in the visual sensor's coordinate system, please refer to the relevant descriptions, which will not be repeated here.
[0089] (3) Based on the visual three-dimensional geometric representation in the visual sensor coordinate system and the first extrinsic matrix, the visual three-dimensional geometric representation in the target reference coordinate system is obtained.
[0090] Specifically, through the rotation matrix of the first extrinsic parameter matrix Translation vector The visual 3D geometric representation in the visual sensor coordinate system is mapped to a unified target reference coordinate system. For any 3D point in the visual 3D geometric representation... Its corresponding three-dimensional point in the target reference coordinate system It can be obtained through the following coordinate transformation relationship:
[0091]
[0092] in, This is a three-dimensional coordinate vector in the visual sensor coordinate system. This represents a three-dimensional coordinate vector in the target reference coordinate system. By performing the above coordinate transformation on each three-dimensional point in the visual three-dimensional geometric representation, the entire visual three-dimensional geometric representation is mapped to a unified target reference coordinate system.
[0093] Optionally, in one possible implementation, converting the tactile perception data into a corresponding three-dimensional geometric representation in the same coordinate system includes:
[0094] (1) Calculate the depth information of the tactile perception data based on the tactile perception data.
[0095] Specifically, the tactile perception data is acquired by a tactile sensor mounted on the robot's end effector during the grasping contact process. In this embodiment, the GelSight Mini tactile sensor is used, which can acquire high-resolution RGB images reflecting the deformation of the tactile surface. Since this tactile sensor does not have the ability to directly output depth information in a real environment, the system needs to calculate the depth information corresponding to the tactile contact area based on the tactile perception data in order to recover the local three-dimensional geometry.
[0096] Specifically, to obtain the depth information corresponding to the tactile perception data, the system performs depth restoration processing on the tactile RGB image to estimate the local three-dimensional geometry of the tactile contact area. In this embodiment, a depth estimation model is used for depth restoration, thereby converting the tactile perception data containing only two-dimensional color information into a corresponding local depth map. In addition, the depth restoration processing can also be implemented through physically constrained inversion methods, learned regression models, or a combination of these methods; this embodiment does not limit the specific implementation.
[0097] Specifically, the depth estimation model can be constructed with reference to the model structure in the NeuralFeels framework, and trained or parameter-adapted based on a dataset containing tactile RGB images and corresponding depth annotations to adapt to the imaging characteristics of the GelSight Mini tactile sensor in terms of gel deformation patterns and changes in lighting conditions. The parameter adaptation process may include fine-tuning an existing model or retraining model parameters based on the target application scenario; the specific implementation method does not constitute a limitation on the scope of protection of this application. Preferably, the tactile depth estimation model can be combined with simulation-real-domain randomization or conventional data augmentation methods for parameter adaptation during actual deployment to improve its prediction stability in real grasping scenarios. Through the above processing, the system can obtain depth information corresponding to the tactile contact area based on tactile perception data.
[0098] (2) Based on the tactile perception data, the depth information and the tactile intrinsic parameter matrix, the tactile three-dimensional geometric representation in the coordinate system of the tactile sensor is obtained.
[0099] Specifically, combined with the tactile intrinsic parameter matrix The depth value of each pixel in the 2D tactile image is back-projected onto the coordinate system of the tactile sensor to obtain the corresponding 3D coordinates. For example, for any pixel in the tactile image... Its corresponding depth is Then the three-dimensional coordinates in the tactile sensor coordinate system It can be calculated using the following formula:
[0100]
[0101] in, The tactile intrinsic parameter matrix is obtained by performing the above back projection on each pixel in the tactile image, which is usually expressed in the form of a 3D point cloud.
[0102] (3) Based on the tactile three-dimensional geometric representation in the coordinate system of the tactile sensor, and combined with the second extrinsic matrix, the tactile three-dimensional geometric representation in the target reference coordinate system is obtained.
[0103] Specifically, through the rotation matrix of the second extrinsic parameter matrix. Translation vector The tactile 3D geometric representation in the tactile sensor coordinate system is mapped to a unified target reference coordinate system. For any 3D point in the tactile 3D geometric representation... Its corresponding three-dimensional point in the target reference coordinate system This can be obtained through the following coordinate transformation relationship:
[0104]
[0105] in, and These originate from the second extrinsic parameter matrix. The rotation matrix and translation vector are used to perform the above coordinate transformation on all three-dimensional points in the tactile three-dimensional geometric representation, thereby mapping the entire tactile three-dimensional geometric representation to a unified target reference coordinate system, providing an accurate spatial basis for the fusion with the visual three-dimensional geometric representation and the construction of multimodal three-dimensional representation.
[0106] In summary, the system has transformed visual perception data and tactile perception data into a unified target reference coordinate system, achieving alignment of multimodal information in the same three-dimensional space, and providing an accurate spatial basis for subsequent construction of multimodal three-dimensional representations and evaluation of grasping stability.
[0107] S104. Based on the three-dimensional geometric representation, construct a multimodal three-dimensional representation of the target object.
[0108] Specifically, the visual 3D geometric representation and the tactile 3D geometric representation are fused according to their spatial correspondence. The visual 3D geometric representation provides the overall geometric outline of the target object, while the tactile 3D geometric representation supplements the local details and deformation information of the contact area. Through multimodal fusion, a complete 3D geometric representation is formed. This multimodal 3D representation not only has complete object shape information but also preserves the local features of the contact surface.
[0109] Optionally, in one possible implementation, constructing a multimodal 3D representation of the target object based on the 3D geometric representation includes:
[0110] (1) Extract the features of the visual three-dimensional geometric representation and construct a visual feature representation; extract the features of the tactile three-dimensional geometric representation and construct a tactile feature representation.
[0111] Specifically, in order to reduce the computational load of subsequent processing and ensure the preservation of key spatial features, the system generates a visual 3D geometric representation. Random downsampling to a fixed scale yields a visual feature representation. Besides random sampling, voxelization or feature encoding can also achieve the same effect; this is not a specific limitation here. The random downsampling function... Generate visual feature representations:
[0112]
[0113] in, Represented by visual features, The visual feature representation, while ensuring key spatial features, reduces the size of the point cloud and optimizes the computational load, providing a unified input format for subsequent fusion processing.
[0114] Specifically, the system uniformly downsamples the tactile 3D geometric representation to a fixed size, and then processes it through a function. Generate tactile feature representation:
[0115]
[0116] in, Represented by tactile features, For the tactile three-dimensional geometric representation, the downsampling process is used to constrain the scale of the tactile three-dimensional geometric representation, so that the tactile point cloud participates in subsequent feature encoding and multimodal fusion processing with a fixed or controlled data scale.
[0117] (2) The visual feature representation and the tactile feature representation are fused to obtain the multimodal three-dimensional representation.
[0118] Specifically, based on the spatial coordinates of the visual and tactile feature representations, the local contact point clouds in the tactile feature representation are embedded into the corresponding spatial positions in the visual feature representation to form a unified three-dimensional geometric representation. Weighted fusion strategies or neighborhood search algorithms can be introduced during the fusion process; no specific limitations are imposed here, to ensure that the local details of the tactile point cloud are fully reflected in the overall three-dimensional representation, while maintaining the overall shape and contour of the visual point cloud without distortion.
[0119] Optionally, in one possible implementation, after constructing the tactile feature representation based on the tactile three-dimensional geometric representation, the method further includes:
[0120] Based on the tactile perception data, contact identification information is added to the points in the tactile feature representation.
[0121] Specifically, based on the spatial relationship of the gel layer region in the tactile perception data, it is determined whether the point belongs to the area in contact with the target object, thereby generating the corresponding contact marker. Through the above processing, the tactile feature representation not only contains the geometric information of the local contact surface, but also carries the contact state information. Accordingly, the system generates two sub-representations according to different uses: (1) tactile feature representation (retaining the gel-covered area), which contains complete original tactile gel points, representing the contact and non-contact parts of the object on the tactile surface, used for subsequent grasping stability assessment; (2) contact point sub-representation (removing the gel area), which only retains the points that actually contact the target object. The number of points in this part changes dynamically with the size of the contact area, and is used to splice with the visual feature representation of the target object's shape.
[0122] S105. Based on the multimodal three-dimensional representation, reason about the missing region of the target object to obtain the complete three-dimensional geometric representation of the target object.
[0123] Optionally, in one possible implementation, obtaining the completed 3D geometric representation of the target object based on the multimodal 3D representation includes:
[0124] (1) Based on the multimodal three-dimensional representation, a modal-free shape representation is obtained.
[0125] Specifically, the multimodal 3D representation includes a visual 3D geometric representation and a tactile 3D geometric representation located in a unified target reference coordinate system. In this embodiment, this is specifically a visual point cloud and a tactile point cloud, where each point contains corresponding 3D spatial coordinate information. The system performs unified processing on the visual point cloud and the tactile point cloud to construct a modal-free shape representation for characterizing the geometric structure of the observed area of the target object. Specifically, the processing does not distinguish the modal origin of the points; it only processes the multimodal point cloud based on the geometric distribution characteristics of the points in 3D space, thereby eliminating the influence of modal differences on the subsequent geometric modeling process and retaining only the 3D geometric coordinate information of the points to characterize part of the geometric shape of the target object.
[0126] Preferably, in one possible implementation, spatial equalization sampling is performed on the multimodal 3D representation before obtaining the completed 3D geometric representation.
[0127] Specifically, to meet the requirement of a fixed input scale for the point cloud completion model, while maintaining a balanced contribution of visual, tactile, and missing region features in space, the system performs spatial equalization sampling on the multimodal point cloud. Specifically, the spatial distribution range of the multimodal point cloud is estimated under a unified target reference coordinate system. This spatial distribution range characterizes the coverage scale or geometric expansion of the point cloud in three-dimensional space. In one embodiment, the spatial distribution range can be obtained through axis-aligned bounding boxes, minimum bounding spheres, or other geometric scale measures. In this embodiment, axis-aligned bounding boxes are preferably used to obtain the geometric scale information of the point cloud along each coordinate axis. Based on the spatial distribution range estimation results, the system performs a resampling operation on the multimodal point cloud, ensuring that the resampled point set maintains a balanced overall spatial distribution, thereby obtaining a modal-free shape representation. The modal-free shape representation contains only three-dimensional geometric coordinate information, which is used as input for subsequent steps to complete the three-dimensional geometric representation generation.
[0128] (2) Based on the modal shape representation, and combined with the relative positional relationship between the observed region and the missing region in three-dimensional space, geometric reasoning is performed on the missing region to generate complete shape features.
[0129] Specifically, after completing the spatial equalization sampling, the system uses the obtained modal-free shape representation as input data and feeds it into the point cloud completion model for processing. The point cloud completion model performs overall modeling of the spatial distribution characteristics of the observed regions in the input point cloud, performs geometric reasoning on the unobserved missing regions of the target object, and thus generates a completion result containing the geometric information of the missing regions. This process can be uniformly represented as follows:
[0130]
[0131] in, This is a modal-free shape representation; in this embodiment, it is a modal-free shape point cloud. The geometric inference mapping function corresponding to the point cloud completion model; This is the completed 3D geometric representation of the target object, which, in addition to the original observed area, further includes the set of geometric points of the missing region obtained by model inference.
[0132] Optionally, in one possible implementation, the point cloud completion model can be AdaPoinTr, and its corresponding geometric reasoning process can be represented as:
[0133]
[0134] in, This refers to the geometric inference function corresponding to the point cloud completion model, used to perform geometric inference on unobserved missing regions in the target object based on the input modal-free shape point cloud. It should be noted that this embodiment does not limit the specific network structure or internal implementation of the point cloud completion model; any model capable of geometrically completing missing regions based on locally observed point clouds can be used as the model. The implementation form is as follows. In summary, the completed 3D geometric representation not only includes the geometric information observed by vision and touch, but also completes the missing parts caused by viewpoint limitations or gripper occlusion through reasoning, providing complete geometric input for subsequent grasping stability assessment.
[0135] S106. Based on the completed three-dimensional geometric representation, the tactile three-dimensional geometric representation, and the pre-trained grasping stability evaluation model, evaluate the grasping stability of the target object and output the evaluation result.
[0136] Specifically, the completed 3D geometric representation is input into the grasping stability assessment model for inference calculations. This model comprehensively analyzes the force balance, contact distribution, and geometric constraints of the target object under current grasping conditions, thereby outputting an assessment result characterizing the degree of grasping stability. This assessment result can be used to indicate whether the current grasping meets stability requirements, or as a basis for subsequent grasping adjustments and decisions.
[0137] Optionally, in one possible implementation, the training method of the pre-trained grasping stability evaluation model includes:
[0138] (1) Obtain training samples of at least one training object under different grasping postures.
[0139] Specifically, to collect high-density tactile data for each object, a large number of potential grasping postures need to be covered. This embodiment utilizes existing object-level metadata that provides high-density grasping postures, obtaining training sample data covering diverse grasping postures without the need for additional grasping candidate generation. This metadata can originate from public datasets (such as GraspNet-1 Billion), self-collected grasping simulation data, or other information sources that provide 6-DOF grasping postures. By calling the corresponding interfaces or generation processes, training samples covering a large number of grasping postures can be obtained efficiently.
[0140] (2) The training samples are screened through offline simulation experiments to obtain effective training samples, which contain multiple training representations.
[0141] Specifically, to ensure the effectiveness of training samples and the reliability of tactile data, the system filters training sample data through offline simulation. For each grasping posture, the system evaluates whether the gripper makes unintended contact with the object before the grasping action begins using simulation or other analysis methods, and detects whether non-fingerpoint contact or failure of the gripper to grasp the object occurs during the grasping process. Specifically, this embodiment uses the PyBullet physics simulation environment to perform collision detection and grasping effectiveness evaluation for each grasping posture. By eliminating grasping postures with initial collisions, accidental collisions, and grasping failures, only training representations that can provide effective tactile feedback during the grasping action are retained.
[0142] (3) Obtain the training perception data corresponding to the effective training sample, and convert the training perception data into a three-dimensional training representation in the same coordinate system based on the transformation relationship.
[0143] Specifically, for each valid training sample, the robot performs a grasping action according to the corresponding grasping posture, and collects training perception data corresponding to that grasping action during the grasping process. The training perception data includes visual training data acquired by the vision sensor on the grasping system and the corresponding tactile training data. Both the visual training data and the tactile training data correspond to the same grasping action. For the specific implementation principle and method of collecting training perception data, please refer to the relevant description, which will not be repeated here.
[0144] Specifically, based on pre-obtained coordinate system transformation relationships, the system performs spatial transformations on both visual and tactile training data, unifying them under the same target reference coordinate system to obtain corresponding 3D training representations. These 3D training representations include a visual 3D training representation obtained from the visual training data and a tactile 3D training representation obtained from the tactile training data, both having a consistent spatial reference under the same target reference coordinate system. For the specific implementation principles and methods of the transformation process, please refer to the relevant descriptions, which will not be repeated here.
[0145] (4) Based on the three-dimensional training representation, construct the multimodal training representation of the training object.
[0146] Specifically, the visual 3D training representation and the tactile 3D training representation are combined to form a unified multimodal training representation containing multi-source geometric information, which is used to characterize the overall geometric state and contact characteristics of the training object under the corresponding grasping posture. The construction of the multimodal training representation is the same as the construction of the multimodal 3D representation. For the specific implementation principle and method, please refer to the relevant description, which will not be repeated here.
[0147] (5) Obtain the stability results of each training representation in the grasping experiment, and generate the corresponding stability label for each multimodal training representation.
[0148] Specifically, the stability result of each training representation in the grasping experiment is obtained, and a corresponding stability label is generated for each multimodal training representation. Specifically, in the simulation environment, the grasping action corresponding to each training sample is executed, and the constraints are removed after the grasping is completed. Gravity is enabled and the simulation is continued for a certain period of time (2 seconds in this embodiment) to determine whether the object falls. Based on the state of the object after the grasping is completed, the corresponding training representation is labeled as stable or unstable, and a stability label is generated for that training representation. , where 1 indicates stable crawling and 0 indicates unstable crawling.
[0149] (6) Based on the multimodal training representation and the stability label, train the grasping stability evaluation model.
[0150] Specifically, during training, each multimodal training representation and its corresponding stability label are input into the grasping stability evaluation model. Supervised learning is used to train the model. A standard loss function, such as cross-entropy loss or mean squared error loss, is used to calculate the error between the model output and the corresponding stability label. The model parameters are then iteratively updated based on this error, enabling the grasping stability evaluation model to learn the mapping relationship from multimodal geometric representations to grasping stability. In this embodiment, a binary cross-entropy loss function is used for training.
[0151]
[0152] in, For batch size, To capture stability tags, This represents the stability score predicted by the model. After model training is complete, the trained grasping stability assessment model can be used for subsequent grasping stability assessments.
[0153] Optionally, in one possible implementation, before training the grasping stability evaluation model, a completed 3D geometric representation of the training object is calculated based on the multimodal training representation, and the grasping stability evaluation model is trained using the completed 3D geometric representation of the training object. For the specific implementation principle and method of the completed 3D geometric representation, please refer to the relevant description, which will not be repeated here.
[0154] Optionally, in one possible implementation, before obtaining the training sensing data corresponding to the valid training samples, the method further includes:
[0155] Spatial transformation is performed on the grasping posture of each sample in the effective training samples to expand the grasping posture configuration in the effective training samples.
[0156] Specifically, to increase the diversity of grasping postures relative to the direction of gravity and spatial orientation, a spatial transformation is performed on the grasping postures corresponding to each valid training sample. This spatial transformation includes applying a spatial pose transformation to the grasping posture while maintaining the relative geometric relationship between the gripper's pre-grasping posture and the target object, thereby changing the spatial configuration of the grasping relative to the reference coordinate system or the direction of gravity. In one optional implementation, by applying a posture adjustment around a horizontal axis to the gripper posture, the gripper is made to approach in an approximately horizontal state during the pre-grasping stage, thus covering grasping configurations under different gravity reference directions. After completing the above pose adjustment, a planar rotation transformation can also be applied to the gripper posture around an axis perpendicular to the gripper's approach direction to generate multiple grasping orientation configurations. By recording the transformed gripper pre-grasping posture and the corresponding object posture as a set of associated poses, each valid training sample can correspond to multiple sets of grasping postures with different configurations, thereby expanding the grasping posture configurations in the valid training samples and improving the coverage and diversity of the training data in spatial distribution.
[0157] Specifically, to address the perceptual sparsity and alignment error issues in visual-tactile grasping, this application constructs a large-scale visual-tactile grasping dataset, 3DA-VTG (3D Alignment Visual-Tactile Grasping Dataset), based on the aforementioned method. The 3DA-VTG dataset uses 88 object classes from the GraspNet-1 Billion dataset as its base objects. These objects include: (1) everyday household objects from the YCB dataset (YCB: Everyday Household Baseline Objects Dataset); (2) novel 3D printed objects from the DexNet 2.0 dataset (DexNet 2.0: 3D Printed Objects Dataset); and (3) other object categories further selected by the GraspNet-1 Billion dataset creators to enhance the diversity and complexity of geometric shapes.
[0158] Specifically, during data acquisition, all objects are first subjected to mesh simplification and convex decomposition using V-HACD (Volumetric Hierarchical Approximate Convex Decomposition algorithm) to improve collision detection efficiency in the simulation environment. Subsequently, a FrankaResearch 3 three-finger gripper model is constructed in the simulation environment, with GelSight Mini tactile sensor models configured at the ends of the two fingers to simulate the acquisition of visual and tactile signals during real-world grasping. To simulate the unavoidable structural uncertainties in real-world assembly, center offset errors of 0.25mm and 0.5mm are artificially introduced in the URDF model of the tactile sensor along the adhesive surface plane to simulate the impact of tactile sensor installation errors on visual-tactile alignment accuracy. Compared to random noise, this type of assembly offset error more closely resembles real-world sensing scenarios, thereby enhancing the realism and research value of the dataset.
[0159] Specifically, under different grasping postures, corresponding visual perception data, tactile perception data, and grasping posture parameters were collected. The sampling points of the visual and tactile perception data were then aligned in 3D space to form a unified 3D geometric representation structure. Simultaneously, the grasping stability result was labeled for each grasping action, generating positive and negative sample labels, with a positive-to-negative sample ratio of approximately 6:4. Ultimately, the constructed 3DA-VTG dataset contains approximately 440,000 grasping annotation data points, averaging about 5,000 samples per object. Statistical analysis was performed on the alignment errors between the visual and tactile modalities, with an average visual alignment error of approximately 0.90 mm and an average tactile alignment error of approximately 2.24 mm.
[0160] Figure 3A visualization of a portion of the 3DA-VTG dataset shown as an exemplary embodiment of this application. (Refer to...) Figure 3 , Figure 3 (a) shows a 3D point cloud recovered by the method of this application, wherein gray dots represent the real geometric model of the object, colored dots represent the visual point cloud acquired by a visual sensor, and red dots represent the tactile point cloud acquired by a tactile sensor. Figure 3 (b) shows the corresponding raw sensor data, including RGB images acquired by the visual sensor and RGB images acquired by the tactile sensor. This diagram visually demonstrates the correspondence and fusion effect of visual and tactile modal data in three-dimensional space.
[0161] Optionally, in one possible implementation, the evaluation of the grasping stability of the target object includes:
[0162] (1) Based on the completed three-dimensional geometric representation and the tactile three-dimensional geometric representation, the feature encoding is obtained.
[0163] Specifically, tactile feature representations are obtained based on the tactile 3D geometric representation. The point clouds containing the completed 3D geometric representation and tactile feature representations are then input into a point cloud encoding network for feature extraction. The tactile input includes additional contact identification information to distinguish between contact and non-contact areas. The point cloud encoding network can employ a Dynamic Graph Convolutional Neural Network (DGCNN) structure, utilizing edge convolution (EdgeConv) to aggregate local neighborhood features and combining hierarchical farthest point sampling (FPS) to progressively reduce point cloud density, thereby capturing multi-scale geometric patterns and generating downsampled point coordinates and aggregated features. The formula is as follows:
[0164]
[0165]
[0166] in, To complete the set of point cloud coordinates for a three-dimensional geometric representation, To represent the set of point cloud coordinates for tactile features, A point coding network for feature extraction from an input point cloud. , These are the sets of coordinates for the downsampled, completed 3D geometric representation point cloud and the tactile feature representation point cloud. , These are the feature vector sets of the downsampled, completed 3D geometric representation and tactile feature representation points, respectively. Then, the aggregated features are processed through linear mapping and positional encoding.
[0167]
[0168]
[0169] in, , The features are obtained after linear mapping and positional encoding. To map aggregated features to a linear mapping layer of the required dimension of the encoder input, To encode the spatial coordinates of the point cloud into a position encoding function with the same dimension as the feature vector, the two branches share the same position encoding parameters to ensure consistency of cross-modal spatial information.
[0170] Furthermore, to enhance the ability of point cloud features to model local geometric structures and spatial relationships, a geometric-aware self-attention module is introduced. This module, based on the standard self-attention mechanism, fuses the relative geometric relationships between a point and its local neighbors. Its core geometric modeling method can be expressed as:
[0171]
[0172] in, The feature vector corresponding to the query point. It is the feature vector of its local neighborhood's neighboring points. The relative geometric difference features between adjacent points and the query point. For query point of Nearest neighbor set For the queried point's neighborhood, the first Spatial coordinates of the nearest neighbor points This represents a multilayer perceptron used for nonlinear mapping of concatenated features. Based on the geometric relationship modeling method described above, a geometrically perceptual self-attention operator (GeoSelfAttn) can be constructed. It can also be extended to construct a geometrically perceptual cross-attention operator (GeoCrossAttn) to achieve spatial alignment and enhancement of cross-modal features.
[0173] Furthermore, the initial encoded features are processed by the geometry-aware self-attention operator (GeoSelfAttn) to obtain an enhanced geometric feature encoding representation:
[0174]
[0175]
[0176] in, To complete the geometric enhancement feature encoding corresponding to the 3D geometric representation, The geometric enhancement feature encoding is used to encode the tactile 3D geometric representation, and the geometric enhancement feature encoding is used for subsequent cross-modal feature fusion and grasping stability evaluation.
[0177] (2) Based on the feature encoding, high-level fusion features are obtained.
[0178] Specifically, the cross-modal fusion module can adopt a phased, hierarchical structure to achieve joint encoding of shape feature encoding and tactile feature encoding. Each phase includes three complementary components: geometric aggregation downsampling, geometric perception cross-modal attention, and standard cross-modal attention. Geometric aggregation downsampling operates only on the shape feature encoding branch, aggregating local geometric features through farthest point sampling (FPS) combined with edge convolution (EdgeConv) to gradually reduce the spatial resolution of the shape feature encoding branch while preserving the complete resolution of the tactile feature encoding branch, thus maintaining fine spatial information of local contact features. In this embodiment, geometric aggregation downsampling, geometric perception cross-modal attention, and standard cross-modal attention can be implemented with reference to existing technologies, such as using the FPS+EdgeConv aggregation method to achieve local geometric feature downsampling. This embodiment does not exclude other cross-modal fusion methods, such as single-stage fusion or convolution-based fusion structures.
[0179] Specifically, assuming the first The input features for this stage are shape enhancement feature encodings. Encoding with haptic enhancement features The corresponding coordinates are and Then the geometric aggregation downsampling operation can be expressed as:
[0180]
[0181] in, This represents the number of target points after downsampling. This involves an aggregation operation of far-point sampling plus edge convolution. Furthermore, a geometry-aware cross-modal attention module is used to re-establish the spatial correspondence between shape feature encoding and tactile feature encoding, enabling tactile query features to effectively perceive local geometric structures. Its operation can be represented as:
[0182]
[0183] in, For the tactile feature encoding branch, For shape feature encoding branch, These are the corresponding coordinates. This is a cross-modal attention mechanism with geometric perception capabilities. Subsequently, a standard cross-modal attention module is used to further fuse information from both modalities, achieving comprehensive information interaction. The formula is expressed as:
[0184]
[0185] in, To illustrate the cross-modal attention mechanism, the three components mentioned above can be executed sequentially in a fixed order and iteratively stacked through multiple stages. In each stage, the tactile feature encoding branch undergoes local neighborhood aggregation, nonlinear mapping, and residual connections to gradually converge from local contact features to global geometric information. Neighborhood aggregation at each stage can employ radius or KNN pooling operations, nonlinear mapping is implemented through a feedforward fully connected network or a multilayer perceptron (MLP), and residual connections ensure stable feature flow propagation. After the final stage... After processing, the output features of the tactile feature encoding branch This can be considered a high-level integration feature. Meanwhile, the corresponding shape feature encoding branch is also enhanced during the fusion process. This high-level fusion feature fully reflects the multimodal spatial relationship and local-global geometric pattern of vision and touch, and can be directly used for subsequent grasping stability assessment, providing reliable feature support for grasping decisions.
[0186] (3) Based on the high-level fusion features, the grasping stability of the target object is evaluated.
[0187] Specifically, the high-level fused features are input into the prediction layer of the grasping stability classification network. The network generates grasping stability scores through layer-by-layer feature mapping and information fusion processing. The score represents the probability of a successful grasping action and can be used to guide the selection of subsequent grasping strategies. The feature mapping and information fusion processing can be implemented using multi-layer feedforward networks, fully connected networks, convolutional networks, or other equivalent algorithms. This embodiment employs a multi-layer feedforward classification network, which generates a grasping stability score for the target object through layer-by-layer nonlinear mapping and feature convergence. In each layer, the system performs a linear transformation on the input features and incorporates a non-linear activation function to enhance the representational power of high-level features. Simultaneously, residual connections maintain the stability of the feature flow, ensuring the effective transfer of local and global information. The final output of the classification network is mapped to a grasping stability score using either a sigmoid or softmax function. This represents the probability of a successful grasping action. It should be noted that the classification network structure and activation function form are not limited, as long as they can output a grasping stability evaluation result based on the high-level fusion features. The grasping stability score can be used to guide subsequent grasping decisions or strategy selection, such as selecting a grasping posture with a higher stability score to execute the actual grasping action. Through the above steps, the system can complete the grasping stability evaluation of the target object based on the input multimodal fusion high-level fusion features, achieving feasible grasping decision support.
[0188] Specifically, the output evaluation result can be a continuous grasping stability score, such as a probability value between 0 and 1, or a binary label, such as "stable" or "unstable". The output result can be displayed in digital form on the control interface or transmitted to the upper controller to guide subsequent grasping strategy selection, action planning, or alarm prompts.
[0189] Furthermore, to verify the effectiveness of the crawling stability evaluation method provided in this application, relevant verification experiments are also provided, as follows:
[0190] This study compared and analyzed the effects of different functional modules in the grasping stability assessment method proposed in this application through ablation experiments. During the experiment, while keeping other conditions consistent, the shape completion module, multimodal fusion method, and visual-tactile spatial alignment mechanism in the method were removed or replaced to evaluate the impact of each module on the grasping stability recognition performance, thereby verifying the effectiveness of the technical solution of this application.
[0191] Specifically, multiple comparison schemes were constructed in the experiment for performance verification. Comparison Scheme 1 is the implementation method using the complete 3D ground truth of the object: in the algorithm presented in this paper, the completed 3D geometric representation is replaced with the complete 3D ground truth of the target object, while the rest of the process remains the same. Comparison Scheme 2 is the implementation method without a shape completion module: in the algorithm presented in this paper, the shape completion module is removed, and the grasping stability evaluation is directly based on the observed visual and tactile 3D geometric representations, while the rest of the modules remain the same. Comparison Scheme 3 is the implementation method without cross-modal attention: in the algorithm presented in this paper, the cross-modal attention module is removed, and only feature concatenation is used to fuse shape feature encoding and tactile feature encoding, while the rest of the modules remain the same. Comparison Scheme 4 is the implementation method without a spatial alignment mechanism: in the algorithm presented in this paper, the spatial alignment mechanism for visual and tactile perception is removed, and the two types of perceptual data are not unified to the same reference coordinate system, while the rest of the modules remain the same. Comparison Scheme 5 is the algorithm presented in this paper. Statistical analysis was performed on the accuracy (Accuracy, Acc) and F1 score of each scheme for grasping stability classification under the same test dataset.
[0192] Table 1 illustrates the impact of different methods on the performance of grasping stability assessment. The table shows that the grasping stability assessment using the complete 3D ground truth object method outperforms other comparative methods, indicating that shape integrity has a positive impact on grasping stability judgment. In contrast, the algorithm presented in this paper also achieves high accuracy and F1 score under the condition of using visual and tactile observations to reconstruct the completed shape, demonstrating that the multimodal completion method can effectively infer the complete geometric information of the target object even with missing observations, thereby improving the grasping stability assessment capability. When using a shape-less completion scheme, using only partial visual and tactile information without completion processing, the accuracy drops to 77.1% and the F1 score drops to 78.9%, indicating that the shape completion module can enhance the geometric consistency between vision and touch, improving the grasping state discrimination ability. When using a cross-modal attention scheme, the accuracy is 77.9% and the F1 score is 77.0%, indicating that the cross-modal attention module can effectively align and fuse visual and tactile features, thereby improving the expressive ability of multimodal information and the grasping stability discrimination performance. When using a spatial alignment mechanism, i.e., without unifying visual and tactile data to the same reference coordinate system for processing, the accuracy drops to 71.6% and the F1 score drops to 75.6%, indicating that the spatial alignment module plays a key role in reducing multimodal perception errors and improving the reliability of grasping stability assessment.
[0193] Table 1 Comparison of the impact of each method on algorithm performance
[0194] Complete 3D True Value of an Object 80.5 82.7 Shapeless completion 77.1 78.9 No cross-modal attention 77.9 77.0 No spatial alignment mechanism 71.6 75.6 This paper's algorithm 79.8 80.7
[0195] As can be understood from the results in the table, the visual-tactile grasping stability evaluation method provided in this application, by introducing a shape completion module, a cross-modal attention module, and a visual-tactile spatial alignment mechanism, can effectively improve the accuracy and robustness of grasping stability recognition in complex grasping scenarios. Compared with existing technical solutions, the method provided in this application has at least the following beneficial effects: First, the shape completion module enhances multimodal geometric consistency and reduces discrimination errors caused by incomplete perception; second, the cross-modal attention module achieves spatial alignment and effective fusion of visual and tactile features, improving multimodal information expression capabilities and grasping stability discrimination performance; third, the spatial alignment mechanism reduces spatial deviation between vision and touch, improving the reliability of grasping stability evaluation; fourth, without increasing system complexity, it achieves a significant improvement in overall grasping stability recognition performance and has good engineering application value.
[0196] Corresponding to the aforementioned embodiment of a gripping stability assessment method, this application also provides an embodiment of a gripping stability assessment device.
[0197] Figure 4This is a schematic diagram of the structure of Embodiment 1 of the gripping stability evaluation device provided in this application. Please refer to... Figure 4 The crawling stability evaluation device provided in this embodiment includes an acquisition module 401, a construction module 402, an inference module 403, and an evaluation module 404, wherein:
[0198] The acquisition module 401 is used to calibrate the grasping system and obtain the transformation relationship between multiple coordinate systems of the grasping system;
[0199] The acquisition module 401 is used to acquire visual perception data of the target object and tactile perception data corresponding to the contact when the robot grasps and makes contact with the target object.
[0200] The construction module 402 is used to convert the visual perception data and the tactile perception data into corresponding three-dimensional geometric representations in the same coordinate system based on the conversion relationship. The three-dimensional geometric representations include a visual three-dimensional geometric representation obtained from the visual perception data and a tactile three-dimensional geometric representation obtained from the tactile perception data.
[0201] The construction module 402 is used to construct a multimodal three-dimensional representation of the target object based on the three-dimensional geometric representation;
[0202] The inference module 403 is used to evaluate the grasping stability of the target object based on the completed three-dimensional geometric representation, the tactile three-dimensional geometric representation and the pre-trained grasping stability evaluation model, and output the evaluation result.
[0203] The evaluation module 404 is used to input the completed three-dimensional geometric representation into a pre-trained grasping stability evaluation model, evaluate the grasping stability of the target object, and output the evaluation result.
[0204] The apparatus of this embodiment can be used to perform... Figure 1 The steps of the method embodiment shown are similar in principle and process, and will not be repeated here.
[0205] The specific implementation process of the functions and roles of each unit in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.
[0206] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0207] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A method for evaluating the stability of grasping, characterized in that, include: The grasping system is calibrated to obtain the transformation relationships between multiple coordinate systems of the grasping system; When the robot grasps and makes contact with the target object, it acquires the visual perception data of the target object and the tactile perception data corresponding to the contact. Based on the transformation relationship, the visual perception data and the tactile perception data are respectively converted into corresponding three-dimensional geometric representations in the same coordinate system. The three-dimensional geometric representations include the visual three-dimensional geometric representations obtained from the visual perception data and the tactile three-dimensional geometric representations obtained from the tactile perception data. Based on the aforementioned three-dimensional geometric representation, a multimodal three-dimensional representation of the target object is constructed; Based on the multimodal 3D representation, a complete 3D geometric representation of the target object is obtained; Based on the completed 3D geometric representation, the tactile 3D geometric representation, and the pre-trained grasping stability evaluation model, the grasping stability of the target object is evaluated, and the evaluation results are output.
2. The method according to claim 1, characterized in that, The construction of a multimodal 3D representation of the target object based on the 3D geometric representation includes: Extract the features of the visual three-dimensional geometric representation to construct a visual feature representation; extract the features of the tactile three-dimensional geometric representation to construct a tactile feature representation. The visual feature representation and the tactile feature representation are fused to obtain the multimodal three-dimensional representation.
3. The method according to claim 2, characterized in that, The process of obtaining a complete 3D geometric representation of the target object based on the multimodal 3D representation includes: Based on the multimodal 3D representation, a modal-free shape representation is obtained; Based on the modal-free shape representation, and combined with the relative positional relationship between the observed region and the missing region in three-dimensional space, a complete three-dimensional geometric representation is generated.
4. The method according to claim 1, characterized in that, The calibration and grasping system obtains the transformation relationships between multiple coordinate systems of the grasping system; including: Collect calibration data from different observation perspectives; Based on the predetermined imaging model and the calibration data, the transformation relationship between the image coordinate system and the visual sensor coordinate system, and the transformation relationship between the image coordinate system and each tactile sensor coordinate system are calculated respectively, to obtain the visual intrinsic parameter matrix and multiple tactile intrinsic parameter matrices; Based on the calibration data and the visual intrinsic parameter matrix, the transformation relationship between the visual sensor coordinate system and the target reference coordinate system is calculated to obtain the first extrinsic parameter matrix; Based on the installation pose information of each of the tactile sensors, the transformation relationship between the coordinate system of each tactile sensor and the target reference coordinate system is calculated to obtain multiple second extrinsic parameter matrices; Based on the visual intrinsic parameter matrix and the first extrinsic parameter matrix, as well as each of the tactile intrinsic parameter matrices and their corresponding second extrinsic parameter matrices, a transformation relationship between multiple coordinate systems of the grasping system is established.
5. The method according to claim 4, characterized in that, The process of converting visual perception data into a corresponding three-dimensional geometric representation in the same coordinate system includes: Obtain predetermined reference pixels, and based on the reference pixels and the visual perception data, obtain the two-dimensional image region of the target object; Based on the two-dimensional image region, the visual perception data, and the visual intrinsic parameter matrix, a visual three-dimensional geometric representation in the visual sensor coordinate system is obtained. Based on the visual three-dimensional geometric representation in the visual sensor coordinate system and the first extrinsic parameter matrix, the visual three-dimensional geometric representation in the target reference coordinate system is obtained.
6. The method according to claim 4, characterized in that, The process of converting tactile perception data into a corresponding three-dimensional geometric representation in the same coordinate system includes: Based on the tactile perception data, the depth information of the tactile perception data is calculated; Based on the tactile perception data, the depth information, and the tactile intrinsic parameter matrix, a three-dimensional geometric representation of tactile sensation in the coordinate system of the tactile sensor is obtained; Based on the tactile three-dimensional geometric representation in the coordinate system of the tactile sensor, and combined with the second extrinsic parameter matrix, the tactile three-dimensional geometric representation in the target reference coordinate system is obtained.
7. The method according to claim 2, characterized in that, After constructing the tactile feature representation based on the tactile three-dimensional geometric representation, the method further includes: Based on the tactile perception data, contact identification information is added to the points in the tactile feature representation.
8. The method according to claim 1, characterized in that, The training method for the pre-trained grasping stability evaluation model includes: Obtain training samples of at least one training object under different grasping postures; The training samples are screened through offline simulation experiments to obtain effective training samples, which contain multiple training representations; Obtain the training perception data corresponding to the effective training samples, and based on the transformation relationship, convert the training perception data into a three-dimensional training representation in the same coordinate system; Based on the three-dimensional training representation, a multimodal training representation of the training object is constructed; Obtain the stability result of each training representation in the crawling experiment, and generate a corresponding stability label for each multimodal training representation; The grasping stability evaluation model is trained based on the multimodal training representation and the stability label.
9. The method according to claim 1, characterized in that, The evaluation of the stability of grasping the target object includes: Based on the completed 3D geometric representation and the tactile 3D geometric representation, feature encoding is obtained; Based on the aforementioned feature encoding, high-level fusion features are obtained; Based on the aforementioned high-level fusion features, the grasping stability of the target object is evaluated.
10. A device for evaluating grasping stability, characterized in that, The device includes an acquisition module, a construction module, a reasoning module, and an evaluation module, wherein: The acquisition module is used to calibrate the grasping system and obtain the transformation relationship between multiple coordinate systems of the grasping system; The acquisition module is used to acquire visual perception data of the target object and tactile perception data corresponding to the contact when the robot grasps and makes contact with the target object. The construction module is used to convert the visual perception data and the tactile perception data into corresponding three-dimensional geometric representations in the same coordinate system based on the conversion relationship. The three-dimensional geometric representations include a visual three-dimensional geometric representation obtained from the visual perception data and a tactile three-dimensional geometric representation obtained from the tactile perception data. The construction module is used to construct a multimodal three-dimensional representation of the target object based on the three-dimensional geometric representation; The reasoning module is used to obtain a complete three-dimensional geometric representation of the target object based on the multimodal three-dimensional representation; The evaluation module is used to evaluate the grasping stability of the target object based on the completed three-dimensional geometric representation, the tactile three-dimensional geometric representation, and the pre-trained grasping stability evaluation model, and output the evaluation results.