Visual and touch combined multi-mode manipulator grabbing performance testing method

By generating multimodal data combined with visual touch in a simulation environment, training neural network models to predict the crawling performance, solving the problems of high computing costs and unstable crawling in the existing technology, achieving more efficient crawling performance prediction and robot crawling stability.

CN120533749APending Publication Date: 2025-08-26SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510654734.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The prior art relies on mechanical models and geometric constraints in the inference of crawling performance, and is computationally expensive and unsuitable for dynamic environments. The lack of tactile feedback of crawling methods based on visual data leads to instability in crawling and cannot handle complex environmental changes.

Method used

In the simulation environment, a multimodal object grab data that combines visual and touch is generated. The performance inference matrix is ​​trained by neural networks, and the performance prediction of the grabbing performance is combined with visual and tactile information. The scene cloud is sampled using arrayed tactile sensing units and simulated grabbing to train the neural network model.

Benefits of technology

It improves the accuracy and reliability of crawling performance prediction, enhances the practicality of robot crawling technology, and supports the application of actual crawling scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120533749A_ABST
    Figure CN120533749A_ABST
Patent Text Reader

Abstract

A visual touch combined multi-mode mechanical arm grabbing performance testing method comprises the steps that after a clamping jaw comprising a plurality of arrayed touch sensing units is used in a simulation environment and sampling is carried out to obtain scene point cloud and grabbing poses of the scene point cloud, grabbing is further simulated through a two-finger clamping jaw in the simulation environment, and visual touch combined multi-mode data is recorded to serve as a training set; training the constructed neural network; and through the trained neural network, in an actual grabbing scene, according to the input object point cloud, the clamping jaw grabbing pose and the grabbing performance reasoning matrix, a grabbing performance quantitative index is predicted and obtained. According to the method, a large amount of vision and touch combined multi-mode object grabbing data is generated in a simulation environment, a grabbing performance reasoning matrix is calculated according to scene point cloud and contact mechanical information, and a neural network model is trained to predict the grabbing performance level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technology in the field of robot control, specifically a method for testing the grasping performance of a multi-modal manipulator combining vision and touch. Background Art

[0002] In the field of robotic automation, grasping stability is crucial for subsequent operations, so accurate grasping performance reasoning is crucial. Existing technologies primarily rely on mechanical models and geometric constraints, using metrics such as shape closure and force closure to reason about grasping performance. This is computationally expensive and unsuitable for handling changes in dynamic environments. Their reliability and generalization are significantly limited. Data-driven machine learning methods generate datasets through simulated environments or physical grasping, and techniques that combine neural networks to predict grasping performance typically rely solely on visual data. The lack of tactile feedback and quantitative stability metrics leads to unstable grasping execution and is unable to address the complex environmental changes encountered during the grasping process. Summary of the Invention

[0003] In response to the above-mentioned deficiencies in the prior art, the present invention proposes a method for testing the grasping performance of a multimodal manipulator combining vision and touch. By generating a large amount of multimodal object grasping data combining vision and touch in a simulation environment, the grasping performance inference matrix is ​​calculated based on the scene point cloud and contact mechanics information, and a neural network model is trained to predict the grasping performance level.

[0004] The present invention is achieved through the following technical solutions:

[0005] The present invention relates to a method for testing the grasping performance of a multimodal manipulator combining vision and touch. After using a gripper containing several arrayed tactile sensing units to sample and obtain a scene point cloud and its grasping posture in a simulation environment, a two-finger gripper is used to simulate grasping in the simulation environment and record the multimodal data combining vision and touch as a training set to train a constructed neural network. The trained neural network is used to predict quantitative indicators of grasping performance in an actual grasping scenario based on the input object point cloud, gripper grasping posture, and grasping performance inference matrix.

[0006] The neural network includes: a vision-guided tactile perception model, an object center of mass prediction model, and a 3D convolutional neural network that predicts quantitative indicators of grasping performance through a grasping performance inference matrix, wherein: the tactile perception model obtains the contact force and clamping jaw closing width of each sensing unit based on the scene point cloud and the gripping posture of the gripper; the object center of mass prediction model obtains the center of mass position of the object to be grasped based on the scene point cloud and the gripping posture of the gripper; and the 3D convolutional neural network obtains quantitative indicators of grasping performance based on the prediction of the grasping performance inference matrix.

[0007] The tactile perception model includes: PointNet module, convolution module, maximum pooling module and output head module, wherein: PointNet module is based on the input gripper grasping posture and scene point cloud After extracting the feature vector, it is spliced ​​with the grasping posture. The spliced ​​vector passes through three convolution modules and a maximum pooling module in sequence, and outputs the contact force of each sensor unit through an output head module containing three fully connected layers and a ReLU activation function. and jaw closing width .

[0008] The object center of mass prediction model includes: Transformer module, convolution module, maximum pooling module and output head module, wherein: Transformer module is based on the input gripper grasping posture and scene point cloud After extracting the feature vector, it is spliced ​​with the grasping posture. The spliced ​​vector passes through two convolution modules and a maximum pooling module in sequence, and the center of mass position of the object to be grasped is output through an output head module containing two fully connected layers and a ReLU activation function. .

[0009] The 3D convolutional neural network includes: a first convolution module, a second convolution module, a third convolution module and an output head module, wherein: the first convolution module infers the matrix M according to the input grasping performance G ] , the feature vector is obtained by convolution, batch normalization, ReLU activation function and maximum pooling. The second and third convolution modules perform the same processing on the feature vector in turn, and the output of the output head module containing two fully connected layers, ReLU activation function and discard module outputs the quantitative index of the crawling performance. .

[0010] The training set is obtained by simultaneously creating multiple simulated grasping scenes and recording scene point clouds. After repeatedly performing the grasping action, the corresponding gripper grasping posture, tactile information, center of mass position of the grasped object and grasping performance quantitative indicators are recorded. The training set specifically includes: grasping scene point clouds, a set of effective grasping postures and their corresponding gripper grasping postures, gripper closing width, contact force of each sensor unit, center of mass position of the grasped object, grasping performance inference matrix and grasping performance quantitative indicators.

[0011] The grasping performance inference matrix refers to: the force component and torque component of each tactile sensing unit in the direction of the three-dimensional coordinate axis.

[0012] The crawling performance quantitative indicators ,in: It is the force applied to the grasped object simultaneously along the x, y, and z coordinate directions, with the unit being Newton.

[0013] The present invention relates to a multimodal manipulator grasping performance testing system for implementing the above-mentioned method, comprising: a grasping performance inference matrix calculation module, a training set generation module and a grasping performance quantitative index prediction neural network, wherein: the grasping performance inference matrix calculation module calculates the grasping performance inference matrix based on the geometric information of the target object and the interaction information between the gripper and the object; the training set generation module performs voxel sampling on the three-dimensional space of the grasping scene object and simulates grasping in the IsaacSim simulation environment to generate a visual-tactile combined multimodal data training set; the grasping performance quantitative index prediction neural network generates corresponding tactile data based on the scene point cloud and the gripping posture of the gripper, and predicts the grasping performance quantitative index.

[0014] Technical Effects

[0015] This invention generates a large-scale visual-tactile multimodal grasping dataset and simultaneously integrates visual and tactile information into a grasping performance quantification prediction neural network. This allows for grasping performance inference based solely on scene point clouds and gripper grasping poses. Compared to existing technologies, this invention addresses the lack of tactile force information in existing grasping datasets, enhancing the comprehensiveness of grasping data. It also improves the accuracy and reliability of grasping performance predictions, supports application in real-world grasping scenarios, and enhances the practicality of robotic grasping technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 Flowchart of the present invention;

[0017] Figure 2 Schematic diagram of the gripper and arrayed tactile sensing unit;

[0018] Figure 3 Schematic diagram of randomly generated grasping scenes and recorded point clouds;

[0019] Figure 4 Visualization result diagram of the contact force of the sensing unit;

[0020] Figure 5 Schematic diagram of the force component calculation method;

[0021] Figure 6 Schematic diagram of the overall neural network structure;

[0022] In the figure: FDN represents the tactile perception model, OCN represents the object center of mass prediction model, and SAN represents the grasping performance quantitative index prediction model;

[0023] Figure 7 The relationship between the sparsity of the crawling performance reasoning matrix and the quantitative indicators of crawling performance;

[0024] Figure 8 Schematic diagram of the grasping scene and tactile information for the physical experiment. DETAILED DESCRIPTION

[0025] like Figure 1 As shown, this embodiment relates to a method for testing the grasping performance of a multi-modal manipulator combining vision and touch, including:

[0026] Step S1: Start the IssacSim simulation environment, configure the two-finger gripper model, and configure 8 rows and 5 columns of arrayed tactile sensor units on its left and right gripping surfaces respectively. Each sensor unit provides real-time feedback on the magnitude, direction, and position of the force, such as Figure 2 As shown. At the same time, the randomly generated grasping scene point cloud is recorded , and initialize the pose of the target object based on the scene point cloud, such as Figure 3 shown.

[0027] Step S2: Get the maximum and minimum values ​​of the object to be grasped in the three coordinate axis directions, perform voxel sampling on the object's three-dimensional space, generate a set of potential grasping poses, and discretize and sample each position in the Euler angle space to generate a set of candidate grasping poses. .

[0028] Step S3: Use a two-finger gripper to simulate grasping of all sampled grasping postures, specifically including:

[0029] 3.1 In the IssacSim environment, transform the two-finger gripper to the grasping posture sampled in step S2 , simulate the closing of the gripper, and stop closing when the left and right gripping surfaces of the two-finger gripper are in contact with the object and cannot close further;

[0030] 3.2 Apply external force interference to the grasping to obtain the grasping performance quantitative index: After the gripper is closed, slowly move it up 1 meter to detect whether the grasped object falls. If it falls, the grasping performance quantitative index If it is 0, then if it does not fall, a gradually increasing force will be applied to the grasped object along the x, y, and z coordinate directions at the same time. , the unit is cow, until the object falls, grasping performance quantitative index ;

[0031] 3.3 After screening, the effective grasping pose set in the world coordinate system is obtained , record the corresponding gripper grasping posture , jaw closing width , the contact force of each sensing unit , the center of mass position of the grasped object And quantitative indicators of crawling performance The contact force can be visualized as Figure 4 shown.

[0032] Step S4: The gripping posture of the gripper obtained in step S3 , jaw closing width , the contact force of each sensing unit , the center of mass position of the grasped object Calculate the force component of each sensor unit [ ] and moment components [ ], combined into the crawling performance reasoning matrix M G ],like Figure 5 As shown, specifically including:

[0033] 4.1 The position of each sensor unit is calculated from the clamping jaw closing width w and the geometric model , obtain the force f of any sensing unit a and its direction vector da= ;

[0034] 4.2 Calculate the force f separately a Angle with the coordinate axis x: ,in: is the inverse sine function;

[0035] 4.3 Calculate the force component in the x-direction, specifically: ,in: is the force-torque harmonic coefficient, is the internal angle of the friction cone, A small positive constant is introduced to prevent the force Calculation error when approaching zero;

[0036] In this embodiment, the force-torque harmonic coefficient is 5. Take 0.1.

[0037] 4.4 Repeat steps 4.1-4.3 to obtain the force components in the y and z directions , combined to obtain the force component of each sensing unit [ ];

[0038] 4.5 Project the spatial contact surface S onto the YoZ plane and calculate the moment component in the x direction ,in: represents the contact surface between the gripper and the object, are the unit vectors in the directions of the x, y, and z coordinate axes, is the force direction vector at the contact point, is the direction component of the vector from the contact point to the object's center of mass in the YoZ plane;

[0039] 4.6 Repeat step 4.4 to obtain the force components in the y and z directions , combined to obtain the torque component of each sensor unit [ ].

[0040] Step S5: Create N grabbing scenes in the IssacSim simulation environment. Each scene executes the contents of steps S1-S4 to create a data set D, which includes: grabbing scene point cloud , effective grasping pose set And its corresponding gripper grasping posture , jaw closing width , the contact force of each sensing unit , the center of mass position of the grasped object , crawling performance reasoning matrix M G And quantitative indicators of crawling performance .

[0041] Step S6: construct a tactile perception model, an object center of mass prediction model, and a grasping performance quantitative index prediction model, and train them respectively using the dataset D created in step S5, specifically including:

[0042] 6.1 Input a scene point cloud of shape (M, N, 3) and a gripper grasp pose of shape (M, 7) into the tactile perception model, where M is the batch size (M=12) and N is the number of points in the scene point cloud. First, the scene point cloud is converted into a high-dimensional feature of (M, 1024) using a PointNet variant. This feature is then concatenated with the gripper grasp pose of (M, 7). Finally, the contact force magnitude and gripper width (M, 1) for each sensor unit are output in a shape of (M, 2, 8, 5, 3).

[0043] The tactile perception model uses the Adam optimizer to iteratively update the network parameters, and its loss function is: ,in: and To adjust the hyperparameters as needed, Usually set to ten times , is the cross entropy loss at the contact point, is the mean square error loss at the contact point, specifically: , ,in: represents the predicted contact probability of the i-th sensor, is the activation function.

[0044] 6.2 Input a scene point cloud of shape (M, N, 3) and a gripper grasp pose of shape (M, 7) into the object center of mass prediction model, where M is the batch size (M=12) and N is the number of points in the scene point cloud. A Transformer variant converts the scene point cloud into a high-dimensional feature of shape (M, 1024) and concatenates it with the gripper grasp pose of shape (M, 7). The final output is the center of mass position of the object to be grasped, of shape (M, 3).

[0045] The object center of mass prediction model uses the Adam optimizer to iteratively update the network parameters and uses the mean square error to calculate the loss together with the object center of mass position in the dataset.

[0046] 6.3 Input the five-dimensional grasping performance inference matrix of shape (M, 2, 8, 5, 6) into the grasping performance quantitative indicator prediction model, where M is the batch size (M=8). At the beginning of the network, the matrix is ​​rearranged to (M, 6, 2, 8, 5). The neural network finally outputs a vector of shape (M, 9), where the index of the maximum value in the second dimension is the predicted grasping performance quantitative indicator.

[0047] The grasping performance quantitative index prediction model calculates the loss by using the cross entropy loss function together with the output grasping performance quantitative index and the grasping stability label in the data set, and uses the Adam optimizer to iteratively update the network parameters.

[0048] Step S7: Input the scene point cloud under actual conditions and gripper grasping posture The tactile perception model trained in step S6 and the object center of mass prediction model are used to predict the closing width of the gripper. , the contact force of each sensor unit and the center of mass of the grasped object , and then calculate the crawling performance reasoning matrix M through the method of step S4 G Finally, the grasping performance quantitative index prediction model trained in step S6 is used to predict the final grasping performance quantitative index , the overall network structure is as follows Figure 6 shown.

[0049] Through virtual simulation experiments, the relationship between the sparsity of the grasping performance inference matrix and the grasping performance quantitative index is analyzed. Sparsity refers to the proportion of matrix units less than 0.001. The final results are as follows Figure 7 The grasping performance inference matrix of high-stability grasping has low sparsity, indicating that most sensor units contribute to grasping; in contrast, the grasping performance inference matrix of low-stability grasping has high sparsity, indicating that most sensor units cannot generate effective grasping margin, so the grasp is weak in resisting external interference.

[0050] As shown in Table 1, joint experiments were conducted on seen / mixed / novel datasets generated in a virtual environment to verify the overall performance of the network. Finally, the prediction accuracy of the performance indicators on the seen dataset can reach 75%, and the prediction accuracy of the performance indicators on the mixed and novel datasets can reach 70%.

[0051] Table 1 Network performance verification experiment results

[0052] After specific practical experiments, in a real scene, the scene point cloud obtained by the sensor scan and the pre-generated gripper grasping posture are used as input, and step S7 is executed to obtain the corresponding grasping performance inference matrix and grasping performance quantitative index. The final result is as follows Figure 8 shown.

[0053] Compared with the existing technology, this method generates a large-scale vision-tactile multimodal grasping dataset through simulated grasping, which makes up for the lack of tactile force information in the existing grasping dataset and enhances the comprehensiveness of the grasping data; designs a neural network to predict the quantitative indicators of grasping performance, combines visual information and tactile information, and improves the accuracy of grasping performance prediction; grasping performance is inferred only based on the scene point cloud and the gripping posture of the gripper, supports the application of actual grasping scenarios, and enhances the practicality of robot grasping technology.

[0054] The above-mentioned specific implementation can be partially adjusted in different ways by those skilled in the art without departing from the principles and purpose of the present invention. The scope of protection of the present invention shall be based on the claims and shall not be limited by the above-mentioned specific implementation. All implementation schemes within its scope shall be subject to the constraints of the present invention.

Claims

1. A method for testing the grasping performance of a multimodal manipulator combining vision and touch, characterized in that: After using a gripper containing several arrayed tactile sensing units to sample the scene point cloud and its grasping posture in a simulation environment, the constructed neural network is trained by further simulating grasping with a two-finger gripper in the simulation environment and recording the visual-tactile combined multimodal data as a training set. The trained neural network is then used to predict quantitative indicators of grasping performance in actual grasping scenarios based on the input object point cloud, gripper grasping posture, and grasping performance inference matrix. The grasping performance inference matrix refers to: the force component and torque component of each tactile sensing unit in the direction of the three-dimensional coordinate axis; The training set is obtained by simultaneously creating multiple simulated grasping scenes and recording scene point clouds. After repeatedly performing the grasping action, the corresponding gripper grasping posture, tactile information, center of mass position of the grasped object and grasping performance quantitative indicators are recorded. The training set specifically includes: grasping scene point clouds, a set of effective grasping postures and their corresponding gripper grasping postures, gripper closing width, contact force of each sensor unit, center of mass position of the grasped object, grasping performance inference matrix and grasping performance quantitative indicators.

2. The method for testing the grasping performance of a multi-modal manipulator combining vision and touch according to claim 1 is characterized in that: The neural network includes: a vision-guided tactile perception model, an object center of mass prediction model, and a 3D convolutional neural network that predicts quantitative indicators of grasping performance through a grasping performance inference matrix, wherein: the tactile perception model obtains the contact force and clamping jaw closing width of each sensing unit based on the scene point cloud and the gripping posture of the gripper; the object center of mass prediction model obtains the center of mass position of the object to be grasped based on the scene point cloud and the gripping posture of the gripper; and the 3D convolutional neural network obtains quantitative indicators of grasping performance based on the prediction of the grasping performance inference matrix.

3. The method for testing the grasping performance of a multi-modal manipulator combining vision and touch according to claim 2 is characterized in that: The tactile perception model includes: PointNet module, convolution module, maximum pooling module and output head module, wherein: PointNet module is based on the input gripper grasping posture and scene point cloud After extracting the feature vector, it is spliced ​​with the grasping posture. The spliced ​​vector passes through three convolution modules and a maximum pooling module in sequence, and outputs the contact force of each sensor unit through an output head module containing three fully connected layers and a ReLU activation function. and jaw closing width .

4. The method for testing the grasping performance of a multi-modal manipulator combining vision and touch according to claim 2 is characterized in that: The object center of mass prediction model includes: Transformer module, convolution module, maximum pooling module and output head module, wherein: Transformer module is based on the input gripper grasping posture and scene point cloud After extracting the feature vector, it is spliced ​​with the grasping posture. The spliced ​​vector passes through two convolution modules and a maximum pooling module in sequence, and the center of mass position of the object to be grasped is output through an output head module containing two fully connected layers and a ReLU activation function. .

5. The method for testing the grasping performance of a multi-modal manipulator combining vision and touch according to claim 2 is characterized in that: The 3D convolutional neural network includes: a first convolution module, a second convolution module, a third convolution module and an output head module, wherein: the first convolution module infers the matrix M according to the input grasping performance G ] , the feature vector is obtained by convolution, batch normalization, ReLU activation function and maximum pooling. The second and third convolution modules perform the same processing on the feature vector in turn, and the output of the output head module containing two fully connected layers, ReLU activation function and drop module is used to output the quantitative index of the crawling performance. .

6. The method for testing the grasping performance of a multi-modal manipulator combining vision and touch according to claim 1 is characterized in that: The crawling performance quantitative indicators ,in: It is the force applied to the grasped object simultaneously along the x, y, and z coordinate directions, with the unit being Newton.

7. The method for testing the grasping performance of a multi-modal manipulator combining vision and touch according to claim 1 is characterized in that: The simulated crawling specifically includes: 3.1 In the IssacSim environment, transform the two-finger gripper to the grasping posture sampled in step S2 , simulate the closing of the gripper, and stop closing when the left and right gripping surfaces of the two-finger gripper are in contact with the object and cannot close further; 3.2 Apply external force interference to the grasping to obtain the grasping performance quantitative index: After the gripper is closed, slowly move it up 1 meter to detect whether the grasped object falls. If it falls, the grasping performance quantitative index If it is 0, then if it does not fall, a gradually increasing force will be applied to the grasped object along the x, y, and z coordinate directions at the same time. , the unit is cow, until the object falls, grasping performance quantitative index ; 3.3 After screening, the effective grasping pose set in the world coordinate system is obtained , record the corresponding gripper grasping posture , jaw closing width , the contact force of each sensing unit , the center of mass position of the grasped object And quantitative indicators of crawling performance .

8. The method for testing the grasping performance of a multi-modal manipulator combining vision and touch according to claim 1, 2 or 5, wherein: The crawling performance reasoning matrix M G ], obtained by: 4.1 The position of each sensor unit is calculated from the clamping jaw closing width w and the geometric model , obtain the force f of any sensing unit a and its direction vector da= ; 4.2 Calculate the force f separately a Angle with the coordinate axis x: ,in: is the inverse sine function; 4.3 Calculate the force component in the x-direction, specifically: ,in: is the force-torque harmonic coefficient, is the internal angle of the friction cone, A small positive constant is introduced to prevent the force Calculation error when approaching zero; 4.4 Repeat steps 4.1-4.3 to obtain the force components in the y and z directions , combined to obtain the force component of each sensing unit [ ]; 4.5 Project the spatial contact surface S onto the YoZ plane and calculate the moment component in the x direction ,in: represents the contact surface between the gripper and the object, are the unit vectors in the directions of the x, y, and z coordinate axes, is the force direction vector at the contact point, is the direction component of the vector from the contact point to the object's center of mass in the YoZ plane; 4.6 Repeat step 4.4 to obtain the force components in the y and z directions , combined to obtain the torque component of each sensor unit [ ].

9. The method for testing the grasping performance of a multi-modal manipulator combining vision and touch according to claim 1, wherein: The training specifically includes: 6.1 Input the scene point cloud of shape (M, N, 3) and the gripper grasping posture of shape (M, 7) into the tactile perception model, where: M is the batch size, M=12, and N is the number of points in the scene point cloud. First, the scene point cloud is converted into a high-dimensional feature of (M, 1024) using a PointNet variant, and then spliced ​​with the gripper grasping posture of (M, 7). Finally, the contact force size and gripper width (M, 1) of each sensor unit are output in the shape of (M, 2, 8, 5, 3); The tactile perception model uses the Adam optimizer to iteratively update the network parameters, and its loss function is: ,in: and To adjust the hyperparameters as needed, Usually set to ten times , is the cross entropy loss at the contact point, is the mean square error loss at the contact point, specifically: , ,in: represents the predicted contact probability of the i-th sensor, is the activation function; 6.2 Input the scene point cloud of shape (M, N, 3) and the gripper grasping posture of shape (M, 7) into the object center of mass prediction model, where: M is the batch size, M=12, and N is the number of points in the scene point cloud. Use the Transformer variant to convert the scene point cloud into a high-dimensional feature of (M, 1024), which is then spliced ​​with the gripper grasping posture of (M, 7). Finally, the center of mass position of the object to be grasped is output with a shape of (M, 3); The object centroid prediction model uses the Adam optimizer to iteratively update the network parameters and calculates the loss using the mean square error together with the object centroid position in the dataset; 6.3 Input the five-dimensional grasping performance inference matrix of shape (M, 2, 8, 5, 6) into the grasping performance quantitative index prediction model, where M is the batch size (M=8). At the beginning of the network, it is rearranged to (M, 6, 2, 8, 5). Finally, the neural network outputs a vector of shape (M, 9). The index of the maximum value in the second dimension is the predicted grasping performance quantitative index. The grasping performance quantitative index prediction model calculates the loss by using the cross entropy loss function together with the output grasping performance quantitative index and the grasping stability label in the data set, and uses the Adam optimizer to iteratively update the network parameters.

10. A multi-modal manipulator grasping performance testing system that implements the method according to any one of claims 1 to 9, characterized in that: include: The grasping performance inference matrix calculation module, the training set generation module and the grasping performance quantitative index prediction neural network are as follows: the grasping performance inference matrix calculation module calculates the grasping performance inference matrix based on the geometric information of the target object and the interaction information between the gripper and the object; the training set generation module performs voxel sampling on the three-dimensional space of the grasping scene objects in the IsaacSim simulation environment and performs simulated grasping to generate a visual-tactile combined multimodal data training set; the grasping performance quantitative index prediction neural network generates corresponding tactile data based on the scene point cloud and the gripping posture of the gripper, and predicts the grasping performance quantitative index.