Accurate object grabbing method of intelligent mechanical arm

Through robotic arm position and RGB-D vision generation, 3D position and contour combined with Flow-Matching and supervised learning data alignment methods, the perception and universality of robotic arm in complex environments is solved, and object grabbing with high precision and high success rate is achieved.

CN120552072APending Publication Date: 2025-08-29ZHEJIANG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510962419.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

When existing robotic arms grasp objects in complex unstructured environments, they lack perception and understanding capabilities and lack universality, and there is a gap in contact mechanical perception between the simulation model and the physical environment, resulting in low crawling success rate and poor robustness.

Method used

The robotic arm position and RGB-D visually generate the 3D position and contour of the target object, combined with Flow-Matching to generate the claw clamp position space, and through the simulation environment construction and supervised learning data alignment method, the authenticity and reliability of simulation evaluation are improved and accurate capture is achieved.

Benefits of technology

It significantly improves the accuracy of 3D pose and contour generation of objects, enhances the universality and diversity of grasping postures, greatly improves the success rate and safety of grabbing, lays the foundation for fine operation, and achieves more than 95% of object grabbing success rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120552072A_ABST
    Figure CN120552072A_ABST
Patent Text Reader

Abstract

The invention discloses a precise object grabbing method of an intelligent mechanical arm, and belongs to the field of robotics. The method comprises the following steps: generating a 3D pose and a contour of a target object by using a mechanical arm pose and RGB-D visual content; the method comprises the following steps: generating a claw clamping pose space through Flow-Matching; a simulation test environment is constructed through a simulation environment, and an optimal grabbing pose is obtained from the obtained claw clamping pose space through a simulation test; according to the data alignment method based on supervised learning, an obtained result in a simulation test environment is aligned with an actual sensing result of physical electronic skin, so that the authenticity and reliability of simulation evaluation are improved, and then accurate grabbing of objects by the intelligent mechanical arm is achieved. The invention aims to improve the universality, the precision and the robustness grabbing capability of the mechanical arm on diversified objects in a complex unstructured environment, especially in a home assistant scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of robotics, specifically to a method for accurately grasping objects using an intelligent robotic arm. Specifically, the invention covers interdisciplinary areas such as robotic grasping, robot vision / computer vision, machine learning / deep learning, and robot simulation and verification. Background Art

[0002] The existing method of grasping objects with a robotic arm has the following drawbacks:

[0003] 1. Insufficient perception and understanding:

[0004] 1) Limited perception in complex environments: Home environments are highly unstructured and dynamically changing, with fluid placement of objects, variable lighting, and occlusions. Traditional 2D vision is limited in its ability to judge distance and shape, making it difficult to accurately capture the three-dimensional geometry of objects.

[0005] 2) Lack of fine manipulation perception: Many household tasks require fine tactile and force perception, such as picking up fragile objects. Existing solutions often struggle to achieve the fine perception capabilities of the human hand.

[0006] 3) Difficulty handling transparent, reflective or textureless objects: Pure vision solutions have difficulty accurately identifying the shape and posture of objects made of these special materials, affecting the success rate of grasping.

[0007] 2. Lack of versatility:

[0008] 1) Limited operational diversity: Objects in the home are diverse, with varying shapes, materials, and sizes, requiring robotic arms to possess strong grasping and manipulation versatility. Current robotic arms excel at repetitive, structured tasks, but their ability to generalize to unknown objects and complex operations needs improvement.

[0009] 2) Insufficient grasping robustness: In actual operations, grasping posture may fail due to slight errors (such as unstable grasp or crushing the object), and there is a lack of real-time detection and correction mechanism for slippage during the grasping process.

[0010] 3) The gap between model predictions and the physical environment: Although large-scale data-driven general grasping solutions (such as GraspNet, or solutions based on large models) have strong generalization capabilities, the grasping posture predictions output by their models are not adjusted by real-time contact feedback in the physical environment. Especially in terms of contact mechanics and fine interaction, pure simulation or pure vision solutions are difficult to completely bridge the gap with the real world, which limits their application in tasks that are sensitive to force or require precise manipulation, such as grasping fragile or soft objects.

[0011] Therefore, there is an urgent need to provide a new method for accurately grasping objects with an intelligent robotic arm. Summary of the Invention

[0012] The present invention aims to overcome the shortcomings of the prior art and address the technical issues of insufficient perception and understanding capabilities and versatility of existing robotic arms in object grasping, particularly in complex, unstructured environments. Furthermore, it addresses the gap in contact mechanics perception between simulation models and physical hardware, and provides a method for precise object grasping using an intelligent robotic arm. The present invention aims to improve the versatility, precision, and robustness of robotic arms in grasping a variety of objects in complex, unstructured environments, particularly in home assistant scenarios.

[0013] The specific technical solutions adopted in the present invention are as follows:

[0014] In a first aspect, the present invention provides a method for accurately grasping an object using an intelligent robotic arm, as follows:

[0015] S1. Generate the 3D pose and outline of the target object using the robot arm pose and RGB-D visual content;

[0016] S2, based on the results obtained in S1, generate the gripper pose space through Flow-Matching;

[0017] S3, constructing a simulation test environment through the simulation environment, and obtaining the optimal grasping posture from the gripper posture space obtained in S2 through simulation testing;

[0018] S4: Using a supervised learning-based data alignment method, the results obtained in S3 in the simulation test environment are aligned with the actual perception results of the physical electronic skin to improve the authenticity and reliability of the simulation evaluation, thereby achieving precise grasping of objects by the intelligent robotic arm;

[0019] A data alignment method based on supervised learning is as follows:

[0020] S41, respectively collecting simulation environment data and physical environment data, and preprocessing the collected data;

[0021] S42. Based on the result obtained in S41, mapping from high-level grasping information to electronic skin contact force is achieved through the first model, and mapping from simulated collision impulse to electronic skin contact force is achieved through the assistance of the second model using the first model;

[0022] S43. Integrate the second model in S42 into the grasping feasibility evaluation module to improve the accuracy of grasping force closure and stability evaluation in the simulation test environment.

[0023] Preferably, the S1 is as follows:

[0024] S11. Use the RGB-D camera to obtain the color image and depth information of the target object. At the same time, use the robot arm's own encoder or inertial navigation unit to obtain the real-time and accurate position and posture of the robot arm's end effector.

[0025] S12, applying a deep learning model to the color image obtained in S11 to detect and segment the target object, thereby obtaining a 2D image contour of the target object;

[0026] S13, combining the depth information obtained in S11, and upgrading the 2D image contour obtained in S12 to 3D point cloud data;

[0027] S14, using a point cloud processing network or in combination with a PnP algorithm, estimating a 6D pose of the target object from the 3D point cloud data obtained in S13; the 6D pose includes a 3D position and a 3D orientation;

[0028] S15: Using the pre-calibrated hand-eye relationship, the target object pose estimated in the RGB-D camera coordinate system in S14 is transformed into the robotic arm base coordinate system. At the same time, the real-time and accurate pose information of the robotic arm in S11 is incorporated into the calculation of the target object pose transformation to further correct and optimize the accuracy of the target object's 3D pose.

[0029] S16. Based on the 3D point cloud data obtained in S13 and the result obtained in S15, a relatively accurate 3D outline of the target object is generated.

[0030] Preferably, the RGB-D camera is Intel RealSense D435, the deep learning model is one of Mask R-CNN, YOLO-V8, and DETR, and the point cloud processing network is PointNet++ or DGCNN;

[0031] In S15, the coordinate system conversion formula is as follows:

[0032]

[0033] in, is the pose of the target object in the robot coordinate system, is the position of the end effector in the robot arm coordinate system, is the pose of the camera in the end effector coordinate system, is the pose of the target object in the camera coordinate system.

[0034] Preferably, the S2 is as follows:

[0035] S21. Build a generative model based on Flow-Matching;

[0036] S22. Using the Flow-Matching model trained in S21, and taking the 3D pose and contour of the target object obtained in S1 as input, generate or sample a series of candidate gripper grasping poses for the target object, and then obtain the gripper pose space; wherein each candidate gripper grasping pose includes the three-dimensional position, posture and clamping opening of the gripper.

[0037] Preferably, the S21 is specifically as follows:

[0038] Through Flow-Matching learning, a dataset containing the 3D pose and contour of the target object and the corresponding successful grasping gripper pose is generated based on the 3D geometric information of the input target object to generate a continuous, high-dimensional gripper grasping pose spatial distribution.

[0039] Preferably, the S3 is as follows:

[0040] S31. Using the selected physical simulator, import the MJCF model of the robotic arm to ensure that the physical parameters of the model are consistent with the real hardware. Then, import the 3D model of the target object and the environment model into the physical simulator to simulate the necessary physical interactions between the robotic arm gripper and the target object, and between the target object and the environment, thereby constructing a simulation test environment.

[0041] S32, for each candidate gripper grasping pose in the gripper pose space generated by S2, performing an iterative evaluation in the simulation test environment constructed by S31 to evaluate the feasibility of the grasping;

[0042] S33. According to the evaluation index in S32, a comprehensive score is assigned to the candidate gripper grasping posture; a particle swarm optimization algorithm or a genetic algorithm is used to search in the gripper posture space of S2, and the grasping posture with the highest comprehensive score is selected as the final optimal grasping posture.

[0043] Preferably, the S32 is specifically as follows:

[0044] S321. Perform self-collision detection between the robot arm and itself and environmental collision detection between the robot arm and surrounding obstacles in the physical simulator to eliminate postures that may cause collision; introduce a collision score as one of the evaluation indicators;

[0045] S322. Performing inverse kinematics of the manipulator for each candidate gripper grasping posture in a physical simulator to determine whether the manipulator can reach the target object with the posture; introducing a reachability score as one of the evaluation indicators;

[0046] S323. Simulate a grasping action in a physical simulator to evaluate the contact force distribution, friction, and overall force closure between the gripper and the target object; introduce force closure measurement or simulated grasping success rate as one of the evaluation indicators;

[0047] The stability of the grasping process is evaluated in a physical simulator by simulating the disturbances imposed on the target object by the gripper; the stability score is introduced as one of the evaluation indicators.

[0048] Preferably, in S41,

[0049] Simulation environment data acquisition involves building a robot model, including the manipulator and end effector, as well as a variety of 3D models of target objects and environmental models, in a high-performance physical simulator. By simulating the manipulator's grasping of the target object, the system systematically collects the gripper's position, object position, object properties, gripper closure, and collision impulses of the contact parts in the simulation test environment.

[0050] Physical environment data acquisition refers to performing actual grasping tests on physical objects corresponding to the simulated test environment on a real robotic arm gripper equipped with an electronic skin sensor, and simultaneously collecting the gripper posture, object posture, object properties, gripper closure degree and electronic skin contact force matrix under physical experimental conditions.

[0051] Preferably, in S42,

[0052] The first model input feature vector is a unified input vector that is obtained by feature engineering the gripper posture, object posture, object attributes, and gripper closure degree information. The output vector is the target electronic skin contact force matrix.

[0053] The network architecture of the first model is divided into:

[0054] The fully connected layer is used to capture the nonlinear relationship between different input modalities. It uses a multi-layer perceptron structure to map high-dimensional input vectors into one-dimensional latent feature vectors.

[0055] The reshaping layer is used to reshape the one-dimensional latent feature vector obtained by the fully connected layer into a low-resolution two-dimensional latent feature map;

[0056] The upsampling layer is used to gradually upsample the two-dimensional latent feature map obtained by the reshaping layer using three layers of transposed convolutional layers to restore it to the resolution of the target electronic skin contact force matrix;

[0057] The output layer is used to compress the number of channels to the same number as the target electronic skin contact force matrix using the last layer of convolution. The activation function can be selectively used to ensure the physical rationality of the output value.

[0058] The training objective of the first model is to minimize the difference between the predicted contact force matrix and the real electronic skin contact force matrix;

[0059] The second model input eigenvector is the collision impulse matrix in the simulation environment, and the output vector is the target electronic skin contact force matrix;

[0060] The second model uses a UNet-like convolutional neural network architecture, which is divided into an encoder, a decoder, and an output layer. The output layer is used to compress the number of channels to the same number as the target electronic skin contact force matrix using the last layer of convolution, and can optionally use an activation function.

[0061] The training data set of the second model consists of the collision impulse matrix and the electronic skin contact force matrix obtained in S41, and the training goal is to minimize the difference between the predicted contact force matrix and the real electronic skin contact force matrix.

[0062] Preferably, the S43 is as follows:

[0063] S431. Using the second model, for each candidate gripper grasping posture generated in the simulation test environment, simulate its grasping action and extract the collision impulse matrix between the gripper and the target object;

[0064] S432. The collision impulse matrix obtained in the simulation test environment in S431 is used as input to the trained second model. The second model can output a predicted electronic skin contact force matrix that is closer to the actual one.

[0065] Compared with the prior art, the present invention has the following beneficial effects:

[0066] By combining robotic arm posture, RGB-D vision, flow-matching models, comprehensive simulation verification, and an innovative simulation and physical data alignment mechanism (based on electronic skin and supervised learning), this invention is expected to achieve the following significant effects and advantages:

[0067] 1) Significantly improve the accuracy of object 3D pose and contour generation:

[0068] By fusing the robot's own high-precision pose data with RGB-D visual information, the limitations of a single visual solution in accurately perceiving objects with complex lighting, partial occlusion, low texture, and even transparent / reflective objects are overcome, making the object's 3D pose and contour information more accurate and robust.

[0069] 2) Enhance the versatility and diversity of grasping postures:

[0070] The introduction of a Flow-Matching generative model intelligently generates a space of high-quality gripper poses encompassing a wide range of possibilities based on the object's 3D geometry. This enables the robot arm to handle new objects of varying shapes, even unknown ones, significantly improving the versatility and generalization of its grasping capabilities, rather than being limited to preset or fixed poses.

[0071] 3) Greatly improve crawling success rate and security:

[0072] Before actual execution, all candidate grasping postures are rigorously tested for collision detection and reachability analysis in a comprehensive, high-fidelity simulation environment. By introducing the physical hardware perception capabilities of electronic skin and using supervised learning methods (neural networks) to align the collision impulse data in the simulation environment with the contact force data of real electronic skin with high precision, the present invention can provide a more accurate and realistic assessment of grasping force closure and stability. This ensures that the optimal grasping posture ultimately selected has an extremely high success rate and reliability, while reducing the risk of damage to the robotic arm and the environment, especially when handling fragile, soft objects or objects that require precise force application.

[0073] Thanks to precise 3D object pose and contour generation, intelligent gripper pose space generation, and rigorous simulation testing and optimization (including alignment of simulation and physical contact perception), the present invention has the strong potential to achieve an object grasping success rate of over 95% (grasping test using YCB Video Models), thereby significantly improving the practicality and reliability of the robotic arm in complex home environments.

[0074] 4) Lay the foundation for fine operation:

[0075] By introducing electronic skin and employing supervised learning methods (such as UNet-like CNN) to align simulation data with physical hardware data, this invention effectively addresses the shortcomings of traditional simulation in contact mechanics modeling, enabling simulation evaluation results to more accurately reflect the contact perception of the real physical world. This not only brings higher precision to grasping force closure and stability assessments, but also provides a more solid and reliable foundation for subsequent precision manipulation using real-world robotic arms in conjunction with tactile feedback (such as adaptive adjustment of gripping force, real-time slip prevention, and object material identification).

[0076] In summary, this paper provides an innovative and efficient robotic arm object grasping method through the fusion of multi-source perception data (robotic arm posture + RGB-D vision), advanced generative models (Flow-Matching), and rigorous simulation verification (combining advanced simulation and physical data alignment mechanisms). This method will enable the robotic arm to demonstrate stronger environmental adaptability, higher task completion rate, and better object manipulation capabilities in complex and unstructured application scenarios such as home assistants. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] Figure 1 is a flow chart of the method of the present invention;

[0078] Figure 2 Schematic diagram of the first model and the second model;

[0079] Figure 3 is the RGB camera perspective image;

[0080] Figure 4 for Figure 3 Depth map of

[0081] Figure 5 To capture the example image. DETAILED DESCRIPTION

[0082] The present invention will be further described and illustrated below with reference to the accompanying drawings and specific embodiments. The technical features of each embodiment of the present invention may be combined accordingly, provided that there is no conflict between them.

[0083] like Figure 1 The figure shows a precise object grasping method for an intelligent robotic arm provided by the present invention. By integrating the robotic arm's posture with RGB-D vision, applying the Flow-Matching model to generate the grasping posture, and conducting comprehensive simulation verification, the method can achieve accurate and robust grasping of diverse objects in complex unstructured environments, thereby solving the problems of insufficient perception and insufficient flexibility and versatility in the existing technology. The specific steps are as follows:

[0084] S1. Generate the 3D pose and outline of the target object using the robot arm pose and RGB-D visual content.

[0085] As a preferred embodiment of the present invention, the purpose of this step is to accurately obtain the position, posture and geometric shape of the target object in three-dimensional space. The specific implementation method is as follows:

[0086] S11, using an RGB-D camera (such as Intel RealSense D435) to obtain a color image (RGB) and depth information (D) of the target object. In a preferred embodiment of the present invention, the color image is such as Figure 3 As shown, the depth information is Figure 4 At the same time, the real-time accurate position T of the end effector of the manipulator is obtained through the encoder or inertial navigation unit of the manipulator itself. EE ∈SE(3); where SE(3) represents the three-dimensional special Euclidean group, so T EE It is a set of geometric transformations with rotation R and translation t, which can be regarded as a 6D posture, a 3D position plus a 3D rotation.

[0087] S12, the color image obtained in S11 (such as Figure 3 As shown in Figure 2, deep learning models (such as Mask R-CNN, YOLO-V8, DETR, etc.) are used to detect and segment target objects, thereby obtaining the 2D image contour C of the target object. 2D .

[0088] S13, combined with the depth information obtained in S11 (such as Figure 4 As shown), the 2D image contour C obtained by S12 2D Upgraded to 3D point cloud data P 3D .

[0089] S14, using point cloud processing network (such as PointNet++, DGCNN, etc.) or combined with traditional PnP algorithm, the 3D point cloud data P obtained in S13 3D Estimate the 6D pose of the target object in, It represents the transformation of coordinate system A relative to coordinate system B, that is, the position and orientation of A with B as the reference system. The 6D posture includes three-dimensional position and three-dimensional orientation.

[0090] S15, through the pre-calibrated hand-eye relationship (For Eye-in-Hand configuration), the estimated target object pose in the RGB-D camera coordinate system in S14 Convert to the robot base coordinate system; at the same time, the real-time accurate posture information of the robot in S11 (This can be obtained by obtaining the rotation angle of each joint based on the rotary encoder of the robot arm joint motor and then performing a forward kinematics solution based on the geometric information of the joint-link) and incorporated into the calculation of the target object's pose transformation to further correct and optimize the accuracy of the target object's 3D pose (especially when dealing with visual blind spots or low-texture objects).

[0091] The coordinate system conversion formula is as follows:

[0092]

[0093] in, is the pose of the target object in the robot coordinate system, is the position of the end effector in the robot arm coordinate system, is the pose of the camera in the end effector coordinate system, is the pose of the target object in the camera coordinate system.

[0094] S16, based on the 3D point cloud data obtained by the depth camera in S13 and the 3D pose obtained in S15, reconstruct the geometric model to generate a relatively accurate 3D outline (polygonal mesh model) of the target object 3D .

[0095] S2. Based on the 3D pose and contour of the target object obtained in step S1, the gripper pose space is generated through Flow-Matching.

[0096] As a preferred embodiment of the present invention, the purpose of this step is to generate a series of high-quality and diverse feasible gripper grasping posture candidates for the target object, forming a "gripper posture space" G. The specific implementation method is as follows:

[0097] S21. Build a generative model based on Flow-Matching.

[0098] The model learns from a large number of data sets containing the 3D poses and contours of objects and the corresponding successful gripper poses. Flow-Matching can learn the complex mapping relationship between data distributions, thereby generating a continuous, high-dimensional gripper grasping pose distribution based on the 3D geometric information of the input object. Specifically, let the 3D geometric information of the input object be The grasping posture is G={P g ,R g ,O g}, where P g is the three-dimensional position of the claw, R g It is the claw grip posture, O g The Flow-Matching model learns a continuous mapping from a simple noise distribution P0(G) to a complex target grasping posture distribution P1(G|X), that is, learning a vector field v t (G, t, X), so that the integral trajectory along this vector field can transform P0 into P1.

[0099] S22, using the Flow-Matching model trained in S21, and taking the 3D pose and contour of the target object obtained in S1 as input, generates or samples a series of candidate gripper grasping poses G for the target object. i ∈G, and then the gripper posture space is obtained. Among them, each candidate gripper grasping posture includes the three-dimensional position, posture (orientation) and clamping opening of the gripper; the sampling process can be expressed as G iFlowMatchingSample(X).

[0100] During the above generation process, the model implicitly or explicitly considers the physical limitations and characteristics of the gripper (such as the EFGC80 gripper), such as its clamping range, clamping method (parallel clamping), and maximum clamping force, to ensure that the generated posture is physically feasible.

[0101] The generative capabilities of the Flow-Matching model enable it to generate reasonable grasping postures for objects of various shapes and sizes, including unknown objects that have never appeared in the training data, thereby improving the versatility and generalization ability of the grasping solution.

[0102] S3. Construct a simulation test environment through the simulation environment. After simulation testing, search for the best grasping posture from the gripper posture space obtained in step S2.

[0103] As a preferred embodiment of the present invention, the purpose of this step is to conduct rigorous virtual verification and evaluation of the generated gripper posture before physical execution, and to select the best grasping posture that is the safest, most robust, and has the highest success rate. The specific implementation method is as follows:

[0104] S31. Using the selected physical simulator (such as MuJoCo), import the MJCF model of the robot model (Elfin5 robotic arm, EFGC80 gripper) to ensure that the model's physical parameters (mass, inertia, joint limits) are consistent with the real hardware; then import the 3D model of the target object and the environment model (desktop, target object, obstacles) into the physical simulator to simulate the necessary physical interactions (including collision, friction, elasticity) between the robotic arm gripper and the target object, and between the target object and the environment, to build a simulation test environment. In a preferred embodiment of the present invention, the constructed simulation test environment is as follows: Figure 5 As shown;

[0105] S32, for each candidate gripper grasping pose G in the gripper pose space g generated by S2, i , iterative evaluation is performed in the simulation test environment constructed by S31 to assess the feasibility of grasping.

[0106] It should be noted that the above method for evaluating the feasibility of crawling is as follows:

[0107] S321, perform self-collision detection between the robot arm and itself and environmental collision detection between the robot arm and surrounding obstacles in the physical simulator to eliminate postures that may cause collision. Introduce collision score S collision (G i )∈{0,1} is used as one of the evaluation indicators, where 1 indicates no collision and 0 indicates collision.

[0108] S322, solve the inverse kinematics (IK) of the manipulator for each candidate gripper grasping posture in the physical simulator to determine whether the manipulator can reach the target object with this posture; introduce the reachability score S reachable (G i )∈{0,1} is used as one of the evaluation indicators, where 1 indicates reachable and 0 indicates unreachable.

[0109] S323. Simulate grasping actions in a physical simulator to evaluate the contact force distribution, friction, and overall force closure between the gripper and the target object; introduce force closure metrics Or the simulated grasping success rate is used as one of the evaluation indicators. In the physical simulator, the stability of the grasping is evaluated by simulating the disturbance imposed on the target object after the claw grasps it; the stability score S is introduced stability (G i )∈{0,1} is used as one of the evaluation indicators.

[0110] S33, according to the evaluation index in S32, the candidate gripper grasping posture G i Assign an overall score S total (G i ).

[0111] The comprehensive score can be expressed as the weighted sum of each score. The specific formula is:

[0112]

[0113] Among them, w1, w2, w3, and w4 are corresponding weights;

[0114] Using heuristic algorithms such as particle swarm optimization (PSO) or genetic algorithm, an efficient search is performed in the gripper pose space of S2, and the grasping pose with the highest comprehensive score is selected as the final optimal grasping pose G best .in In this step, through interaction with the simulation environment, the scores of each item are estimated to obtain a comprehensive score, thereby finding the optimal solution.

[0115] S4. In order to more accurately evaluate the closure and stability of the grasping force, especially after the introduction of electronic skin perception, this step trains the model through supervised learning so that the contact information predicted in the simulation environment can be better aligned with the actual perception results of the physical hardware (electronic skin). This step aims to elaborate on a data alignment method based on supervised learning to bridge the differences in tactile perception, especially contact mechanics, between the robot simulation environment and the physical hardware. This method achieves accurate mapping from simulation data to physical electronic skin contact force data by constructing a pair of neural network models (i.e., the first model and the second model), thereby significantly improving the authenticity and reliability of the simulation evaluation, as follows:

[0116] Through a data alignment method based on supervised learning, the results obtained in step S3 in the simulation test environment are aligned with the actual perception results of the physical electronic skin, so as to improve the authenticity and reliability of the simulation evaluation, thereby realizing the precise grasping of objects by the intelligent robotic arm.

[0117] As a preferred embodiment of the present invention, this step is to train the proposed supervised learning model, which requires synchronously collecting multimodal data in the simulation environment and the physical environment, and performing corresponding preprocessing.

[0118] Among them, a data alignment method based on supervised learning is as follows:

[0119] S41, data collection and preprocessing, collecting simulation environment data and physical environment data respectively, and preprocessing the collected data.

[0120] It should be noted that data collection is divided into simulation environment data collection and physical environment data collection:

[0121] B1) Simulation environment data acquisition involves building an accurate robot model (including the robotic arm and end effector, such as the Elfin5 robotic arm and EFGC80 gripper) and a variety of target object 3D models and environment models in a high-performance physical simulator (such as MuJoCo). By simulating the robotic arm to perform grasping operations on the target object, the gripper position, object position, object properties, gripper closure degree, and collision impulse of the contact parts in the simulation test environment are systematically collected.

[0122] The details are as follows:

[0123] Gripper Pose: The real-time 3D position and posture of the robot end effector (gripper) relative to the base coordinate system, expressed as a homogeneous transformation matrix T gripper ∈SE(3). Considering the large numerical span of the three-dimensional position component, it is necessary to standardize it (i.e., preprocess it). Generally, the following steps can be used: 1. Subtract the mean value; 2. Divide by the standard deviation.

[0124] Object Pose: The real-time 3D position and posture of the grasped target object in the simulation environment, expressed as T object ∈SE(3). Considering the large numerical span of the three-dimensional position component, it is necessary to standardize it. Generally, the following steps can be adopted: 1. Subtract the mean value; 2. Divide by the standard deviation.

[0125] Object Properties: Physical properties of the target object, including but not limited to material (encoded as a one-hot vector), mass, and geometry. Preprocessing: One-hot encode material properties, normalize mass values, and convert geometry information into voxel information.

[0126] Gripper Aperture: The real-time opening value of the gripper during the gripping process. Preprocessing: Perform numerical normalization.

[0127] Contact Impulses: The simulator provides information about the contact area between the gripper and the object. These impulse data are projected and aggregated onto a predefined two-dimensional grid, simulating the layout of the electronic skin sensor to form a discrete collision impulse matrix I. collision ∈R M×N , where M×N is the virtual resolution of the electronic skin. Preprocessing: Normalize the element values ​​of the collision impulse matrix.

[0128] B2) Physical environment data acquisition involves performing actual grasping tests on physical objects corresponding to those in the simulated test environment using a real robotic gripper equipped with an electronic skin (e-skin) sensor. To ensure effective data alignment, the physical experimental conditions must be as consistent as possible with the simulation settings. The gripper pose, object pose, object properties, gripper closure, and e-skin contact force matrix under the physical experimental conditions are simultaneously collected.

[0129] The details are as follows:

[0130] Gripper Pose: The actual gripper pose T obtained by the robot arm encoder or a high-precision external tracking system ′ gripper ∈SE(3). Considering the large numerical span of the three-dimensional position component, it is necessary to standardize it. Generally, the following steps can be adopted: 1. Subtract the mean value; 2. Divide by the standard deviation.

[0131] Object Pose: The actual object pose T′ obtained by an external visual positioning system or high-precision measurement equipment object ∈SE(3). Considering the large numerical span of the three-dimensional position component, it is necessary to standardize it. Generally, the following steps can be adopted: 1. Subtract the mean value; 2. Divide by the standard deviation.

[0132] Object Properties: The physical properties of the object are consistent with those in the simulation. Object property data is preprocessed in the same way as in the simulation.

[0133] Gripper Aperture: The actual degree of gripper opening. Preprocessing: Normalize the values.

[0134] E-skin Contact Force Matrix: The contact force data read in real time from the electronic skin sensor array. This data is usually output in the form of a two-dimensional matrix, representing the pressure or force value of different sensing units on the claw surface, denoted as F e-skin ∈R M×N . Preprocessing: Normalize the pressure values ​​of the perception unit.

[0135] S42, supervised learning model training, based on the results obtained in S41, the first model is used to realize the mapping from high-level captured information to electronic skin contact force, and the first model is used to assist the second model in realizing the mapping from simulated collision impulse to electronic skin contact force (since it is difficult to obtain training data for the second model, a large amount of data can be generated with the help of the first model to provide to the second model for training).

[0136] This step includes two neural network models, whose general architecture is as follows Figure 2 As shown, it aims to achieve accurate mapping from high-level grasping information to electronic skin contact force and from simulated collision impulse to electronic skin contact force.

[0137] It should be noted that the first model is a high-level information to electronic skin contact force prediction network, which aims to learn to predict the actual contact force distribution of electronic skin in the physical world from abstract grasping configuration information.

[0138] The first model above predicts the feature vector X of the network input high This involves feature engineering the gripper pose, object pose, object attributes, and gripper closure information (step B2, physical environment data collection) and concatenating them into a unified input vector. For example, a 6D pose can be represented as a 6- or 7-dimensional vector, and object attributes can be one-hot encoded or embedded.

[0139] The output vector of this prediction network is the target e-skin contact force matrix.

[0140] Specifically, the network architecture of the first model can be divided into:

[0141] Fully connected layer (Feature Embedding): uses a multi-layer perceptron (MLP) structure to map the high-dimensional input vector into a low-dimensional but information-rich latent feature vector V latent These fully connected layers are responsible for capturing the nonlinear relationships between different input modalities.

[0142] Reshape layer: transform the one-dimensional latent feature vector V latent Reshape into a low-resolution two-dimensional latent feature map H∈R h×w×C , where h×w is the initial two-dimensional spatial dimension and C is the number of feature channels. The dimension of the reshape operation needs to match the design of the subsequent transposed convolutional layer.

[0143] Upsampling Layer (Transposed Convolutional Decoder): Utilizing three layers of transposed convolution (also known as deconvolution), the latent feature map H is progressively upsampled to restore the resolution of the target electronic skin contact force matrix, M×N. Each layer of transposed convolution can be followed by an activation function (such as ReLU) and a batch normalization layer to enhance the network's nonlinear representation capabilities and training stability.

[0144] Output layer: The last convolution layer compresses the number of channels to the same as F e-skin The same number of channels (usually 1, representing the pressure value), and optionally using Sigmoid or ReLU as the activation function to ensure the physical rationality of the output value (such as non-negativity).

[0145] The training objective is to minimize the predicted contact force matrix Contact force matrix F with real electronic skin e-skin The difference between them is used for training. Commonly used loss functions include Mean Squared Error (MSE)

[0146] The second model is a simulation impulse to electronic skin contact force alignment network. This network is the key to aligning simulation data with physical hardware. It learns to convert collision impulses in the simulation environment into contact forces that match the real electronic skin perception.

[0147] The second model predicts that the input feature vector of the network is the collision impulse matrix I in the simulation environment collision ∈R M×N , the output vector is the target electronic skin contact force matrix F e-skin ∈R M×N .

[0148] The second model uses a UNet-like convolutional neural network architecture, which is good at processing image-to-image mapping tasks, can effectively capture local features and fuse multi-scale information. Specifically, the network architecture of the second model can be divided into:

[0149] Encoder (Downsampling Path): Consists of three layers of convolutional blocks, each containing one or more convolutional layers (e.g., 3×3 kernels), an activation function (ReLU), and batch normalization. Each convolutional block is followed by a downsampling operation (such as max pooling or strided convolution), gradually reducing the feature map resolution while increasing the number of feature channels to extract multi-scale semantic features. Global average pooling or a fully connected layer can be introduced after the deepest encoder feature layer to capture global context.

[0150] The decoder (Upsampling Path) consists of three layers of transposed convolutional layers, which progressively upsample the features extracted by the encoder back to the original input resolution. Feature maps of corresponding sizes from the encoder path are directly concatenated to the input of the decoder path. This ensures that high-resolution local detail information is passed directly from the encoder to the decoder, effectively preserving spatial details and alleviating information bottlenecks. The concatenated feature maps are then fused and compressed in the channel dimension through additional convolutional layers.

[0151] Output layer: The last convolution layer compresses the number of channels to the number of channels of the target contact force matrix and optionally uses an activation function.

[0152] The training data set of this network consists of the collision impulse matrix I collected in the simulation collision And the electronic skin contact force matrix F collected in physics e-skin The training objective is to minimize the predicted contact force matrix With real F e-skin The difference between them, the loss function is

[0153] S43, application in simulation evaluation, integrates the second model in S42 into the grasping feasibility evaluation module to improve the accuracy of grasping force closure and stability evaluation in the simulation test environment.

[0154] As a preferred embodiment of the present invention, the second neural network (i.e., the second model) trained in the above steps will be integrated into the grasping feasibility assessment module to improve the accuracy of grasping force closure and stability assessment in the simulation environment. Through this data-driven alignment method, this step can effectively bridge the gap in contact perception between the simulation environment and the physical world, providing a solid foundation for achieving high-precision and high-robustness robotic grasping. The details are as follows:

[0155] S431, using the second model, for each candidate gripper grasping posture G generated in the simulation test environment i , simulate its grasping action and extract the collision impulse matrix I between the claw and the target object collision (G i );

[0156] S432, the collision impulse matrix I obtained in S431 under the simulation test environment collision (G i ) as input and fed into the trained second model (simulation impulse to electronic skin contact force alignment network), the neural network will output a predicted electronic skin contact force matrix Based on this more realistic predicted contact force matrix, the gripping force closure can be calculated. and stability S stability (G i ) for more detailed and accurate quantitative assessments. For example, the uniformity of force distribution, points of maximum pressure, and the presence of potential slip areas can be analyzed (by monitoring changes in the predicted force matrix with small perturbations).

[0157] The present invention achieves accurate generation of the 3D pose and contour of objects by: 1) combining the robot arm pose information and RGB-D visual content, thus overcoming the limitations of traditional vision in the perception of complex environments and objects of specific materials; 2) introducing an advanced generative model (Flow-Matching) to intelligently generate a high-quality gripper pose space based on the 3D pose and contour of the object, thereby improving the versatility and diversity of grasping postures; 3) utilizing an efficient and real-time simulation test environment, combining the physical perception capabilities of electronic skin with the data alignment method of supervised learning, to virtually verify and optimize the generated gripper pose space, thereby searching for the optimal grasping posture, effectively reducing the risk of real machine deployment, and improving the grasping success rate and safety. In particular, it improves the assessment accuracy of grasping force closure and stability, thus bridging the gap between simulation and real contact perception.

[0158] The embodiment described above is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Persons skilled in the art may make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, any technical solution obtained by equivalent substitution or equivalent transformation falls within the scope of protection of the present invention.

Claims

1. A method for accurately grasping an object by an intelligent robotic arm, characterized in that: The details are as follows: S1. Generate the 3D pose and outline of the target object using the robot arm pose and RGB-D visual content; S2, based on the results obtained in S1, generate the gripper pose space through Flow-Matching; S3, constructing a simulation test environment through the simulation environment, and obtaining the optimal grasping posture from the gripper posture space obtained in S2 through simulation testing; S4: Using a supervised learning-based data alignment method, the results obtained in S3 in the simulation test environment are aligned with the actual perception results of the physical electronic skin to improve the authenticity and reliability of the simulation evaluation, thereby achieving precise grasping of objects by the intelligent robotic arm; A data alignment method based on supervised learning is as follows: S41, respectively collecting simulation environment data and physical environment data, and preprocessing the collected data; S42. Based on the result obtained in S41, mapping from high-level grasping information to electronic skin contact force is achieved through the first model, and mapping from simulated collision impulse to electronic skin contact force is achieved through the assistance of the second model using the first model; S43. Integrate the second model in S42 into the grasping feasibility evaluation module to improve the accuracy of grasping force closure and stability evaluation in the simulation test environment.

2. The method for accurately grasping an object with an intelligent robotic arm according to claim 1, characterized in that: The S1 is specifically as follows: S11. Use the RGB-D camera to obtain the color image and depth information of the target object. At the same time, use the robot arm's own encoder or inertial navigation unit to obtain the real-time and accurate position and posture of the robot arm's end effector. S12, applying a deep learning model to the color image obtained in S11 to detect and segment the target object, thereby obtaining a 2D image contour of the target object; S13, combining the depth information obtained in S11, and upgrading the 2D image contour obtained in S12 to 3D point cloud data; S14, using a point cloud processing network or in combination with a PnP algorithm, estimating a 6D pose of the target object from the 3D point cloud data obtained in S13; the 6D pose includes a 3D position and a 3D orientation; S15: Using the pre-calibrated hand-eye relationship, the target object pose estimated in the RGB-D camera coordinate system in S14 is transformed into the robotic arm base coordinate system. At the same time, the real-time and accurate pose information of the robotic arm in S11 is incorporated into the calculation of the target object pose transformation to further correct and optimize the accuracy of the target object's 3D pose. S16. Based on the 3D point cloud data obtained in S13 and the result obtained in S15, a relatively accurate 3D outline of the target object is generated.

3. The method for accurately grasping an object with an intelligent robotic arm according to claim 2, characterized in that: The RGB-D camera is Intel RealSense D435, the deep learning model is one of Mask R-CNN, YOLO-V8, and DETR, and the point cloud processing network is PointNet++ or DGCNN; In S15, the coordinate system conversion formula is as follows: in, is the pose of the target object in the robot coordinate system, is the position of the end effector in the robot arm coordinate system, is the pose of the camera in the end effector coordinate system, is the pose of the target object in the camera coordinate system.

4. The method for accurately grasping an object with an intelligent robotic arm according to claim 1, characterized in that: The S2 is specifically as follows: S21. Build a generative model based on Flow-Matching; S22. Using the Flow-Matching model trained in S21, and taking the 3D pose and contour of the target object obtained in S1 as input, generate or sample a series of candidate gripper grasping poses for the target object, and then obtain the gripper pose space; wherein each candidate gripper grasping pose includes the three-dimensional position, posture and clamping opening of the gripper.

5. The method for accurately grasping an object with an intelligent robotic arm according to claim 4, characterized in that: The S21 is specifically as follows: Through Flow-Matching learning, a dataset containing the 3D pose and contour of the target object and the corresponding successful grasping gripper pose is generated based on the 3D geometric information of the input target object to generate a continuous, high-dimensional gripper grasping pose spatial distribution.

6. The method for accurately grasping an object with an intelligent robotic arm according to claim 1, characterized in that: The S3 is as follows: S31. Using the selected physical simulator, import the MJCF model of the robotic arm to ensure that the physical parameters of the model are consistent with the real hardware. Then, import the 3D model of the target object and the environment model into the physical simulator to simulate the necessary physical interactions between the robotic arm gripper and the target object, and between the target object and the environment, thereby constructing a simulation test environment. S32, for each candidate gripper grasping pose in the gripper pose space generated by S2, performing an iterative evaluation in the simulation test environment constructed by S31 to evaluate the feasibility of the grasping; S33. According to the evaluation index in S32, a comprehensive score is assigned to the candidate gripper grasping posture; a particle swarm optimization algorithm or a genetic algorithm is used to search in the gripper posture space of S2, and the grasping posture with the highest comprehensive score is selected as the final optimal grasping posture.

7. The method for accurately grasping an object with an intelligent robotic arm according to claim 6, characterized in that: The S32 is specifically as follows: S321. Perform self-collision detection between the robot arm and itself and environmental collision detection between the robot arm and surrounding obstacles in the physical simulator to eliminate postures that may cause collision; introduce a collision score as one of the evaluation indicators; S322. Performing inverse kinematics of the manipulator for each candidate gripper grasping posture in a physical simulator to determine whether the manipulator can reach the target object with the posture; introducing a reachability score as one of the evaluation indicators; S323. Simulate a grasping action in a physical simulator to evaluate the contact force distribution, friction, and overall force closure between the gripper and the target object; introduce force closure measurement or simulated grasping success rate as one of the evaluation indicators; The stability of the grasping process is evaluated in a physical simulator by simulating the disturbances imposed on the target object by the gripper; the stability score is introduced as one of the evaluation indicators.

8. The method for accurately grasping an object with an intelligent robotic arm according to claim 1, characterized in that: In the S41, Simulation environment data acquisition involves building a robot model, including the manipulator and end effector, as well as a variety of 3D models of target objects and environmental models, in a high-performance physical simulator. By simulating the manipulator's grasping of the target object, the system systematically collects the gripper's position, object position, object properties, gripper closure, and collision impulses of the contact parts in the simulation test environment. Physical environment data acquisition refers to performing actual grasping tests on physical objects corresponding to the simulated test environment on a real robotic arm gripper equipped with an electronic skin sensor, and simultaneously collecting the gripper posture, object posture, object properties, gripper closure degree and electronic skin contact force matrix under physical experimental conditions.

9. The method for accurately grasping an object with an intelligent robotic arm according to claim 8, characterized in that: In the above S42, The first model input feature vector is a unified input vector that is obtained by feature engineering the gripper posture, object posture, object attributes, and gripper closure degree information. The output vector is the target electronic skin contact force matrix. The network architecture of the first model is divided into: The fully connected layer is used to capture the nonlinear relationship between different input modalities. It uses a multi-layer perceptron structure to map high-dimensional input vectors into one-dimensional latent feature vectors. The reshaping layer is used to reshape the one-dimensional latent feature vector obtained by the fully connected layer into a low-resolution two-dimensional latent feature map; The upsampling layer is used to gradually upsample the two-dimensional latent feature map obtained by the reshaping layer using three layers of transposed convolutional layers to restore it to the resolution of the target electronic skin contact force matrix; The output layer is used to compress the number of channels to the same number as the target electronic skin contact force matrix using the last layer of convolution. The activation function can be selectively used to ensure the physical rationality of the output value. The training objective of the first model is to minimize the difference between the predicted contact force matrix and the real electronic skin contact force matrix; The second model input eigenvector is the collision impulse matrix in the simulation environment, and the output vector is the target electronic skin contact force matrix; The second model uses a UNet-like convolutional neural network architecture, which is divided into an encoder, a decoder, and an output layer. The output layer is used to compress the number of channels to the same number as the target electronic skin contact force matrix using the last layer of convolution, and can optionally use an activation function. The training data set of the second model consists of the collision impulse matrix and the electronic skin contact force matrix obtained in S41, and the training goal is to minimize the difference between the predicted contact force matrix and the real electronic skin contact force matrix.

10. The method for accurately grasping an object with an intelligent robotic arm according to claim 1, characterized in that: The S43 is specifically as follows: S431. Using the second model, for each candidate gripper grasping posture generated in the simulation test environment, simulate its grasping action and extract the collision impulse matrix between the gripper and the target object; S432. The collision impulse matrix obtained in the simulation test environment in S431 is used as input to the trained second model. The second model can output a predicted electronic skin contact force matrix that is closer to the actual one.

Citation Information

Cited By

  • Underwater operation robot based on rapid reloading and control method

    CN121553332A