A robot seven-degree-of-freedom grasping method based on generative adversarial learning

By introducing generative adversarial learning and enhanced grasping evaluation methods, combined with a multi-channel associative attention network, the stability and accuracy of robot grasping are improved. This solves the problem of insufficient grasping success rate and robustness of traditional six-degree-of-freedom grasping methods in complex environments, and enables precise grasping of diverse objects.

CN118990500BActive Publication Date: 2025-11-04DONGHUA UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411291302.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-14
Publication Date
2025-11-04
Estimated Expiration
2044-09-14

AI Technical Summary

Technical Problem

Traditional six-degree-of-freedom grasping methods lack stability and accuracy when dealing with objects of complex shapes and diverse materials. Existing methods rely on a single force closure index and ignore the balance of contact point distribution, resulting in low grasping success rate and robustness.

Method used

A seven-DOF grasping method based on generative adversarial learning is adopted, which combines enhanced grasping evaluation method and generative adversarial learning. By comprehensively considering multiple scoring indicators such as contact point distribution balance, centroid distance and contact point non-coplanarity, Hough voting and PnP3D network are used to extract point cloud features, and a discriminator network is added for optimization to improve grasping prediction accuracy and stability.

Benefits of technology

It significantly improves the success rate and robustness of grasping, enabling precise grasping of diverse objects in complex environments, overcoming the limitations of traditional methods, and has broad application prospects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118990500B_ABST
    Figure CN118990500B_ABST
Patent Text Reader

Abstract

The application relates to a robot seven-degree-of-freedom grabbing method based on generative adversarial learning, which comprehensively considers various scoring indexes such as grabbing stability, contact point distribution balance, center of mass distance and non-coplanarity of contact points by introducing an enhanced grabbing evaluation method and generative adversarial learning, and comprehensively evaluates a grabbing pose; a set abstraction layer network based on Hough voting and a PnP3D network feature extractor are combined with a multi-stream channel correlation attention network to finely extract point cloud features and improve grabbing prediction accuracy; the generative adversarial optimization feature extraction result is optimized, so that the model has higher robustness and precision under different environments and object changes, the grabbing success rate is significantly improved, the problems of low grabbing prediction success rate and robustness, low precision and low stability of a traditional six-degree-of-freedom grabbing method are solved, accurate grabbing of diversified objects in a complex environment can be realized, the grabbing pose is closer to the real value, and strong support is provided for the development of intelligent robot grabbing technology.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of robot vision grasping, in particular to a seven-degree-of-freedom grasping method based on generative adversarial learning. BACKGROUND

[0002] In recent years, with the rapid development of robot technology, robot grasping technology has been widely used in industrial automation, logistics sorting, home service and other fields. The core of robot grasping technology is how to realize stable and accurate grasping of objects with different shapes, materials and sizes, which puts high requirements on the perception, planning and control capabilities of robots. In the robot grasping task, the selection of grasping posture and the evaluation of grasping stability are two key problems, which directly affect the success rate of robot grasping and task execution effect.

[0003] Traditional six-degree-of-freedom (6-DoF) grasping method has been widely used in many applications. However, the traditional 6-DoF grasping evaluation method mainly relies on the force closure index, which judges the feasibility of grasping by evaluating whether the contact force between the robot gripper and the object to be grasped can form a stable balance. However, the force closure index may not fully reflect the stability of grasping when dealing with objects with complex shapes and diversified materials, leading to grasping failure. The traditional method often ignores the balance of the distribution of contact points when evaluating the grasping posture, leading to the imbalance of the moment in the grasping process, affecting the stability and safety of grasping. Existing methods mostly only consider the local information of the grasping point, without considering the stress condition and posture change of the object in the whole grasping process, resulting in low success rate and robustness of grasping.

[0004] The traditional grasping prediction method has insufficient generalization ability and robustness when dealing with different environments and diversified objects, and is easily affected by data noise and sample distribution difference. Common methods rely on static feature extraction and scoring mechanism, which cannot dynamically adapt to the changes of the data set, resulting in low precision and stability of grasping prediction. SUMMARY

[0005] In view of the problems of low success rate and robustness, low precision and stability of traditional six-degree-of-freedom grasping method, a seven-degree-of-freedom grasping method based on generative adversarial learning is proposed, taking the seven-degree-of-freedom grasping of robot as the research object, introducing the enhanced grasping evaluation method and generative adversarial learning method, and proposing an efficient and stable robot grasping technology, which can realize accurate grasping of diversified objects in complex environment.

[0006] The technical scheme of the present application is:

[0007] A robot seven-degree-of-freedom grasping method based on generative adversarial learning, comprising the following steps:

[0008] S1, acquire and input the original data set containing point cloud data, perform data enhancement on the original data set by enhancing the grasping evaluation method, use the enhanced grasping evaluation method EGE, introduce a new scoring index, comprehensively consider the distribution balance of the contact point and the stability of the grasping, and improve the existing grasping scoring mechanism as a learning tendency;

[0009] S2, use the improved backbone network Backbone based on Hough voting Votenet and a plug-and-play enhanced feature PnP3D to perform point cloud downsampling and feature extraction;

[0010] S3, the grasping feature extraction network is realized by a multilayer perceptron, extracts the sampled point cloud features, the required position of the grasping, and the candidate set of points to be grasped and the feature vector set;

[0011] S4, obtain seven-degree-of-freedom information of six-degree-of-freedom posture and gripper width degree of freedom through a multi-stream channel associated attention network;

[0012] S5, a discriminator network is added between the feature extraction network and the grasping optimization network, and a generative adversarial network structure is formed by discriminating the intermediate features of the feature extraction network and the real labels of the data set optimized by EGE, to obtain a more robust seven-degree-of-freedom grasping prediction network;

[0013] S6, input the RGB-D image of the object to be grasped into the seven-degree-of-freedom grasping prediction network obtained in step S5 to obtain the seven-degree-of-freedom pose information of the object to be grasped;

[0014] S7, calibrate the camera intrinsic parameters, and use the Tsai-Lenz algorithm to calibrate the camera and the robot hand-eye relationship matrix to obtain the camera coordinate system and the robot base coordinate system, map the seven-degree-of-freedom pose information obtained in step 6 from the camera coordinate system to the robot base coordinate system, and deploy it to the robot arm using Moveit to move to the grasping pose and execute object grasping.

[0015] Further, the method specifically comprises the following steps:

[0016] S11, initialize the grasping pose and force closure index in the original data set, obtain the grasping pose and force closure index score fc as initial data;

[0017] S12, calculate the contact point normal vector and the connecting line vector, find the closest contact points c1, c2 of the object to be grabbed and the gripper in a grabbing pose using the K nearest neighbor (KNN) method, select 15 nearby points to fit a curved surface, calculate the first-order partial derivative of the curved surface at the contact points c1, c2 through the fitted curved surface equation, and obtain the normal vectors of the curved surfaces at the two contact points and the connecting line vector of the two contact points Through the formula

[0018]

[0019] Calculate the cosine similarity of the two normal vectors and the connecting line vector as the torque balance index sore s , θ1, θ2 represent the and angles, respectively;

[0020] S13, calculate the centroid distance index, obtain the contact points c1, c2 through the KNN method of S12, calculate the contact point center coordinates C, and obtain the center of gravity coordinates O of the object point cloud by calculating the average coordinates of the point cloud, and use the formula

[0021] score d = ||OC||2

[0022] Calculate the distance between the object to be grabbed and the center of the gripper, which is used as the centroid distance index score d of the object;

[0023] S14, calculate the contact point distribution balance index, obtain the contact point normal vectors through S12, and calculate the contact point distribution balance index; the formula is as follows

[0024]

[0025] score b is the contact point distribution balance index, which is used to represent the balance degree of the distribution of the contact points between the gripper and the object;

[0026] S15, calculate the contact point non-coplanar index, obtain the contact point center coordinates C and the object centroid coordinates O through step S13, and use the formula

[0027]

[0028] Calculate the contact point non-coplanar index score v ;

[0029] S16, comprehensive evaluation of the grabbing pose, use the comprehensive evaluation formula for the original grabbing pose to update the score of each grabbing pose, and the formula is as follows:

[0030] score = a x score + b x score + g x score + m x (score s + b x score d + g x score fc + m x (score b ) v )

[0031] wherein a, b, g, m are weight coefficients, and score represents the comprehensive score of the grasping;

[0032] S21, split the input point cloud data into Cartesian coordinates and feature descriptors, use multi-layer Hough voting-based point set abstraction and feature enhancement, sample and extract features of the point cloud data through the multi-layer Hough voting point set abstraction layer and the PnP3D module to enhance the feature representation, and then perform feature fusion and upsampling: fuse and upsample the features at different levels to obtain the final feature representation;

[0033] S3, the grasping feature extraction network is implemented through a multi-layer perceptron, extracts the features of the sampled point cloud, and grasps the required position and the candidate point set and feature vector set to be grasped;

[0034] S41, select four grasping feature vectors X1, X2, X3, and X4 of different scales as inputs of the multi-stream channel correlation attention network, and obtain a merged feature map X through add feature fusion;

[0035] S42, perform convolution operations on X through a query convolution layer, a key convolution layer, and a value convolution layer, respectively, to obtain projection query features p q , projection key features p k , and a value matrix p v ;

[0036] S43, calculate a similarity matrix mat s

[0037] mat s = p k ·P q T

[0038] S44, perform a max-pooling operation on mat s obtained in S43 and use a Softmax activation function to obtain a channel correlation matrix mat c ;

[0039] S45, perform a dot product operation on the mat c matrix and the projection value features p v , and weight the comprehensive feature map X to calculate an output transition feature map, and the formula is:

[0040] xm = mat c ·p v + ρX

[0041] x m is the transition feature map, mat c is the channel correlation matrix calculated by S44, and ρ is a learnable weight;

[0042] S46, the four different scale grasping feature vectors X1, X2, X3, X4 of S41 are used to calculate the formula

[0043]

[0044] X o represents the output feature map, X k represents X1, X2, X3, X4;

[0045] the output feature vector X o is obtained o After the multi-layer perceptron, the feature vector X o is converted into the required seven-degree-of-freedom grasping pose parameter set (x, y, z, ψ, θ, φ, w) through the hierarchical structure of the network; wherein (x, y, z) represents the Cartesian coordinates of the gripper in the camera coordinate system, (φ, θ, φ) represents the rotation around the x, y, z coordinate axes, and w represents the width of the gripper;

[0046] S5, a discriminator network is added between step S3 and step S4, and the generated data set obtained in step S3 and the real label are trained through the Wasserstein loss function to optimize the overall neural network parameters; specifically including:

[0047] S51, a discriminator network is built through a multi-layer perceptron;

[0048] S52, the intermediate feature vector output by S3 and the real grasping pose label in the data set are input into the discriminator network, and the discriminator outputs the determination value of the false grasping perspective and the determination value of the real grasping perspective;

[0049] S53, the discriminator loss function is defined as:

[0050]

[0051] L D is the loss function of the discriminator D, represents the expected value of the discriminator output when the input x is sampled from the real data distribution P r , represents the input from the generated data distribution P gdenotes the expected value of the discriminator output, λ denotes a parameter that controls the strength of the penalty term in the loss function, denotes the L2 norm of the gradient of the discriminator with respect to the input ;

[0052] S54, define the seven degrees of freedom grasp prediction network loss function:

[0053]

[0054] L G denotes the loss function of the generator G, N cls denotes the number of categories of the classification task, denotes the classification loss function, c i denotes the predicted category of the model, denotes the true category, t denotes the training round, λ1(t) denotes the time weight coefficient for adjusting the contribution of each term in the loss function, N reg denotes the number of terms related to the regression task, denotes the weight of the regression sample selected, 1(·) is the indicator function, which means 1 if true, 0 otherwise, denotes the regression loss function, s ij is the predicted regression value, is the true regression value, λ2 is the regularization weight, denotes the generator weight;

[0055] S55, calculate the loss of the discriminator and the seven degrees of freedom grasp prediction network respectively during model training, and calculate the parameters of each part of the model through back propagation;

[0056] S6, input the RGB-D image of the object to be grasped into the seven degrees of freedom grasp prediction network obtained in step 5 to obtain the seven degrees of freedom pose information of the object to be grasped.

[0057] S7, camera intrinsic calibration is performed, and Tsai-Lenz algorithm is used for camera and robot hand-eye calibration to obtain the hand-eye relationship matrix of the camera coordinate system and the robot base coordinate system. The seven degrees of freedom pose information obtained in step 6 is mapped from the camera coordinate system to the robot base coordinate system, and is deployed to the robot arm by Moveit to move to the grasping pose to execute object grasping.

[0058] Further, in step S11, the original data set uses the public data set GraspNet-1Billion.

[0059] Further, step S55 specifically:

[0060] In each training cycle, adjust the learning rate and batch normalization momentum, set the model to training mode, perform forward propagation and back propagation through the batch data, calculate the loss and update the network parameters, train the discriminator multiple times to ensure the balance of the generator and the discriminator, calculate the total loss including the generator loss and the grasping loss during the training of the generator, and update the generator parameters.

[0061] The present application has the following advantages:

[0062] Compared with the traditional multi-degree-of-freedom grasping method, the present application has significant advantages in grasping stability, contact point distribution balance and grasping prediction accuracy, can effectively overcome the defects of existing methods, and has wide application prospect and important technical significance.

[0063] The traditional six-degree-of-freedom grasping method mainly relies on the force closure index, the evaluation standard is single and the distribution balance of the contact point is easily ignored, resulting in insufficient grasping stability and accuracy. In addition, the feature extraction process is simple, and the generalization ability and robustness of the model are weak, which is difficult to adapt to complex environments and diversified objects. The present application proposes an enhanced robot grasping method, which comprehensively considers grasping stability, contact point distribution balance, centroid distance and contact point non-coplanarity and other scoring indicators by introducing enhanced grasping evaluation method and generative adversarial learning, and comprehensively evaluates the grasping pose; based on the set abstraction layer network of Hough voting and the PnP3D network feature extractor, combined with the multi-stream channel correlation attention network, the point cloud features are extracted in detail to improve the grasping prediction accuracy; through generative adversarial optimization of the feature extraction result, the model has higher robustness and accuracy under different environments and object changes, significantly improves the grasping success rate, overcomes the limitations of traditional methods such as insufficient precision and grasping robustness, and has wide application prospect and important technical significance. BRIEF DESCRIPTION OF DRAWINGS

[0064] Figure 1 is the overall framework schematic diagram of the present application;

[0065] Figure 2 is the actual deployment diagram of the present application;

[0066] Figure 3 is the multi-stream channel correlation attention network model structure diagram of the present application. DETAILED DESCRIPTION

[0067] The present application will be described in detail below in combination with the drawings and specific embodiments. The present embodiment is implemented on the basis of the technical solution of the present application, and gives detailed implementation mode and specific operation process, but the protection scope of the present application is not limited to the following examples.

[0068] The application provides a robot seven-degree-of-freedom grasping method based on generative adversarial learning. By combining an enhanced grasp evaluation method (EGE), a feature extraction network, and a multi-stream channel associated attention network, and introducing a generative adversarial method, the accuracy and stability of grasp prediction are significantly improved. The method considers the balance of contact point distribution and grasp stability, improves the existing grasp scoring mechanism, and ensures that the grasp pose is closer to the true value. In addition, through camera intrinsic calibration and hand-eye calibration, the grasp parameters output by the network are deployed to the actual environment, realizing efficient grasping operation of the robot in a complex environment. This method not only improves the success rate of grasping, but also enhances the robustness of the robot in actual grasping, providing strong support for the development of intelligent robot grasping technology.

[0069] As Figure 1 shown, the embodiment provides a robot seven-degree-of-freedom grasping method based on generative adversarial learning. The method flow is shown in Figure 1 , and the actual deployment is shown in Figure 2 , including the following steps:

[0070] S1, obtain the original data set containing point cloud data through network, actual measurement, etc. and input, and perform data enhancement on the original data set through the enhanced grasp evaluation method. The enhanced grasp evaluation method (EGE) is used to introduce a new scoring index. The EGE method considers the balance of contact point distribution and grasp stability, and improves the existing grasp scoring mechanism as a learning tendency. The EGE specifically includes:

[0071] S11, initialize the grasp pose and force closure index in the original data set, and obtain the grasp pose and force closure index score fc from the original data set as initial data, wherein the original data set uses a public data set GraspNet-1Billion.

[0072] S12, calculate the contact point normal vector and connecting vector, use the K nearest neighbor (KNN) method to find the contact points c1 and c2 closest to the object to be grasped and the robot gripper under a grasp pose, select 15 nearby points to fit a surface, calculate the first-order partial derivative of the surface at the contact points c1 and c2 through the fitted surface equation, and obtain the normal vectors of the surfaces at the two contact points and the connecting vector of the two contact points The cosine similarity of the two normal vectors and the connecting vector is calculated by the formula

[0073]

[0074] as the torque balance index score s , wherein θ1 and θ2 represent and The included angle.

[0075] S13. Calculate the centroid distance index. Using the contact points c1 and c2 obtained through the KNN method in S12, calculate the center coordinates C of the contact points. Then, obtain the centroid coordinates O of the object's point cloud by calculating the average coordinates of the point cloud. Use the formula...

[0076] score d =||OC||2

[0077] Calculate the distance between the object to be grasped and the center of the gripper; this distance is used as the object's centroid distance score. d .

[0078] S14. Calculate the contact point distribution balance index, using the contact point normal vector obtained in S12. Calculate the contact point distribution balance index. The formula is as follows:

[0079]

[0080] score b It is an index for the balance of contact point distribution, used to indicate the degree of uniformity in the distribution of contact points between the gripper and the object.

[0081] S15. Calculate the non-coplanarity index of the contact point. Using the center coordinates C of the contact point obtained in step 13 and the coordinates O of the object's centroid, use the formula...

[0082]

[0083] Calculate the non-coplanarity index score of contact points v .

[0084] S16. Comprehensive evaluation of grasping poses: Apply the comprehensive evaluation formula to the original grasping poses to update the score of each grasping pose. The formula is as follows:

[0085] score = α × score s +β×score d +γ×score fc +μ×(score b ×score v )

[0086] In the formula, α, β, γ, and μ are weighting coefficients, and score represents the overall score for capturing.

[0087] S2. Point cloud downsampling and feature extraction are performed using a backbone network improved based on Hough voting (Votenet) and a plug-and-play augmented features (PnP3D), specifically including:

[0088] S21, split the input point cloud data into Cartesian coordinates and feature descriptors, use multi-layer Hough voting based point set abstraction and feature enhancement, sample and feature extract the point cloud data through a multi-layer Hough voting point set abstraction layer and a PnP3D module to enhance the feature representation, and then perform feature fusion and upsampling: fuse and upsample the features at different levels to obtain the final feature representation.

[0089] S3, the grasping feature extraction network is implemented through a multi-layer perceptron, extracts the point cloud features after sampling, and grasps the required position and candidate grasping point set and feature vector set;

[0090] S4, obtain the six-degree-of-freedom pose and seven-degree-of-freedom information of the gripper width through a multi-stream channel associated attention network, as shown in Figure 3 , specifically comprising:

[0091] S41, select four grasping feature vectors X1, X2, X3 and X4 of different scales as inputs of the multi-stream channel associated attention network, and obtain a merged feature map X through add feature fusion;

[0092] S42, perform convolution operations on X through a query convolution layer, a key convolution layer and a value convolution layer respectively to obtain a projection query feature p q , a projection key feature p k and a value matrix p v ;

[0093] S43, calculate a similarity matrix mat s between the projection query feature and the projection key feature

[0094] mat s = p k · p q T

[0095] S44, perform a max-pooling operation on mat s obtained in S43 and use a Softmax activation function to obtain a channel association matrix mat c ;

[0096] S45, perform a dot product operation on the mat c matrix and the projection value feature p v , and weight the integrated feature map X to calculate the output transition feature map, the formula being:

[0097] x m = mat c · p v + ρX

[0098] x mis the transition feature map, mat c is the channel correlation matrix calculated by S44, and p is a learnable weight;

[0099] S46, the four different scale feature vectors X1, X2, X3, X4 output by S41 are input into the formula

[0100]

[0101] X o represents the output feature map, X k represents X1, X2, X3, X4.

[0102] The output feature vector X o is obtained by passing through a multi-layer perceptron for nonlinear mapping, and gradually converting the feature vector X o through the hierarchical structure of the network into the required seven-degree-of-freedom grasp pose parameter set (x, y, z, ψ, θ, φ, w) (the grasp pose parameters are not explicitly provided at the input, but they are generated at the output part of the neural network / multi-layer perceptron and are constantly updated to the true value through the optimization algorithm); wherein (x, y) represents the Cartesian coordinates of the gripper in the camera coordinate system, (ψ, θ, φ) represents the rotation around the x, y, z coordinate axes, and w represents the width of the gripper.

[0103] S5, in order to further improve the accuracy of the grasp prediction, the present application adds a discriminator network between the feature extraction network and the grasp optimization network, forms a generative adversarial network structure by discriminating the intermediate features of the feature extraction network and the real labels of the data set optimized by EGE, and obtains a more robust seven-degree-of-freedom grasp prediction network;

[0104] A discriminator network is added between step 3 and step 4, and the feature vector set to be grasped obtained in step 3 and the real label are trained through the Wasserstein loss function to optimize the overall neural network parameters. Specifically, it includes:

[0105] S51, a discriminator network is built through a multi-layer perceptron;

[0106] S52, the intermediate feature vector output by S3 and the real grasp pose label in the data set are input into the discriminator network, and the discriminator outputs the determination value of the false grasp perspective and the determination value of the real grasp perspective;

[0107] S53, define the discriminator loss function:

[0108]

[0109] L D is the loss function of the discriminator D, denotes the expected value of the discriminator output when the input x is sampled from the real data distribution P r denotes the expected value of the discriminator output when the input x is sampled from the real data distribution P denotes the input from the generated data distribution P g denotes the expected value of the discriminator output when the input x is sampled from the real data distribution P denotes the L2 norm of the gradient of the discriminator with respect to the input denotes the L2 norm of the gradient of the discriminator with respect to the input

[0110] S54, define the seven-degree-of-freedom grasp prediction network loss function:

[0111]

[0112] L G denotes the loss function of the generator G, B cls denotes the number of classes for the classification task, denotes the classification loss function, c i denotes the model predicted class, denotes the real class, t denotes the training round, and λ1(t) denotes the time weight coefficient for adjusting the contribution of each term in the loss function, N reg denotes the number of terms related to the regression task, denotes the weight of the regression sample selected, 1(·) is the indicator function, which is 1 when 0 otherwise, denotes the regression loss function, s ij is the predicted regression value, is the real regression value, and λ2 is the regularization weight, denotes the generator weight.

[0113] S55, calculate the loss of the discriminator and the seven-degree-of-freedom grasp prediction network respectively during model training, and calculate the parameters of each part of the model through backpropagation.

[0114] In each training cycle, adjust the learning rate and batch normalization momentum, set the model to training mode, perform forward propagation and backpropagation through batch data, calculate the loss and update the network parameters, and train the discriminator multiple times to ensure the balance between the generator and the discriminator. During the training of the generator, the total loss including the generator loss and the grasp loss is calculated, and the generator parameters are updated.

[0115] S6, input the RGB-D image of the object to be grasped into the seven-degree-of-freedom grasp prediction network obtained in step 5 to obtain the seven-degree-of-freedom pose information of the object to be grasped.

[0116] S7, camera intrinsic calibration is performed, and camera and mechanical arm hand-eye calibration is performed using a Tsai-Lenz algorithm (an algorithm for hand-eye calibration) to obtain a camera coordinate system and robot base coordinate system hand-eye relationship matrix, seven-degree-of-freedom pose information obtained in step 6 is mapped from the camera coordinate system to the robot base coordinate system, and Moveit (an open-source and general robot operating system motion framework) is used to deploy to the mechanical arm to move to a to-be-grabbed pose to perform object grabbing.

[0117] The above-described embodiments only express one embodiment of the present application, which is described in a more specific and detailed manner, but should not be understood as limiting the scope of the patent. It should be noted that, for those skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the scope of protection of the present application. Therefore, the scope of protection of the patent of the present application should be subject to the appended claims.

Claims

1. A seven-DOF grasping method for robots based on generative adversarial learning, characterized in that, Includes the following steps: S1. Obtain and input the original dataset containing point cloud data. Perform data augmentation on the original dataset using the augmented grasp evaluation method. Use the augmented grasp evaluation method EGE to introduce a new scoring index. Take into account the distribution balance of contact points and the stability of grasping. Improve the existing grasp scoring mechanism as a learning tendency. S2. Point cloud downsampling and feature extraction are performed using a backbone network improved by Hough Votenet and a plug-and-play enhanced feature PnP3D. S3. The feature extraction network is implemented through a multilayer perceptron to extract the features of the sampled point cloud, capture the required location, and the candidate set of points to be captured and the feature vector set. S4. Obtain the seven-degree-of-freedom information of six-degree-of-freedom pose and gripper width through a multi-channel association attention network; S5. A discriminator network is added between the feature extraction network and the grasping optimization network. By identifying the intermediate features of the feature extraction network and the real labels of the dataset optimized by EGE, a generative adversarial network structure is formed, resulting in a more robust seven-DOF grasping prediction network. S6. Input the RGB-D image of the object to be grasped into the seven-degree-of-freedom grasping prediction network obtained in step S5 to obtain the seven-degree-of-freedom pose information of the object to be grasped. S7. Perform camera intrinsic parameter calibration, and use the Tsai-Lenz algorithm to perform camera and robotic arm hand-eye calibration to obtain the hand-eye relationship matrix between the camera coordinate system and the robot base coordinate system. Map the seven-degree-of-freedom pose information obtained in step S6 from the camera coordinate system to the robot base coordinate system, and use Moveit to deploy the robotic arm to move to the pose to be grasped to perform object grasping.

2. The robot seven-DOF grasping method based on generative adversarial learning according to claim 1, characterized in that, Specifically, the following steps are included: S11. Initialize the grasp pose and force closure index in the original dataset, and obtain the grasp pose and force closure index score from the original dataset. fe , as initial data; S12. Calculate the contact point normal vector and the connecting vector. Use the K-nearest neighbor (KNN) method to find the closest contact points c1 and c2 between the object to be grasped and the robot gripper in a grasping pose. Select 15 nearby points to fit a surface. Calculate the first-order partial derivatives of the surface at contact points c1 and c2 using the fitted surface equation to obtain the normal vectors of the surfaces at the two contact points. and the vector connecting the two contact points Through formula The cosine similarity between two normal vectors and the connecting vector is calculated as a torque balance index score. s θ1 and θ2 respectively represent and The included angle; S13. Calculate the centroid distance index. Using the contact points c1 and c2 obtained through the KNN method in S12, calculate the center coordinates C of the contact points. Then, obtain the centroid coordinates O of the object's point cloud by calculating the average coordinates of the point cloud. Use the formula... score d =||OC||2 Calculate the distance between the object to be grasped and the center of the gripper; this distance is used as the object's centroid distance score. d ; S14. Calculate the contact point distribution balance index, using the contact point normal vector obtained in S12. Calculate the contact point distribution balance index; the formula is as follows: score b It is an index for the balance of contact point distribution, used to indicate the degree of uniformity in the distribution of contact points between the gripper and the object; S15. Calculate the non-coplanarity index of the contact point. Using the center coordinates C of the contact point obtained in step S13 and the coordinates O of the object's centroid, use the formula... Calculate the non-coplanarity index score of contact points v ; S16. Comprehensive evaluation of grasping poses: Apply the comprehensive evaluation formula to the original grasping poses to update the score of each grasping pose. The formula is as follows: score=α×score s +β×score d +γ×score fc +μ×(score b ×score v ) In the formula, α, β, γ, and μ are weighting coefficients, and score represents the overall score for capture. S21. The input point cloud data is split into Cartesian coordinates and feature descriptors. Multi-layer Hough voting-based point set abstraction and feature enhancement are used. The point cloud data is sampled and features are extracted through multi-layer Hough voting point set abstraction layers and PnP3D module to enhance feature representation. Then feature fusion and upsampling are performed: features at different levels are fused and upsampled to obtain the final feature representation. S3. The feature extraction network is implemented through a multilayer perceptron to extract the features of the sampled point cloud, capture the required location, and the candidate set of points to be captured and the feature vector set. S41. Select four different scale feature vectors X1, X2, X3, and X4 as inputs to the multi-channel association attention network, and obtain the merged feature map X by adding feature fusion. S42. By performing convolution operations on X through the query convolutional layer, the key convolutional layer, and the value convolutional layer respectively, the projected query feature p is obtained. q Projection bond feature p k and value matrix p v ; S43. Calculate the similarity matrix mat between the projection query features and the projection key features. s mat s =p k ·p q T S44, mat obtained from S43 s Perform max pooling and use the Softmax activation function to obtain the channel correlation matrix mat. c ; S45, mat c Matrix and projection value characteristics p v Perform a dot product operation and weight the combined feature map X to calculate the output transition feature map, using the following formula: x m =mat c ·p v +ρX x m It is a transition feature map, mat c It is the channel correlation matrix calculated by S44, and ρ is a learnable weight; S46. Using the formula, extract the four feature vectors X1, X2, X3, and X4 from S41 at different scales. X o X represents the output feature map. k Represents X1, X2, X3, X4; The output feature vector X is obtained. o After passing through a multilayer perceptron, nonlinear mapping is performed, and the feature vector X is gradually transformed through the hierarchical structure of the network. o Convert to the required seven-DOF gripping pose parameter set (x, y, z, ψ, θ, φ, w); where (x, y, z) represents the Cartesian coordinates of the gripper in the camera coordinate system, (ψ, θ, φ) represents the rotation about the x, y, z coordinate axes, and w represents the width of the gripper. S5. A discriminator network is added between steps S3 and S4. The feature vector set to be captured obtained in step S3 and the real labels are used for generative adversarial training through the Wasserstein loss function to optimize the overall neural network parameters; specifically including: S51. Construct a discriminator network using a multilayer perceptron; S52. Input the intermediate feature vector output from S3 and the real grasping pose label in the dataset into the discriminator network. The discriminator outputs the judgment value of the fake grasping viewpoint and the judgment value of the real grasping viewpoint. S53. Define the discriminant loss function: L D It is the loss function of the discriminator D. This indicates that when the input x is from the true data distribution P r During mid-sampling, the expected value output by the discriminator, Indicates input From the generated data distribution P g The expected value of the discriminator output, where λ represents a parameter controlling the strength of the penalty term in the loss function. This indicates that the discriminator is related to the input. The L2 norm of the gradient; S54. Define the loss function for a seven-DOF grasping prediction network: L G Let N represent the loss function of the generator G. cls Indicates the number of task categories. Let c represent the classification loss function. i Indicates the model's predicted category. Let N represent the true class, t represent the training epoch, λ1(t) represent the time weight coefficients used to adjust the contributions of each term in the loss function, and N represent the training epochs. reg This indicates the number of terms relevant to the regression task. This represents the weights of the samples selected for regression, and 1(·) is an indicator function, indicating that... If the value is 1, then the value is 0; otherwise, the value is 0. Let s represent the regression loss function. ij These are the predicted regression values. λ is the true regression value, and λ² is the regularization weight. Indicates the generator weights; S55. During model training, the losses of the discriminator and the seven-degree-of-freedom grasping prediction network are calculated separately, and the parameters of each part of the model are calculated through backpropagation. S6. Input the RGB-D image of the object to be grasped into the seven-degree-of-freedom grasping prediction network obtained in step 5 to obtain the seven-degree-of-freedom pose information of the object to be grasped. S7. Perform camera intrinsic parameter calibration, and use the Tsai-Lenz algorithm to perform camera and robotic arm hand-eye calibration to obtain the hand-eye relationship matrix between the camera coordinate system and the robot base coordinate system. Map the seven-degree-of-freedom pose information obtained in step 6 from the camera coordinate system to the robot base coordinate system, and use Moveit to deploy the robotic arm to move to the pose to be grasped to perform object grasping.

3. The robot seven-DOF grasping method based on generative adversarial learning according to claim 2, characterized in that, In step S11, the original dataset used is the publicly available dataset GraspNet-1Billion.

4. The robot seven-DOF grasping method based on generative adversarial learning according to claim 2, characterized in that, Step S55 is as follows: In each training cycle, the learning rate and batch normalized momentum are adjusted, the model is set to training mode, forward and backward propagation is performed using batch data, the loss is calculated and the network parameters are updated, the discriminator is trained multiple times to ensure the balance between the generator and the discriminator, and during the generator training process, the total loss, including the generator loss and the grasping loss, is calculated and the generator parameters are updated.

Citation Information

Patent Citations

  • Multispecific binding molecules having specificity to dystroglycan and laminin-2

    CN110997714A

  • Point disturbance adversarial attack method for three-dimensional target tracking model

    CN113808165A