Method for estimating manipulator grasping parameters based on attention mechanism generative network

By introducing a grab parameter estimation method based on an attention mechanism generative network in the robotic arm grasping system, the problem of robotic arm autonomously grasping multiple objects in complex environments is solved, the grasping accuracy and efficiency are improved, and real-time grasping of multiple unknown objects is achieved.

CN114782347BActive Publication Date: 2025-06-24HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210387024.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-13
Publication Date
2025-06-24
Estimated Expiration
2042-04-13

AI Technical Summary

Technical Problem

In complex environments, when the robotic arm automatically grabs multiple chaotic stacked objects, it is difficult to effectively identify and grasp, resulting in inaccuracy and inefficient gripping.

Method used

The robotic arm grab parameter estimation method based on the attention mechanism generative network is adopted to obtain the work scene images through the RGB-D camera, and a two-dimensional image group containing motion instruction vectors is generated using the trained generative neural network to predict the grab quality, angle, width and priority information, thereby determining the best grab parameters.

Benefits of technology

It improves the accuracy and efficiency of the robotic arm in complex environments, enhances the ability to perceive effective information in complex environments, and realizes real-time capture of multiple unknown objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114782347B_ABST
    Figure CN114782347B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for estimating robotic arm grasping parameters based on an attention mechanism generative network. The present invention uses an RGB-D camera to capture corresponding job scene images and inputs them into a trained attention mechanism generative network to obtain robotic arm grasping parameter information such as grasping quality, grasping angle, grasping width, and grasping priority. The remaining grasping parameters are screened through the grasping priority, so as to obtain better robotic arm grasping parameters in a complex multi-object environment. The invention can not only well solve the autonomous grasping ability of the robotic arm in a complex stacking environment, but also further deepen the perception ability of the vision system for effective information in the complex stacking environment through the parameter estimation of the grasping priority, enhance the data processing ability of the entire system for multi-dimensional information, and thus improve the grasping accuracy of the robotic arm in the complex stacking environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of manipulator grasping control, and particularly relates to a method for estimating manipulator grasping parameters based on an attention mechanism generative network. Background Art

[0002] Currently, in the research on manipulator autonomous grasping, the grasping technology for a single object in a simple scenario has been very mature; however, in actual situations, multiple objects are often stacked in a messy and disorderly manner in a complex environment, which brings greater challenges to the manipulator autonomous grasping technology. The present invention proposes a method for estimating manipulator grasping parameters based on an attention mechanism generative network. This method effectively estimates the manipulator grasping parameters through the attention mechanism generative network, deepens the perception ability of the visual system for effective information in a complex environment, improves the fusion ability of multi-channel information, realizes the grasping task of multiple objects in a complex environment, and successfully solves the problem of manipulator autonomous grasping in a multi-object stacking environment. Summary of the Invention

[0003] Aiming at the problem of manipulator autonomous grasping in a complex stacking scenario, the present invention proposes a method for estimating manipulator grasping parameters based on an attention mechanism generative network, thereby improving the grasping accuracy of the manipulator for autonomous grasping in a complex stacking environment.

[0004] In order to achieve the above object, the main technical solutions adopted by the present invention include:

[0005] S1 Use an RGB-D camera to obtain the operation scene image of the manipulator in the current state, including the RGB image I rgb and the depth image I depth and the operation scene reference coordinate system;

[0006] S2 Sequentially input each operation scene image I into a trained generative neural network based on the attention mechanism to generate a two-dimensional image group prediction value containing a motion instruction vector. The two-dimensional image group prediction value contains at least one two-dimensional grasping quality image G θ , a two-dimensional grasping angle image A θ , a two-dimensional grasping width image W θ and a two-dimensional grasping priority image O θ , respectively containing the grasping success rate information, grasping angle information, gripper opening width information and grasping order information when the manipulator grasps an object;

[0007] S3 Sort the pixel values of the two-dimensional grasping quality image G θ in the predicted two-dimensional image group, and select n pixel points with the largest pixel values, which are the predicted values with the highest grasping success rate The pixel coordinates corresponding to the predicted value are mapped to the two-dimensional angular image A θ , the two-dimensional grasping width image W θ and the two-dimensional grasping priority image O θ . Then, the predicted value of the grasping angle can be obtained The predicted value of the grasping width and the predicted value of the grasping order where p n represents the pixel coordinates of the nth pixel point with the largest pixel value in the two-dimensional grasping quality image G θ arranged from large to small;

[0008] S4 sorts the predicted values of the grasping order and selects the predicted value with the highest grasping order priority Then the grasping information corresponding to the pixel coordinates is the optimal motion command vector, that is

[0009] S5 parses the obtained optimal motion command vector to obtain the grasping coordinates, grasping angle, and grasping width of the object to be grasped in the base coordinate system of the robotic arm, that is, the grasping parameters of the robotic arm.

[0010] Preferably, in step S2, it includes:

[0011] S21 trains the generative neural network based on the attention mechanism using the existing dataset;

[0012] S22 preprocesses the operation scene image respectively to obtain an operation scene image of 300×300 pixels;

[0013] S23 inputs the image feature vector into the trained generative neural network based on the attention mechanism;

[0014] S24 outputs the predicted values of the two-dimensional image group containing the motion command vector by the generative neural network based on the attention mechanism, which is the parameter predicted value of the grasping success probability.

[0015] Preferably, in step S3, it includes:

[0016] S31 sorts the pixel values of the two-dimensional grasping quality image in the predicted two-dimensional image group, that is, the size of each pixel value in the two-dimensional grasping quality image represents the grasping success rate of the robotic arm centered on this point;

[0017] S32 selects the coordinates of the n largest grasping success rate predicted values as the grasping center coordinates;

[0018] S33 According to the two-dimensional image group obtained by prediction, obtain the grasping angle pixel value, grasping width pixel value, and grasping priority pixel value corresponding to the grasping center coordinates.

[0019] S34 Parse out the grasping angle, grasping width, and grasping order information according to the grasping angle pixel value, grasping width pixel value, and grasping priority pixel value.

[0020] Preferably, in step S5, it includes:

[0021] S51 Parse the obtained optimal motion instruction vector;

[0022] S52 Perform coordinate transformation on the parsed data through the wrist camera coordinate system;

[0023] S53 Perform coordinate transformation on the coordinates obtained after the camera coordinate system transformation through the robotic arm base coordinate system;

[0024] S54 Input the coordinates obtained after the transformation through the robotic arm base coordinate system, the grasping width, and the grasping angle information into the robotic arm control system for grasping.

[0025] Preferably, the training method of the generative neural network based on the attention mechanism is:

[0026] S01 Create a dataset G for training the network based on the existing dataset; train ; The G train dataset includes working scene images, valid grasping box information, and segmentation images containing only the topmost object. Among them, the segmentation image containing only the topmost object in G train is a two-dimensional grasping priority image;

[0027] S02 Map the valid grasping box information into a 300×300 two-dimensional image to obtain a two-dimensional grasping quality image, a two-dimensional grasping angle image, and a two-dimensional grasping width image. Combine the segmentation image containing only the topmost object in G train to construct a two-dimensional image group;

[0028] S03 Build a generative neural network based on the attention mechanism by constructing an attention mechanism module;

[0029] S04 Use the dataset G train and the two-dimensional image group to train the generative neural network based on the attention mechanism. Input the RGB-D image without grasping information and output the two-dimensional image group containing grasping information to obtain the trained generative neural network based on the attention mechanism.

[0030] Preferably, the generative neural network based on the attention mechanism includes:

[0031] Feature extraction part, attention mechanism part, generation network part;

[0032] Feature extraction part:

[0033] The feature extraction network consists of a convolutional layer with a kernel size of 9×9 and two convolutional layers with a kernel size of 4×4. After each convolutional layer at this stage, there is a Batch Normalization layer and a Rectified Linear Unit activation layer;

[0034] The cropped RGB image I of size 300×300 rgb and the depth image I depth are feature-fused to obtain the fused feature map I fusion , and I fusion is input into the feature extraction network for feature extraction to obtain the feature map I output1 ;

[0035] Attention mechanism part:

[0036] The attention mechanism network consists of five attention modules, and each module is composed of a residual part, a Squeeze part, and an Excitation part;

[0037] The residual part is divided into a direct mapping part and a residual mapping part. The direct mapping is obtained by performing a convolution operation on I output1 using a 1×1 convolutional kernel to obtain the direct mapping result h(I output1 ). The residual mapping consists of two convolutional layers with a kernel size of 3×3. After each of the two convolutional layers, there is a Batch Normalization layer. After the first Batch Normalization layer, there is a Rectified Linear Unit activation layer. After passing through the residual mapping, I output1 obtains R(I output1 );

[0038] The Squeeze part is implemented by introducing Global Average Pooling, and its function is to obtain the global information embedding of each channel of the feature map, that is, the feature vector; assuming u c is a feature map with a size of W×H and C channels, then the Squeeze'd feature map is z c ;

[0039]

[0040] The Excitation part learns the weights of each channel through z c and is composed of the gating mechanism of two fully connected layers; the gating unit sc is a feature vector of size 1×1 and number of channels C, then s c The calculation method of is expressed as:

[0041] s c = F ex (z c , w) = σ(g(z, w)) = σ(w2δ(w1z c )) (2)

[0042] where σ is the sigmoid activation function, δ is the ReLU activation function, γ is the number of nodes in the hidden layer;

[0043] Multiply the obtained s c with u c to obtain the predicted value

[0044]

[0045] Input R(I output1 ) into the Squeeze part and the Excitation part in sequence, and the predicted value after passing through the Squeeze part and the Excitation part can be obtained Input and splice the feature maps obtained from the residual part to obtain the output I of the attention module output2 ;

[0046]

[0047] Input I output1 into five attention modules connected linearly, and the output of the attention mechanism network part can be obtained

[0048] Generation network part:

[0049] The generation network consists of two transposed convolution layers with a convolution kernel size of 4×4 and a transposed convolution layer with a convolution kernel size of 9×9. After each of the two transposed convolution layers with a convolution kernel size of 4×4, there is a Batch Normalization layer and a Rectified Linear Unit activation layer;

[0050] Input into the generation network to obtain the predicted value of the two-dimensional image group containing the motion instruction vector.

[0051] The present invention has the following beneficial effects:

[0052] The present invention proposes a method for estimating robotic arm grasping parameters based on an attention mechanism generative network, which can achieve the grasping of various unknown objects in a complex unstructured environment. Among them, the fusion of multi-information channels such as color images and depth images improves the perception ability of the robotic arm in complex environments; the construction of a lightweight attention generative network ensures the real-time performance of the robotic arm when grasping objects; the establishment of grasping priorities improves the accuracy of effective grasping of the robotic arm. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 is an exemplary diagram of the structure of a generative neural network based on an attention mechanism in an embodiment of the technical solution of the present invention;

[0054] Figure 2 is a framework diagram of a robotic arm grasping parameter learning system based on an attention mechanism generative network in an embodiment of the technical solution of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0055] To better explain the present invention for easy understanding, the following combines the attached Figure 1 As shown, through specific embodiments, the present invention is described in detail. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. The present invention is further described in detail below in combination with specific embodiments:

[0056] S1 Use an RGB-D camera to obtain the operation scene image of the robotic arm in the current state, including the RGB image I rgb and the depth image I depth as well as the operation scene reference coordinate system;

[0057] S2 Sequentially input each operation scene image I into a trained generative neural network based on an attention mechanism to generate a two-dimensional image group prediction value containing a motion instruction vector. The two-dimensional image group includes at least one two-dimensional grasping quality image G θ , one two-dimensional angle image A θ , one two-dimensional grasping width image W θ and one two-dimensional grasping priority image O θ , respectively containing the grasping success rate information, grasping angle information, gripper opening width information and grasping order information when the robotic arm grasps an object;

[0058] The generative neural network based on an attention mechanism includes a feature extraction part, an attention mechanism part, and a generative network part;

[0059] Feature extraction part:

[0060] The feature extraction network consists of a convolutional layer with a convolution kernel size of 9×9 and two convolutional layers with a convolution kernel size of 4×4. After each convolutional layer at this stage, there is a Batch Normalization layer and a Rectified Linear Unit activation layer in sequence.

[0061] The cropped RGB image I with a size of 300×300 rgb and the depth image I depth are subjected to feature fusion to obtain the fused feature map I fusion , and I fusion is input into the feature extraction network for feature extraction to obtain the feature map I output1 .

[0062] Attention mechanism part:

[0063] The attention mechanism network consists of five attention modules, and each module is composed of a residual part, a Squeeze part, and an Excitation part;

[0064] The residual part can be divided into a direct mapping part and a residual mapping part. The direct mapping is obtained by performing convolution operation on I output1 using a 1×1 convolution kernel to obtain the direct mapping result h(I output1 ). The residual mapping is composed of two convolutional layers with a convolution kernel size of 3×3. After each of the two convolutional layers, there is a Batch Normalization layer in sequence. After the first Batch Normalization layer, there is a Rectified Linear Unit activation layer. After passing through the residual mapping, I output1 obtains R(I output1 );

[0065] The Squeeze part is implemented by introducing Global Average Pooling (GAP), and its function is to obtain the global information embedding of each channel of the feature map, that is, the feature vector. Assume that u c is a feature map with a size of W×H and C channels, then the feature map after Squeeze is z c ;

[0066]

[0067] The Excitation part learns the weights of each channel through z c and is composed of the gate mechanism of two fully connected layers. The gating unit s c is a feature vector with a size of 1×1 and C channels, then the calculation method of s c is expressed as:

[0068] s c = F ex (z c , w) = σ(g(z, w)) = σ(w2δ(w1z c )) (2)

[0069] where σ is the sigmoid activation function, δ is the ReLU activation function, γ is the number of nodes in the hidden layer;

[0070] Multiply the obtained s c with u c to obtain the predicted value

[0071]

[0072] Input R(I output1 ) into the Squeeze part and the Excitation part in sequence to obtain the predicted value after passing through the Squeeze part and the Excitation part Input and the feature map obtained from the residual part into a concatenation operation to obtain the output I output2 ;

[0073]

[0074] Input I output1 into five linearly connected attention modules to obtain the output of the attention mechanism network part

[0075] Generation network part:

[0076] The generation network consists of two transposed convolutional layers with a kernel size of 4×4 and one transposed convolutional layer with a kernel size of 9×9. After each of the two transposed convolutional layers with a kernel size of 4×4, there is a Batch Normalization layer and a Rectified Linear Unit activation layer;

[0077] Input into the generation network to obtain the predicted value of the two-dimensional image group containing the motion instruction vector.

[0078] S3 sorts the pixel values of the two-dimensional grasping quality image G θ in the predicted two-dimensional image group, selects ten pixels with the largest pixel values, which are the predicted values with the highest grasping success rate According to the pixel coordinates of the predicted value, map them to the two-dimensional angle image A θ, the two-dimensional grasping width image W θ and the two-dimensional grasping priority image O θ , the grasping angle prediction value can be obtained the grasping width prediction value and the grasping order prediction value

[0079] S4 sorts the grasping order prediction value and selects the prediction value with the highest grasping order priority Then the grasping information corresponding to the pixel point coordinates is the optimal motion command vector, that is

[0080] S5 parses the obtained optimal motion command vector. The parsed data first undergoes coordinate transformation by the wrist camera, then coordinate transformation between the wrist and the base of the robotic arm, and finally obtains the grasping coordinates, grasping angle, and grasping width of the object to be grasped in the base coordinate system of the robotic arm, that is, the robotic arm grasping parameters.

[0081] S6 repeats steps S1 - S5 until all objects are grasped.

[0082] As a specific preference for the implementation of the technical solution of the present invention, before step S2, the method includes:

[0083] S01 creates a dataset G for training the network based on the existing dataset train ; the G train dataset includes working scene images, effective grasping box information, and segmentation images containing only the top - layer objects;

[0084] S02 maps the effective grasping box information into a 300×300 two - dimensional image to obtain a two - dimensional grasping quality image, a two - dimensional grasping angle image, and a two - dimensional grasping width image, and combines with the segmentation image containing only the top - layer objects in G train to construct a two - dimensional image group;

[0085] S03 builds a generative neural network based on the attention mechanism by constructing an attention mechanism module;

[0086] S04 uses the dataset G train and the two - dimensional image group to train the generative neural network based on the attention mechanism, inputs the RGB - D image without grasping information, outputs the two - dimensional image group containing grasping information, and obtains the trained generative neural network based on the attention mechanism;

[0087] As a specific preference for the implementation of the technical solution of the present invention, as Figure 2 shown, a robotic arm grasping parameter learning system based on an attention - mechanism generative network includes:

[0088] Offline learning, through dataset G train Continuously train a generative neural network based on the attention mechanism to obtain a robotic arm grasping parameter prediction model;

[0089] Online learning, through on-site perception of the actual operation scenario, obtain the operation scenario image in the real situation and input it into the robotic arm grasping parameter prediction model to obtain the grasping parameters of the robotic arm in the real scenario, so as to realize the grasping of the robotic arm in the real scenario.

[0090] It should be understood that the above description of the specific embodiments of the present invention is only for explaining the technical route and features of the present invention, and its purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. However, the present invention is not limited to the above specific embodiments. Any changes or modifications made within the scope of the claims of the present invention should be covered by the protection scope of the present invention.

Claims

1. A method for estimating the grasping parameters of a robotic arm based on an attention mechanism generative network, characterized in that Including: S1 Use an RGB-D camera to obtain the operation scene image of the robotic arm in the current state, including the RGB image I rgb and the depth image I depth as well as the reference coordinate system of the operation scene; S2 sequentially inputs each job scenario image I into the trained generative neural network based on the attention mechanism to generate a two-dimensional image group prediction value containing a motion instruction vector, and the two-dimensional image group prediction value includes at least one two-dimensional grasping quality image G θ , a two-dimensional grasping angle image A θ , a two-dimensional grasping width image W θ and a two-dimensional grasping priority image O θ , respectively containing the grasping success rate information, grasping angle information, gripper opening width information and grasping order information when the robotic arm grasps an object; The generative neural network based on the attention mechanism includes: A feature extraction part, an attention mechanism part, and a generation network part; Feature extraction part: The feature extraction network consists of a convolutional layer with a kernel size of 9×9 and two convolutional layers with a kernel size of 4×4. In this stage, each convolutional layer is followed by a Batch Normalization layer and a Rectified Linear Unit activation layer; The cropped RGB image I of size 300×300 rgb and the depth image I depth are feature-fused to obtain the fused feature map I fusion . Then, I fusion is input into the feature extraction network for feature extraction to obtain the feature map I output1 ; Attention mechanism part: The attention mechanism network consists of five attention modules, and each module is composed of a residual part, a Squeeze part, and an Excitation part; The residual part is divided into a direct mapping and a residual mapping. The direct mapping is obtained by convolving I with a 1×1 convolutional kernel output1 to get the direct mapping result h(I output1 ). The residual mapping consists of two convolutional layers with a kernel size of 3×3. After each of the two convolutional layers, there is a Batch Normalization layer immediately following. After the first Batch Normalization layer, there is a Rectified Linear Unit activation layer. After passing through the residual mapping, I output1 obtains R(I output1 ); The Squeeze part is implemented by introducing Global Average Pooling, whose function is to obtain the global information embedding of each channel of the feature map, that is, the feature vector; assuming u c is a feature map with size W×H and C channels, then the feature map after Squeeze is z c ; The Excitation part learns the weights of each channel through z c and is composed of the gate mechanism of two fully connected layers; the gating unit s c is a feature vector of size 1×1 and number of channels C, then s c is calculated as follows: s c = F ex (z c , w) = σ(g(z, w)) = σ(w2δ(w1z c )) (2) where, σ is the sigmoid activation function, δ is the ReLU activation function, γ is the number of nodes in the hidden layer; Multiply the obtained s c by u c to obtain the predicted value Input R(I output1 ) into the Squeeze part and the Excitation part in sequence, and the predicted values after passing through the Squeeze part and the Excitation part can be obtained Input and the feature map obtained from the residual part to splice and obtain the output I of the attention module output2 ; Input I output1 Input it into five linearly connected attention modules, and the output of the attention mechanism network part can be obtained Generation network part: The generation network consists of two transposed convolutional layers with a kernel size of 4×4 and one transposed convolutional layer with a kernel size of 9×9. The two transposed convolutional layers with a kernel size of 4×4 are each followed by a Batch Normalization layer and a Rectified Linear Unit activation layer; Input into the generation network to obtain the predicted values of the two-dimensional image group containing the motion instruction vectors; Input into the generation network to obtain the predicted values of the two-dimensional image group containing the motion instruction vectors; S3 sorts the pixel values of the two-dimensional grasping quality image G in the predicted two-dimensional image group θ in descending order, and selects n pixel points with the largest pixel values, which are the predicted values with the highest grasping success rate According to the pixel point coordinates of the predicted value, it corresponds to the two-dimensional angle image A θ , the two-dimensional grasping width image W θ and the two-dimensional grasping priority image O θ to obtain the predicted grasping angle value the predicted grasping width value and the predicted grasping order value where p n represents the pixel point coordinates of the nth largest pixel value of the two-dimensional grasping quality image G θ in descending order; S4 sorts the predicted values of the grasping order and selects the predicted value with the highest priority of the grasping order Then the grasping information corresponding to the pixel point coordinates is the optimal motion instruction vector, that is S5 parses the obtained optimal motion instruction vector to obtain the grasping coordinates, grasping angle, and grasping width of the target object to be grasped in the base coordinate system of the robotic arm, that is, the grasping parameters of the robotic arm.

2. The robotic arm grasping parameter estimation method based on an attention mechanism generative network according to claim 1, wherein, In step S2, it includes: S21 trains the generative neural network based on the attention mechanism using the existing dataset; S22 preprocesses the operation scene images respectively to obtain operation scene images of 300×300 pixels; S23 inputs the image feature vectors into the trained generative neural network based on the attention mechanism; S24 outputs the two-dimensional image group prediction value containing the motion instruction vector by the generative neural network based on the attention mechanism, which is the parameter prediction value of the grasping success possibility.

3. The robotic arm grasping parameter estimation method based on the attention mechanism generative network according to claim 1 or 2, wherein, In step S3, it includes: S31 sorts the pixel values of the two-dimensional grasping quality image in the predicted two-dimensional image group. That is, the size of each pixel value in the two-dimensional grasping quality image represents the grasping success rate of the robotic arm centered on this point; S32 selects the coordinates of n maximum grasping success rate prediction values as the grasping center coordinates; S33 obtains the grasping angle pixel value, grasping width pixel value, and grasping priority pixel value corresponding to the grasping center coordinates according to the predicted two-dimensional image group, S34 parses the grasping angle, grasping width, and grasping order information according to the grasping angle pixel value, grasping width pixel value, and grasping priority pixel value.

4. The robotic arm grasping parameter estimation method based on the attention mechanism generative network according to any one of claims 1, wherein, In step S5, it includes: S51 parses the obtained optimal motion instruction vector; S52 performs coordinate transformation on the parsed data through the wrist camera coordinate system; S53 performs coordinate transformation on the coordinates obtained after the camera coordinate system transformation through the base coordinate system of the robotic arm; S54 inputs the coordinates obtained after the transformation through the base coordinate system of the robotic arm and the grasping width and grasping angle information into the robotic arm control system for grasping.

5. The method for estimating the robotic arm grasping parameters based on the attention mechanism generative network according to claim 2, characterized in that: The training method of the generative neural network based on the attention mechanism is: S01 creates a dataset G for training a network based on an existing dataset train ; the dataset G train includes working scenario images, valid grasping box information, and segmentation images containing only the topmost object, where the segmentation image train containing only the topmost object in G is a two-dimensional grasping priority image; S02 maps the effective grasping frame information into a 300×300 two-dimensional image to obtain a two-dimensional grasping quality image, a two-dimensional grasping angle image, and a two-dimensional grasping width image, and combines G train The segmentation image containing only the topmost object is used to construct a two-dimensional image group; S03 constructs a generative neural network based on the attention mechanism by building an attention mechanism module; S04 Use dataset G train Train the generative neural network based on the attention mechanism with the two-dimensional image group. Input the RGB-D image without grasping information and output the two-dimensional image group with grasping information to obtain the trained generative neural network based on the attention mechanism.

Citation Information

Patent Citations

  • Residual error network deep learning method for mechanical arm grabbing pose estimation

    CN109934864A

  • Mechanical arm grabbing control method based on machine vision and depth learning

    CN110125930A