A Multi-Scale Fusion Robot Grasping Detection Method Based on Attention Mechanism
Through a multi-scale fusion method based on attention mechanism, a crawling detection model is constructed, which solves the shortcomings in real-time and accuracy of robot crawling detection, and realizes a stable crawling operation in complex scenarios.
Patent Information
- Application Number
- CN202211385821.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-07
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-11-07
AI Technical Summary
The existing robot crawling detection methods are insufficient in real-time and accuracy, making it difficult to achieve stable crawling operations in complex scenarios.
A multi-scale fusion method based on attention mechanism is adopted to build a capture detection model, and features are extracted using super-large convolution modules, residual modules and attention modules, combined with Cornell and Jacquard data sets for training, so as to improve detection accuracy and speed through lightweight network design.
It improves the accuracy and real-timeness of robot grabbing detection, and can quickly and accurately predict the grab position and posture of objects in real scenes, meeting the requirements of real-timeness.
Smart Images

Figure CN115909197B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of image processing, deep learning, and robot grasping control, and particularly relates to a multi-scale fusion robot grasping detection method based on an attention mechanism. Background Art
[0002] Grasp Detection is a technology for obtaining a grasping solution that can be used for actual grasping operations for a specified robot gripper. In home and industrial scenarios, grasping an object from a table is a very important and challenging step for a robot when operating independently or performing a human-robot collaboration task. Generally, robot grasping can be divided into three steps: grasp detection, trajectory planning, and execution. Grasp detection means that the robot obtains the visual information of the target through an RGB or RGBD camera, and then uses this visual information to predict a grasping model to guide the robotic arm and gripper to perform the grasping task.
[0003] The grasping force of robots lags far behind that of humans and is an unsolved problem in the field of robotics. When people see novel objects, they can quickly and easily grasp any unknown object instinctively based on their experience. In recent years, many works related to robot grasping and manipulation have been carried out, but real-time grasp detection remains a challenge.
[0004] Grasp detection methods are mainly divided into two categories: one is the analytical method, and the other is the empirical method. The analytical method refers to defining the grasping pose by designing force closure constraints that meet conditions such as stability and flexibility based on various parameters of the manipulator. This method can be understood as a solution and optimization to a constraint problem based on dynamics and geometry. When the grasping pose satisfies the force closure condition, the object is clamped by the fixture, and under the action of static friction, the object no longer undergoes displacement or rotation, thus maintaining the stability of the grasp. The grasping poses generated by the analytical method can ensure the successful grasping of the target object, but this method is usually only applicable to simple ideal models. The variability of the actual scene, the randomness of object placement, and the noise of the image sensor, etc., on the one hand, increase the computational complexity, and on the other hand, the computational accuracy cannot be guaranteed. The empirical method is to use the information in the knowledge base to detect the grasping pose and judge its rationality. Starting from the characteristics of the object, classification and pose estimation are carried out using similarity, so as to achieve the purpose of grasping. It does not require parameters such as the friction coefficient of the target object as in the analytical method and has better robustness. However, the empirical method usually cannot achieve both accuracy and real-time performance. Summary of the Invention
[0005] In order to overcome the deficiencies of the above-mentioned prior art, the present invention provides a multi-scale fusion robot grasping detection method based on an attention mechanism, which takes into account both the real-time performance and accuracy of robotic arm grasping.
[0006] The technical solution adopted by the present invention is as follows:
[0007] A multi-scale fusion robot grasping detection method based on the attention mechanism, comprising:
[0008] Construct a grasping detection model;
[0009] Collect a grasping dataset, including RGB images and corresponding annotation information, and depth information; perform data augmentation on the dataset, including scale transformation, translation, flipping, and rotation, to expand the dataset, and calibrate the object regions contained in the images;
[0010] Perform data preprocessing on the dataset, including image data preprocessing and annotation parameter preprocessing, and randomly divide the expanded dataset into a training set and a validation set according to a ratio;
[0011] Use the training set data to train the proposed grasping detection model, and adopt the backpropagation algorithm and the optimization algorithm based on the standard gradient to optimize the gradient of the objective function, so as to minimize the difference between the detected grasping box and the true value; at the same time, use the validation set to test the grasping detection model to adjust the learning rate during the training process of the grasping detection model and avoid overfitting of the grasping detection model to a certain extent;
[0012] According to the trained grasping detection model, use the preprocessed real image data as the network input, and the grasping configuration and the five-dimensional representation of the grasping box as the output of the grasping detection model, and finally map them to the real-world coordinates.
[0013] In the above technical solution, the grasping detection model includes a front-end feature extractor and a back-end grasping predictor; among them, the feature extractor includes a super-large convolutional module, a residual and multi-scale module, and an attention module connected in sequence. Specifically, the grasping detection model may include:
[0014] Channel feature extraction layer, using a super-large convolutional kernel and a depthwise separable construction method to extract the features of RGBD respectively, reduce the number of parameters of the model, and fuse the features of the four channels of RGBD; then perform two sparse convolutional downsamplings to further extract features;
[0015] RBF multi-scale receptive layer, composed of several residual modules and RBF modules, using residual modules to avoid the problem of gradient disappearance, and the RBF module uses multi-branch convolutional layers to expand and merge and dilated convolutions of different sizes to simulate the human receptive field;
[0016] Attention encoding layer, adopting a combination of spatial attention and regularization attention, and then performing upsampling to restore the size of the feature map to the input size;
[0017] The grasping generation layer uses the upsampled feature map to obtain the regression result through a multi-branch output mode.
[0018] Further, the grasping dataset may include the currently publicly available Cornell grasping detection dataset and Jacquard grasping detection dataset; and, the categories and contour position information of the items included in the images in the above two datasets are calibrated.
[0019] Further, the image data preprocessing includes cropping the image to intercept the central part of the original data to convert the size of the input image to meet the requirements of the model. Secondly, the RGBD four-channel image data is normalized to accelerate the training of the network. Finally, the normalized RGBD data is stitched together to obtain the final data as the input of the network model.
[0020] Further, the annotation parameter processing includes: the labels of the grasping dataset include a series of grasping poses, and each pose is respectively converted into the form of a rectangular box to describe, that is, the five-dimensional grasping representation {x, y, θ, w, h}, and the label is converted into the form of {G, Θ, W}, where G represents the graspable area, and the part of the center 1 / 3 along the length direction of each rectangle is selected as the encoding of the graspable position, the graspable position encoding is 1, and the non-graspable position encoding is 0; Θ represents the angle of the graspable position. To solve the problem of periodic change of the angle, sin2θ and cos2θ are used to represent the angle, and the graspable positions are respectively encoded as sin2θ and cos2θ; W represents the width of the grasp, and the graspable part is encoded as the width value h of this pose and is normalized to facilitate the convergence of the network; all the label poses are combined to form the final label maps of P, SIN2θ, COS2θ, and W, where P represents the grasp position.
[0021] Further, during the training process, the loss function is defined as:
[0022] L total =L Q +L sin2Θ +L cos2Θ +L width
[0023] where L Q is the grasping quality score loss, L sin2Θ 、L cos2Θ is the angle prediction loss, and L width is the width prediction loss.
[0024] Compared with the prior art, the beneficial effects of the present invention are
[0025] 1. The method of the present invention utilizes the attention mechanism to guide the grasping detector to focus on the features of the target object itself, enabling the robot to predict the most reasonable grasping position of the object based on the category, structure, and texture features of the target object itself.
[0026] 2. In the method of the present invention, the super-large-depth separable convolution can extract the RGBG four-channel global features with a super-large receptive field, and the multi-stage parallel dilated convolutions with different scales are used to extract the feature maps with multi-scale feature fusion. The grasping model can effectively integrate the global and local feature information, effectively improving the detection accuracy and accelerating the training speed at the same time.
[0027] 3. The present invention uses a lightweight network design method to reduce the computational amount while ensuring the accuracy, solving the problem that it is difficult to ensure real-time performance in the real scene for the grasping detection method. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. As shown by the drawings, the above-mentioned and other objects, features, and advantages of the present invention will become clearer. The same reference numerals indicate the same parts in all the drawings. The drawings are not deliberately drawn to scale in actual size, and the focus is on showing the gist of the present invention.
[0029] Figure 1 Schematic diagram of the overall structure of the grasping detection model based on the attention mechanism in the embodiment of the present invention;
[0030] Figure 2 Grasping configuration representation method provided by the present invention;
[0031] Figure 3 Flow chart of a multi-scale fusion robot grasping detection method based on the attention mechanism provided by the present invention;
[0032] Figure 4 Test results of the grasping detection model based on the attention mechanism on actual objects in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0034] As Figure 1 shown in the schematic diagram of the overall structure of the grasping detection model based on the attention mechanism in the embodiment of the present application, this model uses the target feature position and the target feature channel as the attention, which can guide the grasping detection to only focus on the features related to the target grasping; the overall model consists of four parts: the channel feature extraction layer, the multi-scale fusion layer, the attention encoding layer, and the grasping detection generation layer. The channel feature extraction layer uses a super-large convolution kernel and a depthwise separable construction method to extract the features of RGBD respectively, and at the same time obtains a larger receptive field. The depthwise separable design further reduces the number of parameters of the model; the RBF multi-scale receptive layer obtains features of different scales and increases the receptive field by simulating the human receptive field method; the attention encoding layer uses a combination of spatial attention and regularization attention to find the feature regions of the target object and the channels with effective features in the input image, and makes more effective use of the prominent significant features. After two upsamplings, the size of the feature map is restored to the input size. The grasping generation layer uses the upsampled feature map to obtain the regression results P, sin 2θ, cos 2θ, and W through a multi-branch output mode, and finally obtains the grasping position, grasping angle, and grasping width. Specifically, as Figure 1 shown, the grasping detection network in the embodiment of the present invention adopts an encoder-decoder structure, which mainly includes four parts {c1, c2, c3, c4}:
[0035] c1 is the channel feature extraction layer, which uses a super-large convolution kernel and a depthwise separable construction method to extract the features of RGBD respectively. In the DW stage of the depthwise separable convolution, the convolution kernel is set to be larger (set to 31*31 in this embodiment), and a larger convolution kernel is used to perform per-channel convolution on the four channels of RGBD. Next, pointwise convolution is performed on it. Using a large convolution kernel can obtain a larger receptive field and effectively obtain global features; the depthwise separable design further reduces the number of parameters of the model and fuses the features of the four channels of RGBD; after two sparse convolution downsamplings, features are further extracted;
[0036] c2 is the RBF multi-scale receptive layer, which consists of several residual modules and RBF modules. Using the residual module can effectively avoid the problem of gradient disappearance and further extract deep features; the RBF module uses the method of expanding and merging multi-branch convolutional layers and dilated convolutions of different sizes to simulate the human receptive field, which is used to obtain features of different scales and capture information in a larger area, and maintain a low number of parameters;
[0037] c3 is the attention encoding layer, which uses a combination of spatial attention and regularized attention. The feature layer of the previous step passes through two paths. One path passes through maximum pooling and average pooling, one convolution, and then passes through a sigmoid activation function. The other path passes through a layer of BN, then passes through a sigmoid activation function, and finally the two paths are multiplied accordingly. This layer searches for the characteristic areas of the target objects in the input image and the channels with effective features, and makes more effective use of the prominent features. After the c3 layer, it is upsampled twice to restore the size of the feature map to the input size.
[0038] c4 is the grasping generation layer, which uses the upsampled feature map to obtain the regression results P, sin 2θ, cos 2θ and W through the multi-branch output mode. The final result of θ is obtained by the following formula. The final output maps P, θ and W represent the grasping position, grasping angle and grasping width respectively.
[0039]
[0040] The model proposed in this invention adopts Python 3.7 to write the model structure and runs on the Pytorch deep learning framework. The training and verification environment of the present invention is configured under Ubuntu 20.04, the CPU is Intel (R) Xeon (R) CPU E5-2699C v4 @ 2.20GHz, and the GPU is NVIDIA TITAN RTX.
[0041] like Figure 2 The figure shows a schematic diagram of a grasping position representation method of an embodiment of the present application, which is applicable to a parallel plate grasper. Where (x, y) represents the pixel coordinates of the center point of the parallel plate; w represents the size of the opening of the parallel plate; θ represents the angle between the opening direction of the parallel plate and the horizontal direction; and h represents the width of the parallel plate.
[0042] See also Figure 3 , is a flow chart of a multi-scale fusion robot grasping detection method based on an attention mechanism exemplarily shown in an embodiment of the present application, the method comprising the following steps:
[0043] Step S1. Collect the captured data sets (Cornell captured data set and Jacquard captured data set) and preprocess the data sets, the data sets include RGB images and corresponding annotation information and depth information; perform data enhancement by scaling, translation, flipping and rotation on the data sets, and expand the data sets; divide the data sets into training sets and test sets. In this example, the data sets are divided into training sets and test sets in a ratio of 9:1.
[0044] In step S2, data preprocessing operations are performed on the dataset to make the processed data meet the input and output requirements of the model, including specifically the processing of image data and the processing of annotation parameters.
[0045] Among them, the processing of image data includes image cropping, intercepting the central part of the original data to convert the size of the input image into 300*300 to meet the requirements of the model. Secondly, the image data of four channels of RGBD is normalized to accelerate the training of the network. Finally, the normalized RGBD data is spliced to obtain the final data as the input of the network model.
[0046] The processing of labels includes: The labels of the Cornell and Jacquard datasets include a series of grasping poses. Each pose is respectively converted into the form of a rectangular box to describe, that is, a five-dimensional grasping representation {x, y, θ, w, h}. Then, the label is further converted into the form of {G, Θ, W}. Among them, G represents the graspable area. The part of the center 1 / 3 along the length direction of each rectangle is selected as the encoding of the graspable position. The encoding of the graspable position is 1, and the encoding of the non-graspable position is 0; Θ represents the angle of the graspable position. To solve the problem of periodic change of the angle, sin2θ and cos2θ are used to represent the angle, and the graspable positions are respectively encoded as sin2θ and cos2θ; W represents the width of the grasp. The graspable part is encoded as the width value h of this pose and is normalized to facilitate the convergence of the network; all the label pose combinations form the final label maps of P, Sin2θ, Cos2θ, and W.
[0047] In step S3, a grasping detection network model is constructed. The grasping detection network in the present invention adopts an encoder-decoder structure, mainly including four parts. The channel feature extraction layer uses an ultra-large convolutional kernel and a depthwise separable construction method to extract the features of RGBD respectively, and at the same time obtains a larger receptive field. The depthwise separable design further reduces the number of parameters of the model; the RBF multi-scale receptive layer obtains features of different scales and increases the receptive field by simulating the human receptive field; the attention encoding layer adopts a combination of spatial attention and regularization attention to find the feature regions of the target objects in the input image and the channels with effective features, and makes more effective use of the prominent significant features. Then, after two upsamplings, the size of the feature map is restored to the input size; the grasping generation layer uses the upsampled feature map to obtain the regression results P, sin 2θ, cos 2θ, and W through a multi-branch output mode. Finally, the output maps P, Θ, and W are obtained, which represent the grasping position, grasping angle, and grasping width respectively. The Gaussian convolution kernel is used to smooth the output results P, Θ, and W. Finally, the optimal grasping position of the object is determined through the output map result of the grasping position P, and the corresponding grasping angle and grasping width information are obtained from the grasping angle map Θ and width map W according to the current position information, that is, the final inference result is obtained.
[0048] Step S4 uses the Cornell grasping dataset and the Jacquard dataset for model training and testing. The lightweight grasping detection model based on the attention mechanism proposed by the present invention adopts the output results of multiple branches. During the training process, the loss function includes position regression loss, angle regression loss, and width regression loss. The total loss function of the grasping branch of the present invention is defined as follows:
[0049] L total = L Q + L sin2Θ + L cos2Θ + L width
[0050] where L Q is the grasping quality score loss, L sin2Θ , L cos2Θ is the angle prediction loss, and L width is the width prediction loss. L Q , L sin2Θ , L cos2Θ , L width all adopt the smooth l1 loss function, and the smooth l1 loss function is defined as follows:
[0051]
[0052] During the model training process, Adam is used as the optimizer of the model, and the Kaiming method is used to initialize some convolutional layers, so as to optimize the parameters of each layer of the model according to the gradient of the loss. In the present invention, the learning rate of the optimizer is set to 0.001, and the learning rate decay mode of simulated annealing is adopted for training. Step S5 uses the Cornell grasping dataset and the Jacquard grasping dataset to test the model and verify the effectiveness of the present invention. The test index standard is the commonly used rectangular metric index, specifically:
[0053] (1) The difference between the predicted grasping angle and the angle of the label data does not exceed 30°;
[0054] (2) The Jacquard coefficient of the predicted grasping rectangle and the grasping rectangle marked in the dataset is greater than 25%;
[0055] The Jacquard coefficient is as follows:
[0056]
[0057] Among them, B is the grasping rectangle marked in the dataset, which is the ground-truth; A is the grasping rectangle predicted by the model. A∩B is the intersection of the predicted value and the ground-truth; A∪B is the union of the predicted value and the ground-truth.
[0058] In addition, when the distance between the center of the predicted rectangle and the center of the labeled rectangle exceeds 15, the test effectiveness is greatly reduced. By setting a threshold, the efficiency of subsequent tests can be improved.
[0059] Specifically, according to the test results, the accuracies on the Cornell grasping dataset and the Jacquard grasping dataset are 97.73% and 92.79% respectively, and a fast inference speed of 6 ms is achieved with a parameter of 0.7M.
[0060] Step S7 uses the trained grasping detection network to test the actual grasping effect on a real robotic arm. The specific steps include: The preparation link includes camera intrinsic parameter calibration and hand-eye calibration. Further, an RGBD image is obtained, the depth map is aligned with the color map, the depth map is preprocessed, the RGBD data is fused, and the size is cropped to 300*300 and normalized. Further, the obtained image data is input into the trained model to obtain the grasping configuration in the image space and the depth information of the grasping point. The grasping position in the world coordinate system is obtained from the coordinate transformation matrix obtained by hand-eye calibration.
[0061] As Figure 4 shown, the prediction results of the model proposed in the present invention on actual objects. It can be seen from the experimental results that the algorithm proposed in the present invention has higher accuracy and efficiency compared with other algorithms, and shows good results in real scenarios. This scheme
[0062] Utilizes the attention mechanism to guide the grasping detector to focus on the features of the target object itself. Enabling the robot to predict the most reasonable grasping position of the object according to the category, structure and texture features of the target object itself; Adopting a super-large depth separable convolution can extract the RGBG four-channel global features with a super-large receptive field, and a feature map with multi-scale feature fusion can be extracted by multiple levels of parallel dilated convolutions of different scales. The grasping model can effectively integrate the global and local feature information, effectively improving the detection accuracy and at the same time accelerating the training speed; And this scheme uses a lightweight network design method to reduce the computational amount while ensuring the accuracy, solving the problem that it is difficult to ensure real-time performance in real scenarios for grasping detection methods.
[0063] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A multi-scale fusion robot grasping detection method based on an attention mechanism, characterized in that Including: Constructing a grasping detection model; Collecting a grasping dataset, which includes RGB images and corresponding annotation information, as well as depth information; Performing data augmentation on the dataset, including scale transformation, translation, flipping, and rotation, to expand the dataset, and calibrating the object regions contained in the images; Performing data preprocessing on the dataset, including image data preprocessing and annotation parameter preprocessing, and randomly dividing the expanded dataset into a training set and a validation set according to a ratio; Training the proposed grasping detection model using the training set data, and using the backpropagation algorithm and the optimization algorithm based on the standard gradient to optimize the gradient of the objective function, so as to minimize the difference between the detected grasping box and the true value; at the same time, using the validation set to test the grasping detection model to adjust the learning rate during the training process of the grasping detection model; According to the trained grasping detection model, using the preprocessed real image data as the network input, and the grasping configuration and the five-dimensional representation of the grasping box as the output of the grasping detection model, and finally mapping to the real-world coordinates; The grasping detection model includes a front-end feature extractor and a back-end grasping predictor; among them, the feature extractor includes a super-large convolutional module, a residual and multi-scale module, and an attention module connected in sequence; The grasping detection model includes: A channel feature extraction layer, which uses a super-large convolutional kernel and a depthwise separable construction method to extract the features of RGBD respectively, and fuse the features of the four channels of RGBD; and then further extract features through two sparse convolutional downsamplings; An RBF multi-scale receptive layer, which is composed of several residual modules and RBF modules, uses the residual module to avoid the problem of gradient disappearance, and the RBF module uses multi-branch convolutional layers to expand and merge and dilated convolutions of different sizes to simulate the human receptive field; An attention encoding layer, which adopts a combination of spatial attention and regularization attention, and then performs upsampling to restore the size of the feature map to the input size; A grasping generation layer, which uses the upsampled feature map to obtain the regression result through a multi-branch output mode; Annotation parameter processing includes: The labels of the captured dataset include a series of grasping poses, and each pose is respectively converted into the form of a rectangular box for description, that is, the five-dimensional grasping representation {x, y, θ, w, h}, where (x, y) represents the pixel coordinates of the center point of the parallel clamping plate; w represents the opening size of the parallel clamping plate; θ represents the angle between the opening direction of the parallel clamping plate and the horizontal direction; h represents the width of the parallel clamping plate. The label is converted into the form of {G, Θ, W}, where G represents the graspable area, and the middle 1 / 3 part of each rectangle along the length direction is selected as the encoding of the graspable position. The graspable position is encoded as 1, and the non-graspable position is encoded as 0; Θ represents the angle of the graspable position, and sin2θ and cos2θ are used to represent the angle, and the graspable positions are respectively encoded as sin2θ and cos2θ; W represents the width of the grasp, and the graspable part is encoded as the width value h of this pose and normalized; all the label poses are combined to form the final label map, and the label map is {P, sin2θ, cos2θ, W}, where P represents the grasp position.
2. The multi-scale fusion robot grasping detection method based on the attention mechanism according to claim 1, wherein, The captured dataset includes the currently publicly available Cornell Grasp Detection Dataset and Jacquard Grasp Detection Dataset; and, the categories and contour position information of the items included in the images in the above two datasets are calibrated.
3. The multi-scale fusion robot grasping detection method based on the attention mechanism according to claim 1, wherein Image data preprocessing includes cropping the image to intercept the central part of the original data to convert the size of the input image to meet the requirements of the model. Secondly, the RGBD four-channel image data is normalized to accelerate the training of the network. Finally, the normalized RGBD data is spliced to obtain the final data as the input of the network model.
4. The multi-scale fusion robot grasping detection method based on the attention mechanism according to claim 1, characterized in that During the training process, the loss function is defined as: L total = L Q + L sin2Θ + L cos2Θ + L width Among which, L Q is the loss of grasping quality score, L sin2Θ , L cos2Θ is the loss of angle prediction, and L width is the loss of width prediction.
Citation Information
Patent Citations
Robot grabbing detection method based on feature attention mechanism
CN114092766A
Systems and methods for depth estimation using convolutional spatial propagation networks
US20200273192A1