A Robot Grasping Detection Method and Device for Multiple Backgrounds
By building an adaptive multi-background suppression grab network and CBAM module, the problem of grab feature recognition of robots in changing backgrounds is solved, the accuracy and adaptability of grabbing detection are improved, and efficient motion planning and automated grabbing are achieved.
Patent Information
- Application Number
- CN202411665529.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-20
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2044-11-20
AI Technical Summary
The existing robot crawling technology is difficult to accurately identify and utilize key crawling features in an unstructured environment with variable backgrounds, and the existing network architecture fails to fully utilize feature information at different levels, resulting in low robustness and crawling accuracy and insufficient adaptability.
A self-adaptive multi-background suppression grabbing network is built, and trained through an adaptive multi-background suppression grabbing network, and channel and spatial attention adjustment are combined with the CBAM module to extract key features and ignore irrelevant features to achieve high-quality grasping pose prediction.
It improves the accuracy and robustness of the robot's grasping and detection in complex environments, enhances the generalization ability of unknown objects, and realizes efficient motion planning and automated grasping.
Smart Images

Figure CN119635626B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robot intelligent control, and particularly to a robot grasping detection method and device for multiple scenarios. Background Art
[0002] With the rapid development of robot technology, intelligence has become the core development direction of the new generation of robots. As a key link in the intelligentization process, robot grasping technology is crucial for machining, assembly, and handling tasks. Visual perception, as a non-contact environmental detection method, stands out for its high efficiency, stability, and rich information content. The vision system can provide key information on the shape, pose, and size of the target object. In particular, the application of deep learning technology has greatly improved the robot's recognition accuracy of the target object and the accurate detection ability of the graspable position. For complex working environments with single and multiple objects, the research on robot grasping detection networks constructed by deep learning methods and end-to-end grasping detection algorithms can promote the robot's autonomous learning ability and enhance its adaptability to unknown environments. Therefore, the grasping detection method based on deep learning is of great significance for the robot's detection and recognition ability in unstructured grasping environments and for promoting the intelligent development of robot grasping technology.
[0003] Although the grasping technology based on visual perception has achieved certain development in this field, in the unstructured environment with variable backgrounds, many problems and deficiencies still emerge in the existing technologies. Currently, most of the robot grasping technologies rely on pre-set environmental information and object models, which are inadequate in the complex and changeable working environment and severely limit the autonomous decision-making ability of the robot. Especially in terms of feature extraction, the existing technologies are not sufficient to accurately extract key grasping features in the environment with complex backgrounds. In the unstructured scenarios with background interference, the robot needs to identify and focus on important grasping areas, but the existing technologies often fail to effectively identify and utilize regional features. There are also several problems in the existing grasping network architectures. In addition, the existing technologies often only stay at the level of linear operations when dealing with features, and do not fully utilize the feature information at different levels, resulting in redundancy in information fusion and unable to make substantial contributions to information gain. In the background interference scenarios, due to factors such as light changes, object stacking, and object diversity, the robustness and grasping accuracy of the existing methods are low, and the extracted features may show large inconsistencies at the scale and semantic levels, severely restricting the quality of the fusion value. The existing methods do not fully utilize the complementary advantages of the rich detailed information contained in the low-level features and the powerful semantic information demonstrated by the high-level features, resulting in insufficient adaptability of the existing methods when dealing with target objects with variable sizes and complex shapes. The above problems not only limit the autonomous grasping ability of the robot in complex environments but also seriously affect the overall accuracy and reliability, which are the technical problems that need to be focused on in future research. Summary of the Invention
[0004] In order to solve the problems of grasping detection of unknown objects (including single objects, cluttered stacked objects, and transparent objects) under the interference of grasping backgrounds with different colors (including different lighting environments) in the existing technologies and the low grasping accuracy of the existing grasping network technologies, the embodiments of the present invention provide a robot grasping detection method and device for multiple backgrounds. The technical solutions are as follows:
[0005] On the one hand, a robot grasping detection method for multiple backgrounds is provided. This method is implemented by a robot grasping detection device for multiple backgrounds, and the method includes:
[0006] S1. Obtain a grasping data set;
[0007] S2. Preprocess the grasping data set to obtain a preprocessed grasping data set;
[0008] S3. Construct an adaptive multi-background suppression grasping network;
[0009] S4. Train the adaptive multi-background suppression grasping network according to the preprocessed grasping dataset to obtain a trained adaptive multi-background suppression grasping network;
[0010] S5. Obtain the RGB-D image data of the object to be grasped;
[0011] S6. Input the RGB-D image data of the object to be grasped into the trained adaptive multi-background suppression grasping network to obtain a grasping quality map, a grasping angle map, and a grasping width map;
[0012] S7. Estimate the grasping pose data according to the grasping quality map, the grasping angle map, and the grasping width map; perform motion planning according to the grasping pose data to enable the robot to execute the grasping task.
[0013] On the other hand, a robot grasping detection device for multiple backgrounds is provided. This device is applied to the robot grasping detection method for multiple backgrounds. The device includes:
[0014] A first acquisition unit for acquiring a grasping dataset;
[0015] A preprocessing unit for preprocessing the grasping dataset to obtain a preprocessed grasping dataset;
[0016] A construction unit for constructing an adaptive multi-background suppression grasping network;
[0017] A training unit for training the adaptive multi-background suppression grasping network according to the preprocessed grasping dataset to obtain a trained adaptive multi-background suppression grasping network;
[0018] A second acquisition unit for acquiring the RGB-D image data of the object to be grasped;
[0019] A third acquisition unit for inputting the RGB-D image data of the object to be grasped into the trained adaptive multi-background suppression grasping network to obtain a grasping quality map, a grasping angle map, and a grasping width map;
[0020] An execution unit for estimating the grasping pose data according to the grasping quality map, the grasping angle map, and the grasping width map; performing motion planning according to the grasping pose data to enable the robot to execute the grasping task.
[0021] On the other hand, a robot grasping detection device for multiple backgrounds is provided. The robot grasping detection device for multiple backgrounds includes: a processor; a memory, and a computer-readable instruction is stored on the memory. When the computer-readable instruction is executed by the processor, any one of the methods in the above-mentioned robot grasping detection method for multiple backgrounds is implemented.
[0022] On the other hand, a computer-readable storage medium is provided, in which at least one instruction is stored, and the at least one instruction is loaded and executed by a processor to implement any one of the above-mentioned robot grasping detection methods in multiple scenarios.
[0023] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include:
[0024] In the embodiments of the present invention, first, a grasping dataset is obtained; the grasping dataset is preprocessed; the preprocessed grasping dataset is obtained; secondly, an adaptive multi-background suppression grasping network is constructed; according to the preprocessed grasping dataset, the adaptive multi-background suppression grasping network is trained to obtain a trained adaptive multi-background suppression grasping network; the RGB-D image data of the object to be grasped is obtained; the RGB-D image data of the object to be grasped is input into the trained adaptive multi-background suppression grasping network to obtain a grasping quality map, a grasping angle map, and a grasping width map; finally, according to the grasping quality map, the grasping angle map, and the grasping width map, the grasping pose data is estimated; and according to the grasping pose data, motion planning is performed to enable the robot to execute the grasping task.
[0025] The proposed adaptive multi-background suppression grasping network in the present invention can accurately identify the importance of each channel in the feature map and accordingly weight-adjust the contributions of each channel in the feature map; the adaptive multi-background suppression grasping network focuses on the features crucial for the grasping task and can ignore irrelevant features at the same time, ensuring excellent grasping detection effects in a variety of different backgrounds. The proposed adaptive multi-background suppression grasping network in the present invention solves the problem of gradient disappearance in the network training process, simplifies the learning process, speeds up the convergence rate of the network, and improves the generalization ability of the network to objects at new positions; through the CBAM module set in the present invention, the channel attention and spatial attention can be dynamically adjusted, which can enhance the self-adaptability and complementary advantages of feature extraction and effectively cope with complex and changeable target objects; the adaptive multi-background suppression grasping network can be applied to the grasping detection tasks of objects with different sizes, colors, shapes, and surface textures; the present invention constructs a ROS-based upper robot system platform, which can monitor the motion state of the robot in real time and integrate the grasping detection and motion planning. The integrated design can achieve a high level of automation and intelligence, can detect the state of the robot in real time, and ensure that the robot is in a safe state. Description of the Drawings
[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0027] Figure 1 It is a flowchart of a robot grasping detection method for multiple backgrounds provided by an embodiment of the present invention;
[0028] Figure 2 It is a schematic structural diagram of an adaptive multi-background suppression grasping network architecture provided by an embodiment of the present invention;
[0029] Figure 3 It is a general structure diagram of a residual module provided by an embodiment of the present invention;
[0030] Figure 4 It is a general structure diagram of a SEGCM module provided by an embodiment of the present invention;
[0031] Figure 5 It is a general structure diagram of a CBAM module provided by an embodiment of the present invention;
[0032] Figure 6 It is a robot intelligent grasping system framework provided by an embodiment of the present invention;
[0033] Figure 7 It is a robot grasping process architecture provided by an embodiment of the present invention;
[0034] Figure 8 It is the grasping detection result on the Cornell and Jacquard datasets provided by an embodiment of the present invention;
[0035] Figure 9 It is the grasping detection result in a single-object scene provided by an embodiment of the present invention;
[0036] Figure 10 It is the grasping detection result in a complex multi-object scene provided by an embodiment of the present invention;
[0037] Figure 11 It is the grasping detection result of single-object and multi-object in multiple backgrounds provided by an embodiment of the present invention;
[0038] Figure 12 It is a block diagram of a robot grasping detection device for multiple backgrounds provided by an embodiment of the present invention;
[0039] Figure 13 It is a schematic structural diagram of a robot grasping detection device for multiple backgrounds provided by an embodiment of the present invention. Detailed Implementation Manner
[0040] The technical solutions in the present invention will be described below with reference to the accompanying drawings.
[0041] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of the word "example" is intended to present concepts in a specific manner. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.
[0042] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same. "of", "corresponding", and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same.
[0043] In the embodiments of the present invention, sometimes subscripts such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meanings they express are the same.
[0044] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0045] The embodiments of the present invention provide a robot grasping detection method for multiple backgrounds. This method can be implemented by a robot grasping detection device for multiple backgrounds, and the robot grasping detection device for multiple backgrounds can be a terminal or a server. As Figure 1 shown in the flowchart of the robot grasping detection method for multiple backgrounds, the processing flow of this method can include the following steps:
[0046] S1. Obtain a grasping data set.
[0047] Among them, the grasping data sets adopted in the embodiments of the present invention are the Cornell data set and the Jacquard data set; among them, the Cornell data set contains 885 images, and is augmented by performing cropping, rotation and scaling operations on the Cornell data set; among them, the Jacquard data set consists of 11,619 different scenes of 54,485 different objects, with a total of 4,967,454 grasping annotations. For each scene, a rendered RGB image, a segmentation mask, two depth images and grasping annotations are provided.
[0048] S2. Preprocess the grabbed dataset to obtain the preprocessed grabbed dataset.
[0049] Optionally, the preprocessing of the grabbed dataset in S2 includes:
[0050] Performing cleaning processing, data transformation processing, and grabbing annotation processing on the grabbed dataset.
[0051] In a feasible implementation, the preprocessing of the grabbed dataset further includes: decompressing the dataset, viewing and understanding the dataset files. By preprocessing the grabbed dataset, the quality and usability of the data can be ensured.
[0052] In a feasible implementation, the preprocessed grabbed dataset is divided into a training set and a test set in a ratio of 1:9.
[0053] S3. Construct an adaptive multi-background suppression grasping network.
[0054] Among them, the adaptive multi-background suppression grasping network can automatically adjust its parameters and behavior according to different environments and conditions to achieve optimal performance. The adaptive multi-background suppression grasping network can handle a variety of different background environments, rather than just a single or specific background. The adaptive multi-background suppression grasping network has the ability to suppress or reduce background interference, which can improve the accuracy of target recognition and grasping.
[0055] Among them, Figure 2 is a schematic diagram of the structure of the adaptive multi-background suppression grasping network architecture provided by the embodiments of the present invention; the adaptive multi-background suppression grasping network takes a 224×224 RGBD image as input, extracts features from the RGBD image through the adaptive multi-background suppression grasping network, predicts high-quality grasping poses for unknown objects, and outputs the final grasping quality map, grasping angle map, and grasping width map.
[0056] Among them, the value of each point on the grasping quality map represents the grasping confidence at that grasping point, that is, the success probability of the grasping pose, providing a basis for the grasping decision.
[0057] Among them, the value of each point on the grasping angle map represents the angle that the end effector of the robot needs to rotate at that grasping point; the grasping angle map can improve the accuracy of the grasping action.
[0058] Among them, the value of each point on the grasping width map represents the size that the end effector of the robot needs to open at that grasping point; the grasping width map can ensure the adaptability and accuracy of the grasping action.
[0059] Optionally, the adaptive multi-background suppression grasping network includes: a convolutional layer, a transposed convolutional layer, 5 residual modules, a SEGCM module, and a CBAM module.
[0060] Among them, three convolutional layers are used to perform in-depth feature extraction on the RGBD image, effectively extracting and encoding the preliminary features of the image. Through convolutional operations, a feature map containing 128 channels is generated. Among them, a normalization processing operation is added after each convolutional layer, which can optimize the training process and improve the convergence speed of the network.
[0061] In a feasible implementation, three transposed convolutional layers are used to perform size amplification processing on the input feature map. Through upsampling operations, the size of the feature map is gradually restored and finally amplified to the same 224×224 pixels as the original input image.
[0062] Among them, Figure 3 is the overall structure diagram of the residual module provided by the embodiment of the present invention.
[0063] Optionally, the residual module includes: two 3×3 convolutional kernels and the Mish activation function;
[0064] The residual module is used to perform convolutional operations on the input feature map, extract local features, and output a fused feature map.
[0065] In a feasible implementation, the residual module improves the accuracy and efficiency of the network for predicting the grasping pose of unknown objects by solving the gradient vanishing problem, improving the training efficiency, enhancing feature transmission, and constructing a deep network structure.
[0066] Among them, Figure 4 is the overall structure diagram of the SEGCM module provided by the embodiment of the present invention.
[0067] Optionally, the SEGCM module includes: three 1×1 convolutional kernels, 3×3 convolutional kernels, the Mish activation function, the Sigmoid activation function, and a fully connected layer;
[0068] The SEGCM module is used to extract the relationship between channels, optimize the feature representation through global context information, and output key features.
[0069] In a feasible implementation, the SEGCM module can accurately identify and strengthen the key information in the feature map by learning the relationship between channels, ignore the unimportant features, enabling the network to more comprehensively extract the features of the image and adaptively generate high-quality grasping poses under variable grasping backgrounds.
[0070] Among them, Figure 5 is the overall structure diagram of the CBAM module provided by the embodiment of the present invention.
[0071] Optionally, the CBAM module includes: a channel attention module and a spatial attention module;
[0072] Channel attention module, which is used to perform pooling operations on the input features to obtain attention-enhanced features;
[0073] Spatial attention module, which is used to perform pooling operations on the attention-enhanced features to obtain spatial attention features.
[0074] In a feasible implementation, the CBAM module can improve the accuracy and efficiency of grasping detection and enhance the interpretability of the network. The CBAM module adaptively learns channel and spatial attention weights to improve the feature expression ability of the convolutional neural network. By combining channel attention and spatial attention, the CBAM module can capture the correlation between features in different dimensions, thereby improving the real-time performance and accuracy of the grasping detection task.
[0075] Among them, using a deconvolution layer for processing can retain the rich feature information extracted from the convolutional layer, ensure the precise correspondence of features in the spatial dimension, and provide a high-resolution feature map for generating high-quality grasping pose predictions.
[0076] S4. Train the adaptive multi-background suppression grasping network according to the preprocessed grasping dataset to obtain a trained adaptive multi-background suppression grasping network.
[0077] Optionally, the specific implementation process of S4 may include S41 - S45:
[0078] S41. Input the preprocessed grasping dataset into the adaptive multi-background suppression grasping network, and perform feature extraction through the convolutional layer to obtain preliminary features;
[0079] S42. Input the preliminary features into the residual module, and process them through a 3×3 convolutional kernel to obtain local features; according to the residual mechanism, process the local features through normalization, the Mish activation function, and a 3×3 convolutional kernel, and output fused features; the fused features include the preliminary features of the original input and the features processed by the convolutional layer;
[0080] Among them, each kernel in the 3×3 convolutional kernel can cover a 3×3 pixel area of the input feature map, and can effectively extract local features.
[0081] Among them, the normalization process normalizes the input of the normalization layer to make the mean of the output close to 0 and the variance close to 1, thereby stabilizing the training process, accelerating the convergence of the network, and improving the generalization ability of the model.
[0082] Among them, the Mish activation function is a non-monotonic activation function that can improve the non-linear expression ability of the network and can be expressed by the following formula (1):
[0083] (1)
[0084] Among them, f(x) represents the result after being processed by the Mish activation function; x represents the input activation value.
[0085] Among them, after being processed by the Mish activation function, further processing with a 3×3 convolution kernel can further refine the features, and the output has the same dimension as the input features.
[0086] In a feasible implementation manner, the residual module adopts a residual connection mechanism. By directly adding the feature maps after two layers of convolution and activation function processing on the input feature map, the problem of gradient disappearance in the network can be alleviated, and the expression ability of the features is increased.
[0087] S43. Input the fused features into the SEGCM module, and process them through two 1×1 convolution kernels, the Mish activation function, and the Sigmoid activation function to obtain a compressed feature vector; according to the compressed feature vector, process it through a 3×3 convolution kernel, the Mish activation function, the Sigmoid activation function, and a fully connected layer to obtain an attention weight map;
[0088] In a feasible implementation manner, the specific implementation process of the SEGCM module may include S431 - S433:
[0089] S431. Input the fused features into the SEGCM module, perform global average pooling on the input fused features, compress the spatial information of each channel into a single value to obtain a compressed feature vector, and use two 1×1 convolution kernels to perform non - linear transformation on the compressed feature vector. Among them, the Mish activation function is used for processing between the two convolution kernels to increase the ability to extract non - linear features, and through the Sigmoid activation function, generate the attention weight of each channel; multiply the generated attention weight of each channel by each channel of the original feature map to obtain a scaled feature map;
[0090] S432. The branched processing of the scaled feature map includes: the first branch and the second branch. Among them, the processing process of the first branch may include:
[0091] (1) Perform global average pooling and global pooling operations on the scaled feature map to obtain the feature vectors of global average pooling and global pooling;
[0092] (2) Process through a 1×1 convolution kernel and the Mish activation function to obtain a processing result;
[0093] (3) Use a 3×3 convolution kernel and the Sigmoid activation function to process the processing result, and output the attention weight of the first branch; among them, the Sigmoid activation function is used to ensure that the output value is between 0 and 1.
[0094] Among them, the processing process of the second branch may include:
[0095] (1) Perform global average pooling on the scaled feature map to obtain the processing result after global average pooling;
[0096] (2) Process the processing result after global average pooling through two fully connected layers to obtain the attention weights of the second branch; among them, the Mish activation function is used after the first fully connected layer; the Sigmoid activation function is used after the second fully connected layer.
[0097] S433. Obtain the attention weight map according to the attention weights of the first branch and the second branch; apply the attention weight map to the original feature map before the first branch and the second branch, which can strengthen important features and suppress unimportant features.
[0098] Among them, through the SEGCM module, the network can not only extract the relationships between channels, but also optimize the representation of heat syndromes by using global context information; the SEGCM module enables the network to accurately identify and focus on key features when processing grasping tasks in complex backgrounds, improving the performance and generalization ability of the model.
[0099] S44. Input the attention weight map into the CBAM module, process it through the channel attention module to obtain the channel features after attention weighting; input the channel features after attention weighting into the spatial attention module to obtain the spatial attention features; generate the attention enhanced features according to the spatial attention features and the channel features after attention weighting;
[0100] In a feasible implementation manner, the specific implementation process of the CBAM module may include S441 - S445:
[0101] S441. Input the attention weight map into the channel attention module, perform max - pooling operation on each channel of the attention weight map to generate the global maximum feature vector of each channel; perform average - pooling operation on each channel of the attention weight map to generate the average feature vector of each channel;
[0102] S442. Input the global maximum feature vector of each channel and the average feature vector of each channel into a 7×7 convolutional kernel, learn the attention weights of each channel, and process them through the Sigmoid activation function to obtain the channel attention weights;
[0103] S443. Multiply the channel attention weights with each channel of the original feature map to obtain the channel feature map after attention weighting;
[0104] S444. In the spatial attention module, the channel feature map after attention weighting is input. Global max pooling and average pooling operations are performed on the channel feature map after attention weighting to obtain two feature maps of 56×56×1. The two feature maps are concatenated in the channel dimension and processed through a fully connected layer for dimensionality reduction to obtain a feature map of 56×56×1. Through the Sigmoid activation function, spatial attention features are generated. The spatial attention features are multiplied by the input feature map to obtain the final spatial attention features.
[0105] S445. The channel feature map after attention weighting and the final spatial attention features are multiplied element by element to obtain attention-enhanced features. Among them, the attention-enhanced features can retain key information and suppress noise and irrelevant information.
[0106] S45. The attention-enhanced features are input into a transposed convolutional layer for processing to obtain a grasping quality map, a grasping angle map, and a grasping width map. The adaptive multi-background suppression grasping network is trained according to the grasping quality map, the grasping angle map, and the grasping width map to obtain a trained adaptive multi-background suppression grasping network.
[0107] S5. Obtain the RGB-D image data of the object to be grasped.
[0108] S6. The RGB-D image data of the object to be grasped is input into the trained adaptive multi-background suppression grasping network to obtain a grasping quality map, a grasping angle map, and a grasping width map.
[0109] Among them, the adaptive multi-background suppression grasping network can adapt to various grasping backgrounds, perform high-quality grasping pose prediction on unknown objects, and while maintaining excellent performance, does not increase the number of network parameters. It enables the adaptive multi-background suppression grasping network to achieve fast inference and high-precision grasping under various backgrounds, achieving an ideal balance between high efficiency and precise grasping.
[0110] S7. Estimate the grasping pose data according to the grasping quality map, the grasping angle map, and the grasping width map. Perform motion planning according to the grasping pose data to make the robot execute the grasping task.
[0111] In a feasible implementation manner, the grasping pose data is estimated according to the grasping quality map, the grasping angle map, and the grasping width map. Among them, the grasping quality map represents the grasping confidence at each pixel point, and the grasping confidence represents the grasping success rate at each pixel point. The grasping angle map is the angle that the end effector of the robot needs to rotate. Assuming the angle to be rotated is a, calculate the values of cosa, sina, and tana, and obtain the rotation angle through arctana. The grasping width map is the size that the robot gripper needs to open. By estimating the position of the pixel point with the optimal grasping confidence, the corresponding grasping angle, and the grasping width, the optimal grasping pose data is obtained. Among them, the grasping pose in the camera coordinate system is converted into the grasping pose based on the robot coordinate system.
[0112] In a feasible implementation manner, during the process of the robot performing the grasping task, the RGBD camera installed on the end effector of the robot is responsible for analyzing the RGB-D image to detect the object and generate an accurate grasping pose. The pose information is initially defined in the camera coordinate system and is unknown to the robot. When the camera completes the recognition and positioning of the object position, this pose information is converted from the camera coordinate system to the base coordinate system of the robot, and the robot can accurately know the position of the target object and perform the final grasping action. To achieve this conversion, the embodiment of the present invention uses the handeye-calib package to calibrate the hand-eye system. Through the calibration process, the conversion matrix from the camera coordinate system to the robot base coordinate system can be obtained, enabling the robot to understand the pose information provided by the camera and apply it to the actual grasping task. Through hand-eye calibration, it is ensured that the robot can accurately execute complex operations according to the data provided by the vision system, improving the accuracy and reliability of the robot operation.
[0113] Among them, Figure 6 is the framework of the robot intelligent grasping system provided by the embodiment of the present invention. In a feasible implementation manner, the RGBD camera collects the image data of the grasping scene, preprocesses the image to obtain the preprocessed image data, inputs the preprocessed image data into the trained grasping network based on adaptive multi-background suppression to generate the grasping quality image, the grasping angle image, and the grasping width image, generates the grasping pose data according to the grasping quality image, the grasping angle image, and the grasping width image, performs inverse kinematics planning through the ROS interface to obtain the grasping path of the robot, and the robot performs the grasping action according to the grasping path.
[0114] Among them, Figure 7It is the robot grasping process architecture provided by the embodiments of the present invention; in a feasible implementation, the embodiments of the present invention build a robot host computer grasping system platform, which consists of a host computer, a Baxter robot and an RGBD camera. The Baxter robot uses an open-source robot operating system based on ROS and runs through the Linux platform. Users can connect to the internal computer of the robot through the network to read information or send instructions; the host computer control system uses the Ubuntu20.04 system, and the computer control platform and the physical robot are connected to the same IP through USB and a router. The entire robot interaction system framework is based on the ROS noetic version of the robot operating system. The entire algorithm is programmed in Python. ROS distributes and processes subtasks through independently running functional nodes and completes information transmission between nodes through various communication mechanisms, improving the algorithm reuse rate and writing efficiency. The entire robot intelligent grasping system consists of an adaptive multi-background suppression grasping network grasping pose detection module and a control grasping and inverse kinematics planning module.
[0115] Among them, the adaptive multi-background suppression grasping network grasping pose detection module is used to collect visual information of the grasping scene by using an RGBD camera, and can obtain RGB-D image data of the object to be grasped; the adaptive multi-background suppression grasping network grasping pose detection module can effectively align the missing parts in the depth map to ensure data integrity. After preprocessing the input image, the preprocessed image is input into the adaptive multi-background suppression grasping network to generate a high-quality grasping pose. Among them, the control grasping and inverse kinematics planning module is used to convert the pose coordinates to the robot base coordinate system after the RGBD camera successfully obtains the grasping pose, perform inverse kinematics planning through the ROS interface, obtain the grasping path of the robot, and the robot executes the grasping action according to the grasping path.
[0116] In a feasible implementation, the present invention provides an implementation case which specifically includes: First, connect the robot to the host computer through USB or Ethernet, and connect the RGBD camera to the host computer through USB. Then start the robot driver program to establish communication between the host computer and the Baxter robot body. Once the host computer system successfully establishes a connection with the robot, the Baxter robot can be enabled. After successful enabling, the grasping routine can be executed. While the grasping task is being executed in real time, the host computer can also obtain the joint state and the position information of the robot end effector in real time through ROS, and these information are crucial for monitoring the safety state of the robot.
[0117] In a feasible implementation, to verify the effectiveness of the VecGNet network and its related modules in the present invention, the following steps were taken: The VecGNet network designed in the present invention was trained on the Cornell dataset and the Jacquard dataset respectively. After the training was completed, the performance of the VecGNet network was visually analyzed on these two datasets to intuitively display its learning results. In addition, to further evaluate the performance of the network in practical applications, a robot grasping experiment was conducted in a real-world environment, which not only confirmed the theoretical superiority of the VecGNet network but also demonstrated its feasibility and effectiveness in practical applications. A total of four groups of experiments were conducted to verify the performance of the present invention: grasping visualization on the dataset, single-object grasping in the real world, multi-object grasping in the real world, and grasping in various real-world backgrounds.
[0118] (1) Grasping on the Cornell and Jacquard datasets: As Figure 8 shown are the visual grasps on the Cornell and Jacquard datasets.
[0119] (2) Single-object scene grasping in the real world: In the single-object scene grasping experiment in the real world, a series of common objects in daily life were selected. These objects not only have different sizes and complex shapes but also have rich texture information. Single-object grasping experiments were conducted on these objects to test the actual effectiveness of the present invention. Among the 500 grasping attempts made, only 11 failures occurred, and the grasping success rate was as high as 98%, fully demonstrating the excellent performance of the present invention in dealing with single-object grasping scenarios. As Figure 9 shown, a successful case of grasping detection of a single object in the real-world single-object scene is presented.
[0120] (3) Multi-object scene grasping in the real world: In the complex multi-object environment of the real world, 50 groups of scenes were designed, and each group of scenes contained multiple objects with different sizes, shapes, and colors. The objects in the scenes vary in size, color, shape, and surface texture to ensure the diversity and challenge of the test environment. As Figure 10 shown, the above scenes reflect the complexity of multi-object grasping in the real world. After testing, the average grasping success rate of the present invention in complex scenes reached 95%, which not only confirmed the high efficiency of the proposed grasping detection network but also demonstrated its excellent ability to handle multi-object grasping tasks in the face of complex grasping environments.
[0121] (4)Grasping in different grasping backgrounds in the real world: Under different colored backgrounds and different lighting conditions, the present invention has conducted extensive detection experiments specifically for single-object and multi-object scenarios, including transparent objects. In the experiments, common challenges in the real world, such as different colored backgrounds, low-light and high-light environments, were simulated. The system was tested in 30 groups of carefully designed complex scenarios, where the objects in each group differed in size, color, shape, and surface texture to ensure the diversity and challenge of the test environment. After testing, the average grasping success rate of the present invention in complex scenarios reached 94%, and the detection system of the present invention demonstrated excellent performance. In summary, the detection system shows efficient and accurate performance in detecting single objects, multi-objects, and transparent objects under different colored backgrounds, including low-light and high-light conditions, fully demonstrating its applicability and robustness in different background environments. As Figure 11 shown are single-object and multi-object grasps under various backgrounds.
[0122] In the embodiment of the present invention, first, a grasping dataset is obtained; the grasping dataset is preprocessed; a preprocessed grasping dataset is obtained; secondly, an adaptive multi-background suppression grasping network is constructed; according to the preprocessed grasping dataset, the adaptive multi-background suppression grasping network is trained to obtain a trained adaptive multi-background suppression grasping network; RGB-D image data of the object to be grasped is obtained; the RGB-D image data of the object to be grasped is input into the trained adaptive multi-background suppression grasping network to obtain a grasping quality map, a grasping angle map, and a grasping width map; finally, according to the grasping quality map, the grasping angle map, and the grasping width map, grasping pose data is estimated; and motion planning is performed according to the grasping pose data to enable the robot to execute a grasping task.
[0123] The proposed adaptive multi-background suppression grasping network can accurately identify the importance of each channel in the feature map and adjust the contributions of each channel in the feature map accordingly; the adaptive multi-background suppression grasping network focuses on the features crucial for the grasping task while being able to ignore irrelevant features, ensuring excellent grasping detection effects in various different backgrounds. The proposed adaptive multi-background suppression grasping network solves the problem of gradient disappearance during network training, simplifies the learning process, speeds up the convergence rate of the network, and improves the generalization ability of the network to objects at new positions; through the set CBAM module, the present invention can dynamically adjust channel attention and spatial attention, enhance the adaptability and complementary advantages of feature extraction, and effectively cope with complex and variable target objects; the adaptive multi-background suppression grasping network can be applied to grasping detection tasks of objects with different sizes, colors, shapes, and surface textures; the present invention constructs a ROS-based upper robot system platform, which can monitor the motion state of the robot in real time and integrate grasping detection and motion planning. The integrated design can achieve high-level automation and intelligence, can detect the state of the robot in real time, and ensure that the robot is in a safe state.
[0124] Figure 12 is a block diagram of a robot grasping detection device for multiple backgrounds shown according to an exemplary embodiment. This device is used for a robot grasping detection method for multiple backgrounds. Refer to Figure 12 , this device includes a first acquisition unit 310, a preprocessing unit 320, a construction unit 330, a training unit 340, a second acquisition unit 350, a third acquisition unit 360, and an execution unit 370. Among them:
[0125] The first acquisition unit 310 is used to acquire a grasping data set;
[0126] The preprocessing unit 320 is used to preprocess the grasping data set to obtain a preprocessed grasping data set;
[0127] The construction unit 330 is used to construct an adaptive multi-background suppression grasping network;
[0128] The training unit 340 is used to train the adaptive multi-background suppression grasping network according to the preprocessed grasping data set to obtain a trained adaptive multi-background suppression grasping network;
[0129] The second acquisition unit 350 is used to acquire RGB-D image data of the object to be grasped;
[0130] The third acquisition unit 360 is used to input the RGB-D image data of the object to be grasped into the trained adaptive multi-background suppression grasping network to obtain a grasping quality map, a grasping angle map, and a grasping width map;
[0131] An execution unit 370, configured to estimate grasping pose data according to the grasping quality map, the grasping angle map, and the grasping width map; and perform motion planning according to the grasping pose data to enable the robot to perform a grasping task.
[0132] Optionally, the preprocessing unit 320 is configured to:
[0133] Clean, transform, and annotate the grasping data set.
[0134] Optionally, the adaptive multi-background suppression grasping network includes: a convolutional layer, a transposed convolutional layer, 5 residual modules, a SEGCM module, and a CBAM module.
[0135] Optionally, the residual module includes: two 3×3 convolutional kernels and a Mish activation function;
[0136] The residual module is configured to perform a convolution operation on the input feature map, extract local features, and output a fused feature map;
[0137] Optionally, the SEGCM module includes: three 1×1 convolutional kernels, a 3×3 convolutional kernel, a Mish activation function, a Sigmoid activation function, and a fully connected layer;
[0138] The SEGCM module is configured to extract the relationship between channels, optimize the feature representation through global context information, and output key features.
[0139] Optionally, the CBAM module includes: a channel attention module and a spatial attention module;
[0140] The channel attention module is configured to perform a pooling operation on the input feature to obtain an attention-enhanced feature;
[0141] The spatial attention module is configured to perform a pooling operation on the attention-enhanced feature to obtain a spatial attention feature.
[0142] Optionally, the training unit 340 is configured to:
[0143] Input the preprocessed grasping data set into the adaptive multi-background suppression grasping network, perform feature extraction through the convolutional layer to obtain preliminary features;
[0144] Input the preliminary features into the residual module, process them through a 3×3 convolutional kernel to obtain local features; according to the residual mechanism, process the local features through normalization, the Mish activation function, and a 3×3 convolutional kernel, and output fused features; the fused features include the preliminary features of the original input and the features processed by the convolutional layer;
[0145] The fused features are input into the SEGCM module and processed through two 1×1 convolutional kernels, the Mish activation function, and the Sigmoid activation function to obtain a compressed feature vector; based on the compressed feature vector, it is processed through a 3×3 convolutional kernel, the Mish activation function, the Sigmoid activation function, and a fully connected layer to obtain an attention weight map;
[0146] The attention weight map is input into the CBAM module and processed through the channel attention module to obtain the channel features after attention weighting; the channel features after attention weighting are input into the spatial attention module to obtain the spatial attention features; based on the spatial attention features and the channel features after attention weighting, the attention enhanced features are generated; the attention enhanced features are input into the deconvolution layer for processing to obtain the grasping quality map, the grasping angle map, and the grasping width map; the adaptive multi-background suppression grasping network is trained according to the grasping quality map, the grasping angle map, and the grasping width map to obtain a trained adaptive multi-background suppression grasping network.
[0147] In the embodiment of the present invention, first, a grasping dataset is obtained; the grasping dataset is preprocessed; the preprocessed grasping dataset is obtained; secondly, an adaptive multi-background suppression grasping network is constructed; according to the preprocessed grasping dataset, the adaptive multi-background suppression grasping network is trained to obtain a trained adaptive multi-background suppression grasping network; the RGB-D image data of the object to be grasped is obtained; the RGB-D image data of the object to be grasped is input into the trained adaptive multi-background suppression grasping network to obtain the grasping quality map, the grasping angle map, and the grasping width map; finally, according to the grasping quality map, the grasping angle map, and the grasping width map, the grasping pose data is estimated; motion planning is performed according to the grasping pose data to enable the robot to execute the grasping task.
[0148] The proposed adaptive multi-background suppression grasping network of the present invention can accurately identify the importance of each channel in the feature map and accordingly weight and adjust the contributions of each channel in the feature map; the adaptive multi-background suppression grasping network focuses on the features crucial for the grasping task while being able to ignore irrelevant features, ensuring excellent grasping detection effects under various different backgrounds. The proposed adaptive multi-background suppression grasping network of the present invention solves the problem of gradient disappearance during network training, simplifies the learning process, speeds up the convergence rate of the network, and improves the generalization ability of the network to objects at new positions; through the set CBAM module, the present invention can dynamically adjust channel attention and spatial attention, enhance the self-adaptability and complementary advantages of feature extraction, and effectively cope with complex and variable target objects; the adaptive multi-background suppression grasping network can be applied to the grasping detection tasks of objects with different sizes, colors, shapes, and surface textures; the present invention constructs a ROS-based upper robot system platform, which can monitor the motion state of the robot in real time and integrate grasping detection and motion planning. The integrated design can achieve high-level automation and intelligence, and can detect the state of the robot in real time to ensure that the robot is in a safe state.
[0149] Figure 13 FIG. is a schematic structural diagram of a robot grasping detection device for multiple backgrounds provided by an embodiment of the present invention, as Figure 13 shown, the robot grasping detection device for multiple backgrounds may include the above-mentioned Figure 12 shown robot grasping detection device for multiple backgrounds. Optionally, the robot grasping detection device 410 for multiple backgrounds may include a first processor 2001.
[0150] Optionally, the robot grasping detection device 410 for multiple backgrounds may further include a memory 2002 and a transceiver 2003.
[0151] Among them, the first processor 2001, the memory 2002, and the transceiver 2003, such as may be connected through a communication bus.
[0152] Next, in combination with Figure 13 specific introductions will be made to the respective components of the robot grasping detection device 410 for multiple backgrounds:
[0153] Among them, the first processor 2001 is the control center of the robot grasping detection device 410 in various scenarios, which can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 can be one or more central processing units (CPUs), or an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. For example, one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).
[0154] Optionally, the first processor 2001 can execute various functions of the robot grasping detection device 410 in various scenarios by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.
[0155] In a specific implementation, as an embodiment, the first processor 2001 can include one or more CPUs, such as Figure 13 the CPU0 and CPU1 shown in
[0156] In a specific implementation, as an embodiment, the robot grasping detection device 410 in various scenarios can also include multiple processors, such as Figure 13 the first processor 2001 and the second processor 2004 shown in
[0157] Each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). Here, the processor can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0158] Optionally, the memory 2002 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 may be integrated with the first processor 2001 or may exist independently and is coupled to the first processor 2001 through an interface circuit ( Figure 13 not shown) of the robot grasping detection device 410 in various scenarios. The embodiments of the present invention do not make specific limitations on this.
[0159] The transceiver 2003 is used to communicate with a network device or with a terminal device.
[0160] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 13 not shown separately). Among them, the receiver is used to implement the receiving function, and the transmitter is used to implement the sending function.
[0161] Optionally, the transceiver 2003 may be integrated with the first processor 2001 or may exist independently and is coupled to the first processor 2001 through an interface circuit ( Figure 13 not shown) of the robot grasping detection device 410 in various scenarios. The embodiments of the present invention do not make specific limitations on this.
[0162] It should be noted that Figure 13 the structure of the robot grasping detection device 410 in various scenarios shown does not constitute a limitation on the router. The actual knowledge structure recognition device may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0163] In addition, the technical effects of the robot grasping detection device 410 in various scenarios may refer to the technical effects of the robot grasping detection method in various scenarios described in the above method embodiments and will not be elaborated here.
[0164] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0165] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM) or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM) and direct rambus RAM (DR RAM).
[0166] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, or magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0167] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context.
[0168] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0169] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0170] Those of ordinary skill in the art will realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0171] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices, apparatuses, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated herein.
[0172] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0173] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0174] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0175] When the above-mentioned function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0176] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A robot grasping detection method for multiple scenarios, characterized in that, The method includes: S1. Obtain a grasping dataset; S2. Preprocess the grasping dataset to obtain a preprocessed grasping dataset; Among them, the preprocessing of the grasping dataset in S2 includes: Performing cleaning processing, data conversion processing, and grasping annotation processing on the grasping dataset; S3. Construct an adaptive multi-background suppression grasping network; Among them, the adaptive multi-background suppression grasping network includes: a convolutional layer, a deconvolutional layer, 5 residual modules, a SEGCM module, and a CBAM module; Among them, the residual module includes: two 3×3 convolutional kernels and a Mish activation function; The residual module is used to perform a convolution operation on the input feature map, extract local features, and output a fused feature map; Among them, the SEGCM module includes: three 1×1 convolutional kernels, a 3×3 convolutional kernel, a Mish activation function, a Sigmoid activation function, and a fully connected layer; The SEGCM module is used to extract the relationship between channels, optimize the feature representation through global context information, and output key features; Among them, the CBAM module includes: a channel attention module and a spatial attention module; The channel attention module is used to perform a pooling operation on the input feature to obtain an attention-enhanced feature; The spatial attention module is used to perform a pooling operation on the attention-enhanced feature to obtain a spatial attention feature; S4. Train the adaptive multi-background suppression grasping network according to the preprocessed grasping dataset to obtain a trained adaptive multi-background suppression grasping network; Among them, the training of the adaptive multi-background suppression grasping network according to the preprocessed grasping dataset in S4 to obtain a trained adaptive multi-background suppression grasping network includes: S41. Input the preprocessed grasping dataset into the adaptive multi-background suppression grasping network, and perform feature extraction through the convolutional layer to obtain preliminary features; S42. Input the preliminary features into the residual module, and process them through a 3×3 convolutional kernel to obtain local features; according to the residual mechanism, process the local features through normalization, the Mish activation function, and a 3×3 convolutional kernel, and output fused features; the fused features include the preliminary features of the original input and the features processed by the convolutional layer; S43. Input the fused features into the SEGCM module, and process them through two 1×1 convolutional kernels, the Mish activation function, and the Sigmoid activation function to obtain a compressed feature vector; according to the compressed feature vector, process it through a 3×3 convolutional kernel, the Mish activation function, the Sigmoid activation function, and a fully connected layer to obtain an attention weight map; S44. Input the attention weight map into the CBAM module, and process it through the channel attention module to obtain channel features after attention weighting; input the channel features after attention weighting into the spatial attention module to obtain spatial attention features; generate attention-enhanced features according to the spatial attention features and the channel features after attention weighting; S45. Input the attention enhancement feature into the deconvolution layer for processing to obtain the grasping quality map, grasping angle map, and grasping width map; train the adaptive multi-background suppression grasping network according to the grasping quality map, grasping angle map, and grasping width map to obtain the trained adaptive multi-background suppression grasping network; S5. Obtain the RGB-D image data of the object to be grasped; S6. Input the RGB-D image data of the object to be grasped into the trained adaptive multi-background suppression grasping network to obtain the grasping quality map, grasping angle map, and grasping width map; S7. Estimate the grasping pose data according to the grasping quality map, grasping angle map, and grasping width map; perform motion planning according to the grasping pose data to enable the robot to execute the grasping task.
2. A robot grasping detection device for multiple backgrounds, the robot grasping detection device for multiple backgrounds is used to implement the robot grasping detection method for multiple backgrounds as described in claim 1, characterized in that, The device includes: A first acquisition unit for acquiring a grasping data set; A preprocessing unit for preprocessing the grasping data set to obtain the preprocessed grasping data set; A construction unit for constructing an adaptive multi-background suppression grasping network; A training unit for training the adaptive multi-background suppression grasping network according to the preprocessed grasping data set to obtain the trained adaptive multi-background suppression grasping network; A second acquisition unit for acquiring the RGB-D image data of the object to be grasped; A third acquisition unit for inputting the RGB-D image data of the object to be grasped into the trained adaptive multi-background suppression grasping network to obtain the grasping quality map, grasping angle map, and grasping width map; An execution unit for estimating the grasping pose data according to the grasping quality map, grasping angle map, and grasping width map; performing motion planning according to the grasping pose data to enable the robot to execute the grasping task.
3. A robot grasping detection device for multiple backgrounds, characterized in that, The robot grasping detection device for multiple backgrounds includes: A processor; A memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the method according to claim 1 is implemented.
4. A computer-readable storage medium, characterized in that, Program code is stored in the computer-readable storage medium, and the program code can be called by the processor to execute the method according to claim 1.