A robot grasping method, apparatus, device, and storage medium
By using a pre-trained grasping feature map generation model, the problems of high computational complexity and insufficient feature fusion in existing robot grasping models are solved, achieving efficient and accurate robot grasping, which is suitable for resource-constrained industrial scenarios.
Patent Information
- Application Number
- CN202511609093.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-05
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2045-11-05
AI Technical Summary
Existing deep learning-based robot grasping models suffer from high computational complexity, insufficient feature fusion, and imbalance in multi-task learning in practical applications, resulting in low grasping accuracy and efficiency, and making them difficult to adapt to resource-constrained industrial grasping scenarios.
A pre-trained grasping feature map generation model is used to generate grasping quality map, grasping angle map, and gripper width map by acquiring images of the object to be grasped. The grasping trajectory of the robotic arm is determined, and the robotic arm is controlled by the robotic arm controller to grasp the object. The model adopts a two-stage optimization strategy and a multi-task adaptive loss function to balance the weights of different generation tasks.
It improves the accuracy and speed of feature map extraction, enhancing the precision and efficiency of robot grasping, especially performing exceptionally well in resource-constrained industrial scenarios.
Smart Images

Figure CN121105028B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robot grasping technology, and in particular to a robot grasping method, apparatus, device, and storage medium. Background Technology
[0002] In the wave of automation and intelligent manufacturing, robotic grasping technology plays an indispensable and crucial role. As a fundamental component of mechanical operation, grasping not only directly affects production efficiency but also directly impacts the robot's adaptability and flexibility in complex environments. Efficient grasping strategies enable robots to handle various objects more accurately and quickly, thereby promoting the widespread application of robotics technology in multiple fields such as logistics, manufacturing, healthcare, and home services.
[0003] However, existing deep learning-based robot grasping models have revealed numerous problems in practical applications, such as high computational complexity, insufficient feature fusion, and imbalance in multi-task learning. These problems not only affect model performance, reducing robot grasping accuracy and efficiency, but also make the models difficult to adapt to resource-constrained industrial grasping scenarios. Summary of the Invention
[0004] This invention provides a robot grasping method, apparatus, device, and storage medium to improve robot grasping accuracy and efficiency.
[0005] According to one aspect of the present invention, a robot grasping method is provided, the method comprising:
[0006] Acquire an image of the object to be captured;
[0007] The image of the object to be grasped is input into a trained grasping feature map generation model to obtain the grasping feature map corresponding to the object; the grasping feature map includes a grasping quality map and a grasping angle. Image, capture angle Diagram and gripper width diagram;
[0008] Based on the grasping feature map, determine the grasping trajectory of the robotic arm to grasp the object to be grasped;
[0009] Based on the grasping trajectory, the robotic arm controller controls the robotic arm to grasp the object to be grasped.
[0010] According to another aspect of the present invention, a robotic grasping device is provided, the device comprising:
[0011] The object to be captured image acquisition module is used to acquire images of the object to be captured;
[0012] The grasping feature map determination module is used to input the image of the object to be grasped into the trained grasping feature map generation model to obtain the grasping feature map corresponding to the object to be grasped; wherein, the grasping feature map includes a grasping quality map and a grasping angle. Image, capture angle Diagram and gripper width diagram;
[0013] The grasping trajectory determination module is used to determine the grasping trajectory of the robotic arm to grasp the object based on the grasping feature map;
[0014] The object-grabbing module is used to control the robotic arm to grasp the object according to the grasping trajectory via the robotic arm controller.
[0015] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0016] At least one processor;
[0017] and a memory communicatively connected to at least one processor; wherein,
[0018] The memory stores a computer program that can be executed by at least one processor, such that the at least one processor is able to perform the robot grasping method of any embodiment of the present invention.
[0019] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the robot grasping method of any embodiment of the present invention.
[0020] According to another aspect of the present invention, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the robot grasping method of any embodiment of the present invention.
[0021] The technical solution of this invention involves acquiring an image of an object to be grasped; inputting the image of the object to be grasped into a trained grasping feature map generation model to obtain a grasping feature map corresponding to the object to be grasped; wherein, the grasping feature map includes a grasping quality map and a grasping angle. Image, capture angle The above technical solution utilizes a pre-trained grasping feature map generation model. This model addresses the problems of high computational complexity, insufficient feature fusion, and multi-task learning imbalance faced by existing deep learning-based robot grasping models in practical applications. It generates more accurate grasping feature maps of the object being grasped, while also improving the speed of feature map determination. This, in turn, enhances the accuracy and speed of the grasping trajectory determined from the feature maps, enabling the robot arm to grasp the object more precisely and quickly, thus improving both the robot's grasping accuracy and efficiency.
[0022] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is a flowchart of a robot grasping method according to Embodiment 1 of the present invention;
[0025] Figure 2A This is a flowchart of a robot grasping method according to Embodiment 2 of the present invention;
[0026] Figure 2B This is a schematic diagram of the structure of a trained feature map generation model according to Embodiment 2 of the present invention;
[0027] Figure 2C This is a schematic diagram of the structure of a visual state space layer according to Embodiment 2 of the present invention;
[0028] Figure 2D This is a schematic diagram of the structure of a scanning layer in a visual state space layer according to Embodiment 2 of the present invention;
[0029] Figure 2E This is a flowchart of the processing of input feature maps by a visual state space layer according to Embodiment 2 of the present invention;
[0030] Figure 2FThis is a schematic diagram of the structure of a fusion submodule in a multi-scale fusion module according to Embodiment 2 of the present invention;
[0031] Figure 2G This is a flowchart of the processing of the input feature map by the fusion submodule in a multi-scale fusion module according to Embodiment 2 of the present invention;
[0032] Figure 3 This is a schematic diagram of the structure of a robot grasping device according to Embodiment 3 of the present invention;
[0033] Figure 4 This is a schematic diagram of the structure of an electronic device that implements the robot grasping method of this invention. Detailed Implementation
[0034] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0035] It should be noted that the terms "target," "first," and "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0036] Example 1
[0037] Figure 1 This is a flowchart of a robot grasping method provided in Embodiment 1 of the present invention. This embodiment is applicable to industrial robot grasping scenarios, especially resource-constrained industrial robot grasping scenarios. The method can be executed by a robot grasping device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:
[0038] S101. Obtain the image of the object to be grabbed.
[0039] The image of the object to be captured refers to an RGB image or an RGB-D image containing the object to be captured. Specifically, the RGB-D image containing the object to be captured is the image obtained by fusing the RGB image containing the object and the depth image containing the object.
[0040] Specifically, images of the object to be captured can be obtained using an RGB camera or an RGB-D camera.
[0041] S102. Input the image of the object to be grasped into the trained grasping feature map generation model to obtain the grasping feature map corresponding to the object to be grasped.
[0042] The feature map to be captured includes the capture quality map and the capture angle. Image, capture angle The image includes a gripper width image and a gripping quality image. The gripping quality image records the success rate of gripping the object by using each pixel in the image as the gripping center; the success rate ranges from [0,1]. Gripping angle... The image is used to record the sine value of the grasping angle at each pixel in the image of the object to be grasped; grasping angle. The image shows the cosine of the gripping angle at each pixel in the image of the object to be gripped by the robotic arm; the gripper width image shows the gripper opening width at each pixel in the image of the object to be gripped. It should be noted that the gripping quality image and gripping angle... Image, capture angle The image dimensions of the image and the gripper width image are the same, for example, the gripping quality image and the gripping angle. Image, capture angle The image sizes for both the graph and the gripper width graph are [missing information]. Where H represents the height of the input image required by the model, for example, H could be 224 pixels; W represents the width of the input image required by the model, for example, W could be 224 pixels.
[0043] The trained grasping feature map generation model was obtained by training a pre-defined robot grasping network model on a sample grasping dataset (including the Cornell and Jacquard datasets). To expand the sample grasping dataset and enhance its diversity, data augmentation techniques such as translation, flipping, rotation, and affine transformations were employed. It's worth noting that the model training process uses a two-stage optimization strategy: initially, the Nadam optimizer is used to accelerate convergence, leveraging its adaptive learning rate to quickly locate the optimal solution region; once the model parameters stabilize, the SGD (Stochastic Gradient Descent) optimizer is switched for fine-tuning, improving generalization performance by updating the gradient sample-by-sample. The initial learning rate is set to 0.001, and cosine annealing is used as a dynamic decay mechanism to make the training process more robust. The model training parameters were configured as follows: a total of 100 training epochs, with 16 samples processed per batch. Five-fold cross-validation was performed using both the Cornell and Jacquard datasets. Systematic data partitioning and evaluation ensured the model's stability and generalization ability across different data distributions. The loss function used during model training is as follows:
[0044] ;
[0045] in, To capture the uncertainty parameters of the quality map generation task; To capture the angle Uncertainty parameters of graph generation tasks; To capture the angle Uncertainty parameters of graph generation tasks; Generate the task's uncertainty parameters for the gripper width map; This represents the basic loss function for the task of generating quality maps. Indicates the grab angle The fundamental loss function for graph generation tasks; Indicates the grab angle The fundamental loss function for graph generation tasks; The basic loss function for generating the gripper width map. It should be noted that... This is a regularization term in the loss function, used to prevent the feature map generation model from growing infinitely. This makes the loss weight of a certain generation task approach zero, thereby ensuring the balance and numerical stability among the generation tasks during model training.
[0046] Understandably, by introducing four learnable parameters ( , , and ), guiding the task of generating a quality image of the grasping process and the grasping angle in the model. Image generation task, capture angle The graph generation task and the gripper width graph generation task are each dynamically weighted according to their uncertainty. That is, when the uncertainty of a certain generation task is high (i.e., When the value is relatively large (i = 1, 2, 3, 4), the corresponding loss weight (i.e.) The uncertainty of a generation task will automatically decrease, thus reducing its impact on the overall optimization objective of the model; conversely, when the uncertainty of a generation task is low (i.e., ...), the uncertainty will decrease. When it is relatively small, its corresponding loss weight (i.e. The value of will automatically increase, thus increasing the impact of the generation task on the overall optimization objective of the model. Using the aforementioned loss function (i.e. This approach effectively solves the imbalance problem among the four generation tasks in the feature map generation model. It considers the differences in importance among the four generation tasks, allowing the generation tasks with lower uncertainty in the feature map generation model to receive higher loss weights. This promotes the feature map generation model to prioritize the optimization of more certain and reliable generation tasks, achieving adaptive balance of multi-task losses. It effectively avoids the subjective bias and sensitivity problems caused by manually setting loss weights, improves the overall convergence of the feature map generation model during training, and enhances the stability of the feature map generation model's performance.
[0047] Specifically, bilinear interpolation can be used to scale the image of the object to be grasped, resulting in a scaled image. The scaled image is then numerically normalized to obtain a normalized image. This normalized image is then input into a trained grasping feature map generation model. After processing by the model, the grasping quality map and grasping angle corresponding to the object to be grasped are obtained. Image, capture angle The image shows the image and the gripper width. It should be noted that the pixel value of each pixel in the normalized image is within the range [0,1].
[0048] Understandably, by performing preprocessing such as image scaling and numerical normalization on the image of the object to be crawled, the image of the object to be crawled can meet the input requirements of the trained crawling feature map generation model, thereby improving the generalization ability of the trained crawling feature map generation model in different crawling scenarios.
[0049] S103. Based on the grasping feature map, determine the grasping trajectory of the robotic arm to grasp the object to be grasped.
[0050] The grasping trajectory includes, but is not limited to, the angle changes at each joint of the robot's robotic arm, the time nodes of the angle changes at each joint, and the gripper opening and closing control commands.
[0051] Specifically, based on the capture success rate of each pixel in the capture quality image, the pixel with the highest capture success rate in the capture quality image is selected as the target capture point of the object to be captured; based on the two-dimensional position coordinates of the target capture point in the capture quality image, the capture angle is determined. The image shows the sine value of the angle corresponding to the target grasping point, obtained from the grasping angle. The cosine value of the angle corresponding to the target gripping point is obtained from the figure. Based on the sine and cosine values of the angle corresponding to the target gripping point, the rotation angle of the robot's gripper relative to the horizontal axis at the target gripping point is determined using the following formula:
[0052] ;
[0053] in, This indicates the rotation angle of the robot's gripper relative to the horizontal axis at the target grasping point; This represents the sine value of the angle corresponding to the target capture point; This represents the cosine value of the angle corresponding to the target grasping point. Simultaneously, based on the two-dimensional position coordinates of the target grasping point in the grasping quality map, the opening width of the robotic arm's gripper at the target grasping point is obtained from the gripper width map. A grasping rectangle is generated based on the two-dimensional position coordinates of the target grasping point in the grasping quality map, the rotation angle of the robot's robotic arm gripper relative to the horizontal axis at the target grasping point, and the opening width of the robot's robotic arm gripper at the target grasping point. Using a pre-defined inverse kinematics algorithm and path planning algorithm, the grasping rectangle is analyzed to obtain the grasping trajectory of the robot's robotic arm in grasping the object to be grasped.
[0054] S104. Based on the grasping trajectory, control the robotic arm to grasp the object to be grasped through the robotic arm controller.
[0055] Specifically, the grasping trajectory is sent to the robotic arm controller in the robot, which then controls the robot's robotic arm to grasp the object according to the grasping trajectory.
[0056] The technical solution of this invention involves acquiring an image of an object to be grasped; inputting the image of the object to be grasped into a trained grasping feature map generation model to obtain a grasping feature map corresponding to the object to be grasped; wherein, the grasping feature map includes a grasping quality map and a grasping angle. Image, capture angle The above technical solution utilizes a pre-trained grasping feature map generation model. This model addresses the problems of high computational complexity, insufficient feature fusion, and multi-task learning imbalance faced by existing deep learning-based robot grasping models in practical applications. It generates more accurate grasping feature maps of the object being grasped, while also improving the speed of feature map determination. This, in turn, enhances the accuracy and speed of the grasping trajectory determined from the feature maps, enabling the robot arm to grasp the object more precisely and quickly, thus improving both the robot's grasping accuracy and efficiency.
[0057] Example 2
[0058] Figure 2A This is a flowchart of a robot grasping method provided in Embodiment 2 of the present invention. Based on the above embodiments, this embodiment further optimizes the step of "inputting the image of the object to be grasped into a trained grasping feature map generation model to obtain the grasping feature map corresponding to the object to be grasped," providing an optional implementation scheme. It should be noted that parts not detailed in this embodiment can be referred to in the relevant descriptions of other embodiments. It should also be noted that... Figure 2B The trained feature map generation model 20 includes an encoder 21, a multi-scale fusion module 22, a decoder 23, and an output head 24. For example... Figure 2A As shown, the method includes:
[0059] S201. Obtain the image of the object to be grabbed.
[0060] S202. Input the image of the object to be captured into the encoder, and the encoder outputs the first enhanced feature map, the second enhanced feature map, the third enhanced feature map and the fourth enhanced feature map.
[0061] Among them, see Figure 2B The encoder 21 includes a block embedding layer 211, a first visual state space layer 212, a first block merging layer 213, a second visual state space layer 214, a second block merging layer 215, a third visual state space layer 216, a third block merging layer 217, and a fourth visual state space layer 218. The first enhanced feature map refers to the feature map output by the first visual state space layer 212; the second enhanced feature map refers to the feature map output by the second visual state space layer 214; the third enhanced feature map refers to the feature map output by the third visual state space layer 216; and the fourth enhanced feature map refers to the feature map output by the fourth visual state space layer 218. It should be noted that if the size of the object image to be grasped in the input trained grasping feature map generation model is... Then the size of the first enhanced feature map is The size of the second enhanced feature map is The size of the third enhanced feature map is The size of the fourth enhanced feature map is .
[0062] Specifically, the image of the object to be captured is input into the encoder 21. The image is then segmented and embedded through the block embedding layer 211 to obtain a block feature sequence. Each block feature in the block feature sequence is enhanced through the first visual state space layer 212 to obtain multiple first enhanced feature maps. The multiple first enhanced feature maps are then merged through the first block merging layer 213 to obtain multiple first merged feature maps. The multiple first merged feature maps are then enhanced through the second visual state space layer 214 to obtain multiple second enhanced feature maps. The multiple second enhanced feature maps are then merged through the second block merging layer 215 to obtain multiple second merged feature maps. The multiple second merged feature maps are then enhanced through the third visual state space layer 216 to obtain multiple third enhanced feature maps. The multiple third enhanced feature maps are then merged through the third block merging layer 217 to obtain multiple third merged feature maps. Finally, the multiple third merged feature maps are enhanced through the fourth visual state space layer 218 to obtain multiple fourth enhanced feature maps.
[0063] Specifically, the block embedding layer 211 performs block embedding processing on the image of the object to be captured to obtain a block feature sequence. This can be achieved by dividing the image of the object to be captured into a predetermined number of non-overlapping block images using the block embedding layer 211; embedding processing is then performed on each block image to obtain the embedding features of each block image, thereby obtaining a block feature sequence composed of the embedding features of each block image. The predetermined number can be set in advance according to actual business needs or the expert experience of those skilled in the art, and is not specifically limited in this embodiment of the invention.
[0064] Specifically, a first block merging layer 213 merges multiple first enhanced feature maps to obtain multiple first merged feature maps. Specifically, based on the positional relationship between the multiple first enhanced feature maps, the first block merging layer 213 performs pairwise feature merging on adjacent first enhanced feature maps to obtain multiple first merged feature maps. Similarly, a second block merging layer 215 merges multiple second enhanced feature maps to obtain multiple second merged feature maps. Specifically, based on the positional relationship between the multiple second enhanced feature maps, the second block merging layer 215 performs pairwise feature merging on adjacent second enhanced feature maps to obtain multiple second merged feature maps. A third block merging layer 217 merges multiple third enhanced feature maps to obtain multiple third merged feature maps. Specifically, based on the positional relationship between the multiple third enhanced feature maps, the third block merging layer 217 performs pairwise feature merging on adjacent third enhanced feature maps to obtain multiple third merged feature maps.
[0065] The first visual state space layer 212, the second visual state space layer 214, the third visual state space layer 216, and the fourth visual state space layer 218 have the same structure, each consisting of a first normalization layer, a first linear transformation layer, a second linear transformation layer, a depthwise convolutional layer, a scanning layer, a second normalization layer, an element-wise multiplication layer, a third linear transformation layer, and an element-wise addition layer. The connection relationships between these structures are detailed below. Figure 2C . Figure 2C In the visual state space layer 30, the first normalization layer 31 is connected to the first linear transformation layer 32 and the second linear transformation layer 33; the first linear transformation layer 32 is connected to the depth convolution layer 34; the depth convolution layer 34 is connected to the scan layer 35; the scan layer 35 is connected to the second normalization layer 36; the second normalization layer 36 is connected to the element-wise multiplication layer 37; the second linear transformation layer 33 is connected to the element-wise multiplication layer 37; the element-wise multiplication layer 37 is connected to the third linear transformation layer 38; and the third linear transformation layer 38 is connected to the element-wise addition layer 39. See also... Figure 2DThe scanning layer 35 includes a cross-scanning block 351, an S6 block 352, and a cross-merging block 353; the cross-scanning block 351 is connected to the S6 block 352; the S6 block 352 is connected to the cross-merging block 353. The processing procedure of the input feature map by the scanning layer 35 is as follows: The four-way scanning strategy in the cross-scanning block 351 performs four-way scanning processing on the input feature map, obtaining four one-dimensional feature sequences; the S6 block 352 adjusts the weights of the four one-dimensional feature sequences respectively, obtaining four new one-dimensional feature sequences; the cross-merging block 353 first reassembles the four new one-dimensional feature sequences respectively, obtaining four two-dimensional feature maps; then, the Cross-Merge strategy is used to perform multi-directional information fusion processing on the four two-dimensional feature maps, obtaining a two-dimensional global feature map. It can be understood that the visual state space layer 30, through the scanning layer 35, achieves four-way scanning and feature integration of the input feature map, realizing long-distance information propagation in the two-dimensional structure and enhancing the perception ability of the grasping region in the input feature map.
[0066] It should be noted that since the structures of the first visual state space layer 212, the second visual state space layer 214, the third visual state space layer 216, and the fourth visual state space layer 218 are identical, the processing flow of the input feature map for these layers is the same. For details, please refer to [link to relevant documentation]. Figure 2E Specifically, for the input feature map in the input visual state space layer 30, the first normalization layer 31 performs a first normalization process on the input feature map to obtain a first normalized feature map; the first linear transformation layer 32 performs a first linear transformation process on the first normalized feature map to obtain a first linearly transformed feature map; simultaneously, the second linear transformation layer 33 performs a second linear transformation process on the first normalized data to obtain a second linearly transformed feature map; the first linearly transformed feature map is then subjected to a first feature extraction process by the deep convolutional layer 34 to obtain a first spatial feature map; and the first feature map is then subjected to a scanning layer 35 to further refine the first spatial feature map. The spatial feature map undergoes a second feature extraction to obtain a global view feature map. A second normalization layer 36 then normalizes the global view feature map to obtain a second normalized feature map. An element-wise multiplication layer 37 performs element-wise multiplication on the second normalized feature map and the second linearly transformed feature map to obtain a multiplied feature map. A third linear transformation layer 38 performs a third linear transformation on the multiplied feature map to obtain a third linearly transformed feature map. Finally, an element-wise addition layer 39 performs element-wise addition on the third linearly transformed feature map and the input feature map to obtain an output feature map.
[0067] Understandably, the encoder 21 in the trained grasping feature map generation model 20 achieves layer-by-layer feature enhancement of the image of the object to be grasped, while also reducing the spatial resolution of the image of the object to be grasped layer by layer. Compared with existing deep learning-based robot grasping models, it significantly reduces the computational complexity of the model while retaining the global perception capability of the image of the object to be grasped, thereby improving the accuracy and determination speed of subsequent grasping feature maps.
[0068] S203. The first enhanced feature map, the second enhanced feature map, the third enhanced feature map and the fourth enhanced feature map output by the encoder are processed through the multi-scale fusion module and the decoder to obtain the feature map to be analyzed.
[0069] It should be noted that the feature levels of the first, second, third, and fourth enhanced feature maps are as follows: first enhanced feature map < second enhanced feature map < third enhanced feature map < fourth enhanced feature map.
[0070] Among them, see Figure 2B The multi-scale fusion module 22 includes a first fusion submodule 221, a second fusion submodule 222, and a third fusion submodule 223; the decoder 23 includes a fifth visual state space layer 231, a sixth visual state space layer 232, a seventh visual state space layer 233, and a block expansion layer 234. The first fusion submodule 221, the second fusion submodule 222, and the third fusion submodule 223 have the same structure, each consisting of a first depthwise separable convolution module, a bilinear interpolation module, a first segmentation module, a second segmentation module, a first stitching module, a first normalization processing module, a second depthwise separable convolution module, a second stitching module, a second normalization processing module, and a third depthwise separable convolution module. See the detailed structure for further details. Figure 2F It should be noted that the first depthwise separable convolution module, the bilinear interpolation module, and the first segmentation module are used to process the input high-level feature map, while the first segmentation module is used to process the input low-level feature map. It should also be noted that... Figure 2F In the fusion submodule, the first depthwise separable convolution module 001 is connected to the bilinear interpolation module 002; the bilinear interpolation module 002 is connected to the first segmentation module 003; the first segmentation module 003 is connected to the first stitching module 005; the second segmentation module 004 is connected to the first stitching module 005; the first stitching module 005 is connected to the first normalization processing module 006; the first normalization processing module 006 is connected to the second depthwise separable convolution module 007; the second depthwise separable convolution module 007 is connected to the second stitching module 008; the second stitching module 008 is connected to the second normalization processing module 009; and the second normalization processing module 009 is connected to the third depthwise separable convolution module 010.
[0071] It should also be noted that since the first fusion submodule 221, the second fusion submodule 222, and the third fusion submodule 223 have the same structure, their processing flow for the input feature map is the same. For details, please refer to [link to relevant documentation]. Figure 2G . Figure 2G middle, This represents the high-level feature map of the input. This represents the low-level feature map of the input. DOConv indicates a depthwise separable convolution operation, Bilinear Interpolate indicates a bilinear interpolation operation, Split indicates a segmentation operation, Concatation indicates a concatenation operation, and LayerNorm indicates a normalization operation. This represents the output feature map of the fusion submodule. It should be noted that the high-level feature map input to the third fusion submodule 223 is the fourth enhanced feature map output by encoder 21, and the low-level feature map input is the third enhanced feature map output by encoder 21; the high-level feature map input to the second fusion submodule 222 is the first residual feature map; the low-level feature map input is the second enhanced feature map output by encoder 21; and the high-level feature map input to the first fusion submodule 221 is the second residual feature map, and the low-level feature map input is the first enhanced feature map output by encoder 21.
[0072] The first visual state space layer 212, the second visual state space layer 214, the third visual state space layer 216, the fourth visual state space layer 218, the fifth visual state space layer 231, the sixth visual state space layer 232, and the seventh visual state space layer 233 have the same structure and the same processing flow for the input feature map.
[0073] Specifically, the third fusion submodule 223 performs feature fusion processing on the fourth and third enhanced feature maps output by the encoder 21 to obtain a first fused feature map; the fifth visual state space layer 231 performs feature enhancement processing on the first fused feature map to obtain a first residual feature map; the second fusion submodule 222 performs feature fusion processing on the first residual feature map and the second enhanced feature map output by the encoder 21 to obtain a second fused feature map; the sixth visual state space layer 232 performs feature enhancement processing on the second fused feature map to obtain a second residual feature map; the first fusion submodule 221 performs feature fusion processing on the second residual feature map and the first enhanced feature map output by the encoder 21 to obtain a third fused feature map; the seventh visual state space layer 233 performs feature enhancement processing on the third fused feature map to obtain a third residual feature map; and the block expansion layer 234 performs size restoration processing on the third residual feature map to obtain the feature map to be analyzed.
[0074] Specifically, the third fusion submodule 223 performs feature fusion processing on the fourth enhanced feature map and the third enhanced feature map output by the encoder 21 to obtain the first fused feature map. This can be achieved by: using the first depthwise separable convolution module in the third fusion submodule 223 to reduce the dimensionality of the fourth enhanced feature map according to the dimension of the third enhanced feature map, resulting in a dimensionality-reduced fourth enhanced feature map with the same dimension as the third enhanced feature map; using the bilinear interpolation module in the third fusion submodule 223 to perform bilinear interpolation to adjust the size of the dimensionality-reduced fourth enhanced feature map, resulting in a size-adjusted fourth enhanced feature map with the same size as the third enhanced feature map; using the first segmentation module in the third fusion submodule 223 to segment the size-adjusted fourth enhanced feature map along the channel dimension according to a preset number of segmentation blocks, resulting in a preset number of fourth enhanced feature map blocks; and using the second segmentation module in the third fusion submodule 223 to segment the third enhanced feature map along the channel dimension according to a preset number of segmentation blocks, resulting in a preset number of segmentation blocks. Several third enhanced feature map patches are generated. The first stitching module in the third fusion submodule 223 stitches together the fourth enhanced feature map patches and the third enhanced feature map patches located at the same position according to their relative positions in the feature map before segmentation, resulting in several stitched images of the preset segmentation block. The first normalization processing module in the third fusion submodule 223 normalizes each of the several stitched images of the preset segmentation block, resulting in several normalized stitched images of the preset segmentation block. The second depthwise separable convolution module in the third fusion submodule 223 further normalizes each of the preset segmentation blocks. The normalized stitched image with a preset number of blocks is optimized to obtain an optimized stitched image with a preset number of segments. The optimized stitched image with the preset number of segments is then stitched along the channel dimension using the second stitching module in the third fusion submodule 223 to obtain a stitched feature map. The stitched feature map is then normalized using the second normalization module in the third fusion submodule 223 to obtain a normalized stitched feature map. Finally, the normalized stitched feature map is optimized using the third depthwise separable convolution module in the third fusion submodule 223 to obtain a first fused feature map. The preset number of segments can be pre-set according to actual business needs; for example, the preset number of segments can be 4. This embodiment of the invention does not impose a specific limitation on this number.
[0075] Understandably, the multi-scale fusion module 22 in the trained grasping feature map generation model 20 achieves feature fusion between high-level and low-level feature maps, that is, it realizes the interaction between high-level and low-level information, so that the features in the image of the object to be grasped are fully fused, enhancing the model's ability to integrate features in the image of the object to be grasped; the group fusion strategy improves the multi-scale contextual understanding of the feature map output by the encoder 21, thereby improving the model's ability to discriminate grasping regions in complex images of the object to be grasped (such as images of the object to be grasped with occlusion or background interference, etc.), and thus improving the accuracy of subsequent grasping feature maps.
[0076] S204. The output head analyzes and processes the feature map to be analyzed to obtain the grasping feature map corresponding to the object to be grasped.
[0077] Among them, see Figure 2B The output head 24 includes a quality map generation layer 241, a first angle map generation layer 242, a second angle map generation layer 243, and a width map generation layer 244. The first angle map generation layer 242 is used to generate the grasping angle corresponding to the object to be grasped. Figure; Second angle image generation layer 243 is used to generate the grasping angle corresponding to the object to be grasped. Figure. It should be noted that the quality image generation layer 241, the first angle image generation layer 242, the second angle image generation layer 243, and the width image generation layer 244 in the output head 24 are independent of each other.
[0078] Specifically, the quality map generation layer 241 performs a grasping success rate analysis on each pixel in the feature map to be analyzed, obtaining a grasping quality map corresponding to the object to be grasped; the first angle map generation layer 242 performs a first grasping angle analysis on each pixel in the feature map to be analyzed, obtaining a grasping angle corresponding to the object to be grasped. Figure 243 uses the second angle map generation layer to perform second grasping angle analysis on each pixel in the feature map to be analyzed, thus obtaining the grasping angle corresponding to the object to be grasped. Figure 244 uses a width map generation layer to analyze the gripper opening width of each pixel in the feature map to be analyzed, thus obtaining the gripper width map corresponding to the object to be grasped. Specifically, the first gripping angle analysis is used to analyze the sine value of the gripping angle of the robot's arm at each pixel in the feature map to grasp the object; the second gripping angle analysis is used to analyze the cosine value of the gripping angle of the robot's arm at each pixel in the feature map to grasp the object.
[0079] S205. Based on the grasping feature map, determine the grasping trajectory of the robotic arm to grasp the object to be grasped.
[0080] S206. Based on the grasping trajectory, the robotic arm controller controls the robotic arm to grasp the object to be grasped.
[0081] The technical solution of this invention involves: acquiring an image of an object to be grasped; inputting the image into an encoder, which outputs a first enhanced feature map, a second enhanced feature map, a third enhanced feature map, and a fourth enhanced feature map; processing the first, second, third, and fourth enhanced feature maps output by the encoder using a multi-scale fusion module and a decoder to obtain a feature map to be analyzed; analyzing the feature map to be analyzed using an output head to obtain a grasping feature map corresponding to the object to be grasped; determining the grasping trajectory of the robotic arm to grasp the object based on the grasping feature map; and controlling the robotic arm to grasp the object based on the grasping trajectory using a robotic arm controller. The above technical solution inputs the image of the object to be grasped into a trained grasping feature map generation model. The encoder in the model performs layer-by-layer feature enhancement and spatial resolution reduction on the image, significantly reducing the model's computational complexity while preserving the global perception capability of the object image. Then, a multi-scale fusion module in the model achieves feature fusion between high-level and low-level feature maps, ensuring sufficient integration of features in the object image and enhancing the model's ability to integrate features. Finally, the decoder in the model restores the spatial resolution of the image layer by layer, and the output head of the model outputs the grasping quality map and grasping angle corresponding to the object to be grasped. Image, capture angle The image and gripper width image improve the accuracy and speed of grasping feature map determination, which in turn improves the accuracy and speed of grasping trajectory determination based on grasping feature map. This allows the robot's robotic arm to grasp the object to be grasped more accurately and quickly, thus improving the robot's grasping accuracy and grasping efficiency.
[0082] Example 3
[0083] Figure 3 This is a schematic diagram of a robot grasping device according to Embodiment 3 of the present invention. This embodiment is applicable to industrial robot grasping scenarios, especially resource-constrained industrial robot grasping scenarios. The device can be implemented in hardware and / or software and can be configured in electronic devices. Figure 3 As shown, the device includes:
[0084] The object to be grabbed image acquisition module 301 is used to acquire the image of the object to be grabbed;
[0085] The grasping feature map determination module 302 is used to input the image of the object to be grasped into the trained grasping feature map generation model to obtain the grasping feature map corresponding to the object to be grasped; wherein, the grasping feature map includes a grasping quality map and a grasping angle. Image, capture angle Diagram and gripper width diagram;
[0086] The grasping trajectory determination module 303 is used to determine the grasping trajectory of the robotic arm to grasp the object based on the grasping feature map;
[0087] The object-to-be-grabbed module 304 is used to control the robotic arm to grasp the object to be grasped according to the grasping trajectory via the robotic arm controller.
[0088] The technical solution of this invention involves acquiring an image of an object to be grasped; inputting the image of the object to be grasped into a trained grasping feature map generation model to obtain a grasping feature map corresponding to the object to be grasped; wherein, the grasping feature map includes a grasping quality map and a grasping angle. Image, capture angle The above technical solution utilizes a pre-trained grasping feature map generation model. This model addresses the problems of high computational complexity, insufficient feature fusion, and multi-task learning imbalance faced by existing deep learning-based robot grasping models in practical applications. It generates more accurate grasping feature maps of the object being grasped, while also improving the speed of feature map determination. This, in turn, enhances the accuracy and speed of the grasping trajectory determined from the feature maps, enabling the robot arm to grasp the object more precisely and quickly, thus improving both the robot's grasping accuracy and efficiency.
[0089] Optionally, the trained feature map generation model includes an encoder, a multi-scale fusion module, a decoder, and an output head;
[0090] Feature map extraction and determination module 302 includes:
[0091] The enhanced feature map determination unit is used to input the image of the object to be grasped into the encoder, and the encoder outputs a first enhanced feature map, a second enhanced feature map, a third enhanced feature map and a fourth enhanced feature map;
[0092] The feature map determination unit is used to process the first enhanced feature map, the second enhanced feature map, the third enhanced feature map and the fourth enhanced feature map output by the encoder through the multi-scale fusion module and the decoder to obtain the feature map to be analyzed.
[0093] The grasping feature map determination unit is used to analyze and process the feature map to be analyzed through the output head to obtain the grasping feature map corresponding to the object to be grasped.
[0094] Optionally, the encoder includes a block embedding layer, a first visual state space layer, a first block merging layer, a second visual state space layer, a second block merging layer, a third visual state space layer, a third block merging layer, and a fourth visual state space layer.
[0095] The enhanced feature map determination unit is specifically used for:
[0096] The block embedding layer performs block embedding processing on the image of the object to be captured, resulting in a block feature sequence;
[0097] The first visual state space layer is used to perform feature enhancement processing on each block feature in the block feature sequence to obtain multiple first enhanced feature maps.
[0098] Multiple first enhanced feature maps are merged through the first block merging layer to obtain multiple first merged feature maps;
[0099] Multiple first merged feature maps are enhanced by a second visual state space layer to obtain multiple second enhanced feature maps;
[0100] The second block merging layer merges multiple second enhanced feature maps to obtain multiple merged second feature maps.
[0101] Multiple second merged feature maps are enhanced by a third visual state space layer to obtain multiple third enhanced feature maps;
[0102] Multiple third-enhanced feature maps are merged through a third block merging layer to obtain multiple third-merged feature maps;
[0103] Multiple third-merged feature maps are enhanced by applying a fourth visual state space layer to obtain multiple fourth-enhanced feature maps.
[0104] Optionally, the multi-scale fusion module includes a first fusion sub-module, a second fusion sub-module, and a third fusion sub-module; the decoder includes a fifth visual state space layer, a sixth visual state space layer, a seventh visual state space layer, and a block expansion layer;
[0105] The feature map to be analyzed is used to determine the unit, specifically for:
[0106] The third fusion submodule performs feature fusion processing on the fourth enhanced feature map and the third enhanced feature map output by the encoder to obtain the first fused feature map;
[0107] The first fused feature map is enhanced by the fifth visual state space layer to obtain the first residual feature map;
[0108] The second fusion submodule performs feature fusion processing on the first residual feature map and the second enhanced feature map output by the encoder to obtain the second fused feature map.
[0109] The second fused feature map is enhanced by the sixth visual state space layer to obtain the second residual feature map;
[0110] The first fusion submodule performs feature fusion processing on the second residual feature map and the first enhanced feature map output by the encoder to obtain the third fused feature map;
[0111] The third fused feature map is enhanced by the seventh visual state space layer to obtain the third residual feature map;
[0112] The third residual feature map is sized by performing a block-based expansion layer to obtain the feature map to be analyzed.
[0113] Optionally, the output header includes a quality map generation layer, a first angle map generation layer, a second angle map generation layer, and a width map generation layer;
[0114] The feature map extraction and determination unit is specifically used for:
[0115] The quality map generation layer performs a capture success rate analysis on each pixel in the feature map to be analyzed, and obtains the capture quality map corresponding to the object to be captured.
[0116] The first angle map generation layer performs a first grasping angle analysis on each pixel in the feature map to be analyzed, thereby obtaining the grasping angle corresponding to the object to be grasped. picture;
[0117] The second angle map generation layer performs second grasping angle analysis on each pixel in the feature map to be analyzed, thereby obtaining the grasping angle corresponding to the object to be grasped. picture;
[0118] By generating a width map layer, the gripper opening width is analyzed for each pixel in the feature map to be analyzed, and the gripper width map corresponding to the object to be gripped is obtained.
[0119] Optionally, the first visual state space layer, the second visual state space layer, the third visual state space layer, and the fourth visual state space layer have the same structure, each consisting of a first normalization layer, a first linear transformation layer, a second linear transformation layer, a depthwise convolutional layer, a scanning layer, a second normalization layer, an element-wise multiplication operation layer, a third linear transformation layer, and an element-wise addition operation layer.
[0120] Optionally, the loss function used during the training of the feature map extraction and generation model is as follows:
[0121] ;
[0122] in, To capture the uncertainty parameters of the quality map generation task; To capture the angle Uncertainty parameters of graph generation tasks; To capture the angle Uncertainty parameters of graph generation tasks; Generate the task's uncertainty parameters for the gripper width map; This represents the basic loss function for the task of generating quality maps. Indicates the grab angle The fundamental loss function for graph generation tasks; Indicates the grab angle The fundamental loss function for graph generation tasks; The basic loss function for generating the gripper width map.
[0123] The robot grasping device provided in the embodiments of the present invention can execute the robot grasping method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing each robot grasping method.
[0124] According to embodiments of the present invention, the present invention also provides an electronic device, a readable storage medium, and a computer program product.
[0125] Example 4
[0126] Figure 4 A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0127] like Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded into the RAM 13 from storage unit 18. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0128] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0129] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as robotic grasping methods.
[0130] In some embodiments, the robotic grasping method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the robotic grasping method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the robotic grasping method by any other suitable means (e.g., by means of firmware).
[0131] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0132] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0133] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0134] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0135] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0136] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0137] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0138] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A robot grasping method, characterized in that, include: Acquire an image of the object to be captured; The image of the object to be grasped is input into a trained grasping feature map generation model to obtain a grasping feature map corresponding to the object to be grasped; wherein, the grasping feature map includes a grasping quality map and a grasping angle. Image, capture angle The trained grasping feature map generation model includes an encoder, a multi-scale fusion module, a decoder, and an output head. The encoder includes a block embedding layer, a first visual state space layer, a first block merging layer, a second visual state space layer, a second block merging layer, a third visual state space layer, a third block merging layer, and a fourth visual state space layer. The multi-scale fusion module includes a first fusion submodule, a second fusion submodule, and a third fusion submodule. The decoder includes a fifth visual state space layer, a sixth visual state space layer, a seventh visual state space layer, and a block expansion layer. The output head includes a quality map generation layer, a first angle map generation layer, a second angle map generation layer, and a width map generation layer. Based on the grasping feature map, the grasping trajectory of the robotic arm to grasp the object to be grasped is determined; According to the grasping trajectory, the robotic arm controller controls the robotic arm to grasp the object to be grasped; The step of inputting the image of the object to be grasped into a trained grasping feature map generation model to obtain the grasping feature map corresponding to the object to be grasped includes: The image of the object to be captured is input into the encoder, and the encoder outputs a first enhanced feature map, a second enhanced feature map, a third enhanced feature map, and a fourth enhanced feature map; The first enhanced feature map, the second enhanced feature map, the third enhanced feature map, and the fourth enhanced feature map output by the encoder are processed by the multi-scale fusion module and the decoder to obtain the feature map to be analyzed. The output head analyzes and processes the feature map to be analyzed to obtain the grasping feature map corresponding to the object to be grasped; The loss function used in the training process of the feature map generation model is as follows: ; in, To capture the uncertainty parameters of the quality map generation task; To capture the angle Uncertainty parameters of graph generation tasks; To capture the angle Uncertainty parameters of graph generation tasks; Generate the task's uncertainty parameters for the gripper width map; This represents the basic loss function for the task of generating quality maps. Indicates the grab angle The fundamental loss function for graph generation tasks; Indicates the grab angle The fundamental loss function for graph generation tasks; The basic loss function for generating the gripper width map.
2. The method according to claim 1, characterized in that, The step of inputting the image of the object to be captured into the encoder, and having the encoder output a first enhanced feature map, a second enhanced feature map, a third enhanced feature map, and a fourth enhanced feature map, includes: The block embedding layer is used to perform block embedding processing on the image of the object to be captured to obtain a block feature sequence. The first visual state space layer performs feature enhancement processing on each block feature in the block feature sequence to obtain multiple first enhanced feature maps. The first block merging layer merges the multiple first enhanced feature maps to obtain multiple first merged feature maps. The multiple first merged feature maps are enhanced by the second visual state space layer to obtain multiple second enhanced feature maps; The multiple second enhanced feature maps are merged through the second block merging layer to obtain multiple second merged feature maps; The third visual state space layer is used to perform feature enhancement processing on the multiple second merged feature maps to obtain multiple third enhanced feature maps; The third block merging layer merges the multiple third enhanced feature maps to obtain multiple third merged feature maps. The fourth visual state space layer is used to perform feature enhancement processing on the multiple third merged feature maps to obtain multiple fourth enhanced feature maps.
3. The method according to claim 1, characterized in that, The process involves using the multi-scale fusion module and the decoder to process the first, second, third, and fourth enhanced feature maps output by the encoder to obtain the feature map to be analyzed, including: The third fusion submodule performs feature fusion processing on the fourth enhanced feature map and the third enhanced feature map output by the encoder to obtain the first fused feature map; The first fused feature map is enhanced by the fifth visual state space layer to obtain the first residual feature map; The second fusion submodule performs feature fusion processing on the first residual feature map and the second enhanced feature map output by the encoder to obtain the second fused feature map. The second fused feature map is enhanced by the sixth visual state space layer to obtain the second residual feature map; The first fusion submodule performs feature fusion processing on the second residual feature map and the first enhanced feature map output by the encoder to obtain a third fused feature map. The third fused feature map is enhanced by the seventh visual state space layer to obtain the third residual feature map; The third residual feature map is subjected to size restoration processing through the block expansion layer to obtain the feature map to be analyzed.
4. The method according to claim 1, characterized in that, The step of analyzing and processing the feature map to be analyzed through the output head to obtain the grasping feature map corresponding to the object to be grasped includes: The capture success rate of each pixel in the feature map to be analyzed is analyzed by the quality map generation layer to obtain the capture quality map corresponding to the object to be captured. The first angle map generation layer performs a first grasping angle analysis on each pixel in the feature map to be analyzed, and obtains the grasping angle corresponding to the object to be grasped. picture; The second angle map generation layer performs a second grasping angle analysis on each pixel in the feature map to be analyzed, thereby obtaining the grasping angle corresponding to the object to be grasped. picture; The width map generation layer performs gripper opening width analysis on each pixel in the feature map to be analyzed, thereby obtaining the gripper width map corresponding to the object to be gripped.
5. The method according to claim 1, characterized in that, The first visual state space layer, the second visual state space layer, the third visual state space layer, and the fourth visual state space layer have the same structure, each consisting of a first normalization layer, a first linear transformation layer, a second linear transformation layer, a depthwise convolution layer, a scanning layer, a second normalization layer, an element-wise multiplication operation layer, a third linear transformation layer, and an element-wise addition operation layer.
6. A robotic grasping device, characterized in that, include: The object to be captured image acquisition module is used to acquire images of the object to be captured; The grasping feature map determination module is used to input the image of the object to be grasped into a trained grasping feature map generation model to obtain the grasping feature map corresponding to the object to be grasped; wherein, the grasping feature map includes a grasping quality map and a grasping angle. Image, capture angle The trained grasping feature map generation model includes an encoder, a multi-scale fusion module, a decoder, and an output head. The encoder includes a block embedding layer, a first visual state space layer, a first block merging layer, a second visual state space layer, a second block merging layer, a third visual state space layer, a third block merging layer, and a fourth visual state space layer. The multi-scale fusion module includes a first fusion submodule, a second fusion submodule, and a third fusion submodule. The decoder includes a fifth visual state space layer, a sixth visual state space layer, a seventh visual state space layer, and a block expansion layer. The output head includes a quality map generation layer, a first angle map generation layer, a second angle map generation layer, and a width map generation layer. The grasping trajectory determination module is used to determine the grasping trajectory of the robotic arm grasping the object to be grasped based on the grasping feature map. The object-to-be-grabbed module is used to control the robotic arm to grasp the object to be grasped according to the grasping trajectory via the robotic arm controller. The feature map determination module includes: An enhanced feature map determination unit is used to input the image of the object to be captured into the encoder, and the encoder outputs a first enhanced feature map, a second enhanced feature map, a third enhanced feature map, and a fourth enhanced feature map; The feature map determination unit is used to process the first enhanced feature map, the second enhanced feature map, the third enhanced feature map and the fourth enhanced feature map output by the encoder through the multi-scale fusion module and the decoder to obtain the feature map to be analyzed. The grasping feature map determination unit is used to analyze and process the feature map to be analyzed through the output head to obtain the grasping feature map corresponding to the object to be grasped; The loss function used in the training process of the feature map generation model is as follows: ; in, To capture the uncertainty parameters of the quality map generation task; To capture the angle Uncertainty parameters of graph generation tasks; To capture the angle Uncertainty parameters of graph generation tasks; Generate the task's uncertainty parameters for the gripper width map; This represents the basic loss function for the task of generating quality maps. Indicates the grab angle The fundamental loss function for graph generation tasks; Indicates the grab angle The fundamental loss function for graph generation tasks; The basic loss function for generating the gripper width map.
7. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the robot grasping method according to any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the robot grasping method according to any one of claims 1-5.