Multi-target capturing method and system based on deep learning
Through the multi-objective crawling method based on deep learning, the improved YOLOv5 and GRCNN networks and camera calibration algorithms are used to solve the problems of insufficient target recognition accuracy, inaccurate crawling point calculations and lack of real-time performance in the complex environment of traditional robot crawling systems, and high-precision and high-efficiency multi-objective real-time crawling are achieved.
Patent Information
- Application Number
- CN202510653639.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-06-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional robot grasping systems have problems such as insufficient target recognition accuracy, inaccurate calculation of grab points and lack of real-time in complex multi-object environments.
Using a multi-objective capture method based on deep learning, the image acquisition device is used to obtain the image and depth information of the target to be captured, and the object detection model and the capture point prediction model are constructed through the improved YOLOv5 and GRCNN networks. The mapping model between the robot arm and the image acquisition device is established in combination with the camera calibration algorithm to achieve accurate conversion from image coordinates to robot arm coordinates.
In complex environments, high-precision and efficient real-time grabbing of multi-target objects are achieved, which improves the positioning accuracy of grab points, and solves the problems of insufficient target recognition accuracy, inaccurate calculation of grab points and lack of real-timeness.
Smart Images

Figure CN120182383A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a multi - target grasping method and system based on deep learning, belonging to the technical fields of computer vision and robot control. Background Art
[0002] In the fields of industrial production and logistics, with the continuous development of automation technology, the application of robot grasping systems is becoming increasingly widespread. These systems usually rely on traditional vision algorithms for target recognition and positioning. Traditional grasping methods rely on preset feature templates or rules for object recognition, determine the grasping points through geometric analysis, and perform grasping actions. These methods perform well in simple and structured environments and can effectively complete single or small - quantity object grasping tasks. However, in the face of complex multi - object environments, such as object stacking, partial occlusion, or cluttered backgrounds, the limitations of traditional methods become apparent.
[0003] Although traditional robot grasping systems have achieved certain success in specific scenarios, their limitations are becoming increasingly prominent in complex multi - object environments. First, insufficient target recognition accuracy has become a major problem. Traditional vision algorithms often have difficulty efficiently and accurately distinguishing target objects when dealing with complex backgrounds or target object stacking scenarios, resulting in frequent misidentifications or missed identifications. Second, inaccurate calculation of grasping points is also a key factor restricting the grasping success rate. Due to the lack of high - precision algorithm support, traditional methods often rely on experience or simple geometric analysis when determining grasping points, and it is difficult to adapt to objects of different shapes, sizes, and materials, resulting in a relatively high failure rate of grasping actions. In addition, the real - time issue is also a major challenge faced by traditional grasping methods. In complex grasping tasks, the processing speed of algorithms often fails to meet the requirements of real - time grasping, resulting in system response delays and affecting the overall grasping efficiency and accuracy. Summary of the Invention
[0004] The purpose of the present invention is to provide a multi - target grasping method and system based on deep learning, which uses a target grasping detection model to solve the problems of insufficient target recognition accuracy, inaccurate calculation of grasping points, and lack of real - time performance existing in the prior art.
[0005] To solve the above - mentioned technical problems, the present invention is implemented by adopting the following technical solutions:
[0006] In the first aspect, the present invention provides a multi - target grasping method based on deep learning, including:
[0007] Using an image acquisition device to obtain the image and depth information of the target to be grasped;
[0008] Based on the image and depth information of the target to be grasped, and based on a pre-trained target grasping detection model, obtain the coordinates of at least one grasping point, input them into the drive controller of the robotic arm for parsing, and generate a grasping instruction;
[0009] Use the robotic arm to grasp at least one target according to the grasping instruction;
[0010] The target grasping detection model includes a target detection model and a grasping point prediction model respectively improved and constructed from the YOLOv5 network and the GRCNN network;
[0011] Among them, the data processing method of the target grasping detection model includes:
[0012] Input the image and depth information of the target to be grasped into the target detection model to detect the target, and obtain an image containing at least one target;
[0013] Input the image containing at least one target into the grasping point prediction model to determine the grasping point, and obtain the grasping point coordinates in the coordinate system of the image acquisition device;
[0014] Based on the mapping model established in advance between the robotic arm and the image acquisition device through the camera calibration algorithm, convert the grasping point coordinates in the coordinate system of the image acquisition device into the grasping point coordinates in the coordinate system of the robotic arm.
[0015] Further, the method of establishing the mapping model between the robotic arm and the image acquisition device through the camera calibration method includes:
[0016] Collect multi-target images to be grasped from multiple different angles using the image acquisition device and the calibration board to obtain multi-angle images, and record the pose information of the end of the robotic arm. The pose information of the end of the robotic arm includes the position and direction of the end effector;
[0017] Based on the multi-angle images, calculate the internal parameters and external parameters of the image acquisition device through the standard camera calibration algorithm. The internal parameters include the focal length and the optical center, and the external parameters include the position and pose of the camera relative to the coordinate system of the robotic arm;
[0018] Based on the internal parameters and external parameters of the image acquisition device and the pose information of the end of the robotic arm, construct a mapping model between the coordinate system of the image acquisition device and the coordinate system of the robotic arm, which is used to convert points in the coordinate system of the image acquisition device into points in the coordinate system of the robotic arm.
[0019] Further, it also includes, after constructing the mapping model, using the mapping model for robotic arm operation path teaching, including:
[0020] Specify at least one target in the multi-target images to be grasped;
[0021] Use the mapping model to convert the coordinates in the coordinate system of the specified target image acquisition device into the coordinates in the coordinate system of the robotic arm;
[0022] Calculate the motion path of the robotic arm according to the coordinates in the specified coordinate system of the robotic arm;
[0023] According to the motion path of the robotic arm, use inverse kinematics to calculate the motion trajectories and control commands of each joint, input them into the robotic arm control system, and manipulate the robotic arm to complete the grasping action.
[0024] Furthermore, improve the YOLOv5 network and the GRCNN network respectively to construct an object detection model and a grasping point prediction model, including:
[0025] Improve the YOLOv5 network, including replacing the original loss function of the YOLOv5 network with Focal-IOU and introducing the BiFPN module in the Neck part of the YOLOv5 network;
[0026] The Focal-IOU is used to receive the difference between the predicted bounding box and the actual bounding box input by the Head part of the YOLOv5 network for processing, and output a loss value to update the weights of the YOLOv5 network;
[0027] The BiFPN module is used to receive the multi-scale feature maps formed by the image of the target to be grasped and the depth information input by the Neck part of the YOLOv5 network, and generate enhanced multi-scale feature maps through upsampling, downsampling and feature fusion operations, and output them to the Head part of the YOLOv5 network;
[0028] Improve the GRCNN network, including replacing the standard two-dimensional convolutional layer in the GRCNN network with a lightweight C2f structure and introducing the ECA attention mechanism in the GRCNN network;
[0029] The C2f structure is used to receive the feature maps processed from the images containing at least one target input by the convolutional layer or pooling layer in the GRCNN network, and generate lightweight feature maps through lightweight operations, and output them to the next convolutional layer or pooling layer of the GRCNN network;
[0030] The ECA attention mechanism is used to process the lightweight feature maps output by the convolutional layer or pooling layer in the GRCNN network, generate an attention weight map by calculating the correlation between channels, multiply the attention weight map by the original feature map, and obtain a weighted feature map, and output it to the next layer of the GRCNN network.
[0031] Further, input the image and depth information of the target to be grasped into a target detection model to detect the target, and obtain an image containing at least one target, including:
[0032] Using the target detection model, based on the image and depth information of the target to be grasped, obtain the detection box coordinates, height, width, and center point coordinates of the graspable target in the image of the target to be grasped;
[0033] If the number of targets to be grasped is less than 2, taking the center point coordinates as a reference, extend outward and intercept the target image with a resolution exceeding a preset threshold after extension to obtain an image containing one target;
[0034] If the number of targets to be grasped is greater than or equal to 2, taking the center point coordinates as a reference, extend outward and intercept the target image with a resolution exceeding a preset threshold after extension from left to right and from top to bottom to obtain an image containing at least two targets.
[0035] Further, taking the center point coordinates as a reference, extend outward and intercept the target image with a resolution exceeding a preset threshold after extension from left to right and from top to bottom, including:
[0036] If the vertical coordinates of the targets to be grasped are different, arrange them in ascending order of the vertical coordinates of the targets to be grasped, and intercept each target image with a resolution exceeding a preset threshold after extension in sequence;
[0037] If the vertical coordinates of the targets to be grasped are the same, arrange them in ascending order of the horizontal coordinates of the targets to be grasped, and intercept each target image with a resolution exceeding a preset threshold after extension in sequence.
[0038] Further, input the image containing at least one target into a grasping point prediction model to determine the grasping point, and obtain the grasping point coordinates in the coordinate system of the image acquisition device, including:
[0039] If the target image with a resolution exceeding the preset threshold contains one target, input the target image with a resolution exceeding the preset threshold into the grasping point prediction model, obtain the confidence value at each pixel position and calculate the quality sum of all pixel points, and then calculate the two-dimensional pixel coordinates of the center point of a grasping box by the weighted average method;
[0040] If the target image with a resolution exceeding the preset threshold contains at least two targets, input the target image with a resolution exceeding the preset threshold into the grasping point prediction model, obtain the confidence value at each pixel position and calculate the quality sum of all pixel points, and then calculate the two-dimensional pixel coordinates of the center points of at least two grasping boxes by the weighted average method;
[0041] Based on the conversion relationship obtained from camera calibration, the two-dimensional pixel coordinates of the center points of at least two grasping frames and the depth information are mapped into the three-dimensional space, and the grasping point coordinates of at least two targets in the coordinate system of the image acquisition device are calculated.
[0042] Further, the two-dimensional pixel coordinates of the center point are expressed as:
[0043] ;
[0044] In the formula, represents the two-dimensional pixel coordinates of the center point, represents the abscissa of the center point, represents the ordinate of the center point, represents the confidence value of the target image with the two-dimensional pixel coordinates of the center point being and the resolution exceeding a preset threshold, represents the sum of the qualities of all pixel points.
[0045] Further, based on the mapping model established in advance between the robotic arm and the image acquisition device through the camera calibration algorithm, the grasping point coordinates in the coordinate system of the image acquisition device are converted into the grasping point coordinates in the coordinate system of the robotic arm, where the mapping model is expressed as:
[0046] ;
[0047] In the formula, , and respectively represent the abscissa, ordinate and vertical coordinate of the grasping point in the coordinate system of the robotic arm, and respectively represent the rotation matrix and translation vector of the image acquisition device, , and respectively represent the abscissa, ordinate and vertical coordinate of the grasping point in the coordinate system of the image acquisition device.
[0048] In a second aspect, the present invention provides a multi-target grasping system based on deep learning, including:
[0049] An image acquisition terminal, including a D435i depth camera, configured to acquire an image and depth information of a target to be grasped, and transmit the same to an industrial control computer through a wired connection;
[0050] An industrial control computer with an embedded GPU inside is used to input the image of the target to be grasped into a target detection model to detect the target, obtaining an image containing at least one target; inputting the image containing at least one target into a grasping point prediction model to determine the grasping point, obtaining the grasping point coordinates in the coordinate system of the image acquisition device; based on the mapping model established in advance between the robotic arm and the image acquisition device through a camera calibration algorithm, converting the grasping point coordinates in the coordinate system of the image acquisition device into the grasping point coordinates in the coordinate system of the robotic arm and transmitting them to the host computer;
[0051] The host computer with MATLAB and Simulink operating environments inside is used to perform robotic arm operation path teaching based on the mapping model, receive the grasping point coordinates in the coordinate system of the robotic arm, generate the motion trajectories and control commands of each joint of the robotic arm through inverse kinematics calculation, and adjust the motion trajectories and control commands of each joint of the robotic arm according to the real-time execution status of the robotic arm, transmitting the adjusted control commands to the data transmission module.
[0052] The data transmission module includes a PCAN communication module and an Ethernet communication module, which is used to control the PCAN communication module to receive the adjusted control commands and transmit them to the robotic arm through the Ethernet communication module, and at the same time receive the real-time execution status of the robotic arm and feedback it to the host computer to dynamically monitor the actions of the robotic arm;
[0053] The robotic arm with a multi-degree-of-freedom servo drive unit inside is used to execute the motion trajectories of each joint of a complex motion trajectory according to the received control commands, complete the grasping action and feedback the real-time execution status to the host computer through the data transmission module.
[0054] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:
[0055] 1. By using the image acquisition device to obtain the image and depth information of the target to be grasped, and combining with the pre-trained and improved target grasping detection model, the present invention realizes the high-precision and high-efficiency real-time grasping of multi-target objects in a complex environment. By constructing a mapping model between the robotic arm and the image acquisition device, the present invention ensures the accurate conversion from the coordinate system of the image acquisition device to the coordinate system of the robotic arm, improving the positioning accuracy of the grasping point. In addition, the present invention also improves the YOLOv5 network and the GRCNN network to construct a target detection model and a grasping point prediction model, which not only improves the accuracy of target recognition but also optimizes the prediction ability of the grasping point, enabling accurate and rapid recognition of the target and determination of the best grasping position even in complex scenarios such as target object stacking, partial occlusion or cluttered background, solving the problems of insufficient target recognition accuracy, inaccurate grasping point calculation and lack of real-time performance in the prior art.
[0056] 2. The present invention realizes the intelligent teaching of the manipulator operation path through a mapping model. By simply specifying the target in the multi-target image to be grasped, the target coordinates can be automatically converted into the coordinates in the manipulator coordinate system, simplifying the complex process of traditional manipulator teaching. This not only improves the teaching efficiency but also reduces the operation difficulty. Combining the target detection model and the grasping point prediction model can intelligently plan the optimal grasping path according to information such as the position, shape, and size of the target, further improving the success rate and efficiency of the grasping task.
[0057] 3. The present invention can simultaneously identify and locate multiple targets, and intercept the target images with a resolution exceeding the preset threshold according to the center point coordinates and depth information of each target. Using the grasping point prediction model to calculate the optimal grasping point coordinates of each target, in this process, the system can automatically handle complex situations such as occlusion and overlap between targets to ensure that each target can be accurately grasped. In addition, the present invention also reduces the computational complexity and resource consumption while maintaining high precision and high efficiency by introducing a lightweight network structure and an attention mechanism, making the system more suitable for actual application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 is a schematic flowchart of a multi-target grasping method based on deep learning provided by an embodiment of the present invention;
[0059] Figure 2 is a schematic structural diagram of a BiFPN module provided by an embodiment of the present invention;
[0060] Figure 3 is a schematic structural diagram of the network of the target detection model provided by an embodiment of the present invention;
[0061] Figure 4 is a schematic structural diagram of the network of the ECA attention mechanism provided by an embodiment of the present invention;
[0062] Figure 5 is a schematic structural diagram of the network of the grasping point prediction model provided by an embodiment of the present invention;
[0063] Figure 6 is a schematic structural diagram of a multi-target grasping system based on deep learning provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0064] The technical solution of the present invention will be described in detail below through the drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations on the technical solution of the present invention. Without conflict, the technical features in the embodiments of the present invention and the embodiments can be combined with each other.
[0065] The term "and / or" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. Additionally, the character " / " generally indicates that the related objects before and after are in an "or" relationship.
[0066] Embodiment 1:
[0067] As Figure 1 shown, this embodiment introduces a multi-object grasping method based on deep learning, including:
[0068] Step 1: Use an image acquisition device to obtain the image and depth information of the target to be grasped.
[0069] Step 1 is the basis of the grasping process of the present invention, relying on an image acquisition device, such as a depth camera, to obtain the visual information of the target to be grasped. The image acquisition device can not only capture the two-dimensional image information of the target but also obtain the depth information of the target, that is, the distance between the target and the camera, through specific sensor technologies. Based on the image and depth information obtained by the present invention, the shape, size, and position of the target to be grasped can be understood more accurately, thereby improving the accuracy and success rate of grasping.
[0070] Step 2: Based on the image and depth information of the target to be grasped, and based on a pre-trained target grasping detection model, obtain the coordinates of at least one grasping point, and input them into the drive controller of the robotic arm for parsing to generate a grasping instruction.
[0071] Step 2 is the target detection and grasping point prediction process of the present invention. First, the target detection model, that is, the improved YOLOv5 network, is used to identify the target in the image and obtain its position information. Subsequently, the grasping point prediction model, that is, the improved GRCNN network, is used to predict the coordinates of the best grasping point according to the position, shape, and depth information of the target. The grasping point coordinates are then input into the drive controller of the robotic arm, and after parsing, specific grasping instructions are generated.
[0072] Through the pre-trained target grasping detection model, the present invention can quickly and accurately identify the target and predict the best grasping point, improving the efficiency and success rate of grasping while reducing the dependence on manual operation. In addition, since the target grasping detection model provided by the present invention is constructed based on deep learning, it has strong generalization ability and can adapt to targets of different shapes, sizes, and materials.
[0073] Step 3: Use the robotic arm to grasp at least one target according to the grasping instruction.
[0074] Step 3 is the process where the present invention actually performs the grasping operation. The robotic arm precisely moves to the predicted grasping point position according to the grasping instruction generated in Step 2 and executes the grasping action. The motion control of the robotic arm relies on a high-precision drive system and sensor feedback to ensure the accuracy and stability of grasping. The precise motion control of the robotic arm enables the system to accurately perform the grasping task and move the target from the original position to the target position.
[0075] The target grasping detection model includes a target detection model and a grasping point prediction model respectively improved and constructed from the YOLOv5 network and the GRCNN network.
[0076] The improvement of the YOLOv5 and GRCNN networks aims to improve the accuracy of target detection and grasping point prediction. The present invention can make the target detection model and the grasping point prediction model still maintain good performance in complex environments and adapt to targets of different shapes, sizes and materials by introducing technical means such as new loss functions, network structures and attention mechanisms.
[0077] Among them, the data processing method of the target grasping detection model includes:
[0078] Input the image and depth information of the target to be grasped into the target detection model to detect the target, and obtain an image containing at least one target;
[0079] Input the image containing at least one target into the grasping point prediction model to determine the grasping point, and obtain the grasping point coordinates in the coordinate system of the image acquisition device;
[0080] Based on the mapping model established in advance between the robotic arm and the image acquisition device through the camera calibration algorithm, convert the grasping point coordinates in the coordinate system of the image acquisition device into the grasping point coordinates in the coordinate system of the robotic arm.
[0081] Among them, through the camera calibration algorithm, a mapping model between the robotic arm and the image acquisition device is established.
[0082] The present invention establishes a mapping model between the robotic arm and the image acquisition device through the camera calibration algorithm. The mapping model is the key to achieving precise grasping. Through the calibration process, the present invention can obtain the internal parameters and external parameters of the camera, thereby constructing the conversion relationship between the coordinate system of the image acquisition device and the coordinate system of the robotic arm. The internal parameters such as focal length and optical center, and the external parameters such as the position and attitude of the camera relative to the coordinate system of the robotic arm.
[0083] The present invention inputs the image and depth information of the target to be grasped into a target detection model for target recognition, and then inputs the image containing the target into a grasping point prediction model for grasping point prediction. Finally, based on the mapping model, the predicted grasping point coordinates are converted from the image acquisition device coordinate system to the robotic arm coordinate system for the robotic arm to perform the grasping operation, realizing efficient, accurate, and intelligent multi-target grasping, and providing strong technical support for the intelligent upgrading of industrial automation, logistics warehousing, and other fields.
[0084] Embodiment 2:
[0085] Based on the same inventive concept as Embodiment 1, this embodiment introduces the implementation steps of a multi-target grasping method based on deep learning, including:
[0086] Step 1: Use an image acquisition device to obtain the image and depth information of the target to be grasped.
[0087] In this embodiment, an image acquisition device is used to collect 2000 images of different objects stacked on a desktop, and then the images are preprocessed, including:
[0088] First, denoise the 2000 images of different objects stacked on the desktop. Use median filtering to remove image noise, and then apply the adaptive histogram equalization method to enhance the contrast of the image and improve the distinguishability of object features. In addition, to enhance the diversity of the dataset and the generalization ability of the model, data augmentation processing is also performed on each image, including rotation, scaling, cropping, etc., to obtain a dataset after expanding the data. The data volume of the dataset after expanding the data has been tripled, effectively enhancing the robustness and adaptability of the training model.
[0089] Two annotation methods are respectively used to annotate the dataset after expanding the data to meet the requirements of the target detection model and the grasping point prediction model.
[0090] The annotation method is as follows: Use the labelimg annotation tool to annotate each image in the dataset after expanding the data. For the target detection model, the annotation tool identifies each target in the image and uses a bounding box to calibrate its position, specifying the category label of each object to generate an annotation file in YOLO format; for the grasping point prediction model, multiple feasible grasping boxes for each target are annotated. After the annotation is completed, the annotated training sets of the target detection model and the grasping point prediction model are obtained.
[0091] In some embodiments, a mapping model between the robotic arm and the image acquisition device is established through a camera calibration algorithm.
[0092] Step 1.1: Establish a mapping model between the robotic arm and the image acquisition device through a camera calibration method, including:
[0093] Step 1.1.1: Collect multi-target images to be grasped from multiple different angles using an image acquisition device and a calibration board to obtain multi-angle images, and record the pose information of the end of the robotic arm. The pose information of the end of the robotic arm includes the position and direction of the end effector.
[0094] Step 1.1.2: Based on the multi-angle images, calculate the internal parameters and external parameters of the image acquisition device through a standard camera calibration algorithm. The internal parameters include the focal length and the optical center, and the external parameters include the position and pose of the camera relative to the robotic arm coordinate system.
[0095] Step 1.1.3: Based on the internal parameters and external parameters of the image acquisition device and the pose information of the end of the robotic arm, construct a mapping model between the image acquisition device coordinate system and the robotic arm coordinate system for converting points in the image acquisition device coordinate system into points in the robotic arm coordinate system.
[0096] In some embodiments, it further includes, after constructing the mapping model, using the mapping model for teaching the operation path of the robotic arm.
[0097] Step 1.2: Use the mapping model for teaching the operation path of the robotic arm, including:
[0098] Step 1.2.1: Specify at least one target in the multi-target images to be grasped.
[0099] Step 1.2.2: Use the mapping model to convert the coordinates of the specified target in the image acquisition device coordinate system into coordinates in the robotic arm coordinate system.
[0100] Step 1.2.3: Calculate the motion path of the robotic arm according to the coordinates of the specified target in the robotic arm coordinate system.
[0101] Step 1.2.4: According to the motion path of the robotic arm, use inverse kinematics to calculate the motion trajectories and control commands of each joint, input them into the robotic arm control system, and manipulate the robotic arm to complete the grasping action.
[0102] Step 2: Based on the images and depth information of the target to be grasped, and based on a pre-trained target grasping detection model, obtain the coordinates of at least one grasping point, input them into the drive controller of the robotic arm for parsing, and generate a grasping command.
[0103] The target grasping detection model includes a target detection model and a grasping point prediction model constructed by respectively improving the YOLOv5 network and the GRCNN network.
[0104] In some embodiments, the YOLOv5 network and the GRCNN network are respectively improved to construct an object detection model and a grasping point prediction model. Among them, the network structure of the object detection model is as shown in Figure 3 shown.
[0105] Step 2.1: Improve the YOLOv5 network and the GRCNN network respectively, including:
[0106] Step 2.1.1: Improve the YOLOv5 network, including replacing the original loss function of the YOLOv5 network with Focal-IOU and introducing the BiFPN module in the Neck part of the YOLOv5 network; among them, the structure of the BiFPN module is as shown in Figure 2 shown.
[0107] In this embodiment, the Focal-IOU is used to receive the difference between the predicted bounding box and the actual bounding box input by the Head part of the YOLOv5 network for processing, and output a loss value to update the weights of the YOLOv5 network.
[0108] In this embodiment, the BiFPN module is used to receive multiple scale feature maps formed by the image of the target to be grasped and the depth information input by the Neck part of the YOLOv5 network, and generate enhanced multi-scale feature maps through upsampling, downsampling and feature fusion operations, and output them to the Head part of the YOLOv5 network.
[0109] In this embodiment, the Focal-IOU is used to identify stacked, occluded or background-cluttered targets, and the BiFPN module is used to predict grasping points.
[0110] Step 2.1.2: Improve the GRCNN network, including replacing the standard two-dimensional convolutional layer in the GRCNN network with a lightweight C2f structure and introducing the ECA attention mechanism in the GRCNN network.
[0111] In this embodiment, the C2f structure is used to receive the feature map processed by the image containing at least one target input by the convolutional layer or pooling layer in the GRCNN network, and generate a lightweight feature map through lightweight operations, and output it to the next convolutional layer or pooling layer of the GRCNN network.
[0112] In this embodiment, the ECA attention mechanism is used to process the lightweight feature map output by the convolutional layer or pooling layer in the GRCNN network, generate an attention weight map by calculating the correlation between channels, multiply the attention weight map by the original feature map to obtain a weighted feature map, and output it to the next layer of the GRCNN network.
[0113] In this embodiment, the C2f structure is used to improve the running efficiency of the grasping point prediction model, and the ECA attention mechanism is used to improve the feature extraction ability of the grasping point prediction model.
[0114] Among them, the network structure of the ECA attention mechanism is as Figure 5 shown, where C2f represents the C2f convolutional layer structure, BN represents batch normalization, CBAM represents the convolutional block attention module, ReLn represents the rectified linear unit, ResidualBlock represents the residual block, ConvT2D represents the two-dimensional transposed convolution, sinθ and cosθ respectively represent the sine value and cosine value of the grasping box angle. The rotation angle of the grasping box can be calculated through the sine value and cosine value of the grasping box angle. Angle represents the rotation angle of the grasping box calculated from the sine value and cosine value of the grasping box angle, Width represents the width of the grasping box, that is, the size of the opening of the grasping box, and Quality represents the confidence of successfully grasping the target. The value of the confidence ranges between 0 and 1.
[0115] The labeled target detection model training set and the grasping point prediction model training set are respectively trained in the improved YOLOv5 network and the improved GRCNN network. After training, the best training results are taken as the target detection model and the grasping point prediction model.
[0116] Step 2.1: The data processing method of the target grasping detection model, including:
[0117] Step 2.1.1: Input the image and depth information of the target to be grasped into the target detection model to detect the target, and obtain an image containing at least one target;
[0118] In this embodiment, inputting the image and depth information of the target to be grasped into the target detection model to detect the target and obtain an image containing at least one target includes:
[0119] Using the target detection model, based on the image and depth information of the target to be grasped, obtain the detection box coordinates, height, width, and center point coordinates of the graspable target in the image of the target to be grasped;
[0120] If the number of targets to be grasped is less than 2, then based on the center point coordinates, extend outward and intercept the target image with a resolution exceeding a preset threshold after extension to obtain an image containing one target;
[0121] If the number of targets to be grasped is greater than or equal to 2, then based on the center point coordinates, extend outward and intercept the target image with a resolution exceeding a preset threshold from left to right and from top to bottom after extension to obtain an image containing at least two targets.
[0122] In this embodiment, taking the center point coordinates as a reference, extending outward and intercepting the target image with a resolution exceeding a preset threshold from left to right and from top to bottom, including:
[0123] If the vertical coordinates of the targets to be grasped are different, arrange them in ascending order of the vertical coordinates of the targets to be grasped, and sequentially intercept each target image with a resolution exceeding the preset threshold after extension;
[0124] If the vertical coordinates of the targets to be grasped are the same, arrange them in ascending order of the horizontal coordinates of the targets to be grasped, and sequentially intercept each target image with a resolution exceeding the preset threshold after extension.
[0125] In this embodiment, if the number of targets to be grasped is less than 2, the center point coordinates of the target to be grasped are represented as ; if the number of targets to be grasped is greater than or equal to 2, the set of center point coordinates of the targets to be grasped is represented as , where 、 respectively represent the abscissa and ordinate of the center point coordinates of the th target to be grasped, represents the number of targets to be grasped;
[0126] Among them, the boundaries of the target image with a resolution exceeding the preset threshold after extension in the image acquisition device coordinate system are represented as:
[0127] ;
[0128] ;
[0129] ;
[0130] ;
[0131] In the formula, , are respectively the left and right boundaries of the target image with a resolution exceeding the preset threshold; , are respectively the upper and lower boundaries of the target image with a resolution exceeding the preset threshold, 、 respectively represent the abscissa and ordinate of the center point coordinates.
[0132] Step 2.1.2: Input the image containing at least one target into the grasping point prediction model to determine the grasping point, and obtain the grasping point coordinates in the image acquisition device coordinate system.
[0133] In this embodiment, an image containing at least one target is input into a grasping point prediction model to determine the grasping point, and the grasping point coordinates in the coordinate system of the image acquisition device are obtained, including:
[0134] If the target image with a resolution exceeding a preset threshold contains one target, the target image with a resolution exceeding the preset threshold is input into the grasping point prediction model to obtain the confidence value at each pixel position and calculate the quality sum of all pixel points, and then the two-dimensional pixel coordinates of the center point of a grasping box are calculated by the weighted average method;
[0135] If the target image with a resolution exceeding a preset threshold contains at least two targets, the target image with a resolution exceeding the preset threshold is input into the grasping point prediction model to obtain the confidence value at each pixel position and calculate the quality sum of all pixel points, and then the two-dimensional pixel coordinates of the center points of at least two grasping boxes are calculated by the weighted average method;
[0136] Based on the conversion relationship of camera calibration, the two-dimensional pixel coordinates of the center points of at least two grasping boxes and the depth information are mapped into the three-dimensional space, and the grasping point coordinates of at least two targets in the coordinate system of the image acquisition device are calculated.
[0137] In this embodiment, the two-dimensional pixel coordinates of the center point are expressed as:
[0138] ;
[0139] In the formula, represents the two-dimensional pixel coordinates of the center point, represents the abscissa of the center point, represents the ordinate of the center point, represents the confidence value of the target image with a resolution exceeding the preset threshold whose two-dimensional pixel coordinates of the center point are , represents the quality sum of all pixel points.
[0140] Step 2.1.3: Based on the mapping model between the robotic arm and the image acquisition device established in advance through the camera calibration algorithm, the grasping point coordinates in the coordinate system of the image acquisition device are converted into the grasping point coordinates in the coordinate system of the robotic arm.
[0141] In this embodiment, through the internal parameter matrix and the external parameter matrix [R|t] of the image acquisition device, the two-dimensional pixel coordinates of the center point are converted into the three-dimensional coordinates in the coordinate system of the image acquisition device, and the specific conversion formula is expressed as:
[0142] ;
[0143] In the formula, , and respectively represent the abscissa, ordinate and vertical coordinate of the grasping point in the coordinate system of the image acquisition device. Among them, Z c is the depth of the image of the target to be grasped, that is, the depth of the center point of the grasping frame in the coordinate system of the image acquisition device.
[0144] In this embodiment, the three-dimensional coordinates in the coordinate system of the image acquisition device are converted to the position in the coordinate system of the robotic arm, and the transformation is performed using the external parameter matrix:
[0145] Among them, the mapping model is expressed as:
[0146] ;
[0147] In the formula, , and respectively represent the abscissa, ordinate and vertical coordinate of the grasping point in the coordinate system of the robotic arm, and respectively represent the rotation matrix and translation vector of the image acquisition device.
[0148] Step 3: Use the robotic arm to grasp at least one target according to the grasping instruction.
[0149] Embodiment 3:
[0150] Based on the same inventive concept as other embodiments, as Figure 6 shown, this embodiment introduces a multi-target grasping system based on deep learning, including:
[0151] An image acquisition end, including a D435i depth camera, which is used to collect the image and depth information of the target to be grasped, and transmit it to the industrial control computer through a wired connection;
[0152] The industrial control computer is built-in with an embedded GPU, which is used to input the image of the target to be grasped into the target detection model to detect the target, and obtain an image containing at least one target; input the image containing at least one target into the grasping point prediction model to determine the grasping point, and obtain the grasping point coordinates in the coordinate system of the image acquisition device; based on the mapping model established in advance between the robotic arm and the image acquisition device through the camera calibration algorithm, convert the grasping point coordinates in the coordinate system of the image acquisition device into the grasping point coordinates in the coordinate system of the robotic arm and transmit them to the upper computer;
[0153] The upper computer, with MATLAB and Simulink running environments built-in, is used for teaching the operation path of the robotic arm based on the mapping model, receiving the coordinates of the grasping points in the robotic arm coordinate system, generating the motion trajectories and control commands of each joint of the robotic arm through inverse kinematics calculation, and adjusting the motion trajectories and control commands of each joint of the robotic arm according to the real-time execution state of the robotic arm, and transmitting the adjusted control commands to the data transmission module.
[0154] The data transmission module, including a PCAN communication module and an Ethernet communication module, is used to control the PCAN communication module to receive the adjusted control commands and transmit them to the robotic arm through the Ethernet communication module, and at the same time receive the real-time execution state of the robotic arm and feedback it to the upper computer to dynamically monitor the actions of the robotic arm;
[0155] The robotic arm, with a multi-degree-of-freedom servo drive unit built-in, is used to execute the motion trajectories of each joint of the complex motion trajectory according to the received control commands, complete the grasping action and feedback the real-time execution state to the upper computer through the data transmission module.
[0156] Embodiment 4:
[0157] Based on the same inventive concept as other embodiments, this embodiment introduces a computer-readable storage medium, on which computer instructions are stored, and when the computer instructions are executed by a processor, the steps of the method in Embodiment 1 or 2 above are implemented.
[0158] Embodiment 5:
[0159] Based on the same inventive concept as other embodiments, this embodiment introduces a computer program product, including computer instructions, and when the computer instructions are executed by a processor, the steps of the method in Embodiment 1 or 2 above are implemented.
[0160] In summary of the above embodiments, the present invention realizes the high-precision and high-efficiency real-time grasping of multi-target objects in a complex environment by using an image acquisition device to obtain the image and depth information of the target to be grasped and combining a pre-trained and improved target grasping detection model. The present invention ensures the accurate conversion from the image acquisition device coordinate system to the robotic arm coordinate system by constructing a mapping model between the robotic arm and the image acquisition device, improving the positioning accuracy of the grasping points. In addition, the present invention also constructs a target detection model and a grasping point prediction model through the improvement of the YOLOv5 network and the GRCNN network, not only improving the accuracy of target recognition but also optimizing the prediction ability of the grasping points, so that even in complex scenarios such as target object stacking, partial occlusion or cluttered background, the target can be accurately and quickly recognized and the best grasping position can be determined, solving the problems of insufficient target recognition accuracy, inaccurate grasping point calculation and lack of real-time performance in the prior art.
[0161] The present invention realizes the intelligent teaching of the manipulator operation path through a mapping model. By simply specifying the target in the multi-target image to be grasped, the target coordinates can be automatically converted into the coordinates in the manipulator coordinate system, simplifying the complex process of traditional manipulator teaching. This not only improves the teaching efficiency but also reduces the operation difficulty. Combining the target detection model and the grasping point prediction model can intelligently plan the optimal grasping path according to information such as the position, shape, and size of the target, further improving the success rate and efficiency of the grasping task.
[0162] The present invention can simultaneously identify and locate multiple targets, intercept the target images with a resolution exceeding the preset threshold according to the center point coordinates and depth information of each target, and calculate the optimal grasping point coordinates of each target using the grasping point prediction model. In this process, the system can automatically handle complex situations such as occlusion and overlap between targets to ensure that each target can be accurately grasped. In addition, the present invention also reduces the computational complexity and resource consumption while maintaining high precision and high efficiency by introducing a lightweight network structure and an attention mechanism, making the system more suitable for practical application scenarios.
[0163] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0164] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0165] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the functions in the process Figure 1 one process or multiple processes and / or blocksFigure 1 The functions specified in one or more boxes.
[0166] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing the steps of the functions specified in one or more processes and / or boxes Figure 1 One or more processes and / or boxes Figure 1 The steps of the functions specified in one or more boxes.
[0167] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit and scope protected by the claims of the present invention. These all fall within the protection scope of the present invention.
Claims
1. A multi-target grasping method based on deep learning, characterized in that: include: Using image acquisition equipment to obtain images and depth information of the target to be captured; According to the image and depth information of the target to be grasped, based on a pre-trained target grasping detection model, the coordinates of at least one grasping point are obtained, and the coordinates are input into the driving controller of the robot arm for analysis to generate a grasping instruction; Using the robotic arm to grasp at least one target according to the grasping instruction; The target grasping detection model includes a target detection model and a grasping point prediction model respectively constructed by improving the YOLOv5 network and the GRCNN network; The data processing method of the target capture detection model includes: Inputting the image and depth information of the target to be captured into a target detection model to detect the target, thereby obtaining an image containing at least one target; Inputting an image containing at least one target into a grasping point prediction model to determine the grasping point, and obtaining the grasping point coordinates in the image acquisition device coordinate system; Based on the mapping model between the robotic arm and the image acquisition device established in advance through the camera calibration algorithm, the coordinates of the grasping point in the image acquisition device coordinate system are converted into the coordinates of the grasping point in the robotic arm coordinate system.
2. The multi-target grasping method based on deep learning according to claim 1 is characterized in that: The camera calibration method is used to establish a mapping model between the robot arm and the image acquisition device, including: Using an image acquisition device and a calibration plate to acquire multi-target images to be grasped from multiple different angles to obtain multi-angle images, and recording the posture information of the end of the robotic arm, wherein the posture information of the end of the robotic arm includes the position and direction of the end effector; Based on the multi-angle images, the intrinsic parameters and extrinsic parameters of the image acquisition device are calculated by a standard camera calibration algorithm, wherein the intrinsic parameters include focal length and optical center, and the extrinsic parameters include the position and posture of the camera relative to the robot arm coordinate system; Based on the intrinsic and extrinsic parameters of the image acquisition device and the posture information of the end of the robotic arm, a mapping model between the image acquisition device coordinate system and the robotic arm coordinate system is constructed to convert points in the image acquisition device coordinate system into points in the robotic arm coordinate system.
3. The multi-target grasping method based on deep learning according to claim 2 is characterized in that: The method further includes, after constructing the mapping model, using the mapping model to teach the robot arm operation path, including: In the multi-target image to be captured, specify at least one target; The coordinates in the coordinate system of the specified target image acquisition device are converted into the coordinates in the coordinate system of the robot arm by using the mapping model; Calculate the motion path of the robot arm according to the coordinates of the specified target robot arm coordinate system; According to the motion path of the robotic arm, inverse kinematics is used to calculate the motion trajectory and control instructions of each joint, which are input into the robotic arm control system to manipulate the robotic arm to complete the grasping action.
4. The multi-target grasping method based on deep learning according to claim 1 is characterized in that: Improve the YOLOv5 network and GRCNN network respectively to build the target detection model and the grasping point prediction model, including: Improve the YOLOv5 network, including replacing the original loss function of the YOLOv5 network with Focal-IOU, and introducing the BiFPN module in the Neck part of the YOLOv5 network; The Focal-IOU is used to receive the difference between the predicted bounding box and the actual bounding box input by the Head part of the YOLOv5 network, and output the loss value to update the weight of the YOLOv5 network; The BiFPN module is used to receive the image of the target to be captured input by the Neck part of the YOLOv5 network and multiple scale feature maps formed by depth information, and generate enhanced multi-scale feature maps through upsampling, downsampling and feature fusion operations, and output them to the Head part of the YOLOv5 network; Improve the GRCNN network, including replacing the standard two-dimensional convolutional layer in the GRCNN network with a lightweight C2f structure and introducing the ECA attention mechanism in the GRCNN network; The C2f structure is used to receive a processed feature map of an image containing at least one target input by a convolution layer or a pooling layer in the GRCNN network, generate a lightweight feature map through a lightweight operation, and output it to the next convolution layer or pooling layer of the GRCNN network; The ECA attention mechanism is used to receive the lightweight feature map output by the convolution layer or pooling layer in the GRCNN network for processing, generate an attention weight map by calculating the correlation between channels, multiply the attention weight map shown by the original feature map to obtain a weighted feature map, and output it to the next layer of the GRCNN network.
5. The multi-target grasping method based on deep learning according to claim 1 is characterized in that: Inputting the image and depth information of the target to be captured into the target detection model to detect the target, and obtaining an image containing at least one target, including: Using the target detection model, based on the image and depth information of the target to be captured, the detection frame coordinates, height and width, and center point coordinates of the captureable target in the image of the target to be captured are acquired; If the number of targets to be captured is less than 2, the center point coordinates are used as a reference to extend outward and intercept target images whose resolution exceeds a preset threshold after extension, to obtain an image containing one target; If the number of targets to be captured is greater than or equal to 2, the center point coordinates are used as a reference, and the target images whose resolution exceeds a preset threshold after extension are captured from left to right and from top to bottom to obtain an image containing at least two targets.
6. The multi-target grasping method based on deep learning according to claim 5 is characterized in that: Based on the coordinates of the center point, extend outward from left to right and from top to bottom to capture the target image whose resolution exceeds the preset threshold after extension, including: If the vertical coordinates of the objects to be captured are different, the objects to be captured are arranged in ascending order according to their vertical coordinates, and each target image whose resolution after extension exceeds a preset threshold is captured in turn; If the vertical coordinates of the objects to be captured are the same, the objects to be captured are arranged in ascending order according to the horizontal coordinates of the objects to be captured, and each target image whose resolution after extension exceeds a preset threshold is captured in turn.
7. The multi-target grasping method based on deep learning according to claim 5 is characterized in that: Inputting an image containing at least one target into a grasping point prediction model to determine the grasping point, and obtaining the grasping point coordinates in the image acquisition device coordinate system, including: If the target image with a resolution exceeding a preset threshold contains a target, the target image with a resolution exceeding the preset threshold is input into the grasping point prediction model to obtain the confidence value of each pixel position and calculate the quality sum of all pixel points, and then calculate the two-dimensional pixel coordinates of the center point of a grasping box by weighted average method; If the target image with a resolution exceeding a preset threshold contains at least two targets, the target image with a resolution exceeding the preset threshold is input into a grasping point prediction model to obtain a confidence value of each pixel position and calculate the quality sum of all pixel points, and then calculate the two-dimensional pixel coordinates of the center points of at least two grasping frames by a weighted average method; Based on the conversion relationship of camera calibration, the two-dimensional pixel coordinates of the center points of at least two grabbing frames and the depth information are mapped to the three-dimensional space, and the coordinates of the grabbing points of at least two targets in the coordinate system of the image acquisition device are calculated.
8. The multi-target grasping method based on deep learning according to claim 7 is characterized in that: The two-dimensional pixel coordinates of the center point are expressed as: ; In the formula, represents the two-dimensional pixel coordinates of the center point, represents the horizontal coordinate of the center point, represents the vertical coordinate of the center point, The two-dimensional pixel coordinates of the center point are The confidence value of the target image whose resolution exceeds the preset threshold, Represents the quality and % of all pixels.
9. The multi-target grasping method based on deep learning according to claim 8, characterized in that: Based on the mapping model between the manipulator and the image acquisition device established in advance by the camera calibration algorithm, the grabbing point coordinates in the image acquisition device coordinate system are converted into the grabbing point coordinates in the manipulator coordinate system, wherein the mapping model is expressed as: ; In the formula, , and They represent the horizontal, vertical and vertical coordinates of the grasping point in the robot arm coordinate system respectively. and They represent the rotation matrix and translation vector of the image acquisition device, respectively. , and They respectively represent the horizontal, vertical and vertical coordinates of the grabbing point in the image acquisition device coordinate system.
10. A multi-target grasping system based on deep learning, characterized in that: include: The image acquisition end includes a D435i depth camera, which is used to collect images and depth information of the target to be captured and transmit them to the industrial computer through a wired connection; An industrial computer with a built-in embedded GPU is used to input the image of the target to be grasped into a target detection model to detect the target, and obtain an image containing at least one target; input the image containing at least one target into a grasping point prediction model to determine the grasping point, and obtain the grasping point coordinates in the image acquisition device coordinate system; based on a mapping model between the robotic arm and the image acquisition device established in advance by a camera calibration algorithm, the grasping point coordinates in the image acquisition device coordinate system are converted into the grasping point coordinates in the robotic arm coordinate system and transmitted to the host computer; The host computer has a built-in MATLAB and Simulink operating environment, and is used to teach the robot arm operation path based on the mapping model, receive the coordinates of the grasping point in the robot arm coordinate system, generate the motion trajectory and control instructions of each joint of the robot arm through inverse kinematics calculation, and adjust the motion trajectory and control instructions of each joint of the robot arm according to the real-time execution state of the robot arm, and transmit the adjusted control instructions to the data transmission module; The data transmission module includes a PCAN communication module and an Ethernet communication module, which is used to control the PCAN communication module to receive the adjusted control instructions and transmit them to the robotic arm through the Ethernet communication module, and at the same time receive the real-time execution status of the robotic arm and feed it back to the host computer to dynamically monitor the movement of the robotic arm; The robotic arm, with a built-in multi-DOF servo drive unit, is used to execute the motion trajectory of each joint of the complex motion trajectory according to the received control instructions, complete the grasping action and feed back the real-time execution status to the host computer through the data transmission module.
Citation Information
Patent Citations
Target detection method and system based on improved YOLOv11
CN119992048A
Cited By
Grabbing task execution method and device, equipment, storage medium and product
CN120503219A
Multi-mode sensing abalone AI self-adaptive adsorption grabbing control system and method
CN120807938A