A robot vision-guided grasping method
By combining virtual reality technology and monocular vision robot vision guidance method, the image is optimized using cycleGAN and yolov7 algorithm and the GeIOU function is added, the high cost and poor robustness of pose measurement of industrial robots 6DOF is solved, and high-precision and low-cost pose measurement are achieved.
Patent Information
- Application Number
- CN202310704172.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-14
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2043-06-14
AI Technical Summary
The existing industrial robot visual guidance method is costly in 6DOF position measurement, poor measurement robustness, and is susceptible to external ambient light interference.
A robot vision guidance method based on monocular vision is adopted, combined with virtual reality technology image data augmentation algorithm and multi-key point detection model, and image optimization and key point detection are used to use cycleGAN network and yolov7 algorithm for image optimization and key point detection, and the pose is calculated through the EPnP algorithm, and the gradient loss function and GeIOU function are added to improve detection accuracy.
It realizes low-cost and high-precision 6DOF posture measurement, which improves the robustness and measurement accuracy of the detection model and reduces the impact of external ambient light interference.
Smart Images

Figure CN116787432B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent manufacturing of high-end equipment, and in particular to a robot vision-guided grasping method. Background Art
[0002] Machine vision technology is at the core of artificial intelligence, and robotic vision guidance technology based on machine vision is playing an increasingly important role in robotics-related integration and applications. Existing 6DOF robotic vision-guided positioning methods in the industrial sector primarily rely on stereo vision or structured light systems. These methods generally suffer from slow measurement speeds, small measurement areas, and high costs. Existing monocular vision-based 6DOF robotic vision guidance and pose measurement methods also suffer from low positioning accuracy and susceptibility to interference from ambient light.
[0003] This paper addresses the high cost and poor robustness of 6DOF pose measurement when robots grasp objects in industrial environments. We propose a monocular vision-guided robot measurement strategy that achieves high-precision and robust measurement of the target workpiece's 6DOF pose. The proposed method primarily consists of two components: an image data enhancement algorithm based on virtual reality technology and a 6DOF pose measurement algorithm that combines a multi-keypoint detection model with the Epnp algorithm. The former uses image enhancement technology to enhance data for small samples of industrial objects, addressing the problem of poor robustness of detection models due to high image acquisition costs and long acquisition cycles. The latter completes the 6DOF pose measurement of the target workpiece using a single image, achieving low-cost 6DOF pose measurement of the target workpiece using a monocular camera. Summary of the Invention
[0004] In order to solve the above technical problems, the present invention proposes a robot vision-guided grasping method. The technical problem to be solved by the present invention is achieved by the following technical solutions:
[0005] A robot vision-guided grasping method, the method comprising the following steps:
[0006] Step 1: Input the workpiece pattern to be grabbed, and use the virtual engine to generate virtual images of the workpiece in different backgrounds, different environments, and different quantities;
[0007] Step 2: Crop and splice the images collected by the vision system, and use a random algorithm to generate evenly distributed images of the workpiece at different positions;
[0008] Step 3: Use the cycleGAN network to optimize the image generated in the second step, and introduce a gradient loss function and a multi-channel hybrid attention mechanism to improve image clarity and enhance the learning efficiency and generation effect of the detection model;
[0009] Step 4: Use the yolov7 algorithm to predict the key points of the workpiece surface and use the proposed GeIOU function as the loss function to improve the prediction accuracy;
[0010] Step 5: Use the EPnP algorithm to convert the key points on the workpiece surface into position key points, and complete the grasping position calculation through 6DOF pose calculation.
[0011] The gradient loss function formula in the third step is: LossT = |Grad(X)-Grad(Y)|×α;
[0012] Where X is the input image, Y is the output image generated by the network, and α is the weight coefficient of LossT.
[0013] The improved cycleGAN network loss function is Loss=Loss cycle +LossT;
[0014] Among them, Loss cycle is the loss function of the original cycleGAN network.
[0015] The calculation formula for the predicted key points in the fourth step is as follows:
[0016]
[0017] Among them, IOU is the intersection-over-union ratio of the true area of the key point to the predicted area, ρ 2 (A, B) is the Euclidean distance between the center coordinates of the predicted value and the true value, and c is the diagonal distance of the smallest box that encloses them.
[0018] The grasping position calculation in the fifth step is performed by the following steps:
[0019] Step 1: When the industrial robot grabs the workpiece, it uses the formula Calculate, where Grasping pose for industrial robots. T T P is the position of the workpiece in the robot gripper coordinate system;
[0020] When an industrial robot takes a picture of a workpiece, it uses the formula Calculate, where The robot pose when acquiring the image for the vision system. CT T C Already obtained in hand-eye calibration. C T P The position of the workpiece relative to the vision system is directly obtained by the vision system;
[0021] Step 2: Combine the formulas in the first step to get
[0022] Step 3: According to the formula in step 2, when the industrial robot takes pictures of workpieces at different positions, the formula Calculate, where and is the coordinate of the workpiece detected by the vision system, and That is the grasping posture of the target workpiece in the robot coordinate system.
[0023] The present invention has the following beneficial effects: the present invention improves the learning efficiency and generation effect of the network through the gradient loss function and the multi-channel hybrid attention mechanism, while eliminating the grayscale difference between the artifact and the background in the spliced image while ensuring the clarity of the generated image. In the present invention, the accuracy of the network's key point detection is improved by adding GeIOU to the loss function of Yolov7. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The present invention will be further described below with reference to the accompanying drawings and examples.
[0025] Figure 1 It is a schematic diagram of the workflow of the present invention;
[0026] Figure 2 This is a schematic diagram of the traditional yolov7 network structure of the present invention;
[0027] Figure 3 This is a schematic diagram of the improved yolov7 network structure of the present invention. DETAILED DESCRIPTION
[0028] In order to enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be explained more clearly and completely below in conjunction with the drawings in the embodiments. Of course, the described embodiments are only a part of the present invention, not all of it. Based on this embodiment, other embodiments obtained by those skilled in the art without creative work are all within the scope of protection of the present invention.
[0029] like Figures 1 to 3 As shown, a robot vision-guided grasping method comprises the following steps:
[0030] Step 1: Input the image of the workpiece to be captured, and use the virtual engine to generate virtual images of the workpiece in different backgrounds, different environments, and different quantities. Use the virtual engine to directly create small sample images of the target workpiece. Combined with the virtual engine's rendering function, images of the workpiece in different backgrounds and lighting environments can be obtained.
[0031] Step 2: Cropping and splicing the images captured by the visual system, and using a random algorithm to generate evenly distributed images of the workpiece at different positions. By using image cropping and splicing techniques in conjunction with a random distribution algorithm, evenly distributed images of the workpiece at different positions are generated. This enriches the image data while reducing the probability of falling into local minima during deep neural network training.
[0032] Step 3: Optimize the image generated in the second step using the cycleGAN network, and introduce a gradient loss function and a multi-channel hybrid attention mechanism to improve the image clarity of the detection model and improve the learning efficiency and generation effect; there is an obvious rectangular box in the gradient map of the spliced image. In response to this situation, the present invention designs an image gradient loss function. Through the gradient loss function, the grayscale difference between the workpiece and the background in the spliced image is eliminated while ensuring the clarity of the generated image; the learning efficiency and generation effect of the network are improved by combining the multi-channel hybrid attention mechanism with the cycleGAN network; the problem of inconsistent background of the splicing algorithm is solved through the improved cycleGAN image generation technology, thereby improving the robustness and stability of the subsequent target detection network for industrial small sample objects;
[0033] Step 4: Use the yolov7 algorithm to predict the key points of the workpiece surface and use the proposed GeIOU function as the loss function to improve the prediction accuracy; Figures 2 to 3The paper shows an improvement to the YOLOv7 network using three modules: the swin-transformer, PConv, and GAM. The swin-transformer combines the powerful modeling capabilities of the transformer architecture with important visual signals. Compared to traditional convolutional neural network methods, the swin-transformer demonstrates significant advantages in training efficiency. Furthermore, the transformer architecture can be used independently or mixed with conventional convolutional networks, demonstrating excellent scalability. Replacing the first CBS module in the input of the original network with a swin-transformer module improves the model's generalization performance for object recognition. Furthermore, the windowing operation included in the swin-transformer restricts attention to a single window, which not only introduces the locality of conventional CNN convolution operations but also saves computation. By fusing a partial convolution module (PConv) with a global attention mechanism (GAM), an improved ELAN computational module is proposed. This improved ELAN module is named GC-ELAN. Among them, PConv has the advantages of being fast and efficient, and PConv has very little computational complexity, so it can greatly improve the speed of the model during training and reasoning, while the global attention module takes into account the characteristics of channel attention and spatial attention. Compared with most attention mechanisms, GAM can retain most of the feature map detail information, extract more complete information for smaller detection targets, and effectively improve the accuracy in the model. Five GC-ELAN modules are used to replace the ELAN-1 module in the backbone of the original YOLOv7 network. The replaced network will have better accuracy performance and will not introduce too many useless parameters. In addition, in order to better detect key points, the present invention adds 6 branches of key point detection to the output head of yolov7, which predicts multiple key points on the surface of the workpiece while detecting the target object; Figures 2 to 3 As shown in the figure: ELAN-1 represents the computing module of stacking multiple convolutions in the original YOLOv7; SPPCSPC represents the feature pyramid structure; cat represents the feature fusion operation; ELAN-GAM represents the integration of the GAM attention mechanism on the basis of the original ELAN-1; ELAN-Sim represents the integration of the SimAM attention mechanism on the basis of the original ELAN-1.
[0034] Step 5: Use the EPnP algorithm to convert the key points on the workpiece surface into position key points, and complete the grasping position calculation through 6DOF pose calculation.
[0035] The gradient loss function formula in the third step is: LossT = |Grad(X)-Grad(Y)|×α;
[0036] Where X is the input image, Y is the output image generated by the network, and α is the weight coefficient of LossT.
[0037] The improved cycleGAN network loss function is Loss=Loss cycle +LossT;
[0038] Among them, Loss cycle is the loss function of the original cycleGAN network.
[0039] The calculation formula for the predicted key points in the fourth step is as follows:
[0040]
[0041] Among them, IOU is the intersection-over-union ratio of the true area of the key point to the predicted area, ρ 2 (A, B) is the Euclidean distance between the center coordinates of the predicted value and the true value, and c is the diagonal distance of the smallest box that encloses them; IoU is the intersection-over-union ratio of the predicted box and the true value box. The present invention adds the geometric intersection-over-union ratio GeIOU to the loss function of yolov7 to further improve the accuracy of the network's key point detection. GeIOU calculates the minimum circumscribed polygon of the shape composed of the true value and the predicted key points, and calculates IOU from it. Combined with the DIOU function, Ge-IOU is calculated, and the formula is:
[0042] The grasping position calculation in the fifth step is performed by the following steps:
[0043] Step 1: When the industrial robot grabs the workpiece, it uses the formula Calculate, where Grasping pose for industrial robots. T T P is the position of the workpiece in the robot gripper coordinate system;
[0044] When an industrial robot takes a picture of a workpiece, it uses the formula Calculate, where The robot pose when acquiring the image for the vision system. CT T C Already obtained in hand-eye calibration. C T P The position of the workpiece relative to the vision system is directly obtained by the vision system;
[0045] Step 2: Combine the formulas in the first step to get
[0046] Step 3: According to the formula in step 2, when the industrial robot takes pictures of workpieces at different positions, the formula Calculate, where and is the coordinate of the workpiece detected by the vision system, and That is the grasping posture of the target workpiece in the robot coordinate system.
[0047] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and description merely illustrate the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.
Claims
1. A robot vision-guided grasping method, characterized by: The method comprises the following steps: Step 1: Input the workpiece pattern to be grabbed, and use the virtual engine to generate virtual images of the workpiece in different backgrounds, different environments, and different quantities; Step 2: Crop and splice the images collected by the vision system, and use a random algorithm to generate evenly distributed images of the workpiece at different positions; Step 3: Use the cycleGAN network to optimize the image generated in the second step, and introduce the gradient loss function and multi-channel hybrid attention mechanism; Step 4: Use the yolov7 algorithm to predict the key points on the workpiece surface, and use the proposed GeIOU function as the loss function; Step 5: Use the EPnP algorithm to convert the key points on the workpiece surface into pose key points, and complete the grasping position calculation through 6DOF pose calculation; The grasping position calculation in the fifth step is performed through the following steps: S1: When an industrial robot grabs a workpiece, it uses the formula Calculate, where For industrial robots to grasp the pose, is the position of the workpiece in the robot gripper coordinate system; When an industrial robot takes a picture of a workpiece, it uses the formula Calculate, where is the robot pose when the vision system acquires the image, It has been obtained in the hand-eye calibration. The position of the workpiece relative to the vision system is directly obtained by the vision system; S2: Combining the formulas in S1, we can get ; S3: According to the formula in S2, when the industrial robot takes pictures of workpieces at different positions, the formula Calculate, where is the coordinate of the target workpiece detected by the vision system, and That is the grasping posture of the target workpiece in the robot coordinate system.
2. A robot vision-guided grasping method according to claim 1, characterized in that: The gradient loss function formula in the third step is: ; in X is the input image, Y is the output image generated by the network, for LossT The weight coefficient of .
3. The robot vision-guided grasping method according to claim 2, characterized in that: The loss function of the CycleGAN network after introducing the gradient loss function and the multi-channel hybrid attention mechanism is ; in is the loss function of the original cycleGAN network.
Citation Information
Patent Citations
Electric power inspection few-sample image data enhancement method and system
CN111968048A
Method for grabbing target object by mechanical arm based on visual information fusion
CN112171661A
Device for generating prediction image on basis of generator including concentration layer, and control method therefor
US20210326650A1