Visual positioning method, system and device for guiding robot to realize high-precision grabbing
By analyzing the changes in edge points, inflection points and shadows in the image on the conveyor belt in real time, calculating the object existence value and occlusion value, determining the target area and guiding the robot to grab it, the problem of insufficient accuracy in complex dynamic scenarios is solved, and high-precision item grabbing is achieved.
Patent Information
- Application Number
- CN202510629486.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In cold chain logistics transportation, traditional visual positioning technology is difficult to achieve high-precision grabbing in complex dynamic scenarios, especially when the targeted item on the conveyor belt moves and the light changes greatly, resulting in a decrease in positioning accuracy or unrecognition.
By obtaining the images during the transportation of the conveyor belt in real time, obtaining edge points, inflection points and shadow pixel points of each area, combining the difference in the number of edge points, grayscale differences and shadow pixel points, calculate the item existence value and target occlusion value of each area, and then obtain the item grabbing value, determine the target area and guide the robot to grab it.
It improves the visual positioning accuracy of the target item in complex dynamic scenarios, and can accurately identify and grab target items moving on the conveyor belt, especially when the light changes are large.
Smart Images

Figure CN120147430A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of visual positioning technology, and specifically relates to a visual positioning method, system, and device for guiding a robot to achieve high-precision grasping. Background Art
[0002] As the most widely used automation device in industrial robots, the robotic arm plays an extremely important role in scenarios such as dynamic assembly, grasping, and sorting in cold chain logistics transportation. At the same time, when facing a complex working environment, traditional robotic arms are difficult to achieve the optimal path and precise positioning. Therefore, it is necessary to combine vision technology with the robotic arm to achieve high-precision grasping, sorting, etc. in a complex environment. However, traditional target recognition technologies mainly analyze and extract the static features of the target in an ideal scenario, resulting in poor performance in terms of robustness and response time in complex dynamic scenarios.
[0003] Currently, how to improve the accuracy of visual positioning of the target by the robotic arm in a complex scenario has become a research hotspot. In the paper "Research on Robotic Arm Grasping Technology Based on Visual Positioning", an improved YOLOv5 target detection model is used to train and process the collected dataset of leaf springs, and a binocular stereo vision system is introduced to complete the recognition and positioning of the leaf springs. Finally, the grasping simulation of the leaf springs by the robotic arm is completed in the integrated development environment of the Robot Operating System (ROS).
[0004] However, there are still certain defects after processing in the above manner. Cold chain logistics transportation usually uses a conveyor belt to transport goods. Therefore, during the movement of the target item on the conveyor belt, the illumination will change greatly, which will lead to a decrease in positioning accuracy or the problem of not being able to recognize the target item when using the original visual features of the target item in an ideal situation for visual positioning and recognition. Summary of the Invention
[0005] In order to solve the above technical problems, the purpose of this application is to provide a visual positioning method, system, and device for guiding a robot to achieve high-precision grasping. The specific technical solutions adopted are as follows: In the first aspect, an embodiment of this application provides a visual positioning method for guiding a robot to achieve high-precision grasping. The method includes the following steps: Obtain images during the conveyor belt transportation in real time; Obtain the edge points in each image; evenly divide each frame of image into several regions, and obtain the item presence value of each region in the current image according to the difference in the number of edge points and the gray-scale difference between each region in the current frame and the previous frame; obtain the inflection points and shadow pixel points in each frame of image; according to the number of inflection points in each region of the current image and the change in the number of shadow pixel points in each region between the current frame and the previous frame, obtain the target occlusion value of each region in the current image; Obtain the item graspable value of each region in the current image according to the item presence value and the target occlusion value of each region in the current image, and then obtain the target region in the current image; guide the robot to grasp the target item on the target region.
[0006] Preferably, the calculation formula for the item presence value of each region in the current image is: ; where represents the item presence value of the i-th region in the current image, represents the absolute difference in the total number of edge points of the i-th region between the current frame and the previous frame, represents the absolute difference in the average gray scale of the i-th region between the current frame and the previous frame of the image.
[0007] Preferably, the method for obtaining the inflection points is: obtain the curvature of each edge point in each frame of image; record the edge points with a curvature greater than the preset inflection point threshold in each frame of image as the inflection points of each frame of image.
[0008] Preferably, the shadow pixel points are the pixel points with a gray-scale value less than or equal to the preset segmentation threshold in each frame of image.
[0009] Preferably, the calculation formula for the target occlusion value of each region in the current image is: ; where is the target occlusion value of the i-th region in the current image, is the total number of inflection points in the i-th region of the current image, is the absolute value of the difference in the number of shadow pixel points of the i-th region between the current frame and the previous frame, is the first preset constant.
[0010] Preferably, the calculation formula for the item graspable value of each region in the current image is: ; where represents the item graspable value of the i-th region of the current image, represents the item presence value of the i-th region in the current image, is the target occlusion value of the i-th region in the current image, is the second preset constant.
[0011] Preferably, the process of obtaining the target area in the current image is as follows: calculate the graspable value of the items in all areas of the current image. When the maximum value of the graspable value of the items in all areas of the current image is greater than or equal to the preset target threshold, the target area corresponding to the maximum value of the graspable value of the items is used as the target area in the current image.
[0012] Preferably, the specific process of guiding the robot to grasp the target item in the target area is as follows: the ROS integrated development environment issues a control command according to the position of the target area, plans an optimal end-feasible path, and uses joint space trajectory planning to control the joints of the robot's manipulator to move, and finally grasps the item in the target area.
[0013] In a second aspect, an embodiment of the present application provides a visual positioning device for guiding a robot to achieve high-precision grasping. The visual positioning device includes: a data acquisition module, a region analysis module, and a target region acquisition module.
[0014] The data acquisition module is used to obtain images during the conveyor belt transportation process; The region analysis module is used to obtain the item presence value of each region based on the difference in the number of edge points and the gray difference between the current frame and the previous frame of each region; based on the number of inflection points of each region and the change in the number of shaded pixel points between the current frame and the previous frame of each region, obtain the target occlusion value of each region in the current image; The target region acquisition module is used to obtain the target region in the current image and guide the robot to grasp the target item on the target region.
[0015] In a third aspect, an embodiment of the present application further provides a visual positioning system for guiding a robot to achieve high-precision grasping. The system includes a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of the visual positioning method for guiding a robot to achieve high-precision grasping described in any one of the above.
[0016] As can be seen from the above embodiments, the visual positioning method, system, and device for guiding a robot to achieve high-precision grasping provided by the embodiments of the present application at least have the following beneficial effects: When the present application uses the traditional method to perform visual recognition and positioning on the target item on the conveyor belt of cold chain logistics transportation, due to the rapid movement of the item and the stacking of items, the problem that the recognition and positioning using the original visual features of the target item will be unable to recognize or have a low recognition accuracy. By analyzing the illumination features and shadow features when the target item moves in the images collected on the conveyor belt, and combining the feature changes when the items are stacked, the graspable value of the items in each region is obtained, and the target region with the least interference and the most accurate position in the image is obtained, improving the visual positioning accuracy of the target item. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] To more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0018] Figure 1 It is a flowchart of the steps of a visual positioning method for guiding a robot to achieve high-precision grasping provided by an embodiment of the present application; Figure 2 It is a flowchart for obtaining the graspable values of items in each area of the current image provided by an embodiment of the present application; Figure 3 It is a schematic structural diagram of a visual positioning device for guiding a robot to achieve high-precision grasping provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] In order to further elaborate on the technical means and effects adopted by the present application to achieve the intended invention purpose, the following, in combination with the accompanying drawings and preferred embodiments, details the specific implementation manners, structures, features and effects of the visual positioning method, system and device for guiding a robot to achieve high-precision grasping proposed according to the present application. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.
[0020] Unless otherwise specified and limited, terms such as "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a circuit structure, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the article or device including the element. In addition, the term "and / or" used herein includes any and all combinations of one or more of the related listed items. All technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs.
[0021] The following specifically describes the specific solutions of the visual positioning method, system and device for guiding a robot to achieve high-precision grasping provided by the present application in combination with the accompanying drawings.
[0022] Please refer to Figure 1, which shows a flowchart of the steps of a visual positioning method for guiding a robot to achieve high-precision grasping provided by an embodiment of the present application. The method includes the following steps: Step 1: Real-time obtain images during the conveyor belt transportation process.
[0023] A CMOS high-definition camera is set directly above the cold chain conveyor belt at each position where the grasping robot is located to collect image data on the conveyor belt in real time. The interval time for each collection is T. In this embodiment, T is taken as 0.5 s, and the implementer can set it according to the actual transportation speed of the cold chain conveyor belt.
[0024] Step 2: Obtain the edge points in each image; evenly divide each frame of image into several regions, and obtain the item presence value of each region in the current image according to the difference in the number of edge points and the gray level difference of each region in the current frame and the previous frame; obtain the inflection points and shadow pixel points in each frame of image; according to the number of inflection points in each region of the current image and the change in the number of shadow pixel points in each region in the current frame and the previous frame, obtain the target occlusion value of each region in the current image.
[0025] Generally, for the image recognition algorithm of the target item, it will extract the features existing in the item itself, and perform matching or recognition according to the features, and finally locate its position. However, for the target item on the cold chain conveyor belt, since the illumination is fixed, and as the conveyor belt moves continuously, the features of the target item on the conveyor belt will have large differences under different illumination angles and illumination intensities, which will cause the target item to lose its original features, and finally result in deviations in the recognition of the target item. Therefore, when recognizing the target item on the cold chain conveyor belt, it is necessary to consider the dynamic change of its features during movement.
[0026] When there is no target item on the cold chain conveyor belt, although the conveyor belt is running continuously, it appears relatively stationary in the continuously captured images. Therefore, the brightness, edge position, and total number of the conveyor belt and related equipment are also relatively fixed. When the target item starts to enter the image acquisition range, due to the fixed position of the ambient light during cold chain transportation, there will be certain changes in the moving target item on the conveyor belt affected by the light, manifested as changes in the brightness and edge position of the target item as it moves. Therefore, when there are significant changes in the number of edges and brightness in a certain area of the image, it is more likely that the target item is being transported to this area at this time. In addition, due to the fixed ambient light during cold chain transportation, when the moving target item is illuminated at different angles, certain shadows will be generated, and the area near the position of the shadow is usually the area where the target item is located. The shadow is usually the area with the lowest brightness in the captured image, and the shadow area will move as the target item moves. Therefore, when the brightness value in a certain area of the image changes significantly, it is more likely that there is a target item in this area and its adjacent areas.
[0027] To characterize the characteristics of the above-mentioned target item moving with the conveyor belt, each frame of the captured image is used as the input of the Canny edge detection algorithm to obtain all the edge points in the image. Then, the captured image is divided into regions, and the image is evenly divided into n×n regions (in this embodiment, n = 6, and the implementer can adjust according to the actual size of the items transported in the cold chain). The Canny edge detection algorithm is a well-known technology, and the process of this scheme will not be elaborated here.
[0028] As a preferred implementation, the item presence value of each region in the current image is obtained according to the difference in the number of edge points and the gray scale difference between the current frame and the previous frame, which is used to characterize the possibility of the existence of the target item in each region of the current image.
[0029] In this embodiment, the item presence value of the i-th region in the current image is denoted as , and its specific expression is: ; in the formula, represents the item presence value of the i-th region in the current image, represents the absolute difference in the total number of edge points of the i-th region between the current frame and the previous frame, represents the absolute difference in the average gray scale of the i-th region between the current frame and the previous frame of the image.
[0030] The larger the value of
[0031] Furthermore, in cold chain transportation, there may be many target items on the conveyor belt at the same time, that is, there may be more than one item within the grasping range of each robot. Therefore, during the process of collecting images, there may be a situation where there are multiple target items in one image. Since there are many types of items in cold chain transportation, the shapes and sizes of various items may be different. When multiple target items are stacked, some items may be blocked. If only the above method is used for positioning and recognition, multiple targets may be recognized as one target item, resulting in problems such as inaccurate grasping or grasping failure when the robot performs grasping. Therefore, it is necessary to further improve the positioning accuracy.
[0032] When the target items are stacked, due to the different sizes and shapes of the items in cold chain transportation, some target items will be blocked by the rest of the target items themselves or their shadows, resulting in the truncation of the edges of the target items. Specifically, the edges of the blocked and unblocked target items merge into one edge. However, due to the random positions of the target items, the number of inflection points of the superimposed edge increases and the degree of inflection of the inflection points is relatively large. Also, because the cold chain conveyor belt is constantly running, under the condition of fixed ambient light, the position of the shadow generated by the illumination of the target item may shift to a certain extent. Therefore, the blocked part may change from being blocked to unblocked as the conveyor belt runs. To sum up, for each region, the more inflection points there are in each region, and the smaller the change in the area of the shadow region in the current frame compared to the previous frame, the more likely there are blocked target items in the corresponding region.
[0033] Curve fitting is performed on all edge points in the image by the least squares method, and the curvature of each edge point can be obtained according to the fitted curve. Among them, the least squares method is a well-known technology, and the specific process will not be elaborated here. The curvature at the inflection point will be relatively large or even tend to infinity. Therefore, the edge points with a curvature greater than the preset inflection point threshold are recorded as inflection points. In this embodiment, the value of the preset inflection point threshold is 100. To obtain the pixel points of the shadow region in each region, the gray values of all pixel points of each frame of the collected image are used as input, and the Otsu threshold method is used to output the segmentation threshold of the gray value, which is recorded as the preset segmentation threshold. When the gray value of a certain pixel point is less than or equal to the preset segmentation threshold, it is considered that the pixel point is a shadow pixel point. The Otsu threshold method is a well-known technology, and the specific process will not be elaborated here.
[0034] As a preferred implementation manner, according to the number of inflection points in each region of the current image and the change in the number of shadow pixel points in each region between the current frame and the previous frame, the target occlusion value of each region in the current image is obtained, which is used to characterize the degree of occlusion of the target items in each region of the current image.
[0035] In this embodiment, the target occlusion value of the i-th region in the current image is denoted as , and its specific expression is: ; In the formula, is the target occlusion value of the i-th area in the current image, is the total number of inflection points in the i-th area of the current image, is the absolute value of the difference in the number of shaded pixels between the i-th area in the current frame and the previous frame, is the first preset constant to prevent the denominator from being 0. In this embodiment, .
[0036] When the target occlusion value of a single area in the current image is larger, it indicates that the target item in that area is more occluded, and it is more likely to cause problems such as grasping failure or inaccurate grasping when guiding the robot to grasp it.
[0037] Step 3: Obtain the graspable value of the items in each area of the current image based on the item presence value and the target occlusion value of each area in the current image, and then obtain the target area in the current image; guide the robot to grasp the target item on the target area.
[0038] Further, obtain the graspable value of the items in each area of the current image based on the item presence value and the target occlusion value of each area in the current image, which is used to characterize the accuracy of the presence of target items in each area of the current image. Among them, the flowchart for obtaining the graspable value of the items in each area of the current image is as Figure 2 shown.
[0039] In this embodiment, the graspable value of the items in the i-th area of the current image is denoted as , and its specific expression is: ; In the formula, represents the graspable value of the items in the i-th area of the current image, represents the item presence value of the i-th area in the current image, is the target occlusion value of the i-th area in the current image, is the second preset constant for preventing the denominator from being 0. In this embodiment, .
[0040] The meaning of this expression is: when the possibility of the existence of a target item in a certain area of the image is greater, and the degree of occlusion of the target item is smaller, it indicates that the accuracy of the existence of the target item in that area is higher, and the robot can be guided to grasp the target item at that position more.
[0041] Taking the graspable values of items in all regions of the current image as input, and using the cross-validation method, the output is the segmentation threshold of the graspable values of items. Denote this segmentation threshold as the preset target threshold. When the maximum value of the graspable values of items in all regions of the current image is greater than or equal to the preset target threshold, the target region corresponding to the maximum value of the graspable values of items is used as the target region in the current image.
[0042] In this embodiment, a ROS integrated development environment is used to control the robot. The ROS integrated development environment is connected to the CMOS camera and the robot control cabinet through Ethernet communication. After the camera captures an image, the captured image data is uploaded to the ROS integrated development environment. The ROS integrated development environment uses the above method to obtain the target region, and then issues a control command according to the target region position. An optimal feasible path for the end effector is planned using the RRT algorithm, and the joint space trajectory planning is used to control the joints of the robot's manipulator to move, and finally grasp the item in the target region. The RRT algorithm is a well-known technology, and the specific process will not be elaborated here.
[0043] Please refer to Figure 3 , Figure 3 is a schematic structural diagram of a visual positioning device for guiding a robot to achieve high-precision grasping provided by an embodiment of the present application. In this embodiment, each unit included in the terminal is used to execute each step in the corresponding embodiment of the visual positioning method for guiding a robot to achieve high-precision grasping. Refer to Figure 3 , the visual positioning device includes: a data acquisition module, a region analysis module, and a target region acquisition module.
[0044] The data acquisition module is used to obtain images during the conveyor belt transportation process; The region analysis module is used to obtain the item presence values of each region based on the difference in the number of edge points and the gray level difference between the current frame and the previous frame of each region; and obtain the target occlusion values of each region in the current image based on the number of inflection points of each region and the change in the number of shaded pixel points between the current frame and the previous frame of each region; The target region acquisition module is used to obtain the target region in the current image and guide the robot to grasp the target item on the target region.
[0045] Based on the same inventive concept as the above method, an embodiment of the present application also provides a visual positioning system for guiding a robot to achieve high-precision grasping, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of the visual positioning method for guiding a robot to achieve high-precision grasping described in any one of the above.
[0046] Each embodiment in this application is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the differences between each embodiment and other embodiments are emphasized.
[0047] It should be noted that, unless otherwise specified and limited, terms such as "including", "comprising" or any other variants thereof are intended to cover non-exclusive inclusion, so that a circuit structure, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the article or device including the said element. In addition, the term "and / or" used herein includes any and all combinations of any one of the related listed items.
[0048] Those skilled in the art will readily conceive of other embodiments of this application after considering the specification and practicing the invention herein. This application is intended to cover any variations, uses or adaptations of this application, which follow the general principles of this application and include the common general knowledge or conventional technical means in the technical field not invented by this application.
[0049] It should be understood that this application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope.
Claims
1. A visual positioning method for guiding a robot to achieve high-precision grasping, characterized in that: The method comprises the following steps: Real-time acquisition of images during conveyor belt transportation; Obtain edge points in each image; divide each frame of image equally into several regions, and obtain the object existence value of each region in the current image according to the difference in the number of edge points and grayscale difference between the current frame and the previous frame; obtain the inflection points and shadow pixels in each frame of image; obtain the target occlusion value of each region in the current image according to the number of inflection points in each region in the current image and the change in the number of shadow pixels in each region between the current frame and the previous frame; According to the object existence value and target occlusion value of each area in the current image, the object graspable value of each area in the current image is obtained, and then the target area in the current image is obtained; the robot is guided to grasp the target object in the target area.
2. The visual positioning method for guiding a robot to achieve high-precision grasping as claimed in claim 1, characterized in that: The calculation formula for the object existence value of each area in the current image is: ; In the formula, Indicates the existence value of the item in the i-th region in the current image, represents the absolute difference between the total number of edge points of the i-th region in the current frame and the previous frame, It represents the absolute difference between the grayscale mean of the i-th region in the current frame and the previous frame.
3. The visual positioning method for guiding a robot to achieve high-precision grasping as claimed in claim 1, characterized in that: The method for obtaining the inflection point is: obtaining the curvature of each edge point in each frame image; and recording the edge point in each frame image whose curvature is greater than a preset inflection point threshold as the inflection point of each frame image.
4. The visual positioning method for guiding a robot to achieve high-precision grasping as claimed in claim 1, characterized in that: The shadow pixels are pixels in each frame image whose grayscale value is less than or equal to a preset segmentation threshold.
5. The visual positioning method for guiding a robot to achieve high-precision grasping as claimed in claim 1, characterized in that: The calculation formula of the target occlusion value of each area in the current image is: ; In the formula, is the target occlusion value of the i-th region in the current image, is the total number of inflection points in the i-th region of the current image, is the absolute value of the difference in the number of shadow pixels in the i-th region between the current frame and the previous frame, is the first preset constant.
6. The visual positioning method for guiding a robot to achieve high-precision grasping as claimed in claim 1, characterized in that: The calculation formula for the grabbable value of items in each area of the current image is: ; In the formula, Indicates the grabbable value of items in the i-th region of the current image. Indicates the existence value of the item in the i-th region in the current image, is the target occlusion value of the i-th region in the current image, is the second preset constant.
7. The visual positioning method for guiding a robot to achieve high-precision grasping as claimed in claim 1, characterized in that: The process of acquiring the target area in the current image is as follows: calculating the item graspable values of all areas in the current image, and when the maximum value of the item graspable values of all areas in the current image is greater than or equal to a preset target threshold, the target area corresponding to the maximum value of the item graspable value is used as the target area in the current image.
8. The visual positioning method for guiding a robot to achieve high-precision grasping as claimed in claim 1, characterized in that: The specific process of guiding the robot to grab the target object in the target area is as follows: the ROS integrated development environment issues a control instruction according to the position of the target area, plans an optimal terminal feasible path, uses the joint point spatial trajectory planning to control the movement of the robot's mechanical arm joints, and finally grabs the object in the target area.
9. A visual positioning device for guiding a robot to achieve high-precision grasping, characterized in that: Implementing the visual positioning method for guiding a robot to achieve high-precision grasping as described in any one of claims 1 to 8, the visual positioning device comprises: A data acquisition module, used to obtain images during the conveyor belt transportation process; The region analysis module is used to obtain the object existence value of each region based on the difference in the number of edge points and the grayscale difference between the current frame and the previous frame; based on the number of inflection points of each region and the change in the number of shadow pixels of each region between the current frame and the previous frame, obtain the target occlusion value of each region in the current image; The target area acquisition module is used to acquire the target area in the current image and guide the robot to grab the target object in the target area.
10. A visual positioning system for guiding a robot to achieve high-precision grasping, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the visual positioning method for guiding a robot to achieve high-precision grasping as described in any one of claims 1-8 is implemented.