Robot arm vision capturing method and system

By acquiring multiple types of images and using an improved single-stage detection network to identify and capture objects, combined with feature point tracking and positioning technology, the problem of difficulty in accurately identifying and positioning overlapping workpieces or products in the prior art is solved, achieving more efficient and reliable robotic visual capture.

CN119941859AInactive Publication Date: 2025-05-06聊城市检验检测中心
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510078235.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify and position overlapping workpieces or products, resulting in large errors in the robotic hand during grasping, affecting work efficiency and product quality.

Method used

The crawling object is acquired by obtaining infrared images, color images and distance images of the crawling object and identifying the crawling object using an improved single-stage detection network. After identification, feature points are marked on the crawling object and the robot hand, and feature points are tracked and positioned using color images and robot hand moving images to determine the location of the crawling object and the robot hand.

Benefits of technology

It improves the accuracy and reliability of robotic visual capture, reduces the movement error of robotic hand, and improves work efficiency and product quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941859A_ABST
    Figure CN119941859A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, in particular to a robot arm vision capturing method and system, and the method comprises the following steps: continuously obtaining an infrared image, a color image and a distance image of a captured object, and obtaining a robot arm motion image of a robot arm; based on the infrared image, the color image and the distance image, the robot arm uses an improved single-stage detection network to recognize the grabbed object; after the grabbed object is recognized, a plurality of feature points are marked on the grabbed object and the manipulator, and the color image and the manipulator motion image are used for tracking and positioning the feature points; and the positions of the grabbed object and the manipulator are determined by tracking and positioning the feature points. According to the method, the grabbed object is identified based on the infrared image, the color image and the distance image of the grabbed object and the robot arm motion image of the robot arm, so that the accuracy and the reliability of visual capture are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a robot vision capture method and system. Background Art

[0002] With the rapid development of information technology, the research and development of robot technology has become popular and has been widely promoted in some important industrial sites. The structure of the robot also tends to develop in the direction of high speed and high precision, and it can meet the needs of high load-to-weight ratio and complete the expected work tasks in real time. Among them, the robot plays a very important role in the selection of workpieces and the assembly and manufacturing of products. The precise movement of the robot is one of the key factors to improve production efficiency and product quality, and the key to the precise movement of the robot lies in the visual system that guides positioning. Therefore, improving the visual capture accuracy of the robot and then improving the visual system of the robot to reduce the movement error of the robot has naturally become an important topic in the research of robotic arms.

[0003] In the prior art, when a robot arm is used to grab a workpiece or product, the workpiece or product is usually transported to the location of the robot arm through a conveyor belt, and then the conveyor belt stops conveying, so that the robot arm can grab the workpiece or product. This grabbing method has low requirements on the robot arm's visual system, but indirectly reduces the robot arm's work efficiency. In order to solve this problem, researchers install one or more optical sensors on or near the robot arm, and enable the robot arm to visually capture the workpiece or product based on the optical image obtained by the optical sensor. However, the prior art is difficult to accurately identify the workpiece or product, and cannot locate and grab overlapping workpieces or products. Therefore, in real life, many factories still use the method of manually or through a conveyor belt to transport the workpiece or product to the location of the robot arm, and then use the robot arm to grab the workpiece or product, which is very unfavorable to the improvement of the factory's intelligence level and product production efficiency. Summary of the invention

[0004] In view of the defects in the prior art, the present invention provides a robot vision capture method and system.

[0005] In order to achieve the above-mentioned purpose, in the first aspect, the present invention provides a method for visual capture of a robot hand, the method comprising the following steps: continuously acquiring infrared images, color images and distance images of a grasped object, and acquiring a robot hand motion image of the robot hand at the same time; the robot hand identifies the grasped object based on the infrared image, the color image and the distance image using an improved single-stage detection network; after identifying the grasped object, marking multiple feature points on the grasped object and the robot hand, and tracking and locating the feature points using the color image and the robot hand motion image; determining the position of the grasped object and the robot hand based on tracking and locating the feature points. The present invention identifies the grasped object based on the infrared image, color image and distance image of the grasped object and the robot hand motion image of the robot hand, thereby improving the accuracy and reliability of visual capture.

[0006] Optionally, the step of continuously acquiring the infrared image, the color image and the distance image of the grasped object and acquiring the robot hand motion image of the robot hand comprises the following steps: Using an infrared camera to obtain an infrared image of the grasped object; Acquire a distance image of the grasped object using a TOF camera; A color image of the grasped object is acquired through the first camera, and a robot hand motion image of the robot hand is acquired through the second camera.

[0007] Optionally, the single-stage detection network is a YOLOv3 target recognition network; The robot arm uses an improved single-stage detection network to identify the grasped object based on the infrared image, the color image and the distance image, including the following steps: The single-stage detection network is improved to obtain a three-input target recognition network; The infrared image, the color image and the range image are used as inputs of the three-input target recognition network to further identify the grasped object.

[0008] Furthermore, using infrared images, color images, and range images as inputs to a three-input target recognition network to identify grasped objects can combine the advantages of various images to achieve accurate recognition and positioning of stacked grasped objects in dim environments.

[0009] Optionally, the step of improving the single-stage detection network to obtain a three-input target recognition network comprises the following steps: Adding two backbone feature extraction networks to the input end of the single-stage detection network to obtain a first target network; Adding a shape feature optimization algorithm between the feature fusion layer of the first target network and each backbone feature extraction network to obtain a second target network; For the feature fusion layer of the second target network, the features extracted at different stages are fused at the same sub-network depth to obtain the three-input target recognition network.

[0010] Furthermore, the three-input object recognition network can extract and fuse the features of different types of images, thereby realizing the complementary advantages of different types of images and achieving accurate recognition of the grasped objects.

[0011] Optionally, the shape feature optimization algorithm performs the following steps when running: Extracting a shape feature map extracted by the backbone feature extraction network, wherein the shape feature map is a contour feature map extracted based on the infrared image, the color image, and the range image; A graph coordinate system is established on the contour feature map, thereby determining the pixel coordinate position of each pixel point in the object contour on the contour feature map, and at the same time, GCN is used to extract the embedding vector of each pixel point, thereby establishing a deduplication discrimination model; The object contour is refined according to the deduplication discrimination model to obtain a unique contour feature map.

[0012] Furthermore, using a deduplication discriminant model to remove redundant contour features can improve the accuracy and reliability of contour feature maps extracted from different types of images, thereby improving the accuracy and reliability of grasped object recognition.

[0013] Optionally, the step of refining the object contour according to the deduplication discrimination model to obtain a unique contour feature map comprises the following steps: Grouping and labeling the object contours on the same contour feature map, and then pairing the object contours in pairs without duplication to obtain multiple pairs of target object contours; Calculating a deletion discrimination index between the target object contours using the deduplication discrimination model; When the deletion discrimination index is greater than a threshold, a group of the object contours are randomly deleted, and the object contours finally remaining are used as the unique contour feature map.

[0014] Optionally, the deduplication discrimination model satisfies the following relationship: Wherein, P is the deletion discrimination index, N is the minimum value of the total number of pixels in the two groups of target object contours, sigmoid is the restriction function, is the embedding vector of the nth pixel point in the first group of target object contours in the two groups of target object contours, is the embedding vector of the nth pixel point in the second group of target object contours in the two groups of target object contours, is the horizontal coordinate position of the i-th pixel point in the first group of target object contours in the two groups of target object contours in the image coordinate system, is the horizontal coordinate position of the i-th pixel point in the second group of target object contours in the two groups of target object contours in the image coordinate system, is the ordinate position of the i-th pixel point in the first group of target object contours in the two groups of target object contours in the image coordinate system, is the ordinate position of the i-th pixel point in the second group of target object contours in the two groups of target object contours in the image coordinate system.

[0015] Furthermore, the deduplication discrimination model comprehensively considers the node similarity between object contours and the positional relationship between nodes. It is not only simple to calculate, but also can more accurately judge the similarity between contours of different objects, which is conducive to removing redundant contour information and improving the accuracy of grasped object recognition.

[0016] Optionally, after the grasped object is identified, marking a plurality of feature points on the grasped object and the robot hand, and tracking and locating the feature points using the color image and the robot hand motion image comprises the following steps: Constructing a world coordinate system with any fixed point on the robot as a coordinate origin, and determining a first position and a second position of the first camera and the second camera in the world coordinate system; Marking three first-category feature points on each of the grasped objects, and marking one second-category feature point on the robot claw of the robot hand; Tracking the first type of feature points based on the color image, and recording first coordinate positions of the first type of feature points in the world coordinate system in real time; The second type of feature points are tracked according to the robot hand motion image, and the second coordinate positions of the second type of feature points in the world coordinate system are recorded in real time.

[0017] Optionally, the determining the position of the grasped object and the robot hand by tracking and positioning the feature points comprises the following steps: Calculating a unique marking coordinate position of the grasped object using the first coordinate position; The accurate positions of the grasped object and the robot hand are determined in real time by acquiring the unique marking coordinates and the second coordinate position in real time.

[0018] Furthermore, the first coordinate positions of multiple first-category feature points in the world coordinate system are used to determine the unique mark coordinate position of the grasped object, thereby more accurately locating the stacked grasped objects. Using the unique mark coordinate and the second coordinate position to determine the position of the grasped object and the robot arm in real time is beneficial for the robot arm to grasp its relative position with the grasped object in real time.

[0019] In a second aspect, the present invention provides a robot vision capture system, the system uses a robot vision capture method provided by the present invention, and the system includes: a data acquisition device, the data acquisition device is used to continuously acquire infrared images, color images and distance images of the grasped object, and simultaneously acquire the robot motion image of the robot; a data processing device, the data processing device is used for the robot to identify the grasped object based on the infrared image, the color image and the distance image using an improved single-stage detection network; after identifying the grasped object, marking multiple feature points on the grasped object and the robot, and using the color image and the robot motion image to track and locate the feature points; based on the tracking and locating of the feature points, the robot determines whether the grasped object is accurately identified; a data storage device, the data storage device is used to store all data generated in the data processing device; a data output device, the data output device is used to output the data in the data storage device.

[0020] Furthermore, the system provided by the present invention can improve the accuracy and reliability of the robot's visual capture, and can promote the development of the robot towards a more intelligent and efficient direction. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0022] Figure 1 A schematic diagram of a flow chart of a robot vision capture method according to an embodiment of the present invention; Figure 2 A schematic diagram of the framework of a robot vision capture system according to an embodiment of the present invention. DETAILED DESCRIPTION

[0023] The specific embodiments of the present invention will be described in detail below. It should be noted that the embodiments described herein are only for illustration and are not intended to limit the present invention. In the following description, a large number of specific details are set forth in order to provide a thorough understanding of the present invention. However, it is obvious to those of ordinary skill in the art that these specific details do not need to be adopted to implement the present invention. In other examples, in order to avoid confusing the present invention, known circuits, software or methods are not specifically described.

[0024] Throughout the specification, references to "one embodiment," "an embodiment," "an example," or "an example" mean that a particular feature, structure, or characteristic described in conjunction with the embodiment or example is included in at least one embodiment of the present invention. Therefore, the phrases "in one embodiment," "in an embodiment," "an example," or "an example" appearing in various places throughout the specification do not necessarily all refer to the same embodiment or example. In addition, particular features, structures, or characteristics may be combined in one or more embodiments or examples in any suitable combination and / or subcombination. In addition, it should be understood by those of ordinary skill in the art that the figures provided herein are for illustrative purposes and that the figures are not necessarily drawn to scale.

[0025] It should be noted in advance that, in an optional embodiment, except for independent explanations, the same symbols or letters appearing in all formulas have the same meanings and values.

[0026] In an alternative embodiment, see Figure 1 The present invention provides a robot vision capture method, the method comprising the following steps: S1. Continuously acquire infrared images, color images and distance images of the grasped object, and simultaneously acquire the robot hand motion image of the robot hand.

[0027] Wherein, step S1 specifically includes the following steps: S11. Using an infrared camera to obtain an infrared image of the grasped object.

[0028] S12: Acquire a distance image of the grasped object using a TOF camera.

[0029] S13, acquiring a color image of the grasped object through the first camera, and acquiring a robot hand motion image of the robot hand through the second camera.

[0030] Specifically, in this embodiment, the infrared camera used is the LD-SW320-UC infrared camera, the TOF camera model used is the LXPS-DS3221-U / E, the first camera and the second camera used are the AQX-2XE20 high-definition cameras, the infrared camera, the TOF camera and the first camera are installed on the same side of the grasped object, and the second camera is installed just above the robot arm. In other optional embodiments, other models of infrared cameras, TOF cameras, first cameras and second cameras can also be selected, and the specific selection can be determined according to the actual situation, which will not be listed here one by one.

[0031] Furthermore, multiple grasped objects may be stacked, so only a small part of the surface of some grasped objects can be directly observed without contacting other grasped objects. However, the exposed surface of some grasped objects may be difficult to be clearly presented in the color image because the light is blocked by other grasped objects, resulting in the inability to accurately identify these grasped objects using color images, which is a disadvantage of color images. Since the infrared images obtained by infrared cameras are not affected by external ambient light, relatively reliable images of grasped objects can be obtained even under relatively dim conditions, so that grasped objects that are blocked by light can also be clearly presented. Therefore, the use of infrared images can make up for this disadvantage of color images.

[0032] Furthermore, color images contain relatively rich color information and detailed texture information, while range images mainly include the boundary and contour information of the photographed object. Therefore, color images are very useful for identifying object categories, while range image data is very useful for identifying object positions. Therefore, combining color images and range images can accurately identify and locate multiple overlapping grasped objects.

[0033] S2. The robot arm identifies the grasped object using an improved single-stage detection network based on the infrared image, the color image and the range image.

[0034] Wherein, step S2 specifically includes the following steps: S21. Improving the single-stage detection network to obtain a three-input target recognition network.

[0035] Wherein, step S21 specifically includes the following steps: S211. Add two backbone feature extraction networks to the input end of the single-stage detection network to obtain a first target network.

[0036] Specifically, in this embodiment, the single-stage detection network is a YOLOv3 target recognition network, and its input end includes a backbone feature extraction network, which is Darkenet53. In this embodiment, two backbone feature extraction networks are constructed using Darkenet53 at the input end of the single-stage detection network to obtain the first target network. The only difference in structure between the first target network and the single-stage detection network is that the input end of the first target network includes three backbone feature extraction networks, while the input end of the single-stage detection network only has one backbone feature extraction network. Please refer to the prior art for the specific structure of the single-stage detection network.

[0037] Further, the three backbone feature extraction networks included in the input end of the first target network are named as the first backbone feature extraction network, the second backbone feature extraction network and the third backbone feature extraction network in sequence. In addition, the input end of the single-stage detection network mentioned in this embodiment is also directly called the backbone feature extraction network in the prior art. In order to distinguish the single-stage detection network from the first target network, this embodiment names it as the input end.

[0038] S212: Add a shape feature optimization algorithm between the feature fusion layer of the first target network and each backbone feature extraction network to obtain a second target network.

[0039] Specifically, in this embodiment, the backbone feature extraction network can extract the shape features of objects in the image, and the shape features of objects in the image are reflected by the contours of the objects in the image. Since there may be wrinkles at the edges of the grasped object, when identifying multiple stacked grasped objects, the wrinkled part may also be identified as a grasped object, thereby reducing the accuracy and reliability of grasped object identification. Therefore, it is necessary to optimize the shape features extracted by the backbone feature extraction network to improve the recognition accuracy. This embodiment specifically optimizes the shape features through a shape feature optimization algorithm. The following steps are performed when the shape feature optimization algorithm is running: F1. Extracting a shape feature map extracted by the backbone feature extraction network, wherein the shape feature map is a contour feature map extracted based on the infrared image, the color image and the range image.

[0040] F2. Establish a graph coordinate system on the contour feature map, and then determine the pixel coordinate position of each pixel point in the object contour on the contour feature map, and use GCN to extract the embedding vector of each pixel point, and then establish a deduplication discrimination model.

[0041] Specifically, in this embodiment, a graph coordinate system is established with any vertex of the contour feature map as the coordinate origin. The graph coordinate system is a two-dimensional coordinate system, and then the pixel coordinate position of each pixel in the object contour can be determined according to the number of rows and columns of pixels in the contour feature map. Next, GCN can be used to extract the embedding vector of each pixel in the object contour on the contour feature map. This is a prior art, so it will not be described in detail here.

[0042] Furthermore, the pixel coordinate position and the embedded vector are used to establish a duplicate removal discrimination model, which satisfies the following relationship: Among them, P is the deletion discrimination index, N is the minimum value of the total number of pixels in the two groups of target object contours, sigmoid is the restriction function, is the embedding vector of the nth pixel in the first set of target object contours in the two sets of target object contours, is the embedding vector of the nth pixel in the second set of target object contours in the two sets of target object contours, is the horizontal coordinate position of the i-th pixel point in the first set of target object contours in the two sets of target object contours in the image coordinate system, is the horizontal coordinate position of the i-th pixel point in the second set of target object contours in the image coordinate system, is the ordinate position of the i-th pixel point in the first set of target object contours in the two sets of target object contours in the image coordinate system, is the ordinate position of the i-th pixel point in the second group of target object contours in the two groups of target object contours in the image coordinate system.

[0043] Furthermore, the deduplication discrimination model comprehensively considers the node similarity between object contours and the positional relationship between nodes. It is not only simple to calculate, but also can more accurately judge the similarity between different contours, which is conducive to removing redundant contour information and improving the accuracy of grasping object recognition. And finally, the sigmoid function is used to limit the final calculation result to the range of (0, 1), which is conducive to clearly understanding the similarity between different object contours. Using the sigmoid function to limit the calculation result is a prior art.

[0044] F3. Refine the object contour according to the deduplication discrimination model to obtain a unique contour feature map.

[0045] Wherein, step F3 specifically includes the following steps: F31. Group and label the object contours on the same contour feature map, and then pair the object contours in pairs without duplication to obtain multiple pairs of target object contours.

[0046] Specifically, in this embodiment, for the m object contours on the same contour feature map, each object contour is divided into a group and labeled, in order: , and then pair the object contours in pairs to obtain multiple pairs of target object contours. A pair of target object contours can be used To indicate that and When pairing object contours, the principle of non-repetition is followed, that is, .

[0047] F32. Calculate the deletion discrimination index between the contours of the target objects using the deduplication discrimination model.

[0048] Specifically, in this embodiment, for each pair of target object contours on the same contour feature map, their pixel coordinate positions and embedding vectors are substituted into the deduplication discrimination model to calculate the deletion discrimination index.

[0049] Furthermore, when using the deduplication discrimination model to calculate the deletion discrimination index, since the total number of selected pixels is the minimum of the total number of pixels in the two sets of target object contours and the object contours are continuous lines, the deletion discrimination index calculated by sequentially inputting the pixel coordinate position and embedding vector of each pixel point in a clockwise direction may be different from the deletion discrimination index value calculated by sequentially inputting the pixel coordinate position and embedding vector of each pixel point in a counterclockwise direction. Therefore, in order to ensure that redundant contours can be accurately identified and removed, for the two selected sets of target object contours, the pixel coordinate position and embedding vector of each pixel point should be respectively input in a clockwise direction and a counterclockwise direction to calculate the deletion discrimination index, and the maximum value of the two calculated deletion discrimination indices should be retained.

[0050] F33. When the deletion judgment index is greater than a threshold, a group of the object contours are randomly deleted, and the object contours that are finally left are used as the unique contour feature map.

[0051] Specifically, in this embodiment, the threshold is set to 0.75, that is, when the calculated deletion discrimination index is greater than 0.75, it is determined that there is redundant contour information in the first group of target object contours and the second group of target object contours, and at this time, a group of target object contours can be randomly deleted. ,Will and Substitute the pixel coordinate position and embedding vector of into the deduplication discrimination model to calculate the deletion discrimination index. If it is greater than 0.75, then it is judged and There is a set of redundant object contours in the image, and they are randomly deleted. and After calculating the deletion discrimination index of all target object contours and deleting the corresponding object contours, the remaining object contours are used as the unique contour feature map. Using the deduplication discrimination model to remove redundant contour features can improve the accuracy and reliability of contour feature maps extracted from different types of images, thereby improving the accuracy and reliability of grasped object recognition.

[0052] Furthermore, since the shape feature optimization algorithm exists between the feature fusion layer of the first target network and each backbone feature extraction network, unique contour feature maps of the infrared image, color image and range image are obtained respectively.

[0053] S213. For the feature fusion layer of the second target network, the features extracted at different stages are fused at the same sub-network depth to obtain the three-input target recognition network.

[0054] Specifically, in this embodiment, the single-stage detection network mainly uses FPN as the feature fusion layer, which itself includes a top-down convolutional feature extraction path and a bottom-up convolutional feature extraction path. In this embodiment, these two convolutional feature extraction paths are referred to as feature fusion paths. The specific structures of these two convolutional feature extraction paths can refer to the prior art. The operation of this step is actually to replace the FPN feature fusion method in the single-stage detection network with the sub-stage feature fusion method, thereby improving the diversity of features and the receptive field of the single-stage detection network, which is beneficial to improving the recognition accuracy of the three-input target recognition network for the grasped object. Considering that the sub-stage feature fusion method is a prior art, it will not be described in detail here.

[0055] Furthermore, in other optional embodiments, a bottom-up feature fusion path may be added between the feature fusion layer and the output layer of the second target network, which is conducive to further enriching the high-dimensional semantic feature information.

[0056] S22: using the infrared image, the color image and the range image as inputs of the three-input target recognition network to further identify the grasped object.

[0057] Specifically, in this embodiment, before using the three-input target recognition network to identify the grasped object, the three-input target recognition network needs to be trained. A total of 1,200 infrared images, distance images, and color images of different objects are obtained using an infrared camera, a TOF camera, and a first camera to form a model training set, and the model training set is used to complete the training of the three-input target recognition network. The specific training method is the same as the training of the convolutional neural network in the prior art, and will not be described in detail here.

[0058] Furthermore, after completing the training of the three-input target recognition network, the infrared image is used as the input of the first backbone feature extraction network, the color image is used as the input of the second backbone feature extraction network, and the distance image is used as the input of the third backbone feature extraction network, thereby realizing the recognition of the grasped object.

[0059] S3. After the grasped object is identified, a plurality of feature points are marked on the grasped object and the robot arm, and the feature points are tracked and located using the color image and the robot arm motion image.

[0060] Wherein, step S3 specifically includes the following steps: S31: construct a world coordinate system with any fixed point on the robot as a coordinate origin, and determine the first position and the second position of the first camera and the second camera in the world coordinate system.

[0061] Specifically, in this embodiment, the selection of the coordinate origin and the first position and the second position can be set manually, which will not be described in detail here.

[0062] S32: Mark three first-category feature points on each of the grasped objects, and mark one second-category feature point on the robot claw of the robot arm.

[0063] Specifically, in this embodiment, the positions of the first type of feature points on the grasped object are random and non-overlapping. Marking multiple first type feature points on each grasped object is conducive to improving the accuracy of positioning all grasped objects, while marking the second type of feature points on the robot claw of the robot hand is conducive to reflecting the overall movement range of the machine.

[0064] Furthermore, in other optional embodiments, other numbers of first-category feature points may be marked on the grasped object, and second-category feature points may also be marked at other positions of the robot arm, but the positions where the second-category feature points are marked should be at locations where the robot arm has a larger range of motion.

[0065] S33: Track the first type of feature points based on the color image, and record the first coordinate positions of the first type of feature points in the world coordinate system in real time.

[0066] Specifically, in this embodiment, the first type of feature points are tracked and located according to the first coordinate position of the first type of feature points in the world coordinate system, which involves the conversion between the world coordinate system and the camera coordinate system and the conversion between the camera coordinate system and the image coordinate system. Considering that these are existing technologies, they will not be described in detail here.

[0067] S34: Track the second type of feature points according to the robot arm motion image, and record the second coordinate positions of the second type of feature points in the world coordinate system in real time.

[0068] Specifically, in this embodiment, the second type of feature points are tracked and located according to the second coordinate position of the second type of feature points in the world coordinate system, which involves the conversion between the world coordinate system and the camera coordinate system and the conversion between the camera coordinate system and the image coordinate system. Considering that these are existing technologies, they will not be described in detail here.

[0069] S4. Determine the positions of the grasped object and the robot arm by tracking and positioning the feature points.

[0070] Wherein, step S4 specifically includes the following steps: S41. Calculate a unique marking coordinate position of the grasped object using the first coordinate position.

[0071] Specifically, in this embodiment, if the three first coordinate positions of a determined grasping object are , and , then the unique marker coordinate position of the grasped object is Calculating the unique marker coordinate position through the first coordinate position is beneficial to improving the accuracy of positioning all grasped objects.

[0072] S42, determining the accurate positions of the grasped object and the robot hand in real time by acquiring the unique mark coordinates and the second coordinate position in real time.

[0073] Specifically, in this embodiment, the real-time unique mark coordinate position of the grasped object is calculated by acquiring the first coordinate position of the grasped object in real time, and the real-time position of the grasped object is determined by using the unique mark coordinate position. At the same time, the real-time position of the robot hand is determined by the second coordinate position acquired in real time, so that the relative position of the robot hand and the grasped object can be acquired in real time in the determined world coordinate system, which is conducive to improving the intelligence of the robot hand, and further to improving the working efficiency of the robot hand.

[0074] In other optional embodiments, the motion trajectory of the grasped object and the robot arm can be further fitted based on the known unique marker coordinates and the second coordinate position, and then the position where the grasped object may appear in the next period of time can be predicted based on the fitting curve to achieve accurate grasping of the grasped object.

[0075] It should be noted that, in some cases, the actions described in the specification can be performed in a different order and still achieve the desired results. In this embodiment, the order of steps given is only to make the embodiment appear clearer and easier to explain, rather than to limit it.

[0076] In an alternative embodiment, see Figure 2 The present invention also provides a robot vision capture system, which uses a robot vision capture method provided by the present invention. The system includes a data acquisition device A1, a data processing device A2, a data storage device A3 and a data output device A4.

[0077] The data acquisition device A1 is used to continuously acquire infrared images, color images and distance images of the grasped object, and simultaneously acquire the robot hand motion image of the robot hand.

[0078] Specifically, in this embodiment, the data acquisition device A1 includes an infrared camera, a TOF camera, a first camera and a second camera, and the data acquisition device A1 specifically performs the content described in step S1.

[0079] The data processing device A2 is used for the robot hand to identify the grasped object based on the infrared image, the color image and the distance image using an improved single-stage detection network; after identifying the grasped object, multiple feature points are marked on the grasped object and the robot hand, and the feature points are tracked and located using the color image and the robot hand motion image; and the positions of the grasped object and the robot hand are determined based on the tracking and positioning of the feature points.

[0080] Specifically, in this embodiment, the data processing device A2 is connected to the data acquisition device A1. The specific connection method is not limited here, as long as the data processing device A2 can receive the infrared image, color image, distance image and robot arm motion image collected by the data acquisition device A1. For example, the data processing device A2 can be connected to the data acquisition device A1 via a data cable.

[0081] Furthermore, the data processing device A2 specifically executes the contents described in step S2 to step S4.

[0082] The data storage device A3 is used to store all data generated in the data processing device A2.

[0083] Specifically, in this embodiment, the data storage device A3 is connected to the data processing device A2 via a data line, and the data storage device A3 is used to store the position coordinates of the grasped object and the robot arm obtained in the data processing device A2.

[0084] The data output device A4 is used to output the data in the data storage device A3.

[0085] Specifically, in this embodiment, the data output device A4 is connected to the data storage device A3 via a data line, and the data output device A4 is used to output the position coordinates of the grasped object and the robot hand in the data storage device A3.

[0086] In summary, the method provided by the present invention obtains infrared images, distance images, color images and robot motion images through infrared cameras, TOF cameras and cameras respectively, and uses the three-input target recognition network provided by the present invention to combine the advantages of infrared images, distance images and color images to identify the grasped objects of the robot, optimize the effect of robot visual capture, and improve recognition accuracy. At the same time, the three-input target recognition network provided by the present invention uses a shape feature optimization algorithm to optimize the shape features of infrared images, distance images and color images, and then after feature fusion, it can accurately reflect the contour features of the grasped object, further improve the accuracy and reliability of robot visual capture, and then accurately locate the grasped object, and also have a high recognition accuracy when identifying multiple grasped objects stacked together. The positioning of the robot combined with the robot motion image can grasp the accurate positions of the robot and the grasped object in real time, providing a reference for the robot to develop in a more efficient and intelligent direction. In addition, the system provided by the present invention, as a system using the method provided by the present invention, can improve the accuracy and reliability of robot visual capture, and promote the robot to develop in a more intelligent and efficient direction.

[0087] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. These modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be included in the scope of the claims and specification of the present invention.

Claims

1. A robot visual capture method, characterized in that: The steps include: Continuously acquire infrared images, color images, and distance images of the grasped object, and simultaneously acquire the robot hand motion image of the robot hand; The robot arm identifies the grasped object using an improved single-stage detection network based on the infrared image, the color image and the range image; After the grasped object is identified, a plurality of feature points are marked on the grasped object and the robot hand, and the feature points are tracked and located using the color image and the robot hand motion image; The positions of the grasped object and the robot hand are determined by tracking and positioning the feature points.

2. A robot vision capture method according to claim 1, characterized in that: The method of continuously acquiring the infrared image, the color image and the distance image of the grasped object and acquiring the robot hand motion image of the robot hand comprises the following steps: Using an infrared camera to obtain an infrared image of the grasped object; Acquire a distance image of the grasped object using a TOF camera; A color image of the grasped object is acquired through the first camera, and a robot hand motion image of the robot hand is acquired through the second camera.

3. A robot hand vision capture method according to claim 1, characterized in that: The single-stage detection network is a YOLOv3 target recognition network; The robot arm uses an improved single-stage detection network to identify the grasped object based on the infrared image, the color image and the distance image, including the following steps: The single-stage detection network is improved to obtain a three-input target recognition network; The infrared image, the color image and the range image are used as inputs of the three-input target recognition network to further identify the grasped object.

4. A robot hand vision capture method according to claim 3, characterized in that: The step of improving the single-stage detection network to obtain a three-input target recognition network comprises the following steps: Adding two backbone feature extraction networks to the input end of the single-stage detection network to obtain a first target network; Adding a shape feature optimization algorithm between the feature fusion layer of the first target network and each backbone feature extraction network to obtain a second target network; For the feature fusion layer of the second target network, the features extracted at different stages are fused at the same sub-network depth to obtain the three-input target recognition network.

5. A robot hand vision capture method according to claim 4, characterized in that: The shape feature optimization algorithm performs the following steps when running: Extracting a shape feature map extracted by the backbone feature extraction network, wherein the shape feature map is a contour feature map extracted based on the infrared image, the color image, and the range image; A graph coordinate system is established on the contour feature map, thereby determining the pixel coordinate position of each pixel point in the object contour on the contour feature map, and at the same time, GCN is used to extract the embedding vector of each pixel point, thereby establishing a deduplication discrimination model; The object contour is refined according to the deduplication discrimination model to obtain a unique contour feature map.

6. A robot hand vision capture method according to claim 5, characterized in that: The step of refining the object contour according to the deduplication discrimination model to obtain a unique contour feature map comprises the following steps: Grouping and labeling the object contours on the same contour feature map, and then pairing the object contours in pairs without duplication to obtain multiple pairs of target object contours; Calculating a deletion discrimination index between the target object contours using the deduplication discrimination model; When the deletion discrimination index is greater than a threshold, a group of the object contours are randomly deleted, and the object contours finally remaining are used as the unique contour feature map.

7. A robot hand vision capture method according to claim 6, characterized in that: The deduplication discrimination model satisfies the following relationship: Wherein, P is the deletion discrimination index, N is the minimum value of the total number of pixels in the two groups of target object contours, sigmoid is the restriction function, is the embedding vector of the nth pixel point in the first group of target object contours in the two groups of target object contours, is the embedding vector of the nth pixel point in the second group of target object contours in the two groups of target object contours, is the horizontal coordinate position of the i-th pixel point in the first group of target object contours in the two groups of target object contours in the image coordinate system, is the horizontal coordinate position of the i-th pixel point in the second group of target object contours in the two groups of target object contours in the image coordinate system, is the ordinate position of the i-th pixel point in the first group of target object contours in the two groups of target object contours in the image coordinate system, is the ordinate position of the i-th pixel point in the second group of target object contours in the two groups of target object contours in the image coordinate system.

8. A robot hand vision capture method according to claim 2, characterized in that: After the grasped object is identified, a plurality of feature points are marked on the grasped object and the robot hand, and the feature points are tracked and located using the color image and the robot hand motion image, including the following steps: Constructing a world coordinate system with any fixed point on the robot as a coordinate origin, and determining a first position and a second position of the first camera and the second camera in the world coordinate system; Marking three first-category feature points on each of the grasped objects, and marking one second-category feature point on the robot claw of the robot hand; Tracking the first type of feature points based on the color image, and recording first coordinate positions of the first type of feature points in the world coordinate system in real time; The second type of feature points are tracked according to the robot hand motion image, and the second coordinate positions of the second type of feature points in the world coordinate system are recorded in real time.

9. A robot hand vision capture method according to claim 8, characterized in that: The step of determining the position of the grasped object and the robot hand by tracking and positioning the feature points comprises the following steps: Calculating a unique marking coordinate position of the grasped object using the first coordinate position; The accurate positions of the grasped object and the robot hand are determined in real time by acquiring the unique marking coordinates and the second coordinate position in real time.

10. A robot vision capture system, the system using a robot vision capture method according to any one of claims 1 to 9, characterized in that: include: A data acquisition device, the data acquisition device is used to continuously acquire infrared images, color images and distance images of the grasped object, and simultaneously acquire a robot hand motion image of the robot hand; A data processing device, wherein the data processing device is used for the robot hand to identify the grasped object based on the infrared image, the color image and the range image using an improved single-stage detection network; after identifying the grasped object, mark multiple feature points on the grasped object and the robot hand, and track and locate the feature points using the color image and the robot hand motion image; determine the position of the grasped object and the robot hand according to the tracking and positioning of the feature points; A data storage device, the data storage device is used to store all data generated by the data processing device; A data output device is used to output the data in the data storage device.