Robot grabbing control method and device

Through binocular cameras and deep learning technology, the problems of inaccurate target depth information and incorrect grasping point planning in robot grasping control are solved, and a higher crawling success rate and depth information accuracy are achieved.

CN119952710APending Publication Date: 2025-05-0958 INTELLIGENT TECH (HANGZHOU) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510229645.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

In the prior art, in robot grasp control, there are problems such as incorrect planning of target grab points, inaccurate depth information of target center point, and serious depth information errors in target occlusion.

Method used

By combining the target detection network and image segmentation algorithm using a binocular camera, the ROI of the target is extracted, dedistortion processing and feature extraction are performed, pixel disparity is calculated to obtain the depth information of the target, and the grab point is determined based on the depth information and the target outer contour.

Benefits of technology

It improves the accuracy of the depth information of the target object, ensures the accuracy of the grab point, and enhances the success rate of the robot's grabbing of the target object.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119952710A_ABST
    Figure CN119952710A_ABST
Patent Text Reader

Abstract

The invention provides a robot grabbing control method and device, and belongs to the technical field of robots, and the method specifically comprises the steps: putting a target object into a working region, carrying out the feature extraction of an intercepted region, obtaining an ROI region in a right image, carrying out the image processing, comparing with the features of a left image, generating an ROI in the right image, and carrying out the recognition of the ROI region in the right image; the pixel parallax of the left image and the right image is calculated on the basis of the ROI in the left image and the ROI in the right image, the depth information position of a target object with the left image as the reference is obtained on the basis of the pixel parallax of the left image and the right image, the pixel center position of the target object in the left image is determined, and the depth information position of the target object in the right image is determined on the basis of the depth information position and the pixel center position. The space coordinates of the target are obtained, the left and right grabbing points of the target are obtained according to the space coordinates and the outer contour of the target, the left and right grabbing points are used for controlling the robot to grab the target object, the precision of binocular distance measurement is improved, and a favorable guarantee is provided for successfully grabbing the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of robots, and in particular relates to a robot grasping control method and device. Background Art

[0002] In remote mountainous areas or plateaus with harsh environments, the working environment is relatively changeable. This makes it an urgent technical problem to realize the grasping, identification and control of robots and realize that robot equipment can replace border guards. Specifically, in the invention patent application CN202410692264.1 "A 3D positioning and grasping method and system based on 2D camera", a high-precision ranging sensor is combined with a 2D camera, and posture correction is achieved through an algorithm to achieve stable grasping operation of workpieces with irregular surfaces. However, the existing technical solutions have the following technical problems:

[0003] If the grasping point is planned based only on the target detection algorithm, and the position of the target is not facing the camera origin, the grasping point will not be planned on the target body, but will be planned outside the target, and the target cannot be grasped; if the center point of the target is on the upper surface of the target, not on the side facing the camera, there will be a large error in the depth information obtained by the binocular camera, resulting in inaccurate ranging; when the side of the target facing the camera is severely obstructed, the center position of the target will be inaccurate, which will affect the accuracy of the depth information of the target.

[0004] In summary, the shortcomings of the existing technology can be summarized as follows: incorrect planning of target capture points, inaccurate depth information of the target center point, and large depth information errors due to severe target occlusion.

[0005] In view of the above technical problems, the present invention provides a robot grasping control method and device. Summary of the invention

[0006] To achieve the purpose of the present invention, the present invention adopts the following technical solutions:

[0007] According to one aspect of the present invention, a robot grasping control method and device are provided.

[0008] A robot grasping control method, specifically comprising:

[0009] Place the target object in the working area and send the left image of the binocular camera to the target detection network to extract the ROI of the left image;

[0010] The ROI in the left image is processed to obtain the grayscale ROI area in the left image, and the area corresponding to the grayscale ROI area in the right image is intercepted to obtain the intercepted area;

[0011] The intercepted area is subjected to feature extraction to obtain the ROI area in the right image, and after image processing, the ROI in the right image is generated by comparing it with the features of the left image. The pixel disparity between the left image and the right image is calculated based on the ROI in the left image and the ROI in the right image.

[0012] Based on the pixel parallax of the left and right images, the depth information position of the target object based on the left image is obtained, and the pixel center position of the target object in the left image is determined. Based on the depth information position and the pixel center position, the spatial coordinates of the target are obtained. According to the spatial coordinates and the outer contour of the target, the left and right grasping points of the target are obtained, and the left and right grasping points are used to control the robot to perform the grasping task on the target.

[0013] A further technical solution is that before the left image of the binocular camera is sent to the target detection network, the left image needs to be dedistorted.

[0014] A further technical solution is to perform a dedistortion process on the left image, specifically including:

[0015] Establish the conversion relationship between the world coordinate system of a single camera of the binocular camera and the pixel plane coordinate system, and construct a conversion function based on the pixel size and camera focal length;

[0016] The left image is dedistorted using a conversion function.

[0017] A further technical solution is to process the ROI of the left image to obtain the grayscale ROI area in the left image, specifically including:

[0018] The left image is sent to the target detection network to preliminarily extract the rectangular ROI of the object to be detected;

[0019] The rectangular ROI is sent to the image segmentation model to further extract the ROI area, and the ROI area is converted into grayscale to obtain the grayscale ROI area in the left figure.

[0020] The method for determining the depth information position is:

[0021] The left and right camera image epipolar lines of the binocular camera must be on the same horizontal line, so that the pixels to be tested in the right image and the left image are on the same horizontal line to achieve optimal matching;

[0022] The difference between the image to be tested and the pixel to be tested in the horizontal direction is determined by using the pixel disparity of the left image and the right image, and the depth information can be calculated by deriving a formula from the epipolar correction part using the difference, so as to obtain a disparity map with the same size as the ROI of the left image;

[0023] Based on the disparity map, similar triangles are used to infer the depth information position of the target object based on the left image.

[0024] On the other hand, the present application provides a robot grasping control device, which adopts the above-mentioned robot grasping control method, specifically comprising:

[0025] ROI extraction module, interception area acquisition module, pixel time difference evaluation module, grab control module;

[0026] The ROI extraction module is responsible for placing the target object in the working area and sending the left image of the binocular camera to the target detection network to extract the ROI of the left image;

[0027] The clipped region acquisition module is responsible for processing the ROI of the left image to obtain the grayscale ROI region in the left image, and clipping the region corresponding to the grayscale ROI region in the right image to obtain the clipped region;

[0028] The pixel time difference evaluation module is responsible for extracting features from the intercepted area to obtain the ROI area in the right image, and performing image processing to compare the features with the left image to generate the ROI in the right image, and calculating the pixel disparity between the left image and the right image based on the ROI in the left image and the ROI in the right image;

[0029] The grasping control module is responsible for obtaining the depth information position of the target object based on the left image based on the pixel parallax of the left image and the right image, determining the pixel center position of the target object in the left image, and obtaining the spatial coordinates of the target based on the depth information position and the pixel center position, and obtaining the left and right grasping points of the target according to the spatial coordinates and the outer contour of the target, and using the left and right grasping points to control the robot to perform the grasping task on the target.

[0030] The beneficial effects of the present invention are:

[0031] Through the organic combination of deep learning and hardware equipment, this solution can, on the one hand, reduce the dimension of the three-dimensional grasping task and lower the threshold for the successful deployment of subsequent algorithms on the hardware platform. On the other hand, the above solution can further improve the accuracy of image pixels, lay a solid foundation for the subsequent binocular ranging based on pixels to calculate the depth information of the object, improve the accuracy of binocular ranging, and provide favorable guarantees for the successful grasping of the target object.

[0032] Other features and advantages will be described in the following description. The objects and other advantages of the present invention are realized and obtained by the structures particularly pointed out in the description and drawings.

[0033] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The above and other features and advantages of the present invention will become more apparent by describing in detail exemplary embodiments thereof with reference to the attached drawings.

[0035] Figure 1 is a flow chart of a robot grasping control method;

[0036] Figure 2 It is a flow chart of the method for dedistorting the left image;

[0037] Figure 3 It is a flowchart of the specific steps of building an image segmentation model;

[0038] Figure 4 is a flow chart of a method for determining a depth information position;

[0039] Figure 5 It is a framework diagram of a robot grasping control device. DETAILED DESCRIPTION

[0040] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.

[0041] The existing technical solutions only plan the grasping points based on the target detection algorithm. If the position of the target object is not directly facing the camera origin, the grasping point planning will not be on the target object itself, but will be planned outside the target object, and the target object cannot be grasped; if the center point of the target object is on the upper surface of the target object, not on the side facing the camera, the binocular camera will obtain the depth information of the target object with a large error, resulting in inaccurate distance measurement; when the side of the target object facing the camera is severely occluded, the center position of the target object will be inaccurate, thereby affecting the accuracy of the depth information of the target object. In summary, the shortcomings of the existing technology can be summarized as follows: incorrect planning of the target grasping points, inaccurate depth information of the target center point, and large errors in depth information due to severe target occlusion.

[0042] The technical means of the present invention achieves the following technical effects:

[0043] After using the target detection algorithm to initially select the target, the image segmentation algorithm is used to calculate the center point of the image, and then the correct grasping point is planned to improve the accuracy of the target's depth information;

[0044] Use image segmentation algorithm to calculate the center point of image pixels, and calculate the depth information of the target object through the precise center point;

[0045] Use binocular ranging principle to obtain accurate target location information, increasing the possibility of successful capture of the target.

[0046] Technical solution: Collect data on the target object. The continuous video stream data is passed through the target detection algorithm to select the approximate location information of the target object. The area size of the target object is segmented through the image segmentation algorithm, and the center position of the object is calculated. The depth information of the object is calculated through the visual difference of the binocular camera. The above solution can effectively solve the depth information of the target object's position, correctly plan the grasping point of the target object, and increase the possibility of the target object being successfully grasped by the machine equipment.

[0047] ROI stands for Region of Interest, which refers to the region of interest in image processing.

[0048] The following will further explain from two perspectives: Example 1 and Example 2.

[0049] Example 1

[0050] To solve the above problems, according to one aspect of the present invention, Figure 1 As shown, a robot grasping control method is provided, which specifically includes:

[0051] Place the target object in the working area and send the left image of the binocular camera to the target detection network to extract the ROI of the left image;

[0052] The ROI in the left image is processed to obtain the grayscale ROI area in the left image, and the area corresponding to the grayscale ROI area in the right image is intercepted to obtain the intercepted area;

[0053] The intercepted area is subjected to feature extraction to obtain the ROI area in the right image, and after image processing, the ROI in the right image is generated by comparing it with the features of the left image. The pixel disparity between the left image and the right image is calculated based on the ROI in the left image and the ROI in the right image.

[0054] Based on the pixel parallax of the left and right images, the depth information position of the target object based on the left image is obtained, and the pixel center position of the target object in the left image is determined. Based on the depth information position and the pixel center position, the spatial coordinates of the target are obtained. According to the spatial coordinates and the outer contour of the target, the left and right grasping points of the target are obtained, and the left and right grasping points are used to control the robot to perform the grasping task on the target.

[0055] Furthermore, before the left image of the binocular camera is sent to the target detection network, it is also necessary to perform a dedistortion process on the left image.

[0056] Specifically, Figure 2As shown, the left image is subjected to a dedistortion process, specifically comprising:

[0057] Establish the conversion relationship between the world coordinate system of a single camera of the binocular camera and the pixel plane coordinate system, and construct a conversion function based on the pixel size and the camera focal length;

[0058] The left image is dedistorted using a conversion function.

[0059] Because factors such as camera installation and lens distortion will affect the pixel accuracy of the target object, in order to improve the pixel accuracy in the video stream, it is necessary to dedistort the camera accuracy. Camera distortion mainly consists of lateral distortion and tangential distortion. The distortion model can be expressed by the following formula:

[0060]

[0061] Among them: (x0, y0) linear model obtains the image point coordinates (x, y) is the real point coordinates, δ x ,δ y Nonlinear distortion function.

[0062] According to the actual situation, the conversion relationship between the world coordinate system of a single camera and the pixel plane coordinate system is established, which is expressed as follows:

[0063]

[0064] Among them, s is the scale factor, (u, v) is the projection coordinates of the target object on the pixel plane, dX, dY pixel size, f is the camera focal length, R is the rotation matrix, t is the translation matrix, (X w ,Y w ,Z w ) The world coordinates of the target object. Through the above formula, we can obtain the internal and external parameters of the camera and eliminate the influence of distortion on the pixels.

[0065] Further, the ROI of the left image is processed to obtain the grayscale ROI area in the left image, specifically including:

[0066] The left image is sent to the target detection network to preliminarily extract the rectangular ROI of the object to be detected;

[0067] The rectangular ROI is sent to the image segmentation model to further extract the ROI area, and the ROI area is converted into grayscale to obtain the grayscale ROI area in the left figure.

[0068] It should be noted that the target detection network includes three parts: a backbone feature network, an algorithm detection head and an enhanced feature extraction network.

[0069] Specifically, Figure 3 As shown, the specific steps of constructing the image segmentation model are:

[0070] Convolution and pooling are performed on the rectangular ROI, and multiple pooling is performed to obtain features of various sizes. The features of the smallest image size are upsampled or deconvolved to obtain a feature map of the image size corresponding to the next smallest image size. The feature map of the image size corresponding to the next smallest image size is used to perform channel stitching with the features of the next smallest image size to obtain a stitched image.

[0071] The above steps are used to perform channel stitching processing on the stitched image to obtain a stitched processed image with the same size as the input image, and the rectangular ROI frame of the stitched processed image is sent to the segmentation network to obtain the ROI area.

[0072] In one embodiment, in the field of image segmentation, traditional segmentation algorithms are easily affected by the environment, such as watershed algorithms. In order to increase the robustness of the algorithm, the present invention adopts an image segmentation algorithm based on deep learning. The Unet image segmentation algorithm structure first convolves and pools the image, and needs to be pooled 4 times. The original size of the image is 224x224, and after a series of pooling, it becomes 112x112, 56x56, 28x28, and 14x14 four different size features. Then upsample or deconvolve the 14x14 features to get 28x28 features. The 28x28 feature map is concatenated with the 28x28 features obtained by front pooling. Then, the concatenated feature map is convolved and upsampled to get a 56x56 feature map. The 56x56 features obtained by front pooling are concatenated, convolved, and upsampled. After four upsamplings, a 224x224 prediction result with the same size as the input image can be obtained. The first half of the Unet network is feature extraction, and the second half is the upsampling operation. The target detection roughly captures the target box and sends it to the Unet segmentation network to get a new ROI to match with the right image, with a moving step of 1 pixel; the precise matching is to match the new ROI with the target box after segmentation of the right image, with a moving step of 1 pixel. The absolute value of the difference is used for evaluation during the matching process, and the feature ROI extracted by image segmentation is converted into a grayscale image. The SAD function is expressed as follows:

[0073]

[0074] It should be noted that the specific steps of generating the intercepted area are:

[0075] The grayscale ROI area in the left image is template matched with the right image to obtain a matching result, and the matching result is used to cut out the area corresponding to the grayscale ROI area in the right image to obtain a cutout area.

[0076] Furthermore, the template matching includes global matching, semi-global matching and local matching.

[0077] In one embodiment, the specific steps of template matching are:

[0078] Dividing the right image into a plurality of sub-regions using the size of the grayscale ROI region, and determining matching sub-regions among the sub-regions according to the matching results between different sub-regions and the image features of the grayscale ROI region;

[0079] Obtaining a detection frame of the grayscale ROI area, and determining a to-be-detected area of ​​the detection frame based on a grayscale ROI area whose distance from the detection frame is within a preset distance range;

[0080] Determine the matching coefficient of the area to be detected according to the matching result between the image feature of the area to be detected and the matching sub-area, and when the matching coefficient does not meet the requirement, determine the moving step length according to a preset number of pixels, and use the moving step length to expand the matching sub-area to obtain an expanded sub-area;

[0081] Until the matching result between the image features of the extended sub-region and the region to be detected meets the requirement, the extended sub-region is used as the matching region of the grayscale ROI region in the right image, and the matching region is used as the clipping region.

[0082] In another embodiment, the specific steps of determining the matching sub-region are:

[0083] S11 divides the sub-region into a plurality of sub-regions according to the size of the grayscale ROI region, determines the regional feature matching coefficient according to the matching results between different sub-regions and the image features of the grayscale ROI region, and judges whether the regional feature matching coefficient meets the requirements. If so, the sub-region whose regional feature matching coefficient meets the requirements is used as the matching sub-region. If not, proceed to the next step.

[0084] S12 determines whether the regional feature matching coefficient is greater than a preset matching coefficient. If so, the sub-region with the largest regional feature matching coefficient is taken as the target sub-region and the process proceeds to step S14. If not, the process proceeds to the next step.

[0085] S13: taking a sub-region whose regional feature matching coefficient is within a preset matching coefficient interval as a matching sub-region, and dividing the matching sub-region by using a second preset size to obtain divided sub-regions, determining a matching divided region according to the matching results of the image features of different divided regions and the image features of the grayscale ROI region, and taking it as a target sub-region;

[0086] S14 divides the target sub-region into secondary sub-regions, determines secondary matching sub-regions based on the image features in the secondary sub-regions and the matching results of the grayscale ROI region, and combines the secondary matching sub-regions according to the size of the grayscale ROI region to obtain matching sub-regions.

[0087] Furthermore, the method for determining the ROI area in the right figure is:

[0088] The cutout area in the right image is sent to the image segmentation model for feature extraction to obtain the target features in the right image, and the matching results of the target features are used to determine the ROI area in the right image.

[0089] It can be understood that the method for determining the ROI in the right figure is:

[0090] The ROI region in the right image is grayed and binarized to obtain a processed ROI region, and the features of the processed ROI region are compared with those of the left image to generate the ROI in the right image.

[0091] Specifically, the method for determining the pixel disparity between the left image and the right image is:

[0092] The ROI in the left image is template matched with the ROI in the right image, and the pixel disparity between the left image and the right image is determined based on the template matching result.

[0093] It should be noted that if Figure 4 As shown, the method for determining the depth information position is:

[0094] The left and right camera image epipolar lines of the binocular camera must be on the same horizontal line, so that the pixels to be tested in the right image and the left image are on the same horizontal line to achieve optimal matching;

[0095] The difference between the image to be tested and the pixel to be tested in the horizontal direction is determined by using the pixel disparity of the left image and the right image, and the depth information can be calculated by deriving a formula from the epipolar correction part using the difference, so as to obtain a disparity map with the same size as the ROI of the left image;

[0096] Based on the disparity map, similar triangles are used to infer the depth information position of the target object based on the left image.

[0097] In one embodiment, the left and right camera images need to be on the same horizontal line to facilitate subsequent correction of the epipolar lines. The template matching is that the pixels to be tested in the right view and the left view are on the same horizontal line, which is the optimal match. After completing the above optimal match, it is necessary to record the disparity d, that is, the difference xr-x1 between the horizontal direction x1 of the image to be tested and the horizontal direction xr of the matching pixel. The formula derived from the epipolar correction part is: Since f and T are known, the depth Z value can be calculated. Finally, a disparity map D with the same size as the original image can be obtained. The disparity map D obtained through the above matching results can be used to invert the depth map based on the left view using similar triangles.

[0098] Example 2

[0099] On the other hand, Figure 5 As shown, the present application provides a robot grasping control device, which adopts the above-mentioned robot grasping control method, specifically including:

[0100] ROI extraction module, interception area acquisition module, pixel time difference evaluation module, grab control module;

[0101] The ROI extraction module is responsible for placing the target object in the working area and sending the left image of the binocular camera to the target detection network to extract the ROI of the left image;

[0102] The clipped region acquisition module is responsible for processing the ROI of the left image to obtain the grayscale ROI region in the left image, and clipping the region corresponding to the grayscale ROI region in the right image to obtain the clipped region;

[0103] The pixel time difference evaluation module is responsible for extracting features from the intercepted area to obtain the ROI area in the right image, and performing image processing to compare the features with the left image to generate the ROI in the right image, and calculating the pixel disparity between the left image and the right image based on the ROI in the left image and the ROI in the right image;

[0104] The grasping control module is responsible for obtaining the depth information position of the target object based on the left image based on the pixel parallax of the left image and the right image, determining the pixel center position of the target object in the left image, and obtaining the spatial coordinates of the target based on the depth information position and the pixel center position, and obtaining the left and right grasping points of the target according to the spatial coordinates and the outer contour of the target, and using the left and right grasping points to control the robot to perform the grasping task on the target.

[0105] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device, equipment, and non-volatile computer storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0106] The above is a description of a specific embodiment of the present specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0107] The above description is only one or more embodiments of this specification and is not intended to limit this specification. For those skilled in the art, one or more embodiments of this specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of one or more embodiments of this specification shall be included in the scope of the claims of this specification.

Claims

1. A robot grasping control method, characterized in that: Specifically include: Place the target object in the working area and send the left image of the binocular camera to the target detection network to extract the ROI of the left image; The ROI in the left image is processed to obtain the grayscale ROI area in the left image, and the area corresponding to the grayscale ROI area in the right image is intercepted to obtain the intercepted area; The intercepted area is subjected to feature extraction to obtain the ROI area in the right image, and after image processing, the ROI in the right image is generated by comparing it with the features of the left image. The pixel disparity between the left image and the right image is calculated based on the ROI in the left image and the ROI in the right image. Based on the pixel parallax of the left and right images, the depth information position of the target object based on the left image is obtained, and the pixel center position of the target object in the left image is determined. Based on the depth information position and the pixel center position, the spatial coordinates of the target are obtained. According to the spatial coordinates and the outer contour of the target, the left and right grasping points of the target are obtained, and the left and right grasping points are used to control the robot to perform the grasping task on the target.

2. The robot grasping control method according to claim 1, characterized in that: Before the left image of the binocular camera is sent to the target detection network, it is also necessary to perform dedistortion processing on the left image.

3. The robot grasping control method according to claim 2, characterized in that: The left image is subjected to a dedistortion process, specifically comprising: Establish the conversion relationship between the world coordinate system of a single camera of the binocular camera and the pixel plane coordinate system, and construct a conversion function based on the pixel size and the camera focal length; The left image is dedistorted using a conversion function.

4. The robot grasping control method according to claim 1, characterized in that: The ROI in the left image is processed to obtain the grayscale ROI area in the left image, including: The left image is sent to the target detection network to preliminarily extract the rectangular ROI of the object to be detected; The rectangular ROI is sent to the image segmentation model to further extract the ROI area, and the ROI area is converted into grayscale to obtain the grayscale ROI area in the left figure.

5. The robot grasping control method according to claim 1, characterized in that: The specific steps of constructing the image segmentation model are: Convolution and pooling are performed on the rectangular ROI, and multiple pooling is performed to obtain features of various sizes. The features of the smallest image size are upsampled or deconvolved to obtain a feature map of the image size corresponding to the next smallest image size. The feature map of the image size corresponding to the next smallest image size is used to perform channel stitching with the features of the next smallest image size to obtain a stitched image. The above steps are used to perform channel stitching processing on the stitched image to obtain a stitched processed image with the same size as the input image, and the rectangular ROI frame of the stitched processed image is sent to the segmentation network to obtain the ROI area.

6. The robot grasping control method according to claim 1, characterized in that: The specific steps of generating the intercepted area are: The grayscale ROI area in the left image is template matched with the right image to obtain a matching result, and the matching result is used to cut out the area corresponding to the grayscale ROI area in the right image to obtain a cutout area.

7. The robot grasping control method according to claim 1, characterized in that: The specific steps of template matching are: Dividing the right image into a plurality of sub-regions using the size of the grayscale ROI region, and determining matching sub-regions among the sub-regions according to the matching results between different sub-regions and the image features of the grayscale ROI region; Obtaining a detection frame of the grayscale ROI area, and determining a to-be-detected area of ​​the detection frame based on a grayscale ROI area whose distance from the detection frame is within a preset distance range; Determine the matching coefficient of the area to be detected according to the matching result between the image feature of the area to be detected and the matching sub-area, and when the matching coefficient does not meet the requirement, determine the moving step length according to a preset number of pixels, and use the moving step length to expand the matching sub-area to obtain an expanded sub-area; Until the matching result between the image features of the extended sub-region and the region to be detected meets the requirement, the extended sub-region is used as the matching region of the grayscale ROI region in the right image, and the matching region is used as the clipping region.

8. The robot grasping control method according to claim 1, characterized in that: The method for determining the ROI area in the right figure is: The cutout area in the right image is sent to the image segmentation model for feature extraction to obtain the target features in the right image, and the matching results of the target features are used to determine the ROI area in the right image.

9. The robot grasping control method according to claim 1, characterized in that: The method for determining the depth information position is: The left and right camera image epipolar lines of the binocular camera must be on the same horizontal line, so that the pixels to be tested in the right image and the left image are on the same horizontal line to achieve optimal matching; The difference between the image to be tested and the pixel to be tested in the horizontal direction is determined by using the pixel disparity of the left image and the right image, and the depth information can be calculated by deriving a formula from the epipolar correction part using the difference, so as to obtain a disparity map with the same size as the ROI of the left image; Based on the disparity map, similar triangles are used to infer the depth information position of the target object based on the left image.

10. A robot grasping control device, using a robot grasping control method according to any one of claims 1 to 9, characterized in that: Specifically include: ROI extraction module, interception area acquisition module, pixel time difference evaluation module, grab control module; The ROI extraction module is responsible for placing the target object in the working area and sending the left image of the binocular camera to the target detection network to extract the ROI of the left image; The clipped region acquisition module is responsible for processing the ROI of the left image to obtain the grayscale ROI region in the left image, and clipping the region corresponding to the grayscale ROI region in the right image to obtain the clipped region; The pixel time difference evaluation module is responsible for extracting features from the intercepted area to obtain the ROI area in the right image, and performing image processing to compare the features with the left image to generate the ROI in the right image, and calculating the pixel disparity between the left image and the right image based on the ROI in the left image and the ROI in the right image; The grasping control module is responsible for obtaining the depth information position of the target object based on the left image based on the pixel parallax of the left image and the right image, determining the pixel center position of the target object in the left image, and obtaining the spatial coordinates of the target based on the depth information position and the pixel center position, and obtaining the left and right grasping points of the target according to the spatial coordinates and the outer contour of the target, and using the left and right grasping points to control the robot to perform the grasping task on the target.

Citation Information

Patent Citations

  • 3D positioning and capturing method and system based on 2D camera

    CN118269107A