Robot fixed pose grabbing method based on ArUco pseudo two-dimensional code

By combining image statistical prior information and multi-view template matching technology, the problems of ArUco pseudo-QR codes in industrial scenarios are solved, and high-precision and high-rootability robot positioning solution is achieved, ensuring the accurate grasp of the target object.

CN120198505AActive Publication Date: 2025-06-24CHENGDU ZHIXIANG TECHNOLOGY CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510281747.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-24
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

The existing ArUco pseudo-QR code pose solution technology has problems such as light sensitivity, corner misalignment and pose calculation errors in industrial scenarios, resulting in low positioning accuracy and poor system robustness.

Method used

Through image statistical prior information, multi-image thresholds and multi-view template matching, combined with the size, angle and plane fitting information of the ArUco pseudo-QR code, corner point purification and pose calculation are performed to improve the generation accuracy of the robot 6-DOF pose.

Benefits of technology

It significantly improves the robustness of the algorithm, reduces the missed detection and false detection rates caused by uneven lighting, improves the positioning accuracy of the robot's end effector, and ensures accurate capture of the target object.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198505A_ABST
    Figure CN120198505A_ABST
Patent Text Reader

Abstract

The invention provides a robot fixed pose grabbing method based on an ArUco pseudo two-dimensional code, relates to the field of computer vision and image processing, and solves the limitation problems of light sensitivity, angular point dislocation, pose calculation errors and the like easily occurring in the application of the ArUco pseudo two-dimensional code. The method comprises the following steps: acquiring an input image and converting the input image into a grayscale image, calculating a threshold set and then carrying out binaryzation; candidate angular points in the binary image are extracted, the candidate angular points are selected and purified according to prior information of the ArUco pseudo two-dimensional code, and a target angular point and the angular point direction of the target angular point are determined; further determining an information area, matching the information area with the mark template, calculating a pixel 3D coordinate after matching succeeds, and fitting the pixel 3D coordinate into a plane equation corresponding to the information area; and according to the angular point direction and the plane equation, generating 6DOF pose data for the robot to use. According to the method, the generation precision of the 6-DOF pose of the robot is effectively improved, and the robot can accurately grab the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision and image processing, and is applied to a grasping robot. Specifically, it relates to a method for a robot to grasp a fixed pose based on an ArUco pseudo-two-dimensional code. Background Art

[0002] As an important branch in the field of robot perception, the vision positioning technology based on artificial markers plays a key role in scenarios such as industrial automation, intelligent warehousing, and service robots. As a typical encoded marker system among them, the ArUco pseudo-two-dimensional code combines binary encoded information with geometric patterns to construct a composite identification system with both information-bearing and spatial positioning functions. Its working principle is based on feature point detection and perspective projection models in computer vision. By extracting the corner coordinates of the two-dimensional code and establishing the corresponding relationship with a preset template, and combining the internal parameters of the camera, the PnP (Perspective-n-Point) algorithm is used to calculate the three-dimensional pose of the target object relative to the camera. Due to its characteristics such as convenient deployment and high calculation efficiency, this technology is widely used in fields such as robot calibration, scene mapping, and mobile navigation, and has become the mainstream positioning solution especially in the scenario of grasping with a fixed pose; this scenario requires the end effector of the robotic arm to approach the target with a specific spatial pose, and its pose error control is usually within the accuracy range of millimeters and angular grades.

[0003] However, the existing ArUco pseudo-two-dimensional code pose calculation technology has significant limitations in the actual application of industrial scenarios. Traditional corner detection algorithms are highly sensitive to changes in the image perspective. When the angle between the two-dimensional code plane and the camera optical axis exceeds 30 degrees, due to perspective distortion, the deviation of corner coordinate extraction increases, and in extreme cases, corner mis-matching even occurs. Experimental data shows that under the condition of a pitch angle of 45 degrees, the average offset of corner detection can reach 3 to 5 pixels, which will directly cause the translational error of pose calculation to expand to several times the reference value.

[0004] A more severe challenge comes from the dynamic interference of environmental light. Existing algorithms rely on binary processing with a fixed threshold. In low-illumination environments, the boundary of the pseudo-two-dimensional code is easily blurred due to image noise, while in strong specular reflection conditions, pseudo-edge detection phenomena will occur. Tests show that when the environmental illuminance is lower than 100 lux, the success rate of marker detection drops to 72%, and the false detection rate on highly reflective materials such as metal surfaces is as high as 40%. Such optical interference not only affects the positioning accuracy, but also causes the system to trigger an abnormal protection mechanism and interrupt the operation process, severely restricting the general application of the technology. In addition, the common local occlusion problems in industrial sites (such as oil stain adhesion, mechanical component occlusion, etc.) will destroy the geometric integrity of the pseudo-two-dimensional code, and traditional algorithms lack an effective feature compensation mechanism for this, further exacerbating the instability of pose calculation.

[0005] The prior art also overly relies on idealized visual models and fails to effectively integrate multi-dimensional physical constraints and robustness enhancement mechanisms. These technical deficiencies cause problems such as high positioning drift rate, poor system robustness, and weak environmental adaptability in the existing solutions when facing high-precision and strongly constrained fixed pose grasping tasks, seriously restricting the advanced demand for precise grasping technology in the field of intelligent robots.

[0006] In summary, how to achieve high-robustness pose solution through innovative breakthroughs at the algorithm level has become a technical difficulty that urgently needs to be overcome in the field of industrial robot vision positioning. Summary of the Invention

[0007] Based on the current situation in the background technology, the purpose of the present invention is to solve the limitations such as light sensitivity, corner misalignment, and pose calculation errors that are prone to occur in the actual industrial application of ArUco pseudo-two-dimensional codes. Therefore, a robot fixed pose grasping method based on ArUco pseudo-two-dimensional codes is proposed. The present invention solves the above problems through methods such as image statistical prior information, multi-image thresholding, and multi-view template matching, and at the same time combines prior information means such as the size, angle, and plane fitting of ArUco pseudo-two-dimensional codes to perform corner purification, thereby improving the generation accuracy of the robot's 6-DOF pose and enabling the robot to accurately grasp the target object.

[0008] The present invention adopts the following technical solutions to achieve the purpose:

[0009] A robot fixed pose grasping method based on ArUco pseudo-two-dimensional codes, the method comprising the following steps:

[0010] S1. Obtain an input image containing an ArUco pseudo-two-dimensional code, convert the input image into a corresponding grayscale image, and calculate a corresponding threshold set based on the grayscale image;

[0011] S2. Use the threshold set to perform binarization of the grayscale image to obtain a binary image corresponding to the grayscale image, and extract candidate corners corresponding to the ArUco pseudo-two-dimensional code in the binary image;

[0012] S3. According to the prior information of the input image and the ArUco pseudo-two-dimensional code, select and purify the extracted candidate corners to determine the target corners and their corresponding corner directions;

[0013] S4. According to the target corners, determine the information area of the ArUco pseudo-two-dimensional code, and match the information area with a preset marker template;

[0014] S5. After successful matching, calculate the 3D coordinates of the pixels in the information area, and based on the 3D coordinates of the pixels, fit to obtain a plane equation corresponding to the information area;

[0015] S6. Generate 6DOF pose data based on the corner direction and the plane equation, and the robot grasps the target object provided with the ArUco pseudo-two-dimensional code according to the 6DOF pose data.

[0016] Specifically, in step S1, the input image is an RGB image; after converting the RGB image into the grayscale image, calculate the average grayscale and the median grayscale of the grayscale image, and then obtain the corresponding threshold set.

[0017] Furthermore, after calculating the average grayscale and the median grayscale, first determine the starting threshold based on a preset reference value L; then, calculate the threshold step size according to the brightness deviation and the grayscale dynamic range. Take the grayscale level corresponding to the starting threshold as the first grayscale threshold in the threshold set, and then determine the remaining multiple grayscale thresholds based on the first grayscale threshold according to the threshold step size until reaching the maximum grayscale level 255. The multiple grayscale thresholds obtained in this way together constitute the threshold set.

[0018] Specifically, in step S2, there are multiple grayscale thresholds in the threshold set, and the grayscale image is binarized under the action of the multiple grayscale thresholds to obtain multiple binary images; perform edge extraction operations on each binary image, and use the RDP algorithm to fit the extracted edge broken lines; if the polygon obtained by fitting is not a quadrilateral, discard the polygon; if the polygon obtained by fitting is a quadrilateral, take the four corner points of the quadrilateral as the candidate corner points of the ArUco pseudo-two-dimensional code.

[0019] Specifically, in step S3, the prior information includes the side length information and the angle information of the ArUco pseudo-two-dimensional code; after obtaining the prior information in advance, for each quadrilateral obtained by fitting and the extracted candidate corner points in each binary image, sequentially determine whether the angle of each candidate corner point matches the angle information, and further determine whether at least one of the two side lengths corresponding to the candidate corner point matches the side length information when they match; select the optimal candidate corner point from the multiple candidate corner points that match both the angle information and the side length information in the multiple binary images as the target corner point, and the direction corresponding to at least one side length that matches successfully of the target corner point is the corner direction.

[0020] Specifically, in step S4, a perspective transformation matrix is used to transform the information area of the ArUco pseudo-two-dimensional code into a bitmask image with a preset shape. The obtained bitmask image after the transformation is rotated and matched with a preset marker template in multiple directions. The marker template includes multiple preset correct ArUco pseudo-two-dimensional codes. When the matching is successful, it means that the ID of the ArUco pseudo-two-dimensional code corresponding to the bitmask image is correct, and the ArUco pseudo-two-dimensional code is successfully recognized.

[0021] Specifically, in step S5, the depth map information corresponding to the input image is obtained in advance. Based on the depth map information, the 3D coordinates of the pixels in the information area are calculated to obtain a set of 3D coordinates. Based on the set of 3D coordinates, the plane equation corresponding to the information area is fitted.

[0022] Specifically, in step S6, based on the plane equation, the corresponding normal vector n is obtained. According to the vector corresponding to the corner direction, combined with the direction rotation situation during the information area matching, the first direction vector X is determined, and the normal vector n is denoted as the second direction vector Z. Based on the first direction vector X and the second direction vector Z, the right-hand rule is used to calculate the third direction vector Y.

[0023] Since the plane equation represents the space plane where the information area of the successfully matched and recognized ArUco pseudo-two-dimensional code is located, based on the first direction vector X, the second direction vector Z, and the third direction vector Y corresponding to the space plane, a three-dimensional direction vector corresponding to the ArUco pseudo-two-dimensional code is formed, which represents the fixed spatial position of the target object. The robot generates 6DOF pose data according to the preset relative position relationship with the three-dimensional direction vector, and grasps the target object according to the 6DOF pose data.

[0024] In summary, due to the adoption of this technical solution, the beneficial effects of the present invention are as follows:

[0025] By integrating image statistical prior information and multi-image threshold processing technology, the present invention effectively overcomes the technical defect that the traditional ArUco pseudo-two-dimensional code detection is sensitive to lighting conditions, significantly improves the algorithm robustness in complex industrial scenarios, and reduces the occurrence rate of missed detection and false detection caused by uneven lighting. Adopting a corner purification strategy with multi-dimensional prior information collaboration, combined with characteristic parameters such as the physical size, spatial angle, and plane fitting of the two-dimensional code, realizes the precise correction and optimization of the detected corners, and fundamentally solves the problem of corner positioning offset existing in the traditional method.

[0026] The pose calculation and verification system constructed by the present invention can effectively improve the solution accuracy of the 6-DOF pose of the robot, reduce the plane fitting error of the ArUco pseudo-two-dimensional code to the sub-pixel level, and ensure that the end effector of the robot can accurately position and complete the grasping operation of the target object. Through its composite technical path, the present invention enhances the adaptability of the system in a dynamic industrial environment, retains the inherent advantage of fast recognition of the ArUco pseudo-two-dimensional code, and breaks through the original technical bottleneck through algorithm innovation and integration, thus providing a high-precision and highly reliable solution for the visual positioning application in the field of intelligent robot control. Brief Description of the Drawings

[0027] Figure 1 It is a schematic diagram of the overall process of the method of the present invention;

[0028] Figure 2 It is a schematic diagram of the effect of fitting a quadrilateral and extracting candidate corner points in the method of the present invention;

[0029] Figure 3 It is a schematic diagram of the effect of finally obtaining a three-dimensional direction vector in the method of the present invention. Detailed Embodiments

[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations.

[0031] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0032] Embodiment

[0033] A method for a robot to grasp a fixed pose based on an ArUco pseudo-two-dimensional code. The overall process of this method can be referred to Figure 1 for illustration. In this embodiment, the method is first summarized as the following steps:

[0034] S1. Obtain an input image containing an ArUco pseudo-two-dimensional code, convert the input image into a corresponding grayscale image, and calculate a corresponding threshold set based on the grayscale image;

[0035] S2. Use the set of thresholds to perform binarization on the grayscale image to obtain a binary image corresponding to the grayscale image, and extract candidate corner points corresponding to the ArUco pseudo-two-dimensional code in the binary image;

[0036] S3. According to the input image and the prior information of the ArUco pseudo-two-dimensional code, perform selection and purification on the extracted candidate corner points to determine the target corner points and their corresponding corner point directions;

[0037] S4. According to the target corner points, determine the information area of the ArUco pseudo-two-dimensional code, and match the information area with a preset marker template;

[0038] S5. After successful matching, calculate the 3D coordinates of the pixels in the information area, and according to the 3D coordinates of the pixels, fit to obtain the plane equation corresponding to the information area;

[0039] S6. Generate 6DOF pose data according to the corner point direction and the plane equation, and the robot grabs the target object provided with the ArUco pseudo-two-dimensional code according to the 6DOF pose data.

[0040] In this embodiment, the details of each step will be introduced in detail according to the above step sequence.

[0041] In step S1, the input image is an RGB image. After converting the RGB image into a grayscale image, first calculate the average grayscale of the grayscale image as follows:

[0042]

[0043] In the formula, aver_gray represents the average grayscale, and imgGray[i][j] represents the grayscale value of the pixel at the i-th row and the j-th column in the valid area of the grayscale image; after summing the grayscale values of all pixels in the valid area and dividing by the total number of pixels in the valid area, the average grayscale is obtained.

[0044] Continue to calculate the median grayscale of the grayscale image. In this embodiment, first calculate the grayscale probability corresponding to each grayscale level in the grayscale image as follows:

[0045]

[0046] Wherein, gray_prob[I] represents the gray probability corresponding to the gray level I, and I ∈ [0, 255]; gray_pixel_num[I] represents the number of pixels with the gray level I, and pixel_num represents the total number of pixels in the effective area of the gray image. After calculating the gray probability, a total of 256 probability values are obtained, and each probability value corresponds to a specific gray level. For example, when the gray levels are 1, 10, and 100, their respective probability values.

[0047] Subsequently, according to the gray probability, calculate the cumulative probability from 0 to this gray level corresponding to each gray level, and take the smallest gray level I whose cumulative probability exceeds 50% as the median gray level, as shown in the following formula:

[0048] mid_gray = argmin I {gray_distribution[I - 1] ≤ 0.5 < gray_distribution[I]}

[0049] Wherein, mid_gray represents the median gray level, and argmin I represents taking the minimum value among all I values that meet the conditions; gray_distribution[I] represents the cumulative probability from 0 to the gray level I.

[0050] For example, if the probability of the gray level 0 is 1%, the probability of the gray level 1 is 2%, and the probability of the gray level 2 is 4%, then the cumulative probability of the gray level 2 is the sum of all probabilities in the process from 0 to 2, that is, 1% + 2% + 4% = 7%. Therefore, if the cumulative probability of the gray level I exceeds 50% and the cumulative probability of the gray level I - 1 is less than or equal to 50%, then the gray level I is the median gray level to be found.

[0051] In this embodiment, after calculating the average gray level aver_gray and the median gray level mid_gray, the calculation of the threshold set can be carried out, and this part will focus on solving the problem of the recognition error of the ArUco pseudo-two-dimensional code caused by influencing factors such as the illumination environment. First, obtain a preset reference value L, and in this embodiment, its value is taken as 127.5; when the average gray level is greater than or equal to the preset reference value L, it means that the input image is too bright, and then determine the following start threshold start_thres:

[0052] start_thres = 2 × mid_gray - aver_gray

[0053] And when the average gray level is less than the preset reference value L, it means that the input image is too dark, and then determine the following start threshold start_thres:

[0054] start_thres = 2 * aver_gray - mid_gray

[0055] Subsequently, according to the brightness deviation (i.e., the difference from the preset reference value L) and the gray - scale dynamic range (i.e., the difference between the average gray - scale and the median gray - scale), calculate the threshold step stride as follows:

[0056]

[0057] Take the gray - scale level corresponding to the start threshold as the first gray - scale threshold in the threshold set. Subsequently, according to the threshold step, determine the remaining multiple gray - scale thresholds based on the first gray - scale threshold until reaching the maximum gray - scale level 255. The multiple gray - scale thresholds thus obtained together constitute the threshold set.

[0058] In step S2, each gray - scale threshold in the threshold set can binarize the gray - scale image into a corresponding binary image. Therefore, after binarization of the gray - scale image under the action of multiple gray - scale thresholds, multiple binary images can be obtained; for different lighting conditions, this method maximally obtains all situations that can correctly extract and recognize the ArUco pseudo - QR code, making it highly likely to ensure that at least one of the binary images can correctly obtain the ArUco pseudo - QR code during the extraction and recognition process.

[0059] In this embodiment, edge extraction operation is then performed on each binary image, and the Ramer - Douglas - Peucker (RDP) algorithm is used to fit the extracted edge broken line segments. In this embodiment, the shape of the ArUco pseudo - QR code is usually square, which is a kind of quadrilateral. Therefore, during the fitting process, if the fitted polygon is not a quadrilateral, discard the polygon; if a quadrilateral is fitted, take the four corner points of the quadrilateral as the candidate corner points of the ArUco pseudo - QR code.

[0060] In this embodiment, after quadrilateral fitting for each binary image, there may be multiple quadrilaterals. The quadrilaterals presented by different binary images may be different, which may lead to the phenomenon that the corresponding candidate corner points have errors with the actual situation. Therefore, it is necessary to select the target corner point that is most likely to be one of the corner points of the ArUco pseudo - QR code from the multiple candidate corner points of the multiple quadrilaterals in the multiple binary images, and then obtain the corner point direction corresponding to the target corner point.

[0061] Therefore, in step S3, the prior information includes the side length information and the angle information of the ArUco pseudo-two-dimensional code. After obtaining the prior information in advance, for the quadrilaterals fitted in each binary image and the candidate corner points extracted therefrom, it is successively determined whether the angle of each candidate corner point matches the angle information, and when they match, it is further determined whether at least one of the two side lengths corresponding to the candidate corner point matches the side length information. In this way, the optimal candidate corner point can be selected from multiple candidate corner points that successfully match both the angle information and the side length information in multiple binary images as the target corner point, and the direction corresponding to at least one side length that successfully matches of the target corner point is the corner point direction.

[0062] Before matching the angle information and the side length information in this embodiment, the depth map information corresponding to the input image is obtained in advance; according to the depth map information, 3D point cloud transformation is performed on the quadrilaterals fitted in each binary image, and the angle information and the side length information are matched for each quadrilateral and its candidate corner points in the transformed 3D point cloud space, so as to realize the selection and purification of candidate corner points.

[0063] It should be noted here that taking the ArUco pseudo-two-dimensional code of a square as an example, since different binary images are obtained according to different gray thresholds, even after 3D point cloud transformation, there are error situations in the fitted quadrilaterals, for example, the transformed quadrilateral is not a square but a right trapezoid or the four corners are not right angles, etc.; and the side lengths corresponding to the four candidate corner points in such quadrilaterals will also be different. Therefore, the selection and purification of candidate corner points is equivalent to finding the corner point closest to a right angle from multiple quadrilaterals, and at the same time, it can also be considered whether the whole quadrilateral is closest to a square. Then starting from the found corner point, it is checked whether there is at least one side length of its two sides that is closest to the prior side length information. If both the angle and the side length match successfully, it means that the corner point and the corresponding side length are one side of the ArUco pseudo-two-dimensional code. Thus, the corner point and the corresponding corner point direction, that is, the target corner point, can be accurately obtained, and the remaining three candidate corner points of the quadrilateral can form the information area of the subsequent ArUco pseudo-two-dimensional code.

[0064] See Figure 2 For the effect schematic diagram, for the two target objects of the pot lid and the pot handle, the robot needs to perform grasping, and corresponding ArUco pseudo-two-dimensional codes are pasted on both the pot lid and the pot handle. When Figure 2 the perspective is the perspective of the acquired input image, after the image processing process of the above method, the four corner points of the ArUco pseudo-two-dimensional code can be accurately marked, and the center of the ArUco pseudo-two-dimensional code can be determined according to the positions of the four corner points. The corner point determined as the target corner point among the four corner points and its corner point direction will be used in the subsequent determination process of the three-dimensional direction vector.

[0065] After determining the information area of the ArUco pseudo-two-dimensional code in the binary image based on the fitted quadrilateral and its target corner points, in step S4, the perspective transformation matrix can be used to transform this information area into a bitmask image with a preset shape, which in this embodiment is a square image identical to the square ArUco pseudo-two-dimensional code. The perspective transformation matrix is as follows:

[0066]

[0067] In the formula, (x, y) represents the coordinates of the input pixel point in the information area of the ArUco pseudo-two-dimensional code, (x′, y′) represents the coordinates of the pixel point corresponding to this input pixel point in the bitmask image after transformation, ω′ represents the normalization factor, and h (…) represents each matrix coefficient.

[0068] After the bitmask image obtained by the transformation is rotated, it is matched with a preset marker template in the up, down, left, and right directions. The marker template includes multiple preset correct ArUco pseudo-two-dimensional codes. For example, in this embodiment, the ArUco pseudo-two-dimensional code images corresponding to the pot lid and the pot handle in the front view are pre-recorded as the marker template. When the matching is successful, it means that the ID of the ArUco pseudo-two-dimensional code corresponding to the bitmask image is correct, and the ArUco pseudo-two-dimensional code is successfully recognized. At the same time, it is recorded whether the bitmask image has been rotated in any direction. If the matching fails, it indicates that the information area may not be an ArUco pseudo-two-dimensional code, or may not be the ArUco pseudo-two-dimensional code of the target object, or there may be contamination, occlusion, or damage to the ArUco pseudo-two-dimensional code, and it is necessary to re-acquire images from other perspectives or take other measures.

[0069] When the ArUco pseudo-two-dimensional code is correctly recognized due to successful matching, step S5 can be entered. In step S5, the depth map information corresponding to the input image is pre-obtained, and based on this depth map information, the 3D coordinates of the pixels in the information area are calculated to obtain a 3D coordinate set, as follows:

[0070]

[0071]

[0072] In the above formulas, P represents the 3D coordinate set, depth[u, v] represents the pixel coordinates in the depth map information, which are u and v respectively; depth_scalar represents the scalar coefficient for normalizing the data in the depth map information to meters (if the unit of the depth map information is millimeters, the value of depth_scalar can be 1000; conversely, if the unit of the depth map information is meters, the value of depth_scalar can be 1); c xand c y is the principal point coordinate in the camera intrinsic parameters, f x and f y are the focal lengths of the X-axis and Y-axis in the camera intrinsic parameters.

[0073] Immediately following this, based on this set of 3D coordinates, the plane equation corresponding to the information region is fitted as follows:

[0074] ax + by + cz + d = 0

[0075] In the formula, a, b, c, and d are the coefficients of this plane equation, which are calculated after fitting with the set of 3D coordinates.

[0076] Finally, in step S6, based on the obtained plane equation corresponding to the information region, the corresponding normal vector n can be obtained as follows:

[0077] n = (a, b, c)

[0078] In this embodiment, based on the vector corresponding to the corner direction and combining the direction rotation situation during the information region matching, the first direction vector X is determined, and the normal vector n is denoted as the second direction vector Z; during the determination process of the normal vector n in this embodiment, it can always be denoted as facing the outside or the upper side of the image, while when ensuring its perpendicular relationship with the corner direction, the first direction vector X considers whether there is rotation when matching with a preset marking template, and it can always be determined as the vector corresponding to directly above when facing the ArUco pseudo QR code.

[0079] Based on the first direction vector X and the second direction vector Z, the right-hand rule is used to calculate the third direction vector Y, and the third direction vector Y is the vector cross product of X and Z, as follows:

[0080] Y = X × Z

[0081] Since the plane equation represents the spatial plane where the information region of the successfully matched and recognized ArUco pseudo QR code is located, based on the first direction vector X, the second direction vector Z, and the third direction vector Y corresponding to this spatial plane, a three-dimensional direction vector corresponding to the ArUco pseudo QR code is formed, which represents the fixed spatial position of the target object. Combining with being able to determine the center point according to its four corner points after fitting the quadrilateral, after making the vectors X, Y, and Z have a common intersection point and marking this common intersection point at the center point, the Figure 3 shown three-dimensional direction vector rendering effect can be formed. Corresponding to the ArUco pseudo QR codes at the lid and the pot handle, where the green arrow represents the first direction vector X of the two, the blue arrow represents the second direction vector Z, and the red arrow represents the third direction vector Y.

[0082] Combined with Figure 2 and Figure 3As shown in the schematic diagram, it can be seen that the solution of this embodiment has very accurate effects on the corner extraction and three-dimensional direction vector recognition of the ArUco pseudo-two-dimensional code under low-light conditions. Therefore, finally, based on the corresponding three-dimensional direction vector, since the ArUco pseudo-two-dimensional code has corresponding preset fixed position information, there is also a preset standard grasping position for the target object when the robot performs the grasping operation, and this grasping position can be determined according to the relative position relationship with the ArUco pseudo-two-dimensional code; when the three-dimensional direction vector of the ArUco pseudo-two-dimensional code is correctly recognized, the robot can generate its corresponding 6DOF pose data, and based on this 6DOF pose data, grasp the target object.

Claims

1. A robot fixed-pose grasping method based on ArUco pseudo-QR code, characterized in that: The method comprises the following steps: S1. Obtain an input image containing an ArUco pseudo-two-dimensional code, convert the input image into a corresponding grayscale image, and calculate a corresponding threshold set based on the grayscale image; S2, binarizing the grayscale image using the threshold value set to obtain a binary image corresponding to the grayscale image, and extracting candidate corner points corresponding to the ArUco pseudo-two-dimensional code in the binary image; S3, selecting and purifying the extracted candidate corner points according to the prior information of the input image and the ArUco pseudo-two-dimensional code, and determining the target corner point and its corresponding corner point direction; S4, determining the information area of ​​the ArUco pseudo-two-dimensional code according to the target corner point, and matching the information area with a preset marking template; S5. After the matching is successful, the 3D coordinates of the pixels in the information area are calculated, and the plane equation corresponding to the information area is fitted based on the 3D coordinates of the pixels; S6. Generate 6DOF pose data according to the corner point direction and the plane equation. The robot grasps the target object provided with the ArUco pseudo-two-dimensional code according to the 6DOF pose data.

2. The robot fixed-position grasping method according to claim 1, characterized in that: In step S1, the input image is an RGB image; after converting the RGB image into the grayscale image, the average grayscale and median grayscale of the grayscale image are calculated, and then the corresponding threshold value set is obtained.

3. The robot fixed-position grasping method according to claim 2, characterized in that: The calculation formula of the average grayscale is as follows: Wherein, aver_gray represents the average grayscale, and imgGray[i][j] represents the grayscale value of the pixel in the i-th row and j-th column in the effective area of ​​the grayscale image; the average grayscale is obtained by summing the grayscale values ​​of all pixels in the effective area and dividing the sum by the total number of pixels in the effective area; In the process of calculating the median grayscale, the grayscale probability corresponding to each grayscale level in the grayscale image is first calculated as follows: Where gray_prob[I] represents the grayscale probability corresponding to grayscale level I, I∈[0,255]; gray_pixel_num[I] represents the number of pixels with grayscale level I, and pixel_num represents the total number of pixels in the effective area of ​​the grayscale image; then, based on the grayscale probability, the cumulative probability from 0 to the grayscale level corresponding to each grayscale level is calculated, and the minimum grayscale level I with a cumulative probability exceeding 50% is taken as the median grayscale, as shown in the following formula: mid_gray=argmin I {gray_distribution[I-1]≤0.5<gray_distribution[I]} In the formula, mid_gray represents the median grayscale, argmin I represents the minimum value of all I values ​​that meet the conditions; gray_distribution[I] represents the cumulative probability from 0 to gray level I.

4. The robot fixed-position grasping method according to claim 2, characterized in that: After calculating the average grayscale and the median grayscale, first determine the starting threshold as follows: When the average grayscale is greater than or equal to the preset reference value L, the calculation formula for the start threshold is: start_thres=2×mid_gray-aver_gray When the average grayscale is less than the preset reference value L, the calculation formula for the start threshold is: start_thres=2×aver_gray-mid_gray In the above two formulas, start_thres represents the starting threshold, aver_gray represents the average grayscale, and mid_gray represents the median grayscale; then, based on the brightness deviation and the grayscale dynamic range, the threshold step is calculated as follows: In the formula, stride represents the threshold step length, and L represents the preset reference value; The grayscale level corresponding to the starting threshold is taken as the first grayscale threshold in the threshold set, and then the remaining grayscale thresholds are determined based on the first grayscale threshold according to the threshold step until the maximum grayscale level 255 is reached. The multiple grayscale thresholds thus obtained together constitute the threshold set.

5. The robot fixed-position grasping method according to claim 1, characterized in that: In step S2, the threshold value set includes a plurality of grayscale thresholds, and the grayscale image is binarized under the action of the plurality of grayscale thresholds to obtain a plurality of binary images; Performing edge extraction operation on each of the binary images, and fitting the extracted edge polyline segments using the RDP algorithm; if the polygon obtained by fitting is not a quadrilateral, discarding the polygon; If the fitted polygon is a quadrilateral, the four corner points of the quadrilateral are used as candidate corner points of the ArUco pseudo-two-dimensional code.

6. The robot fixed-position grasping method according to claim 5, characterized in that: In step S3, the prior information includes the side length information and angle information of the ArUco pseudo-QR code; after obtaining the prior information in advance, for each quadrilateral fitted in the binary image and the candidate corner points extracted therefrom, determine in turn whether the angle of each candidate corner point matches the angle information, and when matching, further determine whether at least one of the two side lengths corresponding to the candidate corner point matches the side length information; from multiple candidate corner points in the binary images that successfully match both the angle information and the side length information, select the best candidate corner point as the target corner point, and the direction corresponding to at least one side length of the target corner point that successfully matches is the corner point direction.

7. The robot fixed-position grasping method according to claim 6, characterized in that: Before matching the angle information and the side length information, the depth map information corresponding to the input image is obtained in advance; based on the depth map information, a 3D point cloud transformation is performed on each quadrilateral fitted in the binary image, and the angle information and the side length information are matched for each quadrilateral and its candidate corner points in the transformed 3D point cloud space to achieve selection and purification of candidate corner points.

8. The robot fixed-position grasping method according to claim 1, characterized in that: In step S4, the information area of ​​the ArUco pseudo-two-dimensional code is transformed into a bit mask image with a preset shape using a perspective transformation matrix, and the perspective transformation matrix is ​​as follows: Where (x, y) represents the coordinates of the input pixel in the information area of ​​the ArUco pseudo-QR code, (x′, y′) represents the coordinates of the pixel in the bit mask image corresponding to the input pixel after transformation, ω′ represents the normalization factor, and h (...) Represents each matrix coefficient; the bit mask image obtained after the transformation is rotated and matched with a preset marking template in multiple directions. The marking template includes multiple preset correct ArUco pseudo-QR codes. When the match is successful, it means that the ID of the ArUco pseudo-QR code corresponding to the bit mask image is correct, and the ArUco pseudo-QR code is successfully identified.

9. The robot fixed-position grasping method according to claim 1, characterized in that: In step S5, the depth map information corresponding to the input image is obtained in advance, and the 3D coordinates of the pixels in the information area are calculated according to the depth map information to obtain a 3D coordinate set, which is as follows: In the above formulas, P represents a 3D coordinate set, depth[u,v] represents the pixel coordinates in the depth map information, depth_scalar represents the scalar coefficient of the depth map information normalized to meters; c x and c y is the principal point coordinate in the camera intrinsic parameter, f x and f y is the X-axis and Y-axis focal length in the camera intrinsic parameters; According to the 3D coordinate set, the plane equation corresponding to the information area is fitted as follows: ax+by+cz+d=0 Where a, b, c, and d are the coefficients of the plane equation, which are calculated after fitting the 3D coordinate set.

10. The robot fixed-position grasping method according to claim 9, characterized in that: In step S6, based on the plane equation, the corresponding normal vector n is obtained as follows: n=(a,b,c) According to the vector corresponding to the corner point direction, combined with the direction rotation during information area matching, the first direction vector X is determined, and the normal vector n is recorded as the second direction vector Z; based on the first direction vector X and the second direction vector Z, the right-hand rule is used to calculate the third direction vector Y, which is the vector cross product of X and Z, as shown in the following formula: Y=X×Z Since the plane equation represents the spatial plane where the ArUco pseudo-QR code information area is located after successful matching and identification, the three-dimensional direction vector corresponding to the ArUco pseudo-QR code is formed according to the first direction vector X, the second direction vector Z and the third direction vector Y corresponding to the spatial plane, representing the fixed spatial position of the target object; the robot generates 6DOF pose data based on the preset relative position relationship with the three-dimensional direction vector, and grasps the target object based on the 6DOF pose data.

Citation Information

Patent Citations

  • Robot vision positioning method based on two-dimensional code

    CN112364677A

  • Pose positioning method, device and equipment and computer readable storage medium

    CN113538574A

  • Robot autonomous grabbing simulation system and method based on target 6D pose estimation

    CN114912287A

  • Method for refining target pose by robot in home environment

    CN116524011A

  • Robot positioning method based on two-dimensional code road sign

    CN116543048A