Calibration method, device, apparatus and medium

By using a robot's binocular camera to capture multi-view images of the checkerboard pattern on the positioning pen's screen, combined with a multi-step prediction method, the problems of high cost and high feature point requirements in robot-positioning pen calibration were solved, achieving high-precision calibration results.

CN121505051BActive Publication Date: 2026-04-07BEIJING XIAOYU INTELLISYS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-14
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies for calibrating robots and positioning pens suffer from high costs, stringent requirements for environmental feature points, and insufficient calibration accuracy, making it particularly difficult to meet the demands for high-precision calibration in real-world applications.

Method used

The robot's binocular camera is used to capture multiple images of the chessboard displayed on the positioning pen screen from multiple perspectives, obtaining multiple image pairs. The pose of the positioning pen's visual coordinate system relative to the human base coordinate system is determined through a multi-step prediction method. The initial prediction value of the chessboard coordinate system and the feature information of multiple image pairs are used for secondary prediction to improve the accuracy and reliability of calibration.

Benefits of technology

It effectively improves the accuracy and stability of robot and positioning pen calibration, reduces the impact of blind spots and insufficient feature information on prediction accuracy, simplifies the calibration process, and improves calibration efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121505051B_ABST
    Figure CN121505051B_ABST
Patent Text Reader

Abstract

This application proposes a calibration method, apparatus, device, and medium. The method includes: capturing multiple image pairs from multiple perspectives of a chessboard displayed on the screen of a positioning pen using a robot's binocular camera; predicting the value of a target object based on target image pairs among the multiple image pairs to obtain an initial predicted value for the target object; the target object includes a target pose; predicting the value of the target pose based on the multiple image pairs and the initial predicted value to obtain a target predicted value; and determining the pose of the positioning pen's visual coordinate system relative to the human base coordinate system based on the target predicted value to calibrate the robot and the positioning pen. By first performing an initial prediction of the target pose, and then combining multiple image pairs and the initial predicted value with values ​​approximating the true value to perform a secondary prediction, this multi-step prediction method improves the accuracy, stability, and efficiency of target pose prediction, thereby enhancing the accuracy and efficiency of calibration.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of calibration, and in particular to a calibration method, device, apparatus and medium. BACKGROUND

[0002] As a high-precision auxiliary tool, the positioning pen plays a crucial role in the execution of various operation tasks by robots. It can accurately mark target positions or key points, providing robots with clear and explicit spatial coordinate references. Through data interaction with the robot system, the position information captured by the positioning pen can be transmitted in real time to the robot's control unit. Based on these precise data, the robot can adjust its movement path and posture more accurately, thereby achieving precise grasping, placing and assembly of target objects. This auxiliary method with the positioning pen significantly improves the precision and reliability of robot operation, effectively reducing the operation failure rate caused by positioning deviation, and brings higher production efficiency and better operation results to various fields such as industrial production, medical surgery and scientific research.

[0003] However, to achieve efficient collaborative work between robots and positioning pens, which are widely used in industrial and scientific research fields, an accurate calibration process is an indispensable important link. Through this accurate calibration process, systematic errors between devices can be eliminated, ensuring that robots can execute more precise and error-free operation instructions based on the precise position information provided by the positioning pen. SUMMARY

[0004] The present application aims to at least partially solve one of the technical problems in the related art.

[0005] To this end, the first object of the present application is to propose a calibration method.

[0006] The second object of the present application is to propose a calibration device.

[0007] The third object of the present application is to propose an electronic device.

[0008] The fourth object of the present application is to propose a computer-readable storage medium.

[0009] The fifth object of the present application is to propose a computer program product.

[0010] To achieve the above objectives, a first aspect of this application proposes a calibration method, comprising: acquiring multiple image pairs; wherein the multiple image pairs are obtained by capturing a checkerboard pattern displayed on the screen of a positioning pen from multiple perspectives using a robot's binocular camera, the binocular camera including a left-eye camera and a right-eye camera, and any image pair including a left-eye image and a right-eye image; predicting the value of a target prediction object based on a target image pair among the multiple image pairs to obtain an initial predicted value of the target prediction object; wherein the target prediction object includes a target pose, the target pose indicating the pose of the checkerboard coordinate system relative to the human-based coordinate system; predicting the value of the target pose based on the multiple image pairs and the initial predicted value to obtain a target predicted value; and determining the pose of the positioning pen's visual coordinate system relative to the human-based coordinate system based on the target predicted value, so as to calibrate the robot and the positioning pen.

[0011] To achieve the above objectives, a second aspect of this application provides a calibration device, comprising: an acquisition module for acquiring multiple image pairs; wherein the multiple image pairs are obtained by multi-view photography of a checkerboard pattern displayed on the screen of a positioning pen using a robot's binocular camera, the binocular camera including a left-eye camera and a right-eye camera, and any image pair including a left-eye image and a right-eye image; a first prediction module for predicting the value of a target prediction object based on a target image pair among the multiple image pairs, to obtain an initial prediction value of the target prediction object; wherein the target prediction object includes a target pose, the target pose indicating the pose of the checkerboard coordinate system relative to the human-based coordinate system; a second prediction module for predicting the value of the target pose based on the multiple image pairs and the initial prediction value, to obtain a target prediction value; and a determination module for determining the pose of the positioning pen's visual coordinate system relative to the human-based coordinate system based on the target prediction value, so as to calibrate the robot and the positioning pen.

[0012] To achieve the above objectives, a third aspect of this application provides an electronic device, including: a processor and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method described in the first aspect of the application above.

[0013] To achieve the above objectives, a fourth aspect of this application provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, are used to implement the method described in the first aspect of the present application.

[0014] To achieve the above objectives, a fifth aspect of this application provides a computer program product including a computer program that, when executed by a processor, implements the method described in the first aspect of the above-described embodiment.

[0015] The calibration method, apparatus, device, and medium provided in this application acquire multiple image pairs. These image pairs are obtained by capturing images of a checkerboard pattern displayed on the screen of a positioning pen from multiple perspectives using a robot's binocular camera. The binocular camera includes a left-eye camera and a right-eye camera, and each image pair includes a left-eye image and a right-eye image. Based on the target image pairs among the multiple image pairs, the value of the target prediction object is predicted to obtain an initial predicted value for the target prediction object. The target prediction object includes a target pose, which indicates the pose of the checkerboard coordinate system relative to the human-based coordinate system. Based on the multiple image pairs and the initial predicted value, the value of the target pose is predicted to obtain a target predicted value. According to the target predicted value, the pose of the positioning pen's visual coordinate system relative to the human-based coordinate system is determined to calibrate the robot and the positioning pen. Therefore, by first making an initial prediction of the target pose in the chessboard coordinate system relative to the human base coordinate system based on the target image, and approximating the true value of the target pose, a secondary prediction is then performed by combining multiple image pairs and the initial prediction value that approximates the true value. This improves the accuracy and reliability of the target pose prediction. Finally, the pose of the positioning pen's visual coordinate system relative to the human base coordinate system is determined based on the target pose prediction value to complete the calibration of the robot and the positioning pen, effectively improving the accuracy of the calibration. At the same time, by adopting a multi-step prediction method, the impact of problems such as blind spots and insufficient feature information in single-view images on the prediction accuracy is effectively reduced, improving the accuracy and stability of the target pose prediction. Moreover, the initial prediction value can provide a precise initial anchor point for the secondary pose calculation, avoiding problems such as slow iterative convergence and local optimum traps caused by the initial value deviating from the true value during the calculation of high-dimensional data from multiple image pairs. By making full use of the rich feature information of multiple image pairs, a balance between calibration accuracy and calculation efficiency is maintained.

[0016] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0017] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:

[0018] Figure 1 This is a schematic flowchart of a calibration method provided in an embodiment of this application;

[0019] Figure 2A schematic flowchart illustrating a calibration method provided in another embodiment of this application;

[0020] Figure 3 A schematic flowchart illustrating a calibration method provided in another embodiment of this application;

[0021] Figure 4 This is a schematic diagram of the calibration device provided in another embodiment of this application. Detailed Implementation

[0022] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0023] In the technical solutions for calibrating robots and positioning pens, there are mainly two approaches:

[0024] One method involves using a calibration board for calibration. The specific process is to first calibrate the external parameters between the robot's camera and the calibration board, and between the positioning pen's camera and the calibration board. Then, using the calibration board as an intermediary, a series of related calculations are performed to obtain the external parameters between the robot's camera and the positioning pen.

[0025] Secondly, calibration is performed using a co-view method involving both the robot's camera and the positioning pen's camera. In practice, the robot's camera and the positioning pen's camera are used to capture images of feature points in space. Then, 2D matching or 3D matching methods are used to optimize and calculate the captured data, thereby determining the relative positional relationship between the robot and the positioning pen.

[0026] However, both approaches have limitations. The first approach requires an additional calibration board, which not only increases costs but also makes subsequent maintenance of the calibration board more difficult. The second approach requires a high number of feature points in the environment; calibration can only be carried out successfully if there are enough feature points in the environment, but this condition is often difficult to meet in real-world applications.

[0027] In response to at least one of the aforementioned problems, this application proposes a calibration method, apparatus, equipment, and medium.

[0028] The calibration method, apparatus, device, and medium of this application are described below with reference to the accompanying drawings. Before specifically describing the embodiments of this application, commonly used technical terms are first introduced for ease of understanding:

[0029] This application does not limit the form of the robot. For example, the form of the robot includes, but is not limited to, humanoid robots, quadruped robots, multi-legged robots, etc.

[0030] This application does not limit the functions of the robot. For example, the robot includes, but is not limited to: industrial robots (such as welding robots, handling robots, and assembly robots), service robots, medical robots, educational robots, and robots for special environments.

[0031] Figure 1 This is a schematic flowchart of a calibration method provided in an embodiment of this application.

[0032] It should be noted that the calibration method of this application embodiment can be applied to a calibration device. In some possible embodiments, the calibration device can be applied to a robot, electronic device, or chip to enable the robot, electronic device, or chip to perform calibration functions. Additionally, in some possible embodiments, the calibration device can also be software within a robot or electronic device.

[0033] In any embodiment of this application, the chip can be integrated into a robot or electronic device. The chip includes a Central Processing Unit (CPU), Image Signal Processing (ISP), Application-Specific Integrated Circuit (ASIC), Digital Signal Processor (DSP), Field-Programmable Gate Array (FPGA), System-on-Chip (SOC), Reduced Instruction Set Computer (RISC), etc., which will not be listed here.

[0034] Among them, electronic devices can be any device with computing capabilities, such as personal computers (PCs), industrial computers, host computers, mobile terminals, servers, etc. Mobile terminals can be hardware devices with various operating systems, touch screens and / or displays, such as mobile phones, tablets, personal digital assistants, wearable devices, etc.

[0035] like Figure 1 As shown, the method includes the following steps:

[0036] Step S101: Acquire multiple image pairs; wherein, the multiple image pairs are obtained by taking multi-view photos of the chessboard displayed on the screen of the positioning pen using the robot's binocular camera.

[0037] Multiple image pairs can be obtained by capturing the checkerboard pattern displayed on the screen of the positioning pen from multiple perspectives using the robot's binocular camera. It should be noted that any single image pair can be obtained by capturing the checkerboard pattern displayed on the screen of the positioning pen from the robot's binocular camera at the corresponding perspective. It should also be noted that, in this application, the perspective refers to the observation perspective of the binocular camera relative to the checkerboard calibration board.

[0038] It should be noted that this application does not restrict the robot's degrees of freedom; for example, the robot can be a six-degree-of-freedom robot, a five-degree-of-freedom robot, etc.

[0039] It should also be noted that the robot can be equipped with binocular cameras and a positioning pen. The binocular cameras can include a left camera and a right camera, and the screen of the positioning pen can display a chessboard image.

[0040] Each image pair may include a left-eye image and a right-eye image.

[0041] In one embodiment of this application, the calibration device can use a binocular camera configured on a robot to capture multiple images of the checkerboard pattern displayed on the screen of the positioning pen from multiple perspectives, thereby obtaining multiple image pairs. For example, the calibration device can use the binocular camera configured on a robot to capture images of the checkerboard pattern displayed on the screen of the positioning pen from three different perspectives, thereby obtaining images from the corresponding perspectives.

[0042] It should be noted that this application does not limit the number of viewpoints involved in multiple image pairs.

[0043] Step S102: Based on the target image pair among multiple image pairs, predict the value of the target prediction object to obtain the initial prediction value of the target prediction object.

[0044] The target prediction object may include the target pose, which can be used to indicate the pose of the chessboard coordinate system relative to the human base coordinate system.

[0045] The checkerboard coordinate system refers to the reference coordinate system of the checkerboard on the positioning pen screen. For example, this reference coordinate system usually takes the upper left corner of the checkerboard as its origin. The human-based coordinate system, also known as the robot-based coordinate system, is the robot's fixed reference coordinate system.

[0046] It should be noted that in this application, the pose can be composed of a combination of position information and attitude information. The position information is represented by a translation vector t, and the attitude information is represented by a rotation matrix R. For example, the target pose can be the three-dimensional coordinates of the origin of the checkerboard coordinate system in the human-based coordinate system, i.e., the translation vector of the checkerboard coordinate system relative to the human-based coordinate system. And the rotation matrix corresponding to the rotation relationship between the chessboard coordinate system and the human base coordinate system. Combining the target pose It can be represented as:

[0047] (1)

[0048] The target image pair can be any one of multiple image pairs.

[0049] In the embodiments of this application, the value of the target prediction object can be predicted based on the target image pair among multiple image pairs to obtain the initial prediction value of the target prediction object.

[0050] Step S103: Based on multiple image pairs and the initial prediction value, predict the target pose value to obtain the target prediction value.

[0051] In the embodiments of this application, after the initial predicted value of the target object is obtained, the value of the target pose can be predicted based on multiple image pairs and the initial predicted value of the target object, thereby obtaining the target predicted value of the target pose.

[0052] Step S104: Based on the target prediction value, determine the pose of the positioning pen's visual coordinate system relative to the human base coordinate system to calibrate the robot and the positioning pen.

[0053] The visual coordinate system of the positioning pen is a rigid, fixed coordinate system. Its origin and coordinate axis directions are fixedly bound to the physical structure of the positioning pen and will not change their relative relationship with the spatial pose of the positioning pen. It belongs to the right-handed three-dimensional rectangular coordinate system.

[0054] As an example, the calibration device can pre-store coordinate system transformation strategies, which can be constructed based on the cascade operation rules and inverse transformation operation rules of homogeneous transformation matrices. During the calibration process, the calibration device can call the corresponding coordinate system transformation strategy and, according to the coordinate system transformation strategy, perform correlation calculations between the target prediction value of the target pose of the chessboard coordinate system relative to the human base coordinate system and the pose of the pre-calibrated positioning pen's visual coordinate system relative to the chessboard coordinate system, thereby determining the pose of the positioning pen's visual coordinate system relative to the human base coordinate system, and thus realizing the calibration of the relationship between the robot and the positioning pen.

[0055] The calibration method of this application embodiment acquires multiple image pairs. These image pairs are obtained by capturing images of the checkerboard pattern displayed on the screen of the positioning pen from multiple perspectives using the robot's binocular camera. The binocular camera includes a left-eye camera and a right-eye camera, and each image pair includes a left-eye image and a right-eye image. Based on the target image pairs among the multiple image pairs, the value of the target prediction object is predicted to obtain an initial predicted value for the target prediction object. The target prediction object includes a target pose, which indicates the pose of the checkerboard coordinate system relative to the human-based coordinate system. Based on the multiple image pairs and the initial predicted value, the value of the target pose is predicted to obtain a target predicted value. According to the target predicted value, the pose of the positioning pen's visual coordinate system relative to the human-based coordinate system is determined to calibrate the robot and the positioning pen. Therefore, by first making an initial prediction of the target pose in the chessboard coordinate system relative to the human base coordinate system based on the target image, and approximating the true value of the target pose, a secondary prediction is then performed by combining multiple image pairs and the initial prediction value that approximates the true value. This improves the accuracy and reliability of the target pose prediction. Finally, the pose of the positioning pen's visual coordinate system relative to the human base coordinate system is determined based on the target pose prediction value to complete the calibration of the robot and the positioning pen, effectively improving the accuracy of the calibration. At the same time, by adopting a multi-step prediction method, the impact of problems such as blind spots and insufficient feature information in single-view images on the prediction accuracy is effectively reduced, improving the accuracy and stability of the target pose prediction. Moreover, the initial prediction value can provide a precise initial anchor point for the secondary pose calculation, avoiding problems such as slow iterative convergence and local optimum traps caused by the initial value deviating from the true value during the calculation of high-dimensional data from multiple image pairs. By making full use of the rich feature information of multiple image pairs, a balance between calibration accuracy and calculation efficiency is maintained.

[0056] To clearly illustrate how the target image pairs in the above embodiments of this application predict the value of the target object and obtain the initial predicted value of the target object, this application also proposes a calibration method.

[0057] Figure 2 This is a schematic flowchart of a calibration method provided in another embodiment of this application.

[0058] like Figure 2 As shown, the method includes the following steps:

[0059] Step S201: Acquire multiple image pairs; wherein, the multiple image pairs are obtained by taking multi-view photos of the chessboard displayed on the screen of the positioning pen using the robot's binocular camera.

[0060] It should be noted that the explanation of step S201 can be found in the relevant explanations in any embodiment of this application, and will not be repeated here.

[0061] Step S202: Based on the first pose and size scaling factor, any feature point among the multiple feature points of the chessboard is projected onto the image coordinate system corresponding to the target image pair to obtain the first projected coordinates of the feature point.

[0062] The first pose can be used to indicate the pose of the checkerboard coordinate system relative to the right eye camera coordinate system of the right eye camera in a stereo camera setup. The right eye camera coordinate system refers to the three-dimensional coordinate system of the optical center of the right eye camera. It should be noted that the above explanations of pose, checkerboard coordinate system, and stereo camera also apply to this embodiment, and will not be repeated here.

[0063] The size scale factor is used to determine the physical dimensions of the checkerboard grid. The physical dimensions of the checkerboard grid indicate the length (i.e., the side length of the square) between two adjacent interior corners of any square (e.g., a white or black square). It should be noted that because the checkerboard grid is displayed on a relatively small screen, its actual physical dimensions are limited by both screen resolution and display scale. Measuring the physical dimensions of the checkerboard grid using calipers presents numerous challenges. The calipers are difficult to align precisely with the edges of the grid, making accurate measurement extremely difficult. Furthermore, even minor operational deviations during measurement are significantly amplified, thus having a substantial impact on the accuracy of the final measurement result. Given these significant limitations in measurement accuracy, it is currently impossible to accurately obtain the physical dimensions of the checkerboard grid through direct measurement; therefore, the physical dimensions of the checkerboard grid are currently unknown.

[0064] The feature points can be, but are not limited to, the inner corner points of the chessboard, the outer corner points of the chessboard, the center point of the chessboard, etc. This application does not impose any restrictions on them.

[0065] Optionally, in some embodiments, for any feature point among multiple feature points of the checkerboard pattern, the feature point can be projected into the first image coordinate system corresponding to the left eye image in the target image pair and the second image coordinate system corresponding to the right eye image in the target image pair based on the first pose and the size scaling factor, respectively, to obtain the first image coordinates of the feature point in the first image coordinate system and the second image coordinates of the feature point in the second image coordinate system; both the first image coordinates and the second image coordinates are used as the first projection coordinates corresponding to the feature point.

[0066] As an example, suppose the first pose As shown in the following formula:

[0067] (2)

[0068] in, The rotation matrix represents the rotation relationship between the chessboard coordinate system and the right camera coordinate system. This represents the three-dimensional coordinates of the origin of the checkerboard coordinate system in the right-eye camera coordinate system, that is, the translation vector of the checkerboard coordinate system relative to the right-eye camera coordinate system.

[0069] For the i-th feature point (where i is a non-zero natural number) among multiple feature points, the following formula can be used to project the feature point into the first image coordinate system corresponding to the left eye image in the target image pair and the second image coordinate system corresponding to the right eye image in the target image pair, respectively, to obtain the first image coordinates of the i-th feature point in the first image coordinate system and the second image coordinates of the i-th feature point in the second image coordinate system:

[0070] (3)

[0071] (4)

[0072] in, Indicates the coordinates of the second image; p represents the coordinates of the first image; i This represents the coordinates of the i-th feature point in the checkerboard coordinate system, which can be determined based on the physical dimensions of the checkerboard; α represents the size scaling factor. The rotation matrix represents the rotation relationship between the right camera coordinate system and the left camera coordinate system. This represents the three-dimensional coordinates of the origin of the right camera coordinate system in the left camera coordinate system, that is, the translation vector of the right camera coordinate system relative to the left camera coordinate system. It is a projection function. It should be noted that the size scaling factor and the first pose are both unknowns, and the binocular cameras have been pre-calibrated; that is, the pose of the right camera coordinate system relative to the left camera coordinate system is known. and It is known.

[0073] Therefore, the first image coordinates and the second image coordinates corresponding to the i-th feature point can both be determined as the first projection coordinates corresponding to the i-th feature point.

[0074] Step S203: Based on the difference between the first projected coordinates corresponding to the feature point and the actual observed coordinates of the feature point in the image coordinate system corresponding to the target image pair, construct the first residual function corresponding to the feature point.

[0075] Optionally, in some embodiments, the image coordinates of the feature point in the left eye image and the image coordinates of the feature point in the right eye image of the target image pair are used as the actual observation coordinates, that is, the image coordinates of the feature point in the first image coordinate system and the image coordinates of the feature point in the second image coordinate system are used as the actual observation coordinates; based on the difference between the first image coordinates and the actual observation coordinates of the feature point in the left eye image of the target image pair, a first residual function corresponding to the feature point in the left eye camera is constructed, and based on the difference between the second image coordinates and the actual observation coordinates of the feature point in the right eye image of the target image pair, a first residual function corresponding to the feature point in the right eye camera is constructed.

[0076] Using the example above, for the i-th feature point among multiple feature points of the chessboard pattern, the coordinates of the i-th feature point are determined by the second image coordinates and the actual observed coordinates of that feature point in the right eye image of the target image pair. The difference between them can be determined by the following formula, which specifies the first residual function corresponding to the i-th feature point in the right camera:

[0077] (5)

[0078] Based on the second image coordinates and the actual observed coordinates of the feature point in the left eye image of the target image pair... The difference between them can be determined by the following formula, which specifies the first residual function corresponding to the i-th feature point in the left camera:

[0079] (6)

[0080] Step S204: Based on the first residual function corresponding to multiple feature points, predict the value of the target prediction object to obtain the initial prediction value.

[0081] Using the example above, we can predict the value of the target object based on the first residual function corresponding to each feature point in the right camera and the first residual function corresponding to each feature point in the left camera, and obtain the initial predicted value.

[0082] As one possible implementation, if the target prediction object also includes a size scaling factor, the implementation process of step S204 may include the following steps:

[0083] Step S2041: Construct a first objective function based on the first residual function corresponding to multiple feature points.

[0084] As an example, assuming the number of feature points on the chessboard is n, where n is a non-zero natural number, then the right-eye residual vector r, formed by combining the first residual functions corresponding to each feature point in the right-eye camera, is... right for:

[0085] r right =[r right,1 ,……r right,i , ...r right,n ];(7)

[0086] The left-eye residual vector r is formed by combining the first residual functions corresponding to each feature point in the left-eye camera. left for:

[0087] r left =[r left,1 ,……r left,i , ...r left,n ];(8)

[0088] Where, r right,i This represents the value of the first residual function corresponding to the i-th feature point in the right eye camera; r left,i Let i represent the value of the first residual function corresponding to the i-th feature point in the left eye camera; i∈[1,n].

[0089] Therefore, the first objective function F1 can be determined according to the following formula:

[0090] (9)

[0091] It should be noted that the above example of the first objective function is merely exemplary, and in practical applications, it can be other functions as well, which is not limited in this application.

[0092] Step S2042: Obtain the initial estimated size of the physical dimensions of the chessboard grid.

[0093] As an example, a value can be randomly selected from a certain range as the initial estimated size of the physical dimensions of the chessboard. For example, if the range is (0.008~0.012), then 0.01, 0.11, etc. can be selected as the initial estimated size of the physical dimensions of the chessboard.

[0094] Step S2043: Using the initial estimated size as the initial value for iteration, and with the goal of minimizing the value of the first objective function, determine the predicted pose value of the first pose and the predicted scale coefficient of the size scale coefficient.

[0095] As an example, the initial estimated size of the physical dimensions of the checkerboard can be used as the initial value for iteration, and the goal is to minimize the value of the first objective function. The LM (Levenberg-Marquardt) algorithm is used for iterative calculation to determine the optimal value of the first pose and the size scaling factor. The optimal value of the first pose is determined as the predicted pose value of the first pose, and the optimal value of the size scaling factor is determined as the predicted scaling factor of the size scaling factor.

[0096] Step S2044: Determine the initial predicted value of the target object based on the predicted pose value of the first pose and the predicted scale coefficient of the size scale coefficient.

[0097] Optionally, in some embodiments, the initial predicted value of the target pose is determined based on the predicted pose value and the pose of the right eye camera coordinate system relative to the human base coordinate system; the initial predicted value of the physical size is determined based on the predicted scaling factor and the initial estimated size.

[0098] It should be noted that the spatial pose relationship between the right camera coordinate system and the human base coordinate system has been pre-calibrated, meaning that the pose of the right camera coordinate system relative to the human base coordinate system is known.

[0099] As an example, suppose the predicted pose of the first pose is taken as... The pose of the right camera coordinate system relative to the human base coordinate system is: Then, the initial predicted value of the target pose can be determined according to the following formula. for:

[0100] (10)

[0101] Assuming the prediction scaling factor is α and the initial size estimate of the physical dimensions of the checkerboard is s, the initial predicted value of the physical dimensions of the checkerboard can be determined using the following formula:

[0102] s'=α×s;(11)

[0103] Therefore, by deriving the initial predicted value of the target pose based on the predicted pose value and the pose of the right camera coordinate system relative to the human base coordinate system, the initial predicted value of the physical size of the checkerboard pattern is calculated by relying on the predicted scale coefficient and the initial estimated size. This enables precise coupling of the pose relationships between the camera coordinate system, the human base coordinate system, and the checkerboard coordinate system, establishing a parameter transfer link between different coordinate systems and ensuring that the solution of the initial predicted value of the target pose has a clear geometric constraint basis. At the same time, the prediction of physical size and the solution of pose parameters are promoted in a coordinated manner, avoiding the separation between the pose and size parameter solution processes, effectively reducing the impact of single parameter solution error on the overall prediction result, and improving the accuracy and consistency of the initial predicted value of the target object.

[0104] Understandably, by constructing a first objective function based on the first residual function corresponding to multiple feature points, using the initial estimated size of the checkerboard physical dimensions as the initial value for iteration, and minimizing the value of the first objective function as the optimization objective, the predicted pose value of the first pose and the predicted scale coefficient of the size scale coefficient are obtained. This determines the initial predicted value of the target object, transforming the parameter solution of the checkerboard pose and physical dimensions into a standardized numerical optimization problem. By leveraging the minimization constraint of the objective function, effective suppression of imaging distortion and feature point observation errors is achieved, ensuring that the solved pose and scale coefficient parameters are more consistent with the real scene. At the same time, using the initial estimated value of the physical dimensions as the starting point for iteration can reduce the iterative complexity in the parameter optimization process, avoiding the problem of iteration divergence or getting trapped in local optima due to improper selection of initial values, improving the efficiency and accuracy of solving the initial predicted value, and further improving the prediction accuracy of the target pose in subsequent secondary predictions, thereby helping to improve the accuracy and reliability of calibration.

[0105] Step S205: Based on multiple image pairs and the initial prediction value, predict the target pose value to obtain the target prediction value.

[0106] Step S206: Based on the target prediction value, determine the pose of the positioning pen's visual coordinate system relative to the human base coordinate system to calibrate the robot and the positioning pen.

[0107] It should be noted that the explanations of steps S205 to S206 can be found in the relevant explanations in any embodiment of this application, and will not be repeated here.

[0108] Optionally, in some embodiments, the configuration file of the positioning pen can be read to determine the pose of the positioning pen's visual coordinate system relative to the checkerboard coordinate system; based on the pose of the positioning pen's visual coordinate system relative to the checkerboard coordinate system and the target prediction value of the target pose, the pose of the positioning pen's visual coordinate system relative to the human-based coordinate system can be determined.

[0109] The visual coordinate system of the positioning pen, also known as the positioning pen camera coordinate system or positioning pen coordinate system, is a three-dimensional Cartesian coordinate system established with the positioning pen's built-in visual sensor (such as a miniature camera or infrared imaging module) as the core. It is the "local reference system" by which the positioning pen perceives its own relative pose to the chessboard (reference object).

[0110] The configuration file for the positioning pen can record the pose of the positioning pen's visual coordinate system relative to the checkerboard coordinate system.

[0111] As an example, the calibration device can read the positioning pen's configuration file and obtain the positioning pen's pose in the visual coordinate system relative to the checkerboard coordinate system from the configuration file. Assuming the target pose prediction value is Then, the pose of the positioning pen's visual coordinate system relative to the human base coordinate system can be determined using the following formula. :

[0112] (12)

[0113] Furthermore, after determining the pose of the positioning pen's visual coordinate system relative to the human-based coordinate system, the three-dimensional coordinates of the origin of the positioning pen's visual coordinate system in the human-based coordinate system can be extracted from the matrix corresponding to the pose of the positioning pen's visual coordinate system relative to the human-based coordinate system. This yields the translation vector of the positioning pen's visual coordinate system relative to the human-based coordinate system. And the rotation matrix corresponding to the rotation relationship between the visual coordinate system of the positioning pen and the human base coordinate system. This allows for the calibration of the coordinate transformation relationship between the robot and the positioning pen.

[0114] Therefore, by reading the configuration file of the positioning pen to determine its pose relative to the checkerboard coordinate system, and combining it with the target prediction value of the target pose, the pose calculation of the positioning pen's visual coordinate system relative to the human base coordinate system is completed. It can establish a rigid transformation link between the checkerboard coordinate system, the human base coordinate system, and the positioning pen's visual coordinate system by relying on the precise coupling between the inherent parameters of the configuration file and the target prediction value of the pose. This effectively avoids the cumulative error caused by parameter missing or deviation during the multi-coordinate system transformation process. At the same time, this method does not require the introduction of additional complex calibration references or manual intervention processes. It directly achieves pose calculation through the collaborative calculation of existing parameters, which simplifies the calibration operation steps of the robot and the positioning pen, improves the automation level and execution efficiency of the calibration process, and provides stable and reliable pose data support for the high-precision collaborative operation of the robot and the positioning pen.

[0115] The calibration method of this application embodiment projects any one of the multiple feature points of the checkerboard grid onto the image coordinate system corresponding to the target image pair based on the first pose and the size scaling factor, to obtain the first projected coordinates corresponding to the feature point; wherein, the first pose is used to indicate the pose of the checkerboard grid coordinate system relative to the right eye camera coordinate system, and the size scaling factor is used to determine the physical size of the checkerboard grid; based on the difference between the first projected coordinates corresponding to the feature point and the actual observed coordinates of the feature point in the image coordinate system corresponding to the target image pair, a first residual function corresponding to the feature point is constructed; based on the first residual functions corresponding to multiple feature points, the value of the target prediction object is predicted to obtain the initial prediction value. Therefore, by projecting the checkerboard feature points onto the image coordinate system of the target image pair based on the first pose and size scaling factor to obtain the first projected coordinates, and constructing the first residual function in combination with the actual observed coordinates of the feature points, and then relying on the residual functions of multiple feature points to solve the initial prediction value of the target object, the geometric features of the checkerboard can be used as anchor points to transform the coupling relationship between pose parameters and physical size parameters into a quantifiable residual optimization problem. This effectively reduces the interference of factors such as imaging distortion and feature point recognition deviation in the target image pair on the initial prediction, ensuring that the initial prediction value can closely approximate the true value range. At the same time, based on the collaborative optimization logic of multi-feature point residuals, the spatial distribution characteristics of checkerboard feature points can be fully utilized to improve the stability and consistency of the prediction results, providing a high-precision and high-reliability initial anchor point for subsequent secondary pose prediction using multiple image pairs, reducing the number of iterative convergence steps and computational redundancy in the secondary prediction process, and improving the calibration accuracy and efficiency of the robot and positioning pen calibration process.

[0116] To clearly illustrate how the target pose is predicted based on multiple image pairs and initial prediction values ​​in any of the above embodiments of this application, and to obtain the target prediction value, this application also proposes a calibration method.

[0117] Figure 3 This is a schematic flowchart of a calibration method provided in another embodiment of this application.

[0118] like Figure 3 As shown, the method includes the following steps:

[0119] Step S301: Acquire multiple image pairs; wherein, the multiple image pairs are obtained by taking multi-view photos of the chessboard displayed on the screen of the positioning pen using the robot's binocular camera.

[0120] Step S302: Based on the target image pair among multiple image pairs, predict the value of the target prediction object to obtain the initial prediction value of the target prediction object.

[0121] The target prediction object may include the target pose, which can be used to indicate the pose of the chessboard coordinate system relative to the human base coordinate system.

[0122] It should be noted that the explanations of steps S301 to S302 can be found in the relevant explanations in any embodiment of this application, and will not be repeated here.

[0123] Step S303: For any image pair among multiple image pairs, based on the target pose and size scaling factor, project any feature point among multiple feature points of the chessboard grid onto the image coordinate system corresponding to the image pair to obtain the second projection coordinates corresponding to the combination between the feature point and the image pair.

[0124] The size scale factor is used to determine the physical dimensions of the checkerboard grid. The physical dimensions of the checkerboard grid indicate the length (i.e., the side length of the square) between two adjacent interior corners of any square (e.g., a white or black square). It should be noted that because the checkerboard grid is displayed on a relatively small screen, its actual physical dimensions are limited by both screen resolution and display scale. Measuring the physical dimensions of the checkerboard grid using calipers presents numerous challenges. The calipers are difficult to align precisely with the edges of the grid, making accurate measurement extremely difficult. Furthermore, even minor operational deviations during measurement are significantly amplified, thus having a substantial impact on the accuracy of the final measurement result. Given these significant limitations in measurement accuracy, it is currently impossible to accurately obtain the physical dimensions of the checkerboard grid through direct measurement; therefore, the physical dimensions of the checkerboard grid are currently unknown.

[0125] The feature points can be, but are not limited to, the inner corner points of the chessboard, the outer corner points of the chessboard, the center point of the chessboard, etc. This application does not impose any restrictions on them.

[0126] The second projection coordinates corresponding to the combination of the feature point and the image pair are the second projection coordinates of the feature point in the viewpoint corresponding to the image pair.

[0127] Optionally, in some embodiments, for any image pair among multiple image pairs, and for any feature point among multiple feature points of the checkerboard pattern, the feature point can be projected into the third image coordinate system corresponding to the left eye image in the image pair and the fourth image coordinate system corresponding to the right eye image in the image pair, based on the target pose and size scaling factor, to obtain the third image coordinates of the feature point in the corresponding third image coordinate system and the fourth image coordinates in the corresponding fourth image coordinate system; the third image coordinates of the feature point in the corresponding third image coordinate system and the fourth image coordinates in the corresponding fourth image coordinate system are both used as the second projection coordinates corresponding to the combination between the feature point and the corresponding image pair.

[0128] As an example, assume the target pose As shown in the following formula:

[0129] (13)

[0130] in, The rotation matrix represents the rotation relationship between the chessboard coordinate system and the human-based coordinate system. This represents the three-dimensional coordinates of the origin of the checkerboard coordinate system in the human-based coordinate system, that is, the translation vector of the checkerboard coordinate system relative to the human-based coordinate system.

[0131] For an image pair from the j-th viewpoint in a set of multiple image pairs, the extrinsic parameters of the stereo camera are known, i.e., the pose of the right camera coordinate system relative to the left camera coordinate system. Given that the pose of the right camera coordinate system relative to the human base coordinate system at the j-th viewpoint is... Given that:

[0132] (14)

[0133] (15)

[0134] The rotation matrix represents the rotation relationship between the right camera coordinate system and the left camera coordinate system. This represents the three-dimensional coordinates of the origin of the right camera coordinate system in the left camera coordinate system, that is, the translation vector of the right camera coordinate system relative to the left camera coordinate system. This represents the rotation matrix corresponding to the rotation relationship of the right eye camera coordinate system relative to the human base coordinate system at the j-th viewpoint; This represents the three-dimensional coordinates of the origin of the right camera coordinate system in the human-based coordinate system at the j-th viewpoint, which is the translation vector of the origin of the right camera coordinate system relative to the human-based coordinate system at the j-th viewpoint.

[0135] The pose of the checkerboard coordinate system relative to the right camera coordinate system at the j-th viewpoint. It can be represented as:

[0136] (16)

[0137] The pose of the checkerboard coordinate system relative to the left eye coordinate system at the j-th viewpoint It can be represented as:

[0138] (17)

[0139] Furthermore, for the i-th feature point (where i is a non-zero natural number) among multiple feature points, the feature point can be projected into the third image coordinate system corresponding to the left eye image in the image pair and the fourth image coordinate system corresponding to the right eye image in the image pair, respectively, according to the following formula, to obtain the third image coordinates corresponding to the third image coordinate system and the fourth image coordinates corresponding to the fourth image coordinate system:

[0140] (18)

[0141] (19)

[0142] in:

[0143] (20)

[0144] ;(twenty one)

[0145] ;(twenty two)

[0146] ;(twenty three)

[0147] Indicates the coordinates of the third image; Indicates the coordinates of the fourth image; p i,j This represents the coordinates of the i-th feature point in the checkerboard coordinate system from the j-th viewpoint. These coordinates can be determined based on the physical dimensions of the checkerboard from the j-th viewpoint; α represents the size scaling factor. It is a projection function. It should be noted that the target pose and size scale coefficients are both unknowns, and their values ​​need to be predicted.

[0148] Furthermore, the third and fourth image coordinates corresponding to the feature point under the corresponding viewpoint can be determined as the second projection coordinates under the viewpoint corresponding to the image pair of the feature point, that is, the second projection coordinates corresponding to the combination between the feature point and the corresponding image pair are obtained.

[0149] Step S304: Based on the difference between the second projected coordinates and the actual observed coordinates of the feature points in the corresponding image coordinate system of the image pair, construct the second residual function corresponding to the combination.

[0150] Optionally, in some embodiments, the image coordinates of the feature point in the left eye image and the image coordinates of the right eye image in the image pair are used as the actual observed coordinates. That is, the image coordinates of the feature point in the third image coordinate system corresponding to the left eye image and the image coordinates in the fourth image coordinate system corresponding to the right eye image in the image pair are used as the actual observed coordinates. Based on the difference between the third image coordinates and the actual observed coordinates of the feature point in the third image coordinate system corresponding to the left eye image in the image pair, a residual function corresponding to the feature point in the left eye camera at the corresponding viewpoint of the image pair is constructed. Based on the difference between the fourth image coordinates and the actual observed coordinates of the feature point in the fourth image coordinate system corresponding to the right eye image in the image pair, a residual function corresponding to the feature point in the right eye camera at the corresponding viewpoint of the image pair is constructed. Furthermore, the residual function corresponding to the feature point in the left eye camera at the corresponding viewpoint of the image pair and the residual function corresponding to the feature point in the right eye camera at the corresponding viewpoint of the image pair are both used as the second residual function corresponding to the combination between the feature point and the corresponding image pair.

[0151] Using the example above, for an image pair from the j-th viewpoint in multiple image pairs, and for the i-th feature point among multiple feature points of the checkerboard pattern, assume that the actual observed coordinates of this feature point in the third image coordinate system corresponding to the left-eye image in the image pair are... The difference between the third image coordinates of the feature point projected onto the third image coordinate system corresponding to the left eye image in the image pair and the actual observed coordinates of the feature point in the third image coordinate system corresponding to the left eye image in the image pair can be determined by the following formula: The residual function corresponding to the i-th feature point in the left eye camera under the corresponding viewpoint of the image pair is:

[0152] ;(twenty four)

[0153] Assume the actual observed coordinates of this feature point in the fourth image coordinate system corresponding to the right eye image in the image pair are: The difference between the fourth image coordinates of the feature point projected onto the fourth image coordinate system corresponding to the right eye image in the image pair and the actual observed coordinates of the feature point in the fourth image coordinate system corresponding to the right eye image in the image pair can be determined by the following formula: The residual function corresponding to the i-th feature point in the right eye camera under the corresponding viewpoint of the image pair is:

[0154] (25)

[0155] Furthermore, the residual function of the i-th feature point corresponding to the left eye camera in the corresponding view of the image pair and the residual function of the i-th feature point corresponding to the right eye camera in the corresponding view of the image pair can both be determined as the second residual function corresponding to the combination between the i-th feature point and the image pair.

[0156] Step S305: Based on the second residual function corresponding to each combination, predict the value of the target pose to obtain the target predicted value.

[0157] As one possible implementation, step S305 may include the following steps:

[0158] Step S3051: Construct a second objective function based on the second residual function corresponding to each combination.

[0159] It should be noted that the process of constructing the second objective function is similar to the process of constructing the first objective function in step S2041, and will not be described in detail here.

[0160] Step S3052: Using the initial predicted value of the target object as the initial value for iteration, and taking the minimum value of the second objective function as the objective, determine the target predicted value of the target pose.

[0161] As an example, when the target prediction object includes the target pose and the physical dimensions of the checkerboard, the initial predicted value of the target prediction object obtained in step S302 is used as the initial value for iteration. The LM algorithm is used for iterative calculation with the objective of minimizing the value of the second objective function, thereby determining the optimal value of the target pose. This optimal value of the target pose is then determined as the target predicted value of the target pose. Optionally, the optimal value of the size scaling factor can also be determined simultaneously, and the target size of the checkerboard's physical dimensions can be determined based on the optimal value of the size scaling factor and the initial predicted value of the checkerboard's physical dimensions. For example, the value obtained by multiplying the optimal value of the size scaling factor by the initial predicted value of the checkerboard's physical dimensions is used as the target size of the checkerboard's physical dimensions.

[0162] Therefore, by constructing a second objective function based on the second residual function corresponding to each combination (i.e., the combination between feature points and image pairs), the initial predicted value of the target object is used as the initial value for iteration, and the minimization of the value of the second objective function is used as the optimization objective. The target predicted value of the target pose is obtained by solving this problem. It can rely on the accurate anchoring effect of the initial predicted value to compress the iteration convergence cycle of multi-view pose optimization and avoid the problem of iteration divergence or getting trapped in local optima caused by the initial value deviating from the true range. At the same time, with the constraint of the objective function composed of multi-view residuals, the checkerboard feature information under different observation angles can be fully integrated, effectively offsetting the influence of interference factors such as imaging distortion and blind spots of a single viewpoint on the prediction results. This allows the target pose solution process to have multi-dimensional geometric verification basis, improves the accuracy and consistency of the target predicted value, and provides highly reliable data support for the pose calculation of the positioning pen visual coordinate system relative to the human base coordinate system, further maintaining the accuracy and stability of the robot and positioning pen calibration process.

[0163] Step S306: Based on the target prediction value, determine the pose of the positioning pen's visual coordinate system relative to the human base coordinate system to calibrate the robot and the positioning pen.

[0164] It should be noted that the explanation of step S306 can be found in the relevant explanations in any embodiment of this application, and will not be repeated here.

[0165] The calibration method of this application embodiment involves projecting any feature point from multiple feature points of a checkerboard pattern onto the image coordinate system corresponding to the image pair, based on the target pose and a size scaling factor, for any image pair among multiple image pairs, to obtain the second projected coordinates corresponding to the combination of the feature point and the image pair; wherein, the size scaling factor is used to determine the physical size of the checkerboard pattern; based on the difference between the second projected coordinates and the actual observed coordinates of the feature point in the image coordinate system corresponding to the image pair, a second residual function corresponding to the combination is constructed; based on the second residual function corresponding to each combination, the value of the target pose is predicted to obtain the target predicted value. Therefore, by projecting checkerboard feature points onto the corresponding image coordinate system based on the target pose and size scaling factor for any image pair from multiple image pairs to obtain the second projected coordinates, and constructing the second residual function of the image pair by combining the actual observed coordinates of the feature points, the target pose prediction value is solved by relying on the residual function corresponding to the combination of each feature point and each image pair. This fully explores the checkerboard feature information under different observation angles in multiple image pairs, and effectively offsets the imaging limitations and observation biases of single-view images by utilizing the synergistic constraint effect of multi-view residuals. At the same time, by combining pose parameter optimization with feature matching depth of multiple image pairs, the target pose prediction process has multi-dimensional geometric constraints, improving the accuracy and robustness of the target pose prediction results, providing high-precision data support for the pose calculation of the positioning pen's visual coordinate system relative to the human base coordinate system, and further maintaining the accuracy and stability of the robot and positioning pen calibration process.

[0166] To clearly illustrate the calibration method of this application, a detailed explanation is provided below with examples.

[0167] As an example, the following needs to be prepared before implementing the calibration method:

[0168] 1. A robot (e.g., a six-degree-of-freedom robot) may be equipped with a set of binocular cameras. Before being put into use, these binocular cameras have undergone comprehensive and precise calibration to ensure the accuracy of their measurements and imaging. Simultaneously, the hand-eye relationship between the binocular cameras and the robot's end effector has also undergone a meticulous calibration process. This calibration result enables the system to obtain the transformation relationship between the right eye camera coordinate system and the robot's base coordinate system (referred to as the human base coordinate system in this application) in real time and accurately, providing a solid foundation for subsequent precise operations.

[0169] 2. A positioning pen, which can be set on the robot body, is equipped with a display screen that clearly shows a checkerboard pattern. The relationship between its coordinate system and the coordinate system of the built-in camera of the positioning pen has been predetermined and calibrated.

[0170] It should be noted that because the checkerboard pattern is displayed on a relatively small screen, its actual physical size is limited by both screen resolution and aspect ratio. Measuring the physical dimensions of the checkerboard using calipers presents numerous challenges. The calipers are difficult to align precisely with the edges of the checkerboard, making accurate measurement extremely difficult. Furthermore, even minor operational deviations during measurement are significantly amplified, thus having a substantial impact on the accuracy of the final measurement result. Given these significant limitations in measurement accuracy, it is currently impossible to accurately obtain the physical dimensions of the checkerboard through direct measurement; therefore, the physical dimensions of the checkerboard are unknown.

[0171] The calibration method may include the following steps:

[0172] Step 1) Control the robot's binocular camera to capture images of the chessboard displayed on the positioning pen's screen from three different positions or perspectives to obtain corresponding chessboard image pairs (referred to as image pairs in this application).

[0173] It should be noted that this application only uses the example of taking pictures of the chessboard displayed on the screen of the positioning pen from three different positions or perspectives. In practical applications, the chessboard displayed on the screen of the positioning pen can also be taken from four, five or other different perspectives. This application does not limit this.

[0174] It should also be noted that any chessboard image pair includes a left chessboard image (referred to as the left image in this application) and a right chessboard image (referred to as the right image in this application).

[0175] Step 2) Perform the first optimization using the chessboard image corresponding to any one of the three positions.

[0176] Specifically, the first quantity to be optimized is optimized by using a chessboard image pair corresponding to any one of the three positions to obtain the corresponding optimal value.

[0177] The first quantity to be optimized includes:

[0178] A. Pose of the checkerboard coordinate system relative to the right camera coordinate system (Referred to as the first pose in this application);

[0179] B. The size ratio factor α between the physical dimensions of the checkerboard and its estimated dimensions;

[0180] It is necessary to understand the extrinsic parameters of the binocular camera system fixed to the robot. And the pose of the right eye camera coordinate system relative to the robot base coordinate system at that position. The pose of the chessboard coordinate system relative to the left camera coordinate system. for:

[0181] (26)

[0182] The first optimization process can be as follows: For the i-th feature point among multiple feature points of the chessboard, the residual of the feature point in the right camera (denoted as the first residual function in this application) is defined as shown in formula (5), and the residual of the feature point in the left camera is defined as shown in formula (6); based on the residuals of each feature point in the right camera and the residuals of each feature point in the left camera, the first objective function is constructed.

[0183] Furthermore, the initial value of the physical dimensions of the checkerboard (denoted as the initial dimension in this application) is set to 0.01. This initial value is used as the initial value for iterative optimization, with the goal of minimizing the value of the first objective function. The LM algorithm can be used to optimize and solve the relevant nonlinear optimization equations, thereby obtaining the optimal value of the first quantity to be optimized, i.e., the pose of the checkerboard coordinate system relative to the right eye camera coordinate system. And the optimal value of the size scaling factor α. In this application, the pose of the checkerboard coordinate system relative to the right eye camera coordinate system can be determined. The optimal value of the position is denoted as the predicted pose value, and the optimal value of the size scaling factor α is denoted as the predicted scaling factor.

[0184] Step 3) Based on the optimal value of the first quantity to be optimized, determine the initial estimated value (referred to as the initial predicted value) of the target estimation object (referred to as the target prediction object in this application).

[0185] The target estimation objects include: the pose of the chessboard coordinate system relative to the robot base coordinate system (denoted as the target pose in this application) and the physical dimensions of the chessboard.

[0186] In this application, the optimal value of the first quantity to be optimized can be determined, i.e., the pose of the checkerboard coordinate system relative to the right camera coordinate system. The optimal values ​​of the coordinate system and the size scaling factor α are used to determine the initial estimates of the pose of the checkerboard coordinate system relative to the robot's base coordinate system and the initial estimates of the physical dimensions of the checkerboard.

[0187] As an example, the initial estimate of the pose of the chessboard coordinate system relative to the robot base coordinate system can be determined according to formula (10), and the initial estimate of the physical size of the chessboard can be determined according to formula (11).

[0188] Step 4) Optimize again using the chessboard image pairs corresponding to the three positions and the initial estimates.

[0189] Specifically, the second optimization quantity is optimized using the chessboard image pairs corresponding to the three positions and the initial estimate to obtain the corresponding optimal value.

[0190] The second quantity to be optimized includes:

[0191] A. Pose of the checkerboard coordinate system relative to the robot's base coordinate system ;

[0192] B. The size ratio factor α between the physical dimensions of the checkerboard and its estimated dimensions;

[0193] What needs to be understood is the extrinsic parameters of the binocular camera fixed to the robot. And the pose of the right camera coordinate system relative to the robot base coordinate system at the j-th position out of the three positions. Then the pose of the chessboard coordinate system relative to the right camera coordinate system at position j is... Represented as:

[0194] (27)

[0195] The pose of the checkerboard coordinate system relative to the left camera coordinate system at position j. Represented as:

[0196] (28)

[0197] For the i-th feature point among multiple feature points of the chessboard, as shown in formulas (20), (21) and (25), determine the residual of the feature point at the j-th position in the right camera (denoted as the second residual function in this application), as shown in formulas (22), (23) and (24), and determine the residual of the feature point at the j-th position in the left camera; construct the second objective function based on the residuals of each feature point at each position in the right camera and the residuals of each feature point at each position in the left camera.

[0198] Furthermore, using the initial estimated values ​​of the physical dimensions obtained in step 3) and the initial estimated values ​​of the pose of the checkerboard coordinate system relative to the robot base coordinate system as the initial values ​​for iterative optimization, with the goal of minimizing the value of the second objective function, the LM algorithm can be used to optimize and solve the relevant nonlinear optimization equations, thereby obtaining the optimal values ​​of the second optimization quantity, namely the optimal values ​​of the pose of the checkerboard coordinate system relative to the robot base coordinate system and the optimal values ​​of the size scale coefficient;

[0199] Step 5) Based on the optimal pose of the checkerboard coordinate system relative to the robot base coordinate system obtained in Step 4) and the pose of the positioning pen's visual coordinate system relative to the checkerboard coordinate system, determine the pose of the positioning pen's visual coordinate system relative to the robot base coordinate system.

[0200] As an example, the pose of the positioning pen's visual coordinate system relative to the robot's base coordinate system can be determined according to formula (12), based on the optimal value of the pose of the chessboard coordinate system relative to the robot's base coordinate system obtained in step 4) and the pose of the positioning pen's visual coordinate system relative to the chessboard coordinate system.

[0201] The pose of the positioning pen's visual coordinate system relative to the checkerboard coordinate system can be pre-calibrated using the checkerboard grid before the positioning pen leaves the factory and written into the positioning pen's configuration file.

[0202] Therefore, by using the pose of the positioning pen's visual coordinate system relative to the robot's base coordinate system, the relative relationship between the positioning ratio and the robot, or external parameters, can be calibrated.

[0203] The calibration method presented in this application does not rely on specific environmental conditions or additional calibration boards. Calibration can be completed solely through the checkerboard pattern displayed on the positioning pen's screen, eliminating the need for manual measurement of the checkerboard's actual dimensions. This effectively avoids the problem of insufficient dimensional accuracy in checkerboard measurements in related technologies, simplifies the user's operation process, and reduces reliance on operators' professional skills. Furthermore, addressing the pain point of repeated calibration of positioning pens in robot business scenarios due to issues such as loose bases or equipment replacements, this method exhibits extremely high robustness and reliability. Calibration can be completed quickly after each adjustment, simplifying operational complexity, reducing maintenance time, and ensuring stable accuracy of the robot and positioning pen working together. This reduces daily equipment maintenance costs and is more suitable for the high-frequency adjustment needs in actual production operations, improving the overall feasibility and cost-effectiveness of the solution.

[0204] Furthermore, this method optimizes the target pose through step-by-step prediction and precisely couples multi-coordinate system parameters, simplifying operations while maintaining calibration accuracy. It achieves a dual improvement in "operational convenience" and "accuracy reliability," solving the problem of high dependence on environment, tools, and human intervention in related calibration methods. It can also efficiently meet calibration needs in scenarios involving equipment adjustment and replacement, providing strong support for the long-term stable collaborative operation of robots and positioning pens.

[0205] To implement the above embodiments, this application also proposes a calibration device.

[0206] Figure 4 This is a schematic diagram of the calibration device provided in another embodiment of this application.

[0207] like Figure 4 As shown, the calibration device 400 includes: an acquisition module 410, a first prediction module 420, a second prediction module 430, and a determination module 440.

[0208] The acquisition module 410 is used to acquire multiple image pairs; wherein, the multiple image pairs are obtained by taking multi-view pictures of the chessboard displayed on the screen of the positioning pen by the robot's binocular camera. The binocular camera includes a left eye camera and a right eye camera, and any image pair includes a left eye image and a right eye image.

[0209] The first prediction module 420 is used to predict the value of the target prediction object based on the target image pair among multiple image pairs, and obtain the initial prediction value of the target prediction object; wherein, the target prediction object includes the target pose, which is used to indicate the pose of the chessboard coordinate system relative to the human base coordinate system.

[0210] The second prediction module 430 is used to predict the target pose value based on multiple image pairs and initial prediction values ​​to obtain the target prediction value.

[0211] The determination module 440 is used to determine the pose of the positioning pen's visual coordinate system relative to the human base coordinate system based on the target prediction value, so as to calibrate the robot and the positioning pen.

[0212] Further, in one possible implementation of this application embodiment, the binocular camera includes a right eye camera; the first prediction module 420 is used to: project any one of the multiple feature points of the checkerboard grid onto the image coordinate system corresponding to the target image pair based on the first pose and the size scaling factor, to obtain the first projected coordinates corresponding to the feature point; wherein, the first pose is used to indicate the pose of the checkerboard grid coordinate system relative to the right eye camera coordinate system, and the size scaling factor is used to determine the physical size of the checkerboard grid; construct a first residual function corresponding to the feature point based on the difference between the first projected coordinates corresponding to the feature point and the actual observed coordinates of the feature point in the image coordinate system corresponding to the target image pair; and predict the value of the target prediction object based on the first residual functions corresponding to multiple feature points to obtain an initial prediction value.

[0213] Furthermore, in one possible implementation of this application embodiment, the target prediction object further includes the physical dimensions of the chessboard grid; the first prediction module 420 is used to: construct a first objective function based on the first residual function corresponding to multiple feature points; obtain the initial estimated size of the physical dimensions of the chessboard grid; use the initial estimated size as the initial value for iteration and aim to minimize the value of the first objective function to determine the predicted pose value of the first pose and the predicted scale coefficient of the size scale coefficient; and determine the initial prediction value of the target prediction object based on the predicted pose value of the first pose and the predicted scale coefficient of the size scale coefficient.

[0214] Furthermore, in one possible implementation of this application embodiment, the first prediction module 420 is used to: determine the initial predicted value of the target pose based on the predicted pose value and the pose of the right eye camera coordinate system relative to the human base coordinate system; and determine the initial predicted value of the physical size based on the prediction scale factor and the initial estimated size.

[0215] Furthermore, in one possible implementation of this application embodiment, the target prediction object further includes the physical size of the checkerboard; the second prediction module 430 is configured to: for any image pair among multiple image pairs, based on the target pose and size scaling factor, project any feature point among multiple feature points of the checkerboard onto the image coordinate system corresponding to the image pair, to obtain the second projection coordinates corresponding to the combination between the feature point and the image pair; wherein, the size scaling factor is used to determine the physical size of the checkerboard; construct the second residual function corresponding to the combination based on the difference between the second projection coordinates and the actual observed coordinates of the feature point in the image coordinate system corresponding to the image pair; and predict the value of the target pose based on the second residual function corresponding to each combination to obtain the target prediction value.

[0216] Furthermore, in one possible implementation of this application embodiment, the second prediction module 430 is used to: construct a second objective function based on the second residual function corresponding to each combination; use the initial prediction value of the target prediction object as the initial value for iteration, and take the minimum value of the second objective function as the objective to determine the target prediction value of the target pose.

[0217] Furthermore, in one possible implementation of this application embodiment, the determining module 440 is configured to: read the configuration file of the positioning pen to determine the pose of the positioning pen's visual coordinate system relative to the checkerboard coordinate system; and determine the pose of the positioning pen's visual coordinate system relative to the human-based coordinate system based on the pose of the positioning pen's visual coordinate system relative to the checkerboard coordinate system and the target prediction value of the target pose.

[0218] It should be noted that the foregoing explanation of the calibration method embodiment also applies to the calibration device of this embodiment, and will not be repeated here.

[0219] In summary, the calibration device of this application acquires multiple image pairs; based on the target image pair among the multiple image pairs, it predicts the value of the target prediction object to obtain an initial prediction value of the target prediction object; wherein, the target prediction object includes a target pose, which is used to indicate the pose of the chessboard coordinate system relative to the human base coordinate system, the binocular camera includes a left eye camera and a right eye camera, and any image pair includes a left eye image and a right eye image; based on the multiple image pairs and the initial prediction value, it predicts the value of the target pose to obtain a target prediction value; according to the target prediction value, it determines the pose of the positioning pen's visual coordinate system relative to the human base coordinate system to calibrate the robot and the positioning pen. Therefore, by first making an initial prediction of the target pose in the chessboard coordinate system relative to the human base coordinate system based on the target image, and approximating the true value of the target pose, a secondary prediction is then performed by combining multiple image pairs and the initial prediction value that approximates the true value. This improves the accuracy and reliability of the target pose prediction. Finally, the pose of the positioning pen's visual coordinate system relative to the human base coordinate system is determined based on the target pose prediction value to complete the calibration of the robot and the positioning pen, effectively improving the accuracy of the calibration. At the same time, by adopting a multi-step prediction method, the impact of problems such as blind spots and insufficient feature information in single-view images on the prediction accuracy is effectively reduced, improving the accuracy and stability of the target pose prediction. Moreover, the initial prediction value can provide a precise initial anchor point for the secondary pose calculation, avoiding problems such as slow iterative convergence and local optimum traps caused by the initial value deviating from the true value during the calculation of high-dimensional data from multiple image pairs. By making full use of the rich feature information of multiple image pairs, a balance between calibration accuracy and calculation efficiency is maintained.

[0220] To implement the above embodiments, this application also proposes an electronic device, including: a processor and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the calibration method provided in the foregoing embodiments.

[0221] To implement the above embodiments, this application also proposes a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the calibration method provided in the foregoing embodiments.

[0222] To implement the above embodiments, this application also proposes a computer program product, including a computer program that, when executed by a processor, implements the calibration method provided in the foregoing embodiments.

[0223] The collection, storage, use, processing, transmission, provision, and application of user personal information involved in this application all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0224] It should be noted that personal information collected from users should be used for legitimate and reasonable purposes and should not be shared or sold outside of these legitimate uses. Furthermore, such collection / sharing should only be conducted after receiving the user's informed consent, including but not limited to notifying the user to read the user agreement / user notice and sign an agreement / authorization that includes authorization of relevant user information before the user uses the function. In addition, any necessary steps must be taken to protect and safeguard access to such personal information data and ensure that others with access to personal information data comply with their privacy policies and procedures.

[0225] This application is intended to provide an implementation scheme for users to selectively prevent the use or access to their personal information data. Specifically, this application is intended to provide hardware and / or software to prevent or block access to such personal information data. Once personal information data is no longer needed, risks can be minimized by restricting data collection and deleting data. Furthermore, where applicable, such personal information is de-identified to protect user privacy.

[0226] In the foregoing descriptions of the embodiments, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0227] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as controlling or implying relative importance or implicitly specifying the number of technical features controlled. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0228] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.

[0229] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0230] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0231] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0232] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0233] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

Claims

1. A calibration method, characterized in that, The method includes: Multiple image pairs are acquired; wherein, the multiple image pairs are obtained by taking multi-view photos of the chessboard displayed on the screen of the positioning pen using the robot's binocular camera, the binocular camera including a left camera and a right camera, and any image pair includes a left image and a right image. Based on the target image pair among the multiple image pairs, the value of the target prediction object is predicted to obtain the initial prediction value of the target prediction object; wherein, the target prediction object includes a target pose, which is used to indicate the pose of the chessboard coordinate system relative to the human base coordinate system; Based on the multiple image pairs and the initial prediction value, the target pose value is predicted to obtain the target prediction value; Based on the target prediction value, the pose of the positioning pen's visual coordinate system relative to the human base coordinate system is determined to calibrate the robot and the positioning pen.

2. The method according to claim 1, characterized in that, The step of predicting the value of the target object based on the target image pair among the plurality of image pairs to obtain the initial predicted value of the target object includes: Based on the first pose and the size scaling factor, any one of the multiple feature points of the chessboard is projected into the image coordinate system corresponding to the target image pair to obtain the first projection coordinates corresponding to the feature point; wherein, the first pose is used to indicate the pose of the chessboard coordinate system relative to the right eye camera coordinate system, and the size scaling factor is used to determine the physical size of the chessboard. Based on the difference between the first projected coordinates corresponding to the feature point and the actual observed coordinates of the feature point in the image coordinate system corresponding to the target image pair, a first residual function corresponding to the feature point is constructed. Based on the first residual function corresponding to the plurality of feature points, the value of the target prediction object is predicted to obtain the initial prediction value.

3. The method according to claim 2, characterized in that, The target prediction object also includes the physical dimensions of the chessboard grid; the step of predicting the value of the target prediction object based on the first residual function corresponding to the plurality of feature points to obtain the initial prediction value includes: Based on the first residual function corresponding to the multiple feature points, a first objective function is constructed; Obtain the initial estimated size of the physical dimensions of the chessboard grid; Using the initial estimated size as the initial value for iteration, and with the goal of minimizing the value of the first objective function, the predicted pose value of the first pose and the predicted scale coefficient of the size scale coefficient are determined. The initial predicted value of the target object is determined based on the predicted pose value of the first pose and the predicted scale coefficient of the size scale coefficient.

4. The method according to claim 3, characterized in that, The step of determining the initial predicted value of the target prediction object based on the predicted pose value of the first pose and the predicted scale coefficient of the size scale coefficient includes: Based on the predicted pose value and the pose of the right eye camera coordinate system relative to the human base coordinate system, the initial predicted value of the target pose is determined. Based on the predicted scaling factor and the initial estimated size, the initial predicted value of the physical size is determined.

5. The method according to claim 1, characterized in that, The target prediction object also includes the physical dimensions of the chessboard grid; the prediction of the target pose based on the multiple image pairs and the initial prediction value to obtain the target prediction value includes: For any one of the plurality of image pairs, based on the target pose and size scaling factor, any one of the multiple feature points of the chessboard is projected onto the image coordinate system corresponding to the image pair to obtain the second projection coordinates corresponding to the combination between the feature point and the image pair; wherein, the size scaling factor is used to determine the physical size of the chessboard. Based on the difference between the second projected coordinates and the actual observed coordinates of the feature points in the image coordinate system corresponding to the image pair, a second residual function corresponding to the combination is constructed; Based on the second residual function corresponding to each of the combinations, the value of the target pose is predicted to obtain the target predicted value.

6. The method according to claim 5, characterized in that, The step of predicting the target pose value based on the second residual function corresponding to each of the combinations to obtain the target predicted value includes: Based on the second residual function corresponding to each of the aforementioned combinations, a second objective function is constructed; The initial prediction value of the target object is used as the initial value for iteration, and the target prediction value of the target pose is determined with the goal of minimizing the value of the second objective function.

7. The method according to any one of claims 1-6, characterized in that, Determining the pose of the positioning pen's visual coordinate system relative to the human-based coordinate system based on the target prediction value includes: The configuration file of the positioning pen is read to determine the pose of the positioning pen's visual coordinate system relative to the chessboard coordinate system; The pose of the positioning pen's visual coordinate system relative to the human-based coordinate system is determined based on the pose of the positioning pen's visual coordinate system relative to the chessboard coordinate system and the target prediction value of the target pose.

8. A calibration device, characterized in that, The device includes: An acquisition module is used to acquire multiple image pairs; wherein, the multiple image pairs are obtained by taking multi-view photos of the chessboard displayed on the screen of the positioning pen by the robot's binocular camera, the binocular camera includes a left camera and a right camera, and any image pair includes a left image and a right image. The first prediction module is used to predict the value of the target prediction object based on the target image pair among the plurality of image pairs, and obtain the initial prediction value of the target prediction object; wherein, the target prediction object includes a target pose, and the target pose is used to indicate the pose of the chessboard coordinate system relative to the human base coordinate system; The second prediction module is used to predict the value of the target pose based on the multiple image pairs and the initial prediction value, so as to obtain the target prediction value. The determination module is used to determine the pose of the positioning pen's visual coordinate system relative to the human base coordinate system based on the target prediction value, so as to calibrate the robot and the positioning pen.

9. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Monocular visual error measurement system for cooperative target and error limit quantification method

    CN104729534A

  • Binocular calibration method and device, equipment and storage medium

    CN113298885A