Automatic hand-eye calibration method for humanoid robot
By using autonomous robotic arm movement and automated depth camera calibration methods, the problems of low efficiency and unstable accuracy in traditional hand-eye calibration are solved, achieving efficient and stable automated hand-eye calibration, which is suitable for complex application scenarios of humanoid robots.
Patent Information
- Application Number
- CN202511169955.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-11-14
AI Technical Summary
Traditional hand-eye calibration methods rely on human intervention, which is inefficient and has unstable accuracy, making them difficult to adapt to complex application scenarios of humanoid robots.
By using a robotic arm to move autonomously, a depth camera to automatically acquire images, and a computer vision algorithm to automatically identify the coordinates of feature points and combine them with depth information to generate a 3D point cloud, the robotic arm is controlled to automatically adjust its viewing angle for iterative optimization based on the calibration error evaluation results, thus achieving fully automated hand-eye calibration without human intervention.
It significantly improves calibration efficiency and accuracy stability, reduces the operational threshold and maintenance costs, adapts to complex environments, and supports the large-scale deployment and reliable operation of humanoid robots.
Smart Images

Figure CN120941386A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of robot parameter calibration technology, specifically relating to an automated hand-eye calibration method for humanoid robots. Background Technology
[0002] Traditional hand-eye calibration relies on manual intervention, resulting in low efficiency and unstable accuracy. Existing nine-point calibration methods require manual adjustment of camera pose, making them unsuitable for complex humanoid robot applications. This invention addresses these shortcomings through autonomous robotic arm movement, automatic feature point matching, and a closed-loop optimization mechanism. Summary of the Invention
[0003] The purpose of this invention is to provide an automated hand-eye calibration method for humanoid robots to solve the problems mentioned in the background art.
[0004] To achieve the above objectives, the present invention provides the following technical solution: an automated hand-eye calibration method for a humanoid robot, comprising the following steps: Step 1) The robotic arm drives the depth camera to move along a preset trajectory to multiple viewpoints to ensure that the nine feature points on the calibration board are completely covered within the camera's field of view; Step 2) The depth camera acquires RGB images and depth maps of the calibration board at each viewpoint; Step 3) The two-dimensional image coordinates of the nine feature points are automatically detected using a computer vision algorithm, and three-dimensional point cloud coordinates are generated by combining the depth data; Step 4) Based on the two-dimensional-three-dimensional corresponding point pairs of all nine feature points, the intrinsic and extrinsic parameters of the depth camera and the coordinate transformation relationship between the camera and the robotic arm base are calculated; Step 5) According to the calibration error evaluation results, the robotic arm is controlled to automatically adjust the viewpoint and repeat steps 2)-4), iteratively optimizing the calibration parameters until the accuracy threshold is met.
[0005] Preferably, the calibration plate contains nine high-contrast feature points evenly distributed, and the feature point type is circular marker points or checkerboard corner points.
[0006] Preferably, the preset trajectory planning in step 1) needs to meet the following requirements: the range of motion of the robotic arm covers the spatial position of the calibration plate; the motion path avoids obstacles; and each viewpoint fully contains at least one feature point.
[0007] Preferably, the feature point detection in step 3) is implemented using edge detection, Hough transform, or corner detection algorithms.
[0008] Preferably, step 4) uses the Zhang Zhengyou calibration method or the PnP algorithm to calculate the camera's intrinsic and extrinsic parameters.
[0009] Preferably, the coordinate transformation relationship calculation in step 4) is used to fuse the pose data of the robotic arm joint encoder.
[0010] Preferably, the error assessment in step 5) includes at least one of reprojection error and three-dimensional spatial distance error.
[0011] Preferably, the automatic adjustment of the viewing angle in step 5) is specifically done by: replanning the movement trajectory of the robotic arm according to the error distribution, so that the camera focuses on the feature point area with the largest error.
[0012] Preferably, the depth map acquired in step 2) is aligned with the RGB image through time synchronization or spatial registration.
[0013] Preferably, the calibration error evaluation result output in step 5) includes the camera focal length, principal point coordinates, distortion coefficient, and the rotation matrix and translation vector of the camera relative to the robotic arm base.
[0014] Compared with the prior art, the beneficial effects of the present invention are:
[0015] This invention achieves fully automated, unmanned hand-eye calibration through a closed-loop mechanism: a robotic arm autonomously moves along a preset trajectory; a depth camera automatically acquires multi-view images; a computer vision algorithm automatically identifies feature point coordinates and generates a 3D point cloud based on depth information; and the robotic arm automatically adjusts its viewing angle for iterative optimization based on calibration error evaluation results. This significantly improves calibration efficiency and accuracy stability, avoids errors introduced by manual viewing angle adjustment and feature point identification, and ensures that calibration parameters reliably converge to a preset accuracy threshold. Simultaneously, the robotic arm's ability to autonomously and dynamically adjust its viewing angle effectively guarantees complete coverage and reliable detection of feature points in complex application scenarios and under multi-degree-of-freedom robot postures, enhancing the method's adaptability to complex environments. Ultimately, the fully automated process significantly reduces the operational threshold and maintenance costs, enabling high-precision calibration to be completed efficiently without professional personnel intervention. This effectively reduces time consumption and avoids repetitive work caused by human error, providing strong technical support for the large-scale deployment and reliable operation of humanoid robots. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] Example 1:
[0019] An automated hand-eye calibration method for a humanoid robot includes the following steps: Step 1) A robotic arm drives a depth camera to move along a preset trajectory to multiple viewpoints, ensuring that nine feature points on the calibration board are completely covered within the camera's field of view; Step 2) The depth camera acquires RGB images and depth maps of the calibration board at each viewpoint; Step 3) The two-dimensional image coordinates of the nine feature points are automatically detected using a computer vision algorithm, and three-dimensional point cloud coordinates are generated by combining the depth data; Step 4) Based on the two-dimensional-three-dimensional corresponding point pairs of all nine feature points, the intrinsic and extrinsic parameters of the depth camera and the coordinate transformation relationship between the camera and the robotic arm base are calculated; Step 5) Based on the calibration error evaluation results, the robotic arm is controlled to automatically adjust the viewpoint and repeat steps 2)-4), iteratively optimizing the calibration parameters until the accuracy threshold is met. The calibration board contains nine evenly distributed high-contrast feature points, which are either circular markers or checkerboard corner points. The preset trajectory planning in step 1) must meet the following requirements: the robotic arm's movement range covers the spatial position of the calibration board; the movement path avoids obstacles; and each viewpoint completely contains at least one feature point. The feature point detection in step 3) is implemented using edge detection, Hough transform, or corner detection algorithms. Step 4) uses the Zhang Zhengyou calibration method or the PnP algorithm to calculate the camera's intrinsic and extrinsic parameters. The coordinate transformation relationship in step 4) is used to calculate and fuse the pose data of the robotic arm joint encoder. The error evaluation in step 5) includes at least one of reprojection error and 3D spatial distance error. The specific method for automatically adjusting the viewing angle in step 5) is to re-plan the robotic arm's motion trajectory based on the error distribution, so that the camera focuses on the feature point region with the largest error. The depth map acquired in step 2) is aligned with the RGB image through time synchronization or spatial registration. The calibration error evaluation result output in step 5) includes the camera focal length, principal point coordinates, distortion coefficients, and the camera's rotation matrix and translation vector relative to the robotic arm base.
[0020] Through the above technical solution, this invention achieves fully automated hand-eye calibration without human intervention by employing a closed-loop mechanism: a robotic arm autonomously moves along a preset trajectory; a depth camera automatically acquires multi-view images; a computer vision algorithm automatically identifies feature point coordinates and generates a 3D point cloud by combining depth information; and the robotic arm automatically adjusts its viewing angle for iterative optimization based on calibration error evaluation results. This solution significantly improves calibration efficiency and accuracy stability, avoids errors introduced by manual viewing angle adjustment and feature point identification, and ensures that calibration parameters reliably converge to the preset accuracy threshold. Simultaneously, the robotic arm's ability to autonomously and dynamically adjust its viewing angle effectively guarantees complete coverage and reliable detection of feature points in complex application scenarios and under multi-degree-of-freedom robot postures, enhancing the method's adaptability to complex environments. Ultimately, the fully automated process significantly reduces the operational threshold and maintenance costs, efficiently completing high-precision calibration without professional intervention, effectively reducing time consumption and avoiding repetitive work caused by human error, providing strong technical support for the large-scale deployment and reliable operation of humanoid robots.
[0021] Example 2:
[0022] In this embodiment, the robotic arm moves along a preset spiral trajectory, which ensures that the depth camera uniformly samples multiple viewpoints in three-dimensional space. During the movement, the joint angles and end-effector pose of the robotic arm are fed back to the control system in real time. When the calibration plate is detected to be fully within the camera's field of view, the robotic arm pauses its movement and maintains a stable posture. The depth camera synchronously acquires RGB images and depth information of the calibration plate while stationary, with the image resolution set to 1920×1080 and the depth measurement range covering 0.5-3 meters.
[0023] An improved checkerboard corner detection algorithm is used to process RGB images. First, Gaussian filtering is applied to the image to eliminate noise interference. Then, adaptive thresholding is used to extract the checkerboard edges. The algorithm uses a Harris corner detector to locate the sub-pixel coordinates of nine feature points. Simultaneously, it combines the TOF ranging values of the corresponding positions in the depth map to convert the two-dimensional pixel coordinates into three-dimensional point cloud coordinates based on the camera coordinate system. The three-dimensional coordinates of each feature point are calculated by weighted averaging of the depth values of its neighboring regions, effectively suppressing depth measurement noise.
[0024] An optimization equation incorporating feature point observation data from all viewpoints is established. The Levenberg-Marquardt nonlinear optimization algorithm is used to jointly solve for the camera intrinsic matrix, lens distortion coefficients, camera extrinsic parameters for each viewpoint, and hand-eye transformation matrix. The objective function considers both reprojection error and 3D structural consistency constraints. The reprojection error is calculated as the pixel-level deviation between the observed and predicted feature point positions, while the 3D constraint ensures that the reconstructed spatial positions of feature points remain consistent across different viewpoints.
[0025] The system calculates the root mean square (RMS) value of the reprojection error of the current calibration result in real time. When this value exceeds a preset threshold, it autonomously generates a viewpoint adjustment strategy. Based on error distribution analysis, the strategy determines the areas requiring supplementary observation. The robotic arm automatically plans an obstacle avoidance path and moves to a new observation pose according to the strategy. The data acquisition and parameter optimization process is repeated in the new pose until the reprojection error of all feature points is less than 0.3 pixels and the parameter iteration change is less than 1e-6. The final output calibration parameters include camera focal length, principal point coordinates, radial distortion coefficient, tangential distortion coefficient, and the rigid transformation matrix from the camera coordinate system to the robotic arm base coordinate system.
[0026] The entire calibration process runs in real time on the embedded processor. The robotic arm motion control, image acquisition, feature extraction, and parameter optimization modules interact via shared memory. The system has an anomaly detection function; if no complete feature point is detected from three consecutive views, the process is automatically terminated and a prompt to check the calibration board placement is displayed. The calibration results are output in the form of a homogeneous transformation matrix, which can be directly used for robot vision servo control. This method ensures calibration accuracy through a closed-loop optimization mechanism, reducing calibration time by more than 60% compared to traditional methods, and eliminating the need for manual intervention in viewpoint adjustment and feature point selection.
[0027] Example 3:
[0028] This embodiment includes a calibration board with nine high-contrast feature points for automated calibration. The nine feature points on the calibration board are arranged in a uniform distribution, and the feature point type can be either circular markers or checkerboard corner points. Circular markers are represented by a high-contrast black and white concentric circle pattern, while checkerboard corner points are represented by the corner positions of a black and white checkerboard pattern. Both feature point types have distinct visual characteristics, facilitating accurate identification and localization by computer vision algorithms.
[0029] During calibration, the robotic arm moves the depth camera along a pre-programmed trajectory to multiple different viewpoints. The robotic arm's trajectory is pre-planned to ensure that all nine feature points on the calibration board are fully visible within the camera's field of view at each viewpoint. The depth camera simultaneously acquires RGB color images and depth images of the calibration board at each viewpoint. The RGB images are used to extract the two-dimensional coordinate information of the feature points, while the depth images provide the depth information of the feature points.
[0030] The computer vision algorithm processes the acquired RGB image to automatically detect and locate the two-dimensional coordinates of nine feature points within the image. The algorithm first preprocesses the image, including noise reduction and contrast enhancement, and then uses a feature point detection algorithm to accurately identify the center position of each feature point. For circular markers, the algorithm determines the center coordinates through edge detection and circle fitting; for checkerboard corner points, it locates the corner positions using a corner detection algorithm. Simultaneously, the algorithm combines depth image data to convert the two-dimensional image coordinates of each feature point into three-dimensional spatial coordinates, generating corresponding three-dimensional point cloud data.
[0031] Based on the correspondence between the 2D image coordinates and 3D spatial coordinates of all nine feature points, the system establishes a complete set of 2D-3D corresponding point pairs. Using these corresponding point pairs, the intrinsic and extrinsic parameter matrices of the depth camera, as well as the transformation relationship between the camera coordinate system and the robot arm base coordinate system, are calculated through a camera calibration algorithm. The intrinsic parameter matrix includes parameters such as focal length and principal point coordinates, while the extrinsic parameter matrix describes the camera's pose in the world coordinate system.
[0032] The system automatically assesses the accuracy error of the current calibration result. If the error exceeds a preset threshold, the system controls the robotic arm to automatically adjust the camera angle, re-acquire the calibration board image, and repeat the above processing procedure. This closed-loop optimization mechanism gradually improves the calibration accuracy through multiple iterations until the preset accuracy requirements are met. During each iteration, the robotic arm intelligently selects a new viewing angle position based on the previous calibration error to maximize calibration accuracy.
[0033] This method achieves a fully automated hand-eye calibration process through the autonomous movement of a robotic arm, automatic feature point recognition and matching, and a closed-loop optimization mechanism. Compared to traditional methods, it not only improves calibration efficiency but also significantly enhances calibration accuracy and stability. The nine evenly distributed feature points provide sufficient spatial information to ensure accurate calculation of the camera's intrinsic and extrinsic parameters and coordinate transformation relationships. The high-contrast feature point design enhances the robustness of feature point detection, enabling the system to operate reliably under different lighting conditions. The entire calibration process requires no manual intervention, greatly reducing operational difficulty and maintenance costs.
[0034] Example 4:
[0035] When a robotic arm drives a depth camera to perform a calibration task, a motion trajectory that meets specific requirements must first be planned. The core objective of this trajectory planning is to ensure that the robotic arm can drive the camera to completely cover the spatial area where the calibration board is located, while ensuring that at least one feature point can be completely observed from each viewpoint. To achieve this goal, the system will pre-build a spatial position model of the calibration board and calculate the various observation poses that the robotic arm needs to reach based on this model.
[0036] During trajectory planning, the system generates a series of candidate poses that cover the area where the calibration board is located, based on the robotic arm's kinematic model and reachable workspace. These poses must ensure that the robotic arm's end effector can align the camera with the calibration board and that the camera's field of view can completely contain at least one feature point. The system then uses inverse kinematics calculations to convert these poses into angle values for each joint, forming a complete motion path.
[0037] To ensure safety during movement, the system monitors the distribution of obstacles within the robotic arm's workspace in real time. During trajectory planning, the system automatically avoids potential obstacles by combining 3D map information of the environment. Specifically, this is achieved through collision detection algorithms to verify the safety of the planned path, adjusting the path or adding intermediate transition points to bypass obstacle areas if necessary.
[0038] In terms of viewpoint planning, the system ensures that at least one feature point is fully contained at each observation position. This requires calculating the camera's field of view under different poses based on the distribution pattern of feature points on the calibration board. The system predicts the positional distribution of feature points on the camera's imaging plane through projection transformation, ensuring that at least one feature point can be fully displayed in the image without being occluded by other objects under any observation pose.
[0039] As the robotic arm moves along the planned trajectory, the system monitors the motion status of each joint in real time to ensure that the actual motion path matches the planned path. If the original path becomes infeasible due to environmental changes or other factors, the system immediately initiates a replanning mechanism to recalculate the motion trajectory that satisfies all constraints. This dynamic adjustment capability ensures the reliable execution of the calibration process in various complex environments.
[0040] After completing one round of trajectory movement, the system evaluates the quality of data collected from each viewpoint. If some viewpoints fail to meet the requirements for feature point observation, the system automatically adjusts the trajectory planning parameters, adds necessary observation poses, or optimizes the existing path to ensure that all feature points are fully observed. This closed-loop optimization mechanism significantly improves the completeness and reliability of the calibration data.
[0041] The entire trajectory planning process is fully automated, requiring no manual intervention. The system automatically balances multiple constraints, such as motion range coverage, obstacle avoidance requirements, and feature point observation, through built-in algorithms to generate the optimal motion trajectory. This intelligent planning method not only improves calibration efficiency but also ensures high-quality calibration data in various complex environments.
[0042] Example 5:
[0043] In this embodiment, feature point detection is implemented using an edge detection algorithm. Specifically, the RGB image acquired by the depth camera is first preprocessed, including grayscale conversion and image enhancement, to improve the accuracy of subsequent edge detection. Then, the Canny edge detection algorithm is applied, employing steps such as Gaussian filtering to smooth the image, calculating gradient magnitude and direction, non-maximum suppression, and double threshold detection to extract the edge contours of feature points on the calibration board. After morphological processing, the detected edge contours are selected through contour analysis to identify closed contours that match the geometric characteristics of the feature points, ultimately determining the precise two-dimensional image coordinates of the nine feature points.
[0044] In another implementation, feature point detection is achieved using the Hough transform algorithm. This method is particularly suitable for detecting circular feature points on a calibration board. In the process, the RGB image is first preprocessed and edge information is extracted. Then, the Hough circular transform algorithm is applied to accumulate votes in the parameter space to detect circular features in the image. By setting reasonable radius ranges and voting threshold parameters, the algorithm can accurately identify nine circular feature points on the calibration board and calculate their center coordinates as the two-dimensional image positions of the feature points. The Hough transform algorithm has good robustness to noise and partial occlusion, ensuring the stability of feature point detection in complex environments.
[0045] In another implementation, feature point detection is achieved using a corner detection algorithm. Specifically, either the Harris corner detection algorithm or the Shi-Tomasi corner detection algorithm can be selected. The algorithm first calculates the corner response function value for each pixel in the image, and then determines the location of the feature corners through non-maximum suppression and thresholding. For designs with obvious corner features on the calibration board, this method can quickly and accurately locate the coordinates of nine feature points. Corner detection algorithms are computationally efficient and suitable for applications with high real-time requirements.
[0046] During feature point detection, the system automatically adjusts algorithm parameters based on the quality of the detection results. When the number of detected feature points is insufficient or their distribution does not meet expectations, the algorithm automatically lowers the detection threshold or adjusts the filtering parameters to ensure that all nine feature points are reliably detected. Simultaneously, the system correlates the 2D image coordinates with the corresponding depth values in the depth map, generating 3D point cloud coordinates for the feature points through coordinate transformation, providing an accurate data foundation for subsequent camera parameter calibration.
[0047] To improve the robustness of feature point detection, the system also employs a multi-algorithm fusion strategy. Based on the main detection algorithm, other algorithms are used for verification and correction. For example, in a method primarily based on Hough transform, an edge detection algorithm is run simultaneously to verify the accuracy of the detection results. When there is a significant discrepancy between the detection results of the two algorithms, the system automatically selects the more reliable result or triggers a re-detection mechanism to ensure the accuracy of the feature point coordinate data.
[0048] After feature point detection is completed, the system performs a geometric consistency check on the detection results. This is done by calculating the relative positions of the nine feature points and comparing them with the known geometric parameters of the calibration board to verify the reasonableness of the detection results. If the detection results do not meet the geometric constraints, the system will automatically re-execute feature point detection or adjust the camera pose and re-acquire images to ensure the accuracy and reliability of the data input for subsequent calibration calculations.
[0049] Example 6:
[0050] Based on Example 1, this example uses the Zhang Zhengyou calibration method to calculate the camera's intrinsic and extrinsic parameters. Specifically, firstly, based on the two-dimensional image coordinates of the nine feature points obtained in step 3) and their corresponding three-dimensional point cloud coordinates, a two-dimensional-three-dimensional corresponding point pair dataset is constructed. The core principle of the Zhang Zhengyou calibration method is to solve for the camera projection matrix to decompose the camera's intrinsic and extrinsic parameter matrices. The intrinsic parameter matrix includes the focal length, principal point coordinates, and distortion coefficients, while the extrinsic parameter matrix describes the rotation and translation relationship of the camera coordinate system relative to the robot arm's base coordinate system.
[0051] In the calculation process, the initial projection matrix is first solved using the least squares method, which maps 3D spatial points to a 2D image plane. Then, singular value decomposition (SVD) is used to decompose the projection matrix, obtaining initial estimates of the camera's intrinsic and extrinsic parameter matrices. To improve calibration accuracy, a nonlinear optimization algorithm (such as the Levenberg-Marquardt algorithm) is further employed to iteratively optimize the intrinsic and extrinsic parameters, minimizing the reprojection error. That is, the optimized parameters should ensure that the coordinates of the 3D points reprojected onto the image plane are as close as possible to the coordinates of the originally detected 2D feature points.
[0052] In another implementation, step 4) uses the Perspective-n-Point (PnP) algorithm to calculate the camera's extrinsic parameters. The core principle of the PnP algorithm is to use known camera intrinsic parameters and multiple 2D-3D corresponding point pairs to solve for the pose of the camera coordinate system relative to the world coordinate system (i.e., the calibration board coordinate system). In specific implementation, based on the 2D-3D correspondence of the nine feature points obtained in step 3), combined with the pre-calibrated camera intrinsic parameters, efficient solving algorithms such as EPnP or UPnP are used to calculate the camera's extrinsic parameter matrix.
[0053] The advantage of the PnP algorithm lies in its high computational efficiency, making it particularly suitable for scenarios with high real-time requirements. In the calculation process, the 3D points are first represented as a weighted sum of control points, and an initial estimate of the camera pose is obtained by solving a system of linear equations. Subsequently, the initial estimate is nonlinearly optimized using the Gauss-Newton method or the Levenberg-Marquardt algorithm to further improve the accuracy of the extrinsic parameters. The optimization objective is to minimize the reprojection error, ensuring that the calculated extrinsic parameters accurately reflect the spatial relationship between the camera and the calibration board.
[0054] Regardless of whether Zhang Zhengyou's calibration method or the PnP algorithm is used, step 4) will output the camera's intrinsic and extrinsic parameter matrices. The intrinsic parameter matrix describes the camera's optical characteristics, including focal length, principal point coordinates, and distortion coefficients; the extrinsic parameter matrix describes the rotation and translation relationship of the camera coordinate system relative to the robot arm's base coordinate system. These parameters will be used for subsequent robot hand-eye calibration, i.e., calculating the coordinate transformation relationship between the camera and the robot arm's end effector.
[0055] In step 5), the system automatically adjusts the robotic arm pose based on the calibration error evaluation results, re-collects data, and iteratively optimizes the calibration parameters. This closed-loop optimization mechanism ensures the convergence and stability of the calibration process, ultimately enabling the calibration accuracy to meet the preset threshold requirements. Through automated adjustment and optimization, this method significantly improves calibration efficiency and accuracy, making it suitable for high-precision hand-eye coordination tasks for humanoid robots in complex environments.
[0056] Example 7:
[0057] Based on Embodiment 1, in step 4), the coordinate transformation relationship is used to calculate the pose data of the fused robotic arm joint encoder, which is specifically implemented as follows:
[0058] During the robot arm's movement, its joint encoders record the angle changes of each joint in real time and calculate the pose data of the robot arm's end effector using a forward kinematics model. This pose data includes the position and orientation information of the robot arm's end effector in the base coordinate system, which is used to assist in calculating the coordinate transformation relationship between the depth camera and the robot arm's base. During calibration, the 3D coordinates of the calibration board feature points acquired by the depth camera are based on measurements in the camera coordinate system, while the robot arm's motion trajectory data is based on the robot arm's base coordinate system. To establish the transformation relationship between the camera coordinate system and the robot arm's base coordinate system, the pose data of the robot arm's end effector and the feature point data acquired by the camera need to be jointly optimized.
[0059] Specifically, in step 4), the computer vision algorithm first generates three-dimensional point cloud coordinates based on the two-dimensional image coordinates and depth data of nine feature points. These point cloud coordinates are represented in the camera coordinate system. Simultaneously, the pose data recorded by the robotic arm joint encoder is used to calculate the pose matrix of the robotic arm end effector in the base coordinate system through forward kinematics. Since the depth camera is fixedly mounted at the end of the robotic arm, there is a fixed installation offset relationship between it and the robotic arm end effector. Therefore, the transformation matrix between the camera coordinate system and the robotic arm base coordinate system can be solved using an optimization algorithm.
[0060] This optimization process employs the least squares method or nonlinear optimization to match the pose data of the robotic arm's end effector with the 3D coordinates of feature points acquired by the camera, calculating the optimal coordinate transformation relationship. Specifically, the optimization objective is to minimize the error between the 3D coordinates of the feature points in the camera coordinate system and the theoretical projection points in the robotic arm base coordinate system. Through iterative optimization, a high-precision camera extrinsic parameter matrix is ultimately obtained, representing the pose relationship of the camera relative to the robotic arm base.
[0061] Furthermore, the pose data from the robotic arm joint encoders possesses high real-time performance and stability, effectively compensating for camera measurement errors and improving calibration accuracy. Particularly during robotic arm movement, the joint encoder data can be used to dynamically correct camera pose estimation, avoiding the accumulation of calibration errors caused by camera measurement noise. Ultimately, by fusing the pose data from the robotic arm joint encoders, the calibration system can achieve more robust coordinate transformation calculations, ensuring the reliability and stability of calibration results in complex application scenarios.
[0062] This embodiment makes full use of the robotic arm's own sensor data and combines it with computer vision algorithms to achieve high-precision automated hand-eye calibration, avoiding the limitations of traditional methods that rely on manual adjustment and manual measurement, and significantly improving calibration efficiency and accuracy.
[0063] Example 8:
[0064] Based on Embodiment 1, in step 5), the error assessment uses a combination of reprojection error and 3D spatial distance error for comprehensive evaluation. The calculation process for reprojection error is as follows: the 3D point cloud coordinates of the feature points generated in step 3) are reprojected onto the 2D image plane using the camera intrinsic and extrinsic parameters obtained in step 4), resulting in reprojected 2D coordinates; then, the reprojected coordinates are compared with the original 2D image coordinates directly detected in step 3), and the Euclidean distance between the two is calculated as the reprojection error value. This error reflects the consistency of the projection of the calibration parameters onto the image plane and can intuitively assess the accuracy of the camera intrinsic parameters.
[0065] The calculation process for the 3D spatial distance error is as follows: Using the coordinate transformation relationship between the camera and the robotic arm base obtained in step 4), the 3D coordinates of the feature points collected from each viewpoint are uniformly transformed to the coordinate system of the robotic arm base. Since the calibration plate maintains a fixed position during the movement of the robotic arm, theoretically, the 3D coordinates of the same feature point under different viewpoints should coincide in the base coordinate system. In actual calculation, the average value of the transformed coordinates of the same feature point under all viewpoints is taken as the reference value, and then the Euclidean distance between the coordinates of each viewpoint and the reference value is calculated as the 3D spatial distance error. This error directly reflects the positioning accuracy of the hand-eye calibration in 3D space and can effectively evaluate the accuracy of the coordinate transformation relationship.
[0066] During error assessment, the system monitors the trends of both types of errors in real time. When the reprojection error continuously decreases while the 3D spatial distance error fluctuates significantly, it indicates that the camera intrinsic parameter calibration is effective but there is a deviation in the hand-eye transformation relationship. In this case, the control robot arm focuses on adjusting the relative pose of the camera and the end effector. When the 3D spatial distance error stabilizes and converges while the reprojection error remains high, it indicates that the camera intrinsic parameters need optimization, and the system will automatically increase the sampling angles for different tilt angles of the chessboard plane. Through this error correlation analysis mechanism, the system can intelligently identify the main sources of error in the calibration process and adjust the optimization strategy accordingly.
[0067] During the iterative optimization phase, the system establishes an error weighting model, dynamically adjusting the weight ratio of reprojection error and 3D spatial distance error in the total error evaluation based on the current error distribution. When the initial calibration error is large, emphasis is placed on 3D spatial distance error to quickly establish the correct hand-eye transformation framework; as the calibration accuracy improves, the weight of reprojection error is gradually increased to achieve sub-pixel-level fine adjustment. The system sets dual convergence conditions: requiring the mean reprojection error to be less than 0.3 pixels and the mean 3D spatial distance error to be less than 1 mm, and simultaneously requiring that the error decrease does not exceed 5% in three consecutive iterations, at which point the calibration is considered converged.
[0068] To improve the reliability of the evaluation, the system employs a cross-validation mechanism: 20% of the feature point data is reserved and not used in the calibration calculation, specifically for verifying the generalization performance of the calibration results. When the difference between the validation set error and the training set error exceeds 15%, a data re-acquisition process is automatically triggered to ensure that the calibration results have sufficient generalization ability. The entire error evaluation process is fully automated, requiring no manual intervention to complete error analysis, problem diagnosis, and optimization strategy formulation, ultimately outputting calibration parameters that meet the preset accuracy threshold.
[0069] Example 9:
[0070] In step 5) of the automated hand-eye calibration method for humanoid robots, the specific implementation process of automatically adjusting the viewing angle is as follows: After the system completes the initial calibration parameter calculation, the reprojection error of each feature point is statistically analyzed to establish an error distribution model. This model quantifies the calibration accuracy difference of each feature point region by calculating the Euclidean distance between the observed coordinates of each feature point in the image coordinate system and the theoretical coordinates obtained by backprojection based on the current calibration parameters.
[0071] The system then identifies the feature point region with the largest error, which typically corresponds to the edge of the camera's field of view or a location where depth measurement is unstable. Based on the robotic arm's kinematic model and the camera's imaging geometry, the planning module generates a new robotic arm trajectory, causing the depth camera's optical center to move in a directional direction towards this feature point region. Trajectory planning must ensure smooth and continuous movement of all robotic arm joints while maintaining other feature points within the camera's effective field of view.
[0072] As the robotic arm executes a new trajectory, a real-time monitoring system continuously tracks the visibility of feature points. When the camera reaches the predetermined viewing angle, the control system fine-tunes the robotic arm's end effector posture to ensure the target feature point region is precisely located in the center of the camera's focal plane. This active focusing mechanism fully utilizes the depth-of-field characteristics of the depth camera, ensuring the target area receives the highest resolution image sampling.
[0073] After the viewing angle adjustment is completed, the system automatically triggers a new round of image acquisition and feature point detection. Unlike the initial calibration, this acquisition will perform multi-frame oversampling in high-error areas and improve the measurement accuracy of feature point coordinates through temporal filtering. The newly acquired observation data will be fused with the original data to jointly participate in the iterative optimization calculation of calibration parameters.
[0074] This process forms a closed-loop feedback control system, the core of which lies in establishing a mapping relationship between error distribution and robotic arm motion strategy. Each iteration not only optimizes the calibration parameters themselves but also dynamically optimizes the data acquisition strategy. By recording historical adjustment trajectories, the system gradually constructs a precision distribution map of the calibration board in the robot's workspace, providing prior knowledge for subsequent calibrations.
[0075] When the change in the maximum reprojection error is less than a preset threshold in three consecutive iterations, the system determines that the calibration parameters have converged and terminates the automatic adjustment process. The final output calibration parameter matrix integrates observation data from multiple perspectives and working conditions, maintaining uniform high precision across the robot's entire workspace. This active perspective adjustment method based on error feedback effectively solves the problem of local optima in parameters caused by a fixed perspective in traditional calibration.
[0076] Example 10:
[0077] In the implementation of automated hand-eye calibration methods for humanoid robots, the alignment of depth maps with RGB images is a crucial step in ensuring calibration accuracy. This embodiment details the specific implementation process of image alignment through two methods: time synchronization and spatial registration.
[0078] In the time synchronization alignment scheme, the system employs a hardware triggering mechanism to achieve synchronized acquisition of data from the depth sensor and the RGB camera. When the robotic arm moves to the preset calibration pose, the control unit simultaneously sends trigger signals to both the depth sensor and the RGB camera, ensuring that the two image data streams are captured at the same time. The trigger signals are synchronized via a precise hardware clock, with time errors controlled within microseconds. After acquisition, the system automatically verifies the consistency of the timestamps of the two images. If a time deviation exceeds a preset threshold, a re-acquisition process is automatically triggered. This scheme is particularly suitable for scenarios where the robotic arm performs dynamic calibration during movement, effectively eliminating image misalignment caused by robotic arm motion.
[0079] In the spatial registration and alignment scheme, the system first independently acquires depth maps and RGB images, and then achieves spatial alignment of the two images through a feature matching algorithm. Specifically, the system uses nine feature points on a calibration board as registration references, detecting the positions of these feature points in both the depth map and the RGB image. By calculating the spatial transformation relationship between the feature points, a mapping matrix of pixel coordinates between the two images is established. To improve registration accuracy, the system employs a multi-scale feature detection algorithm, first quickly locating the approximate region of feature points on the low-resolution image, and then performing sub-pixel-level precise localization on the high-resolution image. The RANSAC algorithm is also introduced during the registration process to eliminate mismatched points, ensuring the robustness of the transformation matrix. The final generated mapping matrix will be used to accurately project the 3D point cloud data from the depth map onto the corresponding pixel positions in the RGB image.
[0080] For dynamic calibration scenarios, the system can combine two alignment methods to achieve optimal results. During the robotic arm's movement, time-synchronized data acquisition ensures basic alignment accuracy, while spatial registration performs fine-tuning during static pose calibration. The system automatically evaluates the alignment quality of each frame, and when the detected feature point projection error exceeds a threshold, spatial registration is prioritized for correction. This hybrid alignment strategy ensures both the real-time performance of the calibration process and the accuracy of the final calibration parameters.
[0081] During implementation, the system also established a quantitative evaluation mechanism for image alignment quality. The alignment effect is monitored in real time by calculating the reprojection error of feature points in two images. When the alignment error continues to increase, the system automatically adjusts the acquisition parameters or triggers a recalibration process. This evaluation mechanism, together with the calibration parameter optimization process, forms a closed-loop control, ensuring that the final hand-eye transformation matrix has optimal accuracy and stability.
[0082] This embodiment effectively solves the challenge of multimodal image alignment through the aforementioned technical solution, providing a reliable data foundation for subsequent calculation of feature point 3D coordinates. The system can flexibly select the most suitable image alignment method according to the needs of actual application scenarios, ensuring stable calibration results in different working environments. This adaptive alignment capability significantly improves the calibration success rate and accuracy of humanoid robots in complex application scenarios.
[0083] Example 11:
[0084] Based on Example 1, the specific implementation process of outputting the calibration error evaluation result in step 5) including camera focal length, principal point coordinates, distortion coefficients, and the rotation matrix and translation vector of the camera relative to the robotic arm base is as follows:
[0085] After completing the calibration parameter calculation in step 4), the system automatically enters the calibration error evaluation stage. First, based on the currently calculated camera intrinsic parameter matrix, the system extracts the camera's focal length parameter. This parameter includes components in the x and y axes, corresponding to the ratio of the physical size of the image sensor in the horizontal and vertical directions to the pixel size, respectively. The focal length parameter directly reflects the camera's imaging characteristics and is a fundamental parameter for subsequent 3D reconstruction and coordinate transformation.
[0086] The principal point coordinates are evaluated by analyzing the origin offset between the image coordinate system and the camera coordinate system. The system calculates the deviation between the image center point and the actual optical center point in the current calibration result; this deviation value is the camera principal point coordinate. The accuracy of the principal point coordinates directly affects the conversion accuracy of feature point 2D coordinates to 3D coordinates and is one of the important indicators for evaluating calibration quality.
[0087] For the evaluation of distortion coefficients, the system employs a reprojection error analysis method. This involves reprojecting the 3D coordinates of the calibration board feature points onto the image plane and calculating the pixel distance between the projected points and the actually detected feature points. This process evaluates the impact of radial and tangential distortion separately and outputs the corresponding distortion coefficients. The accuracy of the distortion coefficients directly determines the effectiveness of image correction and has a significant impact on the accuracy of subsequent visual measurements.
[0088] The rotation matrix of the camera relative to the robot arm base is evaluated by comparing the difference between the theoretical coordinate system transformation and the actual calculation results. The system compares the forward kinematics solution of the robot arm's current pose with the calibrated camera extrinsic parameter matrix, calculating the rotation angle deviation between the two. The accuracy of this rotation matrix determines the orientation relationship between the camera coordinate system and the robot's base coordinate system, and is a key parameter for hand-eye coordination control.
[0089] The translation vector is evaluated through spatial distance measurement. The system calculates the position of the calibrated camera optical center in the robot arm's base coordinate system and compares it with the actual position of the robot arm's end effector. This translation vector reflects the precise positional relationship between the camera's optical center and the robot arm's end flange, and its accuracy directly affects the spatial positioning accuracy of the robot's grasping and manipulation.
[0090] After evaluating the above parameters, the system calculates the overall calibration error. This error includes multiple dimensions such as reprojection error, coordinate system alignment error, and parameter consistency error. When the evaluation result exceeds a preset threshold, the system automatically generates a new camera pose adjustment strategy, controls the robotic arm to move to a better observation position, re-collects data, and performs iterative optimization. The entire evaluation and optimization process is fully automated, requiring no manual intervention, until all output parameters meet the preset accuracy requirements.
[0091] This embodiment ensures the reliability and accuracy of calibration parameters through an automated error evaluation mechanism. The system not only outputs the final calibration result but also records the parameter change trends and error convergence status for each iteration, providing comprehensive data support for subsequent calibration quality analysis and system performance evaluation. This closed-loop evaluation and optimization mechanism significantly improves the stability and reliability of the calibration process, meeting the high-precision calibration requirements of humanoid robots in complex application scenarios.
[0092] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0093] The above description is only used to illustrate the technical solution of the present invention and is not intended to limit it. Any other modifications or equivalent substitutions made by those skilled in the art to the technical solution of the present invention, as long as they do not depart from the spirit and scope of the technical solution of the present invention, should be covered within the scope of the claims of the present invention.
Claims
1. A method for automated hand-eye calibration of a humanoid robot, characterized in that, Includes the following steps: Step 1) The robotic arm drives the depth camera to move to multiple viewpoints along a preset trajectory to ensure that the nine feature points on the calibration board are completely covered within the camera's field of view; Step 2) The depth camera acquires RGB images and depth maps of the calibration board from various viewpoints; Step 3) Automatically detect the two-dimensional image coordinates of nine feature points using computer vision algorithms, and generate three-dimensional point cloud coordinates by combining them with depth data; Step 4) Based on the two-dimensional-three-dimensional corresponding point pairs of all nine feature points, calculate the intrinsic and extrinsic parameters of the depth camera and the coordinate transformation relationship between the camera and the robot arm base; Step 5) Based on the calibration error evaluation results, control the robotic arm to automatically adjust the viewing angle and repeat steps 2)-4) to iterate and optimize the calibration parameters until the accuracy threshold is met.
2. The method according to claim 1, characterized in that, The calibration board contains nine high-contrast feature points that are evenly distributed, and the feature point types are circular marker points or checkerboard corner points.
3. The method according to claim 1, characterized in that, The preset trajectory planning in step 1) must meet the following requirements: the range of motion of the robotic arm must cover the spatial position of the calibration plate; The movement path avoids obstacles; Each viewpoint must contain at least one complete feature point.
4. The automated hand-eye calibration method for a humanoid robot according to claim 1, characterized in that, The feature point detection in step 3) is implemented using edge detection, Hough transform, or corner detection algorithms.
5. The automated hand-eye calibration method for a humanoid robot according to claim 1, characterized in that, Step 4) uses the Zhang Zhengyou calibration method or the PnP algorithm to calculate the camera's intrinsic and extrinsic parameters.
6. The automated hand-eye calibration method for a humanoid robot according to claim 1, characterized in that, The coordinate transformation relationship calculation in step 4) is used to fuse the pose data of the robotic arm joint encoder.
7. The automated hand-eye calibration method for a humanoid robot according to claim 1, characterized in that, The error assessment in step 5) includes at least one of reprojection error and three-dimensional spatial distance error.
8. The automated hand-eye calibration method for a humanoid robot according to claim 1, characterized in that, The specific method for automatically adjusting the viewing angle in step 5) is as follows: replan the movement trajectory of the robotic arm according to the error distribution, so that the camera focuses on the feature point area with the largest error.
9. The automated hand-eye calibration method for a humanoid robot according to claim 1, characterized in that, The depth map acquired in step 2) is aligned with the RGB image through time synchronization or spatial registration.
10. The automated hand-eye calibration method for a humanoid robot according to claim 1, characterized in that, The calibration error evaluation result output in step 5) includes the camera focal length, principal point coordinates, distortion coefficient, rotation matrix and translation vector of the camera relative to the robot arm base.