Hand-eye calibration method and device, medium, computer device and program product
By filtering and optimizing the hand-eye conversion relationship and utilizing the positioning point information in the target image, the problem of inaccurate camera pose positioning in complex environments was solved, thereby improving the accuracy of camera positioning and the robustness of the system.
Patent Information
- Application Number
- CN202411911068.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2044-12-23
AI Technical Summary
In existing technologies, the pose positioning accuracy of cameras is relatively low, mainly due to the low precision of hand-eye transition caused by complex shooting environments and lighting factors.
By acquiring multiple samples, a subset of samples is randomly selected to determine the initial hand-eye conversion relationship and its quality. The initial hand-eye conversion relationship with the highest quality is selected, and the hand-eye conversion relationship is optimized based on the location information of the localization points in the target image. Localization points affected by noise are filtered out to improve the accuracy of the hand-eye conversion relationship.
It improves the accuracy of camera positioning, enhances the robustness and applicability of the system, reduces the impact of noise on positioning accuracy, and is suitable for complex lighting environments.
Smart Images

Figure CN120070589B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of hand-eye calibration technology, and in particular to a hand-eye calibration method and apparatus, medium, computer equipment and program product. Background Technology
[0002] When using a camera to photograph a target, it is often necessary to determine the camera's pose. Since most cameras lack self-localization capabilities, the common practice is to fix the camera to a positionable rigid body. By positioning this rigid body and using a hand-eye calibration algorithm, the hand-eye translation relationship between the rigid body and the camera is calculated, thus indirectly obtaining the camera's pose. However, the actual shooting environment, including scene and lighting conditions, can be quite complex, leading to lower accuracy in the obtained hand-eye translation relationship and consequently, lower camera positioning accuracy. Summary of the Invention
[0003] In a first aspect, this application provides a hand-eye calibration method for calibrating the hand-eye transformation relationship between a camera and a rigid body fixedly connected to the camera, wherein multiple positioning points are distributed around the rigid body; the method includes: acquiring multiple samples, any one of which includes a camera pose transformation matrix and a rigid body pose transformation matrix, the camera pose transformation matrix representing the relative pose relationship of the camera between two moments, and the rigid body pose transformation matrix representing the relative pose relationship of the rigid body between the two moments; the rigid body pose transformation matrix being determined based on the position information of the multiple positioning points; and determining multiple initial hand-eye transformation relationships and their quality, wherein any initial hand-eye transformation relationship and its quality... The quality is determined as follows: a subset of samples are randomly selected from the plurality of samples; an initial hand-eye conversion relationship is determined based on each selected sample; and the quality of the initial hand-eye conversion relationship is determined based on the matching degree between each selected sample and the initial hand-eye conversion relationship. The initial hand-eye conversion relationship with the highest quality is determined, and each sample used to determine the initial hand-eye conversion relationship with the highest quality is determined as a target sample. The target image corresponding to each target sample is acquired, and the optimized hand-eye conversion relationship is determined based on the position information of the positioning points in the acquired target image. The target image corresponding to the target sample includes the images of the plurality of positioning points acquired at the target time corresponding to the target sample.
[0004] Secondly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the methods described in any embodiment of this application.
[0005] Thirdly, embodiments of this application provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described in any embodiment of this application.
[0006] In this embodiment, after acquiring multiple samples including a camera pose transformation matrix and a rigid body pose transformation matrix, a subset of samples are first selected to determine the initial hand-eye transformation relationship and its quality. Then, the highest quality initial hand-eye transformation relationship is selected from the multiple initial hand-eye transformation relationships. A higher quality initial hand-eye transformation relationship indicates a higher fit between the initial hand-eye transformation relationship and the selected samples, thus indirectly reflecting the high accuracy of the target sample used to determine the initial hand-eye transformation relationship. Since the hand-eye transformation relationship is determined based on the positional information of the localization points in the target image corresponding to the target sample, a higher accuracy of the target sample indicates a higher accuracy of the positional information of the localization points in the target image corresponding to the target sample. Therefore, the optimized hand-eye transformation relationship can be determined based on the positional information of the localization points in the target image. This filters out the positional information of localization points that are significantly affected by noise, and only uses the positional information of highly accurate localization points to determine the hand-eye transformation relationship, thereby improving the accuracy of the determined hand-eye transformation relationship and consequently improving the camera positioning accuracy.
[0007] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application. Attached Figure Description
[0008] The accompanying drawings, which are incorporated in and constitute a part of this application, illustrate embodiments consistent with this application and, together with the description, serve to illustrate the technical solutions of this application.
[0009] Figure 1 This is a schematic diagram of a virtual shooting system according to an embodiment of this application.
[0010] Figure 2 This is a flowchart of the hand-eye calibration method according to an embodiment of this application.
[0011] Figure 3 This is a schematic diagram illustrating the transformation relationship between multiple coordinate systems in an embodiment of this application.
[0012] Figure 4 This is a general flowchart of an embodiment of this application.
[0013] Figure 5 This is a block diagram of a hand-eye calibration device according to an embodiment of this application.
[0014] Figure 6 This is a schematic diagram of a computer device according to an embodiment of this application. Detailed Implementation
[0015] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0016] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items. Additionally, the term “at least one” herein means any combination of at least two of any one or more of a plurality.
[0017] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0018] To enable those skilled in the art to better understand the technical solutions in the embodiments of this application, and to make the above-mentioned objectives, features and advantages of the embodiments of this application more apparent and understandable, the technical solutions in the embodiments of this application will be further described in detail below with reference to the accompanying drawings.
[0019] When using a camera to photograph a target, it is often necessary to determine the camera's pose. In some embodiments, the shooting scene is a virtual shooting scene, and the camera mentioned above is a camera in a virtual shooting system. Figure 1A schematic diagram of a virtual shooting system 20 according to some embodiments is shown. The virtual shooting system 20 includes a camera 202, a display screen 204, and a processing device 206, wherein the display screen 204 is used to display virtual images. The shooting target (such as an actor) can be located in front of the display screen 204. The camera 202 is used to capture images of the shooting target and the virtual images displayed on the display screen 204 to obtain a captured image. The virtual image can be an image within the field of view of the virtual camera. The virtual camera is a mathematical model of the camera 202, used to capture images in a virtual scene as the background for virtual shooting. The pose of the virtual camera can be consistent with the pose of the camera 202. By accurately determining the pose of the camera 202, the pose of the virtual camera can be further determined, thereby rendering a virtual image corresponding to the viewpoint of the virtual camera.
[0020] like Figure 1 As shown, there are three display screens 204. For easy distinction, each display screen is referred to as display screen 2042, display screen 2044, and display screen 2046. The display screens (2042, 2044, and 2046) can display virtual images independently or in conjunction with each other. The display screens 204 used in the virtual shooting system 20 can be physical screens, such as LED displays or LCD displays, and can have curved or flat screen structures. It should be understood that those skilled in the art can customize the type, number, size, resolution, and installation location of the physical screens in the virtual shooting system 20 according to actual needs, and this embodiment does not limit this.
[0021] The processing device 206 is used to control the process of displaying virtual images on the display screen 204, such as controlling the playback frame rate of the virtual images, the parameters of the virtual camera 202, and the switching of virtual images.
[0022] Camera 202 can be fixedly connected to rigid body 208. Camera 202 can be a camera without self-localization function. By positioning rigid body 208 and calculating the hand-eye conversion relationship between rigid body 208 and camera 202 (i.e. the relative pose relationship between rigid body 208 and camera 202) through hand-eye calibration algorithm, the pose of camera 202 can be indirectly obtained.
[0023] In actual shooting scenarios, the shooting environment and lighting may be quite complex, resulting in lower accuracy of the hand-eye conversion relationship and thus lower positioning accuracy of the camera 202.
[0024] Based on this, this application provides a hand-eye calibration method for calibrating the hand-eye conversion relationship between a camera 202 and a rigid body 208 fixedly connected to the camera 202, wherein multiple positioning points are distributed around the rigid body 208; see [link to relevant documentation]. Figure 2 The method includes:
[0025] Step S12: Obtain multiple samples. Each sample includes a camera pose transformation matrix and a rigid body pose transformation matrix. The camera pose transformation matrix is used to represent the relative pose relationship of the camera 202 between two time points, and the rigid body pose transformation matrix is used to represent the relative pose relationship of the rigid body 208 between the two time points. The rigid body pose transformation matrix is determined based on the position information of multiple positioning points.
[0026] Step S14: Determine multiple initial hand-eye conversion relationships and their quality, wherein any initial hand-eye conversion relationship and its quality are determined based on the following method: randomly select a portion of samples from multiple samples, determine the initial hand-eye conversion relationship based on each of the selected samples, and determine the quality of the initial hand-eye conversion relationship based on the matching degree between each of the selected samples and the initial hand-eye conversion relationship.
[0027] Step S16: Determine the highest quality initial hand-eye conversion relationship, and designate each sample used to determine the highest quality initial hand-eye conversion relationship as the target sample;
[0028] Step S18: Obtain the target image corresponding to each target sample, and determine the optimized hand-eye conversion relationship based on the position information of the positioning points in the obtained target image; wherein, the target image corresponding to the target sample includes the images of multiple positioning points collected at the target time corresponding to the target sample.
[0029] This application first selects a subset of samples from multiple samples to determine the initial hand-eye translation relationship and its quality. Then, it selects the highest-quality initial hand-eye translation relationship from among these initial relationships. The quality of the initial hand-eye translation relationship indirectly reflects the accuracy of the target sample used to determine it, and thus the accuracy of the location information of the positioning points. Therefore, the target sample and its corresponding target image can be determined based on the highest-quality initial hand-eye translation relationship, and then the optimized hand-eye translation relationship can be determined based on the location information of the positioning points in the target image. In this way, the location information of positioning points that are significantly affected by noise can be filtered out, and only the location information of positioning points with high accuracy can be used to determine the hand-eye translation relationship, thereby improving the accuracy of the determined hand-eye translation relationship and thus improving the camera positioning accuracy.
[0030] The specific implementation details of this application are illustrated below with reference to the accompanying drawings.
[0031] The rigid body 208 in this application refers to an object whose internal parts maintain their relative positions under the action of external forces. The rigid body 208 can be an object such as a robotic arm, camera bracket, or robot. Multiple positioning points are distributed around the rigid body 208; these positioning points are feature points that can be used to track the rigid body 208. In some embodiments, the multiple positioning points include, but are not limited to, at least one of the following:
[0032] Built-in positioning points on rigid body 208. Built-in positioning points refer to fixed or internal positioning points inherent to rigid body 208 itself. These points are usually determined during the object design and are part of the rigid body 208 structure. These positioning points can be special marks, sensor interfaces, positioning holes, laser points, etc., fixed to the surface of rigid body 208. Built-in positioning points are typically part of the rigid body 208 design, requiring no additional installation and resulting in low deployment costs.
[0033] Positioning points bound to rigid body 208. Positioning points bound to rigid body 208 typically refer to those where positioning devices, markers, or sensors are attached to the surface of rigid body 208 through external methods (such as pasting, fixing, or installation). Common forms include reflective markers, inertial sensors, optical markers, and QR codes. Users can choose appropriate locations to bind positioning points to rigid body 208 according to their needs, and the positions of the positioning points can be flexibly arranged according to motion requirements. Positioning points bound externally do not change the basic structure and function of rigid body 208 and can usually be installed without interfering with the main functions of rigid body 208. Furthermore, when the number of positioning points built into rigid body 208 is insufficient, binding additional positioning points to rigid body 208 can increase the number of positioning points, thereby improving the positioning accuracy of rigid body 208.
[0034] Positioning points on other rigid bodies to which rigid body 208 is connected. In some complex mechanical or robotic systems, a rigid body may be connected to multiple other rigid bodies. Each rigid body can include one or more positioning points. When the number of positioning points on a single rigid body is insufficient, acquiring positioning points on other rigid bodies can increase the number of positioning points, thereby improving the positioning accuracy of rigid body 208.
[0035] In some embodiments, the position information of multiple positioning points can be the position information of multiple positioning points in an image. Multiple positioning points may include points on the rigid body 208 that can emit or reflect infrared light. Based on this, an infrared camera in the motion capture system can acquire images of multiple positioning points on the rigid body 208, obtaining images of the multiple positioning points, and then determining the position information of the multiple positioning points based on the images captured by the infrared camera. Infrared light wavelengths are typically within the range invisible to the naked eye, are unaffected by ambient light, and can operate stably under different lighting conditions, making them particularly suitable for complex lighting environments. Of course, this is only one optional implementation method; in other implementation methods, the positioning points can be of other types, and the images of the positioning points can also be acquired by other devices with image acquisition capabilities.
[0036] In step S12, multiple samples can be acquired. Each sample includes a camera pose transformation matrix and a rigid body pose transformation matrix. The camera pose transformation matrix represents the relative pose relationship of camera 202 between two time points. The rigid body pose transformation matrix represents the relative pose relationship of rigid body 208 between corresponding two time points. The two time points corresponding to the camera pose transformation matrix and the rigid body pose transformation matrix in the same sample are the same. For example, multiple samples include sample I and sample II, where sample I includes camera pose transformation matrix I and rigid body pose transformation matrix I, and sample II includes camera pose transformation matrix II and rigid body pose transformation matrix II. Camera pose transformation matrix I represents the relative pose relationship of camera 202 from time t1 to time t2, and rigid body pose transformation matrix I represents the relative pose relationship of rigid body 208 from time t1 to time t2. The camera pose transformation matrix II represents the relative pose relationship of camera 202 from time t2 to time t3, and the rigid body pose transformation matrix II represents the relative pose relationship of rigid body 208 from time t2 to time t3.
[0037] In some embodiments, the camera pose transformation matrix in any sample can be determined based on the following method: acquiring a first calibration object image obtained by the camera 202 acquiring an image of the calibration object at a first time and a second calibration object image obtained by the camera 202 acquiring an image of the calibration object at a second time different from the first time; detecting the position information of a first feature point on the calibration object from the first calibration object image; detecting the position information of a second feature point on the calibration object corresponding to the first feature point from the second calibration object image; and determining the camera pose transformation matrix between the first time and the second time based on the position information of the first feature point and the position information of the second feature point.
[0038] The calibration object can be a checkerboard, QR code, or calibration ball, etc., used for calibration. The first and second feature points can be corner points on the calibration object, obtained by feature detection on the images of the first and second calibration objects. The poses of the camera 202 at the first and second moments can be different. The camera 202, which is fixedly connected to the rigid body 208, can be moved by controlling the rigid body 208, thereby obtaining the camera pose transformation matrix between the two different moments (i.e., the first and second moments mentioned above) during the movement. After detecting the first and second feature points, a feature point matching algorithm can be used to match the first and second feature points, resulting in multiple feature point pairs. Each feature point pair includes a first feature point and a second feature point, and the first and second feature points in the same feature point pair correspond to the same feature point on the calibration object. Then, based on the position information of the first and second feature points in the multiple feature point pairs, the camera pose transformation matrix between the first and second moments can be determined, for example, by solving the above camera pose transformation matrix using the PnP algorithm.
[0039] In some embodiments, the rigid body pose transformation matrix in any sample is determined based on the following method: acquiring a first positioning point image obtained by image acquisition of multiple positioning points at a first time and a second positioning point image obtained by image acquisition of multiple positioning points at a second time different from the first time; detecting first position information of multiple positioning points from the first positioning point image and detecting second position information of multiple positioning points from the second positioning point image; and determining the rigid body pose transformation matrix between the first time and the second time based on the first position information and the second position information. Both the first and second positioning point images can be acquired by a motion capture system. For example, if the positioning point is a point on the rigid body 208 that can emit or reflect infrared light, the first and second positioning point images can be acquired by an infrared camera in the motion capture system. Of course, this application is not limited to this.
[0040] A pair of positioning points can be determined using a matching algorithm. Each pair includes a positioning point detected from a first positioning point image (hereinafter referred to as the first positioning point) and a positioning point detected from a second positioning point image (hereinafter referred to as the second positioning point). The first and second positioning points in the same pair correspond to the same positioning point on the rigid body 208. Then, based on the first position information of the first positioning point and the second position information of the second positioning point in the multiple positioning point pairs, the rigid body pose transformation matrix between the first and second time moments can be determined. For example, the rigid body pose transformation matrix can be solved using the PnP algorithm.
[0041] In step S14, multiple initial hand-eye transfer relationships and their quality can be determined. The quality of an initial hand-eye transfer relationship represents its degree of matching with each sample used to determine that initial hand-eye transfer relationship. The more samples that match the initial hand-eye transfer relationship, the higher its quality; conversely, the fewer samples that match the initial hand-eye transfer relationship, the lower its quality. Optionally, the quality of the initial hand-eye transfer relationship can be represented by a score. This score can be a percentage score, a five-point score, etc.
[0042] In some embodiments, a subset of samples can be randomly selected from a plurality of samples. After selecting a number of samples each time, an initial hand-eye transition relationship can be determined based on each of the selected samples, and the quality of the initial hand-eye transition relationship can be determined based on the matching degree between each of the selected samples and the initial hand-eye transition relationship. Repeating this process multiple times yields multiple initial hand-eye transition relationships and their qualities. For example, assuming the total number of samples is 10, and 5 samples are selected each time, then in the first execution of the above operation, 5 samples can be randomly selected from the 10 samples, an initial hand-eye transition relationship can be determined based on these 5 samples, and the quality of the first initial hand-eye transition relationship can be determined based on the matching degree between the first 5 selected samples and the first obtained initial hand-eye transition relationship. Similarly, in the second execution of the above operation, 5 samples can be randomly selected again from the 10 samples, an initial hand-eye transition relationship can be determined based on these 5 samples, and the quality of the second initial hand-eye transition relationship can be determined based on the matching degree between the second 5 selected samples and the second obtained initial hand-eye transition relationship. And so on. Assuming the above process is repeated 6 times, a total of 6 initial hand-eye conversion relationships and their quality can be determined.
[0043] In some embodiments, the ideal hand-eye translation relationship between the camera 202 and the rigid body 208 satisfies: the product of the camera pose transformation matrix and the hand-eye translation relationship is equal to the product of the hand-eye translation relationship and the rigid body pose transformation matrix, that is:
[0044] AX = XB
[0045] Where A represents the camera pose transformation matrix, B represents the rigid body pose transformation matrix, and X represents the hand-eye transformation relationship between camera 202 and rigid body 208.
[0046] After selecting a number of samples each time, the selected samples can be substituted into the above formula to obtain multiple equations (i.e., A). 1 X = XB1 A 2 X = XB 2 ...), based on multiple equations, X is estimated through fitting and other related techniques to obtain the initial hand-eye conversion relationship.
[0047] The matching degree between each selected sample and the initial hand-eye conversion relationship can be determined as follows: the residual of each selected sample is determined based on the initial hand-eye conversion relationship; the ratio between the number of samples with residuals greater than a preset threshold and the total number of selected samples is determined; and the matching degree between each selected sample and the initial hand-eye conversion relationship is determined based on this ratio.
[0048] Because noise may exist in the samples, the determined initial hand-eye conversion relationship may not strictly satisfy the condition AX = XB for all samples. For any selected sample, the first product of the camera pose conversion matrix in that sample and the initial hand-eye conversion relationship can be determined, and the second product of the initial hand-eye conversion relationship and the rigid body pose conversion matrix in that sample can be determined. The residual of the sample is determined based on the difference between the first product and the second product. The residual R of the nth sample is... n It can be written as:
[0049] R n =A n X-XB n
[0050] Among them, A n Let B represent the camera pose transformation matrix in the nth sample. n Let represent the rigid body pose transformation matrix in the nth sample.
[0051] Assuming the number of samples used to determine the initial hand-eye transfer relationship is K, and the number of samples with residuals greater than a preset threshold among the K samples is K0, the score of the initial hand-eye transfer relationship can be determined based on the ratio of K0 to K. The quality of the initial hand-eye transfer relationship is inversely correlated with the aforementioned ratio.
[0052] In step S16, the initial hand-eye conversion relationship with the highest quality can be determined, that is, the initial hand-eye conversion relationship with the smallest proportion of samples whose residuals are greater than a preset threshold. Assume there are three initial hand-eye conversion relationships determined in step S14, denoted as initial hand-eye conversion relationship 1, initial hand-eye conversion relationship 2, and initial hand-eye conversion relationship 3. The ratio of the number of samples with residuals greater than the preset threshold determined based on initial hand-eye conversion relationship 1 to the total number of samples used to determine initial hand-eye conversion relationship 1 is β1; the ratio of the number of samples with residuals greater than the preset threshold determined based on initial hand-eye conversion relationship 2 to the total number of samples used to determine initial hand-eye conversion relationship 2 is β2; and the ratio of the number of samples with residuals greater than the preset threshold determined based on initial hand-eye conversion relationship 3 to the total number of samples used to determine initial hand-eye conversion relationship 3 is β3. Since β1 > β2 > β3, initial hand-eye conversion relationship 3 can be determined as the initial hand-eye conversion relationship with the highest quality, and all samples used to determine initial hand-eye conversion relationship 3 can be determined as target samples.
[0053] In step S18, the target images corresponding to each target sample can be obtained. Assume that the samples used to determine the initial hand-eye conversion relationship of the highest quality include samples SP1, SP2, and SP3. Specifically, the camera pose conversion matrix in sample SP1 represents the relative pose relationship of camera 202 at times t11 and t12, and the rigid body pose conversion matrix in sample SP1 represents the relative pose relationship of rigid body 208 at times t11 and t12. That is, the target times corresponding to sample SP1 are times t11 and t12. Similarly, assume that the camera pose conversion matrix in sample SP2 represents the relative pose relationship of camera 202 at times t21 and t12. The relative pose relationships at time t22 are represented by the rigid body pose transformation matrix in sample SP2, which indicates the relative pose relationship of rigid body 208 at times t21 and t22. Therefore, the target times for sample SP2 are t21 and t22. Similarly, it is assumed that the camera pose transformation matrix in sample SP3 represents the relative pose relationship of camera 202 at times t31 and t32, and the rigid body pose transformation matrix in sample SP3 represents the relative pose relationship of rigid body 208 at times t31 and t32. Thus, the target times for sample SP3 are t31 and t32. Therefore, the images of the positioning points acquired at times t11, t12, t21, t22, t31, and t32 can all be used as target images. The optimized hand-eye conversion relationship can be determined based on the positional information of the positioning points in these target images. All of the above target images can be obtained by acquiring images of multiple positioning points using a motion capture system.
[0054] In some embodiments, the target image includes images acquired when the rigid body 208 is in multiple motion modes. These different motion modes can have different velocities, accelerations, and / or motion patterns, including but not limited to rotation and translation. For example, the target images acquired at times t11 and t12 could be images acquired during the rigid body 208's uniform translation at a first velocity; the target images acquired at times t21 and t22 could be images acquired during the rigid body 208's accelerated translation at a second velocity; and the target images acquired at times t31 and t32 could be images acquired during the rigid body 208's rotation at a third velocity. This approach increases data diversity, thereby improving the accuracy and robustness of the calibration results.
[0055] In some embodiments, an initial optimized hand-eye translation relationship can be determined first, and then used as the initial value in the optimization process. A certain optimization method is then employed to optimize the relationship, resulting in an optimized hand-eye translation relationship. Specifically, the positional information of the localization points in the acquired target image can be passed as input parameters to the hand-eye calibration function in the robot vision tool library. This function then solves the mathematical model of the hand-eye translation relationship based on the input parameters to obtain the initial optimized hand-eye translation relationship. The initial optimized hand-eye translation relationship is then iteratively optimized based on a preset optimization step size to obtain the optimized hand-eye translation relationship. The robot vision tool library includes, but is not limited to, OpenCV, PCL (PointCloud Library), and Kalibr. Taking OpenCV as an example, the hand-eye calibration function could be the `calibrateRobotWorldHandEye()` function. This function can obtain a preliminary, less accurate estimate of the matrix X (i.e., the initial optimized hand-eye translation relationship) through singular value decomposition, reducing the influence of the initial value on the final optimization result.
[0056] During the optimization process, a fixed or dynamically changing optimization step size can be used. In the example of dynamically changing optimization step size, as an implementation method, the optimization step size used in any iteration of optimizing the initial hand-eye relationship can be inversely correlated with the accuracy of the initial optimized hand-eye relationship at that iteration. That is, a larger optimization step size can be used at the beginning of the iteration to improve optimization efficiency; as the number of iterations increases, the accuracy of the initial optimized hand-eye relationship also gradually improves. At this point, the optimization step size can be gradually reduced to avoid exceeding the optimal value and improve optimization accuracy.
[0057] In some embodiments, an optimization function can be established based on a first transformation matrix between the world coordinate system and the rigid body coordinate system, a second transformation matrix between the calibration object coordinate system and the camera coordinate system, a third transformation matrix between the world coordinate system and the calibration object coordinate system, and an initial optimized hand-eye transformation relationship. The initial optimized hand-eye transformation relationship and the third transformation matrix in the optimization function are then iteratively optimized based on a preset optimization step size and a preset optimization objective to obtain an optimized hand-eye transformation relationship. The calibration object is used to determine the camera pose transformation matrix in multiple samples.
[0058] Figure 3 A schematic diagram of multiple coordinate systems involved in an embodiment of this application is shown. The coordinates of each positioning point on the rigid body 208 in the world coordinate system can be predetermined, and the coordinates of each positioning point on the rigid body 208 in the rigid body coordinate system can be obtained by a motion capture system. Based on the coordinates of each positioning point in the world coordinate system and the coordinates of each positioning point in the rigid body coordinate system, a first transformation matrix between the world coordinate system and the rigid body coordinate system can be determined.
[0059] The calibration object can be displayed on the screen 204 of the virtual shooting system 20. Therefore, the coordinate system of the calibration object can be the display screen coordinate system corresponding to the screen 204. The coordinates of feature points on the calibration object in the calibration object coordinate system can be obtained. Furthermore, the coordinates of feature points on the calibration object in the camera coordinate system can be determined based on the image of the calibration object captured by the camera 202. Based on the coordinates of the feature points in the calibration object coordinate system and the coordinates of the feature points in the camera coordinate system, the second transformation matrix between the calibration object coordinate system and the camera coordinate system can be determined. In addition, the first transformation matrix between the world coordinate system and the rigid body coordinate system can be determined based on the third transformation matrix between the world coordinate system and the calibration object coordinate system, the second transformation matrix between the calibration object coordinate system and the camera coordinate system, and the transformation matrix between the camera coordinate system and the rigid body coordinate system (i.e., the initial optimized hand-eye transformation relationship). That is, the above-mentioned first transformation matrix, second transformation matrix, third transformation matrix, and initial optimized hand-eye transformation relationship satisfy the following conditions:
[0060]
[0061] This is the transformation matrix from the world coordinate system to the rigid body coordinate system (the first transformation matrix);
[0062] This is the transformation matrix from the camera coordinate system to the rigid body coordinate system (initial optimization of hand-eye transformation relationship);
[0063] This is the transformation matrix (second transformation matrix) from the object coordinate system to the camera coordinate system;
[0064] This is the transformation matrix (third transformation matrix) from the world coordinate system to the calibration object coordinate system.
[0065] in, satisfy:
[0066]
[0067] P d Let P be the coordinates of a point located in the rigid body coordinate system. a These are the coordinates of the point in the world coordinate system.
[0068] An optimization function can be established based on the above formula, and the initial hand-eye transformation relationship and the third transformation matrix can be jointly optimized. The result of optimizing the initial hand-eye transformation relationship is the optimized hand-eye transformation relationship.
[0069] In some embodiments, the original coordinates of the positioning point in the rigid body coordinate system can be obtained, and the coordinates of the positioning point in the world coordinate system can be converted into the coordinates of the positioning point in the rigid body coordinate system through the matrix transformation relationship between the first transformation matrix, the second transformation matrix, the third transformation matrix and the initial optimized hand-eye conversion relationship. The distance between the original coordinates of the positioning point in the rigid body coordinate system and the coordinates of the positioning point in the rigid body coordinate system obtained through coordinate transformation is used as the optimization target, and optimization is performed through optimization methods such as least squares optimization, and finally the optimized hand-eye conversion relationship is obtained.
[0070] The following is combined Figure 4 The overall process of the embodiments of this application will be illustrated by example.
[0071] This application first filters the data. The core idea is to filter out noisy positioning points to eliminate those with large errors, and then perform hand-eye calibration based on the relatively clean positioning points. This process relies on the formula AX = XB to process two sets of key data: A (representing the pose transformation matrix of camera 104 at different times) and B (representing the pose transformation matrix of the rigid body fixedly connected to camera 202 at corresponding times). By effectively filtering these two types of data, the aim is to obtain more accurate hand-eye relationship calibration results. The specific implementation steps are as follows:
[0072] (1) Random sampling: Select M (here set to N-3) samples from N samples (each pair of AB constitutes a sample) as a subset, where each subset consists of a randomly selected pair of AB;
[0073] (2) Estimating the initial hand-eye transition relationship: Using the selected M samples, the initial hand-eye transition relationship is estimated using the given formula AX = XB;
[0074] (3) Interior point recognition: Substitute the initial hand-eye conversion relationship obtained in the above steps into the formula Rn =A n X-XB n The residual R-values for all samples are calculated. A fixed threshold is then set to distinguish which samples can be considered "interiors"—samples with residuals less than or equal to this threshold. The proportion of these interiors to the total number of samples is then calculated and used as a quantitative indicator of data quality in the current iteration.
[0075] (4) Optimal dataset selection: Repeat steps 1 to 3 several times (denoted as N′). Each iteration will generate a new initial hand-eye conversion relationship and its corresponding score. Finally, select the set of samples with the highest score. It is believed that the quality of the set of samples is high and the images corresponding to the set of samples are suitable for further calculation of more accurate optimization of hand-eye conversion relationship.
[0076] By using the above method, this application can not only effectively reduce the impact of outliers on the final calibration accuracy, but also improve the robustness of the entire system to external interference factors.
[0077] This application has the following technical effects:
[0078] (1) Only a portion of the data is considered in each iteration for interior point identification and data filtering. This can filter out noisy data and remove data points with large errors, thereby reducing the impact of noise and outliers on the final calibration accuracy, improving the robustness of the system, and providing stronger fault tolerance.
[0079] (2) Before optimization, the initial optimization hand-eye conversion relationship is determined based on the hand-eye calibration function in the robot vision tool library, which can improve the rationality of the initial optimization value and reduce the impact of the sensitivity of the initial value.
[0080] (3) By using reasonable initialization and step size control, the iteration process can be accelerated and the time consumption can be reduced.
[0081] (4) By adding more positioning points or different motion modes, the robustness and accuracy of the calibration results can be improved.
[0082] (5) By taking the distance between the original coordinates of the positioning point in the rigid body coordinate system and the coordinates of the positioning point in the rigid body coordinate system obtained through coordinate transformation (i.e., spatial registration error) as the optimization target, the error accumulation of intermediate steps is reduced, which can improve the accuracy of calibration and improve the computational efficiency, and is suitable for real-time application scenarios.
[0083] (6) This application does not require the use of camera intrinsic parameter matrix during the optimization process, which reduces the system’s dependence on image quality and feature point detection and enhances the system’s robustness and applicability.
[0084] (7) This application obtains the coordinates of the positioning points on the rigid body through a motion capture system, without relying on a specific calibration target (such as a checkerboard), thus having better adaptability and flexibility, and can be applied to a variety of scenarios.
[0085] like Figure 5 As shown, this application also provides a hand-eye calibration transpose for calibrating the hand-eye conversion relationship between a camera and a rigid body fixedly connected to the camera, wherein multiple positioning points are distributed around the rigid body; the transpose includes:
[0086] The first acquisition module 302 is used to acquire multiple samples, each of which includes a camera pose transformation matrix and a rigid body pose transformation matrix. The camera pose transformation matrix is used to represent the relative pose relationship of the camera between two time points, and the rigid body pose transformation matrix is used to represent the relative pose relationship of the rigid body between the two time points. The rigid body pose transformation matrix is determined based on the position information of the multiple positioning points.
[0087] The first determining module 304 is used to determine multiple initial hand-eye conversion relationships and their quality, wherein any initial hand-eye conversion relationship and its quality are determined based on the following method: randomly selecting a portion of samples from the multiple samples, determining the initial hand-eye conversion relationship based on each of the selected samples, and determining the quality of the initial hand-eye conversion relationship based on the matching degree between each of the selected samples and the initial hand-eye conversion relationship.
[0088] The second determining module 306 is used to determine the highest quality initial hand-eye conversion relationship, and to determine each sample used to determine the highest quality initial hand-eye conversion relationship as the target sample;
[0089] The second acquisition module 308 is used to acquire the target image corresponding to each target sample, and determine the optimized hand-eye conversion relationship based on the position information of the positioning points in the acquired target image; wherein, the target image corresponding to the target sample includes the images of the multiple positioning points acquired at the target time corresponding to the target sample.
[0090] The apparatus provided in this application embodiment has functions or includes modules that can be used to execute the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0091] This application also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method of any embodiment of this application.
[0092] Figure 6This illustration shows a more specific hardware structure diagram of a computer device provided in an embodiment of this application. The device may include: a processor 402, a memory 404, an input / output interface 406, a communication interface 408, and a bus 410. The processor 402, memory 404, input / output interface 406, and communication interface 408 are interconnected internally via the bus 410.
[0093] The processor 402 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The processor 402 may also include a graphics card, such as an Nvidia Titan X graphics card or a 1080Ti graphics card.
[0094] The memory 404 can be implemented in the form of read-only memory (ROM), random access memory (RAM), static storage device, dynamic storage device, etc. The memory 404 can store the operating system and other applications. When the technical solutions provided in the embodiments of this application are implemented by software or firmware, the relevant program code is stored in the memory 404 and is called and executed by the processor 402.
[0095] Input / output interface 406 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touch screens, microphones, various sensors, etc., and output devices may include displays, speakers, vibrators, indicator lights, etc.
[0096] Communication interface 408 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0097] Bus 410 includes a pathway for transmitting information between various components of the device, such as processor 402, memory 404, input / output interface 406, and communication interface 408.
[0098] It should be noted that although the above-described device only shows the processor 402, memory 404, input / output interface 406, communication interface 408, and bus 410, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this application, and not necessarily all the components shown in the figures.
[0099] This application provides a computer program product, including a computer program that, when executed by a processor, implements the methods described in any embodiment of this application.
[0100] This application also provides a computer-readable storage medium storing computer instructions thereon, which, when executed by a processor, implement the addressing method described in any embodiment of this application. The computer-readable storage medium may be phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transmission medium, which can be used to store information accessible by a computing device.
[0101] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that the embodiments of this application can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions of the embodiments of this application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.
[0102] The systems, devices, modules, or units described in the above embodiments can be implemented by computer devices or entities, or by products with certain functions. A typical implementation device is a computer, which can take the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email sending and receiving device, game console, tablet computer, wearable device, or any combination of these devices.
[0103] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate. When implementing the embodiments of this application, the functions of each module can be implemented in one or more software and / or hardware. Alternatively, some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0104] The above description is only a specific implementation of the embodiments of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the embodiments of this application, and these improvements and modifications should also be considered as the protection scope of the embodiments of this application.
Claims
1. A hand-eye calibration method for calibrating the hand-eye conversion relationship between a camera and a rigid body fixedly connected to the camera, wherein multiple positioning points are distributed around the rigid body; the method includes: Multiple samples are acquired, and each sample includes a camera pose transformation matrix and a rigid body pose transformation matrix. The camera pose transformation matrix is used to represent the relative pose relationship of the camera between two time points, and the rigid body pose transformation matrix is used to represent the relative pose relationship of the rigid body between the two time points. The rigid body pose transformation matrix is determined based on the position information of the multiple positioning points. Multiple initial hand-eye transition relationships and their quality are determined, wherein any initial hand-eye transition relationship and its quality are determined based on the following method: a subset of samples are randomly selected from the multiple samples; initial hand-eye transition relationships are determined based on each selected sample; and the quality of the initial hand-eye transition relationship is determined based on the matching degree between each selected sample and the initial hand-eye transition relationship. The more samples used to determine the initial hand-eye transition relationship that match the initial hand-eye transition relationship, the higher the quality of the initial hand-eye transition relationship. The highest quality initial hand-eye transition relationship is determined, and each sample used to determine the highest quality initial hand-eye transition relationship is designated as the target sample; The target image corresponding to each target sample is acquired, and the optimized hand-eye conversion relationship is determined based on the position information of the positioning points in the acquired target image; wherein, the target image corresponding to the target sample includes the images of the multiple positioning points acquired at the target time corresponding to the target sample.
2. The method according to claim 1, further comprising: The residuals of each selected sample are determined based on the initial hand-eye conversion relationship; Determine the ratio between the number of samples with residuals greater than a preset threshold and the total number of all selected samples; The matching degree between each selected sample and the initial hand-eye conversion relationship is determined based on the ratio.
3. The method according to claim 2, wherein determining the residuals of each selected sample based on the initial hand-eye conversion relationship comprises: For any selected sample, determine the first product between the camera pose transformation matrix in the sample and the initial hand-eye transformation relationship, and determine the second product between the initial hand-eye transformation relationship and the rigid body pose transformation matrix in the sample; The residual of the sample is determined based on the difference between the first product and the second product.
4. The method according to claim 1, wherein the plurality of positioning points includes at least one of the following: The rigid body has its own positioning points; Positioning points bound to the rigid body; The positioning points on other rigid bodies connected to the rigid body; and / or The target image includes images acquired when the rigid body is in various motion modes.
5. The method according to claim 1, wherein determining the optimized hand-eye translation relationship based on the position information of the positioning points in the acquired target image includes: The position information of the localization points in the acquired target image is passed as input parameters to the hand-eye calibration function in the robot vision tool library, so that the hand-eye calibration function can solve the mathematical model of the hand-eye conversion relationship based on the input parameters to obtain the initial optimized hand-eye conversion relationship. The initial optimized hand-eye conversion relationship is iteratively optimized based on a preset optimization step size to obtain an optimized hand-eye conversion relationship.
6. The method according to claim 5, wherein iteratively optimizing the initial optimized hand-eye conversion relationship based on a preset optimization step size to obtain an optimized hand-eye conversion relationship includes: An optimization function is established based on the first transformation matrix between the world coordinate system and the rigid body coordinate system, the second transformation matrix between the calibration object coordinate system and the camera coordinate system, the third transformation matrix between the world coordinate system and the calibration object coordinate system, and the initial optimized hand-eye transformation relationship; wherein, the calibration object is used to determine the camera pose transformation matrix in the multiple samples; Based on a preset optimization step size and a preset optimization objective, the initial optimized hand-eye transformation relationship and the third transformation matrix in the optimization function are iteratively optimized to obtain the optimized hand-eye transformation relationship.
7. In the method according to claim 5, the optimization step size used in any iteration of the initial optimized hand-eye conversion relationship is inversely correlated with the accuracy of the initial optimized hand-eye conversion relationship in that iteration.
8. The method according to claim 1, wherein the camera pose transformation matrix in any sample is determined based on the following method: Acquire a first image of the calibration object obtained by the camera at a first moment and a second image of the calibration object obtained by the camera at a second moment different from the first moment; The position information of a first feature point on the calibration object is detected from the first calibration object image, and the position information of a second feature point on the calibration object corresponding to the first feature point is detected from the second calibration object image; The camera pose transformation matrix between the first time point and the second time point is determined based on the position information of the first feature point and the position information of the second feature point.
9. The method according to claim 1, wherein the rigid body pose transformation matrix in any sample is determined based on the following method: Acquire a first positioning point image obtained by image acquisition of the plurality of positioning points at a first moment, and a second positioning point image obtained by image acquisition of the plurality of positioning points at a second moment different from the first moment; First position information of the plurality of positioning points is detected from the first positioning point image, and second position information of the plurality of positioning points is detected from the second positioning point image; The rigid body pose transformation matrix between the first time point and the second time point is determined based on the first position information and the second position information.
10. The method according to claim 1, wherein the plurality of positioning points include points on the rigid body that can emit or reflect infrared light, and the target images corresponding to each target sample are acquired by an infrared camera in the motion capture system.
11. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method of any one of claims 1 to 10.
12. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the method of any one of claims 1 to 10.
13. A computer program product comprising a computer program that, when executed by a processor, implements the method of any one of claims 1 to 10.
Citation Information
Patent Citations
Hand-eye calibration method, device and equipment based on point cloud registration
CN116038720A