Hand-eye calibration method and device, medium, computer equipment and program product
By acquiring multiple samples and optimizing the hand-eye conversion relationship, the problem of low hand-eye calibration accuracy in complex shooting environments is solved, and higher camera positioning accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202411911068.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-12-23
AI Technical Summary
In complex shooting environments, it is difficult to obtain high-precision hand-eye conversion relationships in hand-eye calibration methods, resulting in low camera positioning accuracy.
By obtaining multiple samples, including the camera pose conversion matrix and the rigid pose conversion matrix, the initial hand-eye conversion relationship and its quality are determined, the highest quality initial hand-eye conversion relationship is selected, and the hand-eye conversion relationship is optimized based on the position information of the position point in the target image.
It improves the accuracy of hand-eye conversion relationship, enhances the accuracy and robustness of camera positioning, and reduces the impact of noise on positioning.
Smart Images

Figure CN120070589A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of hand-eye calibration, and in particular, to a hand-eye calibration method, device, medium, computer device, and program product. Background Art
[0002] When using a camera to photograph a target, it is often necessary to determine the pose of the camera. Since most cameras do not have the function of self-positioning, the usual method is to fixedly connect the camera to a positionable rigid body, position the rigid body, and calculate the hand-eye transformation relationship between the rigid body and the camera through a hand-eye calibration algorithm, so as to indirectly obtain the pose of the camera. However, factors such as the scene and lighting of the actual shooting environment may be relatively complex, resulting in a low accuracy of the obtained hand-eye transformation relationship, and thus a low accuracy of camera positioning. Summary of the Invention
[0003] In a first aspect, this application provides a hand-eye calibration method for calibrating the hand-eye transformation relationship between a camera and a rigid body fixedly connected to the camera, and a plurality of positioning points are distributed around the rigid body; the method includes: obtaining a plurality of samples, any one of the samples including a camera pose transformation matrix and a rigid body pose transformation matrix, the camera pose transformation matrix being used to represent the relative pose relationship of the camera between two moments, and the rigid body pose transformation matrix being used to represent the relative pose relationship of the rigid body between the two moments; the rigid body pose transformation matrix is determined based on the position information of the plurality of positioning points; determining a plurality of initial hand-eye transformation relationships and their qualities, wherein any one of the initial hand-eye transformation relationships and their qualities is determined based on the following method: randomly selecting some samples from the plurality of samples, determining an initial hand-eye transformation relationship based on each of the selected samples, and determining the quality of the initial hand-eye transformation relationship based on the matching degree between each of the selected samples and the initial hand-eye transformation relationship; determining the initial hand-eye transformation relationship with the highest quality, and determining all the samples used to determine the initial hand-eye transformation relationship with the highest quality as target samples; obtaining target images corresponding to each of the target samples, and determining an optimized hand-eye transformation relationship based on the position information of the positioning points in the obtained target images; wherein the target image corresponding to the target sample includes an image of the plurality of positioning points collected at the target moment corresponding to the target sample.
[0004] In a second aspect, an embodiment of this application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in any embodiment of this application is implemented.
[0005] In a third aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method described in any embodiment of the present application is implemented.
[0006] In an embodiment of the present application, after obtaining a plurality of samples including a camera pose transformation matrix and a rigid body pose transformation matrix, first select a part of the samples from the plurality of samples to determine an initial hand-eye transformation relationship and its quality, and then determine the initial hand-eye transformation relationship with the highest quality from the plurality of initial hand-eye transformation relationships. The higher the quality of the initial hand-eye transformation relationship, the higher the degree of adaptation of the initial hand-eye transformation relationship to each of the selected samples, thereby indirectly reflecting that the accuracy of the target samples used to determine the initial hand-eye transformation relationship is relatively high. Since the hand-eye transformation relationship is determined based on the position information of the positioning points in the target image corresponding to the target samples, the higher the accuracy of the target samples, the higher the accuracy of the position information of the positioning points in the target image corresponding to the target samples. Therefore, an optimized hand-eye transformation relationship can be determined based on the position information of the positioning points in the target image. In this way, the position information of the positioning points affected by noise and generating large errors can be filtered out, and only the position information of the positioning points with relatively high accuracy is used to determine the hand-eye transformation relationship, thereby improving the accuracy of the determined hand-eye transformation relationship and further improving the camera positioning accuracy.
[0007] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The accompanying drawings herein are incorporated into the specification and constitute a part of this application. These drawings show embodiments consistent with the present application and are used together with the specification to explain the technical solutions of the present application.
[0009] Figure 1 is a schematic diagram of a virtual shooting system according to an embodiment of the present application.
[0010] Figure 2 is a flowchart of a hand-eye calibration method according to an embodiment of the present application.
[0011] Figure 3 is a schematic diagram of the conversion relationship between multiple coordinate systems according to an embodiment of the present application.
[0012] Figure 4 is a general flowchart according to an embodiment of the present application.
[0013] Figure 5 is a block diagram of a hand-eye calibration device according to an embodiment of the present application.
[0014] Figure 6 is a schematic diagram of a computer device according to an embodiment of the present application. Detailed Implementation Modes
[0015] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation modes described in the following exemplary embodiments do not represent all implementation modes consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0016] The terms used in the present application are for the purpose of describing specific embodiments only and are not intended to limit the present application. The singular forms "a", "the", and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. Additionally, the term "at least one" as used herein represents any one of a plurality or any combination of at least two of a plurality.
[0017] It should be understood that although the terms first, second, third, etc. may be used in the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0018] In order to enable those skilled in the art of the present technology to better understand the technical solutions in the embodiments of the present application and make the above-mentioned objects, features, and advantages of the embodiments of the present application more apparent and understandable, the technical solutions in the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.
[0019] When using a camera to photograph a shooting target, it is often necessary to determine the pose of the camera. In some embodiments, the shooting scene is a virtual shooting scene, and the above-mentioned camera is a camera in a virtual shooting system. Figure 1A schematic diagram of a virtual shooting system 20 of some embodiments is shown. The virtual shooting system 20 includes a camera 202, a display screen 204 and a processing device 206, wherein the display screen 204 is used to display a virtual picture. The shooting target (such as an actor) can be located in front of the display screen 204, and the camera 202 is used to shoot the shooting target and the virtual picture displayed on the display screen 204 to obtain a shot picture. The virtual picture can be a picture within the field of view of the virtual camera. The virtual camera is a mathematical model of the camera 202, which is used to collect pictures in the virtual scene as the background of the virtual shooting. The posture of the virtual camera can be consistent with the posture of the camera 202. By accurately determining the posture of the camera 202, the posture of the virtual camera can be further determined, so as to render a virtual picture corresponding to the viewing angle of the virtual camera.
[0020] like Figure 1 As shown, the number of display screens 204 is 3. For the sake of distinction, each display screen is recorded as display screen 2042, display screen 2044 and display screen 2046. The display screens (2042, 2044, 2046) can display virtual images independently or in conjunction. Among them, the display screen 204 used in the virtual shooting system 20 can be a physical screen, which can be a type of LED display screen, liquid crystal display screen, etc., and can be a curved screen or a flat screen. It should be understood that those skilled in the art can customize the type, quantity, size, resolution, installation position, etc. of the physical screen in the virtual shooting system 20 according to actual needs, and this embodiment of the application is not limited to this.
[0021] The processing device 206 is used to control the display process of the virtual picture on the display screen 204, such as controlling the playback frame rate of the virtual picture, the parameters of the virtual camera 202, the switching of the virtual picture, etc.
[0022] The camera 202 may be fixedly connected to the rigid body 208. The camera 202 may be a camera without a self-positioning function, and the position and posture of the camera 202 may be indirectly obtained by positioning the rigid body 208 and calculating the hand-eye conversion relationship between the rigid body 208 and the camera 202 (i.e., the relative position and posture relationship between the rigid body 208 and the camera 202) through a hand-eye calibration algorithm.
[0023] In an actual shooting scene, factors such as the scene and lighting of the shooting environment may be relatively complex, resulting in a low precision of the obtained hand-eye conversion relationship, thereby resulting in a low positioning accuracy of the camera 202 .
[0024] Based on this, the present application provides a hand-eye calibration method for calibrating the hand-eye conversion relationship between the camera 202 and the rigid body 208 fixedly connected to the camera 202, and multiple positioning points are distributed around the rigid body 208; see Figure 2 , the method comprising:
[0025] Step S12: Obtain multiple samples. Any one sample includes a camera pose transformation matrix and a rigid body pose transformation matrix. The camera pose transformation matrix is used to represent the relative pose relationship of the camera 202 between two moments, and the rigid body pose transformation matrix is used to represent the relative pose relationship of the rigid body 208 between the two moments. The rigid body pose transformation matrix is determined based on the position information of multiple positioning points.
[0026] Step S14: Determine multiple initial hand-eye transformation relationships and their qualities. Among them, any one initial hand-eye transformation relationship and its quality are determined based on the following method: Randomly select some samples from multiple samples, determine the initial hand-eye transformation relationship based on each selected sample, and determine the quality of the initial hand-eye transformation relationship based on the matching degree between each selected sample and the initial hand-eye transformation relationship.
[0027] Step S16: Determine the initial hand-eye transformation relationship with the highest quality, and determine all the samples used to determine the initial hand-eye transformation relationship with the highest quality as target samples.
[0028] Step S18: Obtain the target images corresponding to each target sample, and determine the optimized hand-eye transformation relationship based on the position information of the positioning points in the obtained target images. Among them, the target image corresponding to the target sample includes the images of multiple positioning points collected at the target moment corresponding to the target sample.
[0029] In the embodiments of the present application, first select some samples from multiple samples to determine the initial hand-eye transformation relationship and its quality, and then determine the initial hand-eye transformation relationship with the highest quality from multiple initial hand-eye transformation relationships. The quality of the initial hand-eye transformation relationship can indirectly reflect the accuracy of the target samples used to determine the initial hand-eye transformation relationship, and further reflect the accuracy of the position information of the positioning points. Therefore, the target samples and their corresponding target images can be determined based on the initial hand-eye transformation relationship with the highest quality, and then the optimized hand-eye transformation relationship can be determined based on the position information of the positioning points in the target images. In this way, the position information of the positioning points affected by noise and generating large errors can be filtered out, and only the position information of the positioning points with higher accuracy is used to determine the hand-eye transformation relationship, thereby improving the accuracy of the determined hand-eye transformation relationship and further improving the camera positioning accuracy.
[0030] The following takes the accompanying drawings as an example to illustrate the specific implementation details of the present application.
[0031] The rigid body 208 in the present application refers to an object whose internal parts maintain their relative positions unchanged under the action of external forces. The rigid body 208 can be an object such as a robotic arm, a camera bracket, a robot, etc. There are multiple positioning points distributed around the rigid body 208. The positioning point refers to a feature point that can be used to track the rigid body 208. In some embodiments, the multiple positioning points include but are not limited to at least one of the following:
[0032] The positioning points built into the rigid body 208. The built-in positioning points refer to the fixed or internal positioning points that the rigid body 208 itself has. These points are usually determined during the design of the object and are part of the structure of the rigid body 208. Such positioning points can be some special marks fixed on the surface of the rigid body 208, sensor interfaces, positioning holes, laser points, etc. The built-in positioning points are usually part of the design of the rigid body 208 and do not require additional installation, with low deployment costs.
[0033] The positioning points bound to the rigid body 208. The positioning points bound to the rigid body 208 usually refer to the positioning points where positioning devices, marks or sensors are attached to the surface of the rigid body 208 through external means (such as pasting, fixing, installing, etc.). Common forms include reflective marks, inertial sensors, optical marks, QR codes, etc. Users can choose appropriate positions to bind the positioning points to the rigid body 208 according to their needs, and the positions of the positioning points can be flexibly arranged according to the motion requirements. The positioning points bound through external means do not change the basic structure and function of the rigid body 208 and can usually be installed without interfering with the main functions of the rigid body 208. In addition, when the number of built-in positioning points on the rigid body 208 is insufficient, by binding additional positioning points to the rigid body 208, the number of positioning points can be increased, thereby improving the positioning accuracy of the rigid body 208.
[0034] The positioning points on other rigid bodies connected to the rigid body 208. In some complex mechanical systems or robotic systems, one rigid body may be connected to multiple other rigid bodies. Each rigid body can include one or more positioning points. When the number of built-in positioning points on a single rigid body is insufficient, by obtaining the positioning points on other rigid bodies, the number of positioning points can be increased, thereby improving the positioning accuracy of the rigid body 208.
[0035] In some embodiments, the position information of multiple positioning points can be the position information of multiple positioning points in an image. The multiple positioning points can include points on the rigid body 208 that can emit or reflect infrared light. On this basis, an infrared camera in the motion capture system can be used to collect images of multiple positioning points on the rigid body 208 to obtain images of the multiple positioning points, and then the position information of the multiple positioning points can be determined based on the images captured by the infrared camera. The wavelength of infrared light is usually in the range invisible to the naked eye and is not affected by ambient light, and can work stably under different lighting conditions, especially suitable for complex lighting environments. Of course, this is only an optional implementation method. In other implementation methods, the positioning points can also be of other types, and the images of the positioning points can also be collected by other devices with image acquisition functions.
[0036] In step S12, multiple samples can be obtained. Each sample includes a camera pose transformation matrix and a rigid body pose transformation matrix. The camera pose transformation matrix is used to represent the relative pose relationship of the camera 202 between two moments. The rigid body pose transformation matrix is used to represent the relative pose relationship of the rigid body 208 between the corresponding two moments. The two moments corresponding to the camera pose transformation matrix and the rigid body pose transformation matrix in the same sample are the same. For example, the multiple samples include sample I and sample II. Among them, sample I includes the camera pose transformation matrix I and the rigid body pose transformation matrix I, and sample II includes the camera pose transformation matrix II and the rigid body pose transformation matrix II. The camera pose transformation matrix I represents the relative pose relationship of the camera 202 from time t1 to time t2, then the rigid body pose transformation matrix I represents the relative pose relationship of the rigid body 208 from time t1 to time t2. The camera pose transformation matrix II represents the relative pose relationship of the camera 202 from time t2 to time t3, then the rigid body pose transformation matrix II represents the relative pose relationship of the rigid body 208 from time t2 to time t3.
[0037] In some embodiments, the camera pose transformation matrix in any one sample can be determined based on the following method: Obtain the first calibration object image collected by the camera 202 for the calibration object at the first moment and the second calibration object image collected by the camera 202 for the calibration object at the second moment different from the first moment. Detect the position information of the first feature point on the calibration object from the first calibration object image, and detect the position information of the second feature point corresponding to the first feature point on the calibration object from the second calibration object image. Determine the camera pose transformation matrix between the first moment and the second moment based on the position information of the first feature point and the position information of the second feature point.
[0038] Among them, the calibration object can be an object for calibration such as a checkerboard, a QR code, or a calibration sphere. The first feature point and the second feature point can be corner points on the calibration object, and can be obtained by performing feature detection on the first calibration object image and the second calibration object image. The poses of the camera 202 at the first moment and the second moment can be different. The movement of the rigid body 208 can be controlled to drive the movement of the camera 202 fixedly connected to the rigid body 208, and the camera pose transformation matrix between two different moments (i.e., the above-mentioned first moment and second moment) during the movement of the camera 202 can be obtained. After detecting the first feature point and the second feature point, the first feature point and the second feature point can be matched through a feature point matching algorithm to obtain a plurality of feature point pairs. Each feature point pair includes a first feature point and a second feature point, and the first feature point and the second feature point in the same feature point pair correspond to the same feature point on the calibration object. Then, based on the position information of the first feature point and the position information of the second feature point in the plurality of feature point pairs, the camera pose transformation matrix between the first moment and the second moment can be determined. For example, the above camera pose transformation matrix can be solved through the PnP algorithm.
[0039] In some embodiments, the rigid body pose transformation matrix in any one sample is determined based on the following method: obtaining a first positioning point image obtained by collecting images of a plurality of positioning points at a first moment and a second positioning point image obtained by collecting images of a plurality of positioning points at a second moment different from the first moment; detecting the first position information of the plurality of positioning points from the first positioning point image, and detecting the second position information of the plurality of positioning points from the second positioning point image; determining the rigid body pose transformation matrix between the first moment and the second moment based on the first position information and the second position information. Among them, both the first positioning point image and the second positioning point image can be obtained through a motion capture system. For example, when the positioning point is a point on the rigid body 208 that can emit or reflect infrared light, the above first positioning point image and second positioning point image can be obtained through the infrared camera in the motion capture system. Of course, the present application is not limited to this.
[0040] A plurality of positioning point pairs can be determined through a matching algorithm. Among them, each positioning point pair includes a positioning point detected from the first positioning point image (hereinafter referred to as the first positioning point) and a positioning point detected from the second positioning point image (hereinafter referred to as the second positioning point), and the first positioning point and the second positioning point in the same positioning point pair correspond to the same positioning point on the rigid body 208. Then, based on the first position information of the first positioning point and the second position information of the second positioning point in the plurality of positioning point pairs, the rigid body pose transformation matrix between the first moment and the second moment can be determined. For example, the above rigid body pose transformation matrix can be solved through the PnP algorithm.
[0041] In step S14, multiple initial hand-eye transformation relationships and their qualities can be determined. Among them, the quality of an initial hand-eye transformation relationship is used to represent the matching degree between the initial hand-eye transformation relationship and each sample used to determine the initial hand-eye transformation relationship. Among the samples used to determine the initial hand-eye transformation relationship, the more samples that match the initial hand-eye transformation relationship, the higher the quality of the initial hand-eye transformation relationship; conversely, the fewer samples that match the initial hand-eye transformation relationship among the samples used to determine the initial hand-eye transformation relationship, the lower the quality of the initial hand-eye transformation relationship. Optionally, the quality of the initial hand-eye transformation relationship can be represented by a score. The score can be a percentage score, or a five-point scale score, etc.
[0042] In some embodiments, a part of the samples can be randomly selected from multiple samples. After selecting several samples each time, an initial hand-eye transformation relationship can be determined based on each of the samples selected this time, and the quality of the initial hand-eye transformation relationship obtained this time can be determined based on the matching degree between each of the samples selected this time and the initial hand-eye transformation relationship obtained this time. Repeating the above process multiple times can obtain multiple initial hand-eye transformation relationships and their qualities. For example, assuming that the total number of samples is 10 and the number of samples selected each time is 5, then when performing the above operation for the first time, 5 samples can be randomly selected from 10 samples, an initial hand-eye transformation relationship can be determined based on these 5 samples, and the quality of the initial hand-eye transformation relationship obtained for the first time can be determined based on the matching degree between the 5 samples selected for the first time and the initial hand-eye transformation relationship obtained for the first time. Similarly, when performing the above operation for the second time, 5 samples can be randomly selected again from 10 samples, an initial hand-eye transformation relationship can be determined based on these 5 samples, and the quality of the initial hand-eye transformation relationship obtained for the second time can be determined based on the matching degree between the 5 samples selected for the second time and the initial hand-eye transformation relationship obtained for the second time. And so on. Assuming that the above process is repeated 6 times in total, 6 initial hand-eye transformation relationships and their qualities can be determined in total.
[0043] In some embodiments, the ideal hand-eye transformation relationship between the camera 202 and the rigid body 208 satisfies: the product of the camera pose transformation matrix and the hand-eye transformation relationship is equal to the product of the hand-eye transformation relationship and the rigid body pose transformation matrix, that is:
[0044] AX = XB
[0045] where A represents the camera pose transformation matrix, B represents the rigid body pose transformation matrix, and X represents the hand-eye transformation relationship between the camera 202 and the rigid body 208.
[0046] After selecting several samples each time, each of the several samples selected this time can be substituted into the above formula to obtain multiple equations (i.e., A 1 X = XB1 , A 2 X = XB 2 ……), based on multiple equations, estimate X through fitting and other methods in related technologies to obtain an initial hand-eye transformation relationship.
[0047] The matching degree of each selected sample with the initial hand-eye transformation relationship can be determined based on the following method: Determine the residuals of each selected sample based on the initial hand-eye transformation relationship, determine the ratio between the number of samples with residuals greater than a preset threshold and the total number of all selected samples, and determine the matching degree of each selected sample with the initial hand-eye transformation relationship based on this ratio.
[0048] Since there may be noise in the samples, the determined initial hand-eye transformation relationship may not enable all samples to strictly satisfy the condition AX = XB. For any selected sample, the first product of the camera pose transformation matrix in the sample and the initial hand-eye transformation relationship can be determined, and the second product between the initial hand-eye transformation relationship and the rigid body pose transformation matrix in the sample can be determined. The residual of the sample is determined based on the difference between the first product and the second product. The residual R of the nth sample n can be denoted as:
[0049] R n = A n X - XB n
[0050] where A n represents the camera pose transformation matrix in the nth sample, and B n represents the rigid body pose transformation matrix in the nth sample.
[0051] Assume that the number of samples used to determine the initial hand-eye transformation relationship is K, and the number of samples with residuals greater than the preset threshold among the K samples is K0. Then, the score of this initial hand-eye transformation relationship can be determined based on the ratio of K0 to K. The quality of the initial hand-eye transformation relationship is inversely related to the above ratio.
[0052] In step S16, the initial hand-eye transformation relationship with the highest quality can be determined, that is, the initial hand-eye transformation relationship with the smallest proportion of the number of samples whose residuals are greater than the preset threshold is determined. Suppose there are 3 initial hand-eye transformation relationships determined based on step S14, which are respectively denoted as the initial hand-eye transformation relationship 1, the initial hand-eye transformation relationship 2, and the initial hand-eye transformation relationship 3. Among them, the ratio of the number of samples whose residuals are greater than the preset threshold determined based on the initial hand-eye transformation relationship 1 to the total number of samples used to determine the initial hand-eye transformation relationship 1 is β1, the ratio of the number of samples whose residuals are greater than the preset threshold determined based on the initial hand-eye transformation relationship 2 to the total number of samples used to determine the initial hand-eye transformation relationship 2 is β2, and the ratio of the number of samples whose residuals are greater than the preset threshold determined based on the initial hand-eye transformation relationship 3 to the total number of samples used to determine the initial hand-eye transformation relationship 3 is β3, and β1>β2>β3. Then, the initial hand-eye transformation relationship 3 can be determined as the initial hand-eye transformation relationship with the highest quality, and each sample used to determine the initial hand-eye transformation relationship 3 is determined as a target sample.
[0053] In step S18, the target images corresponding to each target sample can be obtained. Suppose each sample used to determine the initial hand-eye transformation relationship with the highest quality includes sample SP1, sample SP2, and sample SP3. Among them, the camera pose transformation matrix in sample SP1 represents the relative pose relationship between the camera 202 at time t11 and time t12, and the rigid body pose transformation matrix in sample SP1 represents the relative pose relationship between the rigid body 208 at time t11 and time t12, that is, the target time corresponding to sample SP1 is time t11 and time t12; similarly, suppose the camera pose transformation matrix in sample SP2 represents the relative pose relationship between the camera 202 at time t21 and time t22, and the rigid body pose transformation matrix in sample SP2 represents the relative pose relationship between the rigid body 208 at time t21 and time t22, that is, the target time corresponding to sample SP2 is time t21 and time t22; and suppose the camera pose transformation matrix in sample SP3 represents the relative pose relationship between the camera 202 at time t31 and time t32, and the rigid body pose transformation matrix in sample SP3 represents the relative pose relationship between the rigid body 208 at time t31 and time t32, that is, the target time corresponding to sample SP3 is time t31 and time t32. Then, the images of the positioning points collected at time t11, time t12, time t21, time t22, time t31, and time t32 can be used as target images, and the optimized hand-eye transformation relationship is determined based on the position information of the positioning points in these target images. The above target images can all be obtained by the motion capture system collecting images of multiple positioning points.
[0054] In some embodiments, the target image includes images acquired when the rigid body 208 is in multiple motion modes. Among them, different motion modes may have different speeds, accelerations, and / or motion manners, and the motion manners include but are not limited to rotation and translation. For example, the target images acquired at the above-mentioned t11 and t12 moments may be images acquired during the uniform translation of the rigid body 208 at the first motion speed, the target images acquired at the above-mentioned t21 and t22 moments may be images acquired during the accelerated translation of the rigid body 208 at the second motion speed, and the target images acquired at the above-mentioned t31 and t32 moments may be images acquired during the rotation of the rigid body 208 at the third motion speed. In this way, the diversity of data can be improved, thereby improving the accuracy and robustness of the calibration result.
[0055] In some embodiments, the initial optimized hand-eye transformation relationship can be determined first, and then the initial optimized hand-eye transformation relationship can be used as the initial value in the optimization process, and a certain optimization method can be adopted for optimization to obtain the optimized hand-eye transformation relationship. Specifically, the position information of the positioning points in the acquired target image can be passed as input parameters to the hand-eye calibration function in the robot vision library, so that the hand-eye calibration function can solve the mathematical model of the hand-eye transformation relationship based on the input parameters to obtain the initial optimized hand-eye transformation relationship, and then perform iterative optimization on the initial optimized hand-eye transformation relationship based on the preset optimization step size to obtain the optimized hand-eye transformation relationship. Among them, the robot vision library includes but is not limited to libraries such as OpenCV, PCL (Point Cloud Library), and Kalibr. Taking OpenCV as an example, the hand-eye calibration function can be the calibrateRobotWorldHandEye() function therein. This function can obtain a matrix X with a relatively low accuracy (i.e., the initial optimized hand-eye transformation relationship) through singular value decomposition, reducing the influence of the initial value on the final optimization result.
[0056] During the optimization process, a fixed or dynamically changing optimization step size can be adopted. In the example where the optimization step size changes dynamically, as an implementation method, the optimization step size adopted for any iteration optimization of the initial optimized hand-eye transformation relationship can be inversely correlated with the accuracy of the initial optimized hand-eye transformation relationship at that iteration. That is, a larger optimization step size can be adopted at the beginning of the iterative optimization to improve the optimization efficiency; as the number of iterations increases, the accuracy of the initial optimized hand-eye transformation relationship of the iteration also gradually increases. At this time, the optimization step size can be gradually reduced to avoid crossing the optimal value and improve the optimization accuracy.
[0057] In some embodiments, an optimization function may be established based on a first transformation matrix between a world coordinate system and a rigid body coordinate system, a second transformation matrix between a calibration object coordinate system and a camera coordinate system, a third transformation matrix between the world coordinate system and the calibration object coordinate system, and an initial optimized hand-eye transformation relationship, and the initial optimized hand-eye transformation relationship and the third transformation matrix in the optimization function may be iteratively optimized based on a preset optimization step size and a preset optimization target to obtain an optimized hand-eye transformation relationship. Among them, the calibration object is used to determine the camera pose transformation matrix in multiple samples.
[0058] Figure 3 FIG. shows a schematic diagram of multiple coordinate systems involved in an embodiment of the present application. The coordinates of each positioning point on the rigid body 208 in the world coordinate system can be determined in advance, and the coordinates of each positioning point on the rigid body 208 in the rigid body coordinate system can be detected by an action capture system. Based on the coordinates of each positioning point in the world coordinate system and the coordinates of each positioning point in the rigid body coordinate system, the first transformation matrix between the world coordinate system and the rigid body coordinate system can be determined.
[0059] The calibration object can be displayed through the display screen 204 of the virtual shooting system 20. Therefore, the calibration object coordinate system can be the display screen coordinate system corresponding to the display screen 204. The coordinates of the feature points on the calibration object in the calibration object coordinate system can be obtained, and the coordinates of the feature points on the calibration object in the camera coordinate system can also be determined according to the image of the calibration object collected by the camera 202. According to the coordinates of the feature points in the calibration object coordinate system and the coordinates of the feature points in the camera coordinate system, the second transformation matrix between the calibration object coordinate system and the camera coordinate system can be determined. In addition, the first transformation matrix between the world coordinate system and the rigid body coordinate system can be determined based on the third transformation matrix between the world coordinate system and the calibration object coordinate system, the second transformation matrix between the calibration object coordinate system and the camera coordinate system, and the transformation matrix between the camera coordinate system and the rigid body coordinate system (i.e., the initial optimized hand-eye transformation relationship). That is, the above first transformation matrix, second transformation matrix, third transformation matrix, and initial optimized hand-eye transformation relationship satisfy the following conditions:
[0060]
[0061] is the transformation matrix from the world coordinate system to the rigid body coordinate system (the first transformation matrix);
[0062] is the transformation matrix from the camera coordinate system to the rigid body coordinate system (the initial optimized hand-eye transformation relationship);
[0063] is the transformation matrix from the calibration object coordinate system to the camera coordinate system (the second transformation matrix);
[0064] The transformation matrix from the world coordinate system to the calibration object coordinate system (the third transformation matrix).
[0065] Among them, Satisfy:
[0066]
[0067] P d is the coordinate of the positioning point in the rigid body coordinate system, and P a is the coordinate of the positioning point in the world coordinate system.
[0068] An optimization function can be established according to the above formula, and the initial hand-eye transformation relationship and the third transformation matrix can be jointly optimized. The optimized result of the initial hand-eye transformation relationship is the optimized hand-eye transformation relationship.
[0069] In some embodiments, the original coordinates of the positioning point in the rigid body coordinate system can be obtained, and through the matrix transformation relationship between the above first transformation matrix, second transformation matrix, third transformation matrix and the initial optimized hand-eye transformation relationship, the coordinates of the positioning point in the world coordinate system are converted into the coordinates of the positioning point in the rigid body coordinate system. The distance between the original coordinates of the positioning point in the rigid body coordinate system and the coordinates of the positioning point in the rigid body coordinate system obtained through coordinate transformation is used as the optimization target, and optimization is carried out through optimization methods such as least squares optimization, and finally the optimized hand-eye transformation relationship is obtained.
[0070] The following combines Figure 4 to illustrate the overall process of the embodiments of the present application by way of example.
[0071] The present application first screens the data. The core idea is to first screen the positions of the positioning points containing noise to eliminate the positions of the positioning points with large errors, and perform hand-eye calibration based on the relatively pure positions of the positioning points. This process relies on the formula AX = XB to process two sets of key data: A (representing the pose transformation matrix between different moments of camera 104) and B (representing the pose transformation matrix between corresponding moments of the rigid body fixedly connected to camera 202). By effectively filtering these two types of data, the aim is to obtain a more accurate hand-eye relationship calibration result. The specific implementation steps are as follows:
[0072] (1) Random sampling: Select M (set here as N - 3) samples from N samples (each pair of A - B constitutes a sample) as a subset, where each subset consists of a randomly selected pair of A - B;
[0073] (2) Estimate the initial hand-eye transformation relationship: Use the selected M samples to estimate the initial hand-eye transformation relationship through the given formula AX = XB;
[0074] (3) Inlier recognition: Substitute the initial hand-eye transformation relationship obtained in the above step into the formula Rn = A n X - XB n Calculate the residual R value for all samples. Subsequently, set a fixed threshold to distinguish which samples can be considered "inliers", that is, samples with a residual less than or equal to this threshold. Count the proportion of these inliers in the total number of samples, and use this proportion as a quantitative indicator of the data quality in the current iteration round;
[0075] (4) Optimal dataset selection: Repeat steps 1 to 3 several times (denoted as N'), and each iteration will generate a new initial hand - eye transformation relationship and its corresponding score. Finally, select the group of samples with the highest score, consider the quality of this group of samples to be high, and the images corresponding to this group of samples are suitable for further calculation of a more accurate optimized hand - eye transformation relationship.
[0076] Through the above - mentioned method, this application can not only effectively reduce the influence of outliers on the final calibration accuracy, but also improve the robustness of the entire system to external interference factors.
[0077] This application has the following technical effects:
[0078] (1) In each iteration, only consider a part of the data for inlier recognition and data filtering, which can screen the data containing noise, eliminate data points with large errors, reduce the influence of noise and outliers on the final calibration accuracy, improve the robustness of the system, and have stronger fault - tolerance ability.
[0079] (2) Before optimization, determine the initial optimized hand - eye transformation relationship based on the hand - eye calibration function in the robot vision tool library, which can improve the rationality of the optimization initial value and reduce the influence of sensitivity to the initial value.
[0080] (3) Through reasonable initialization and step - size control, the iteration process can be accelerated and the time consumption can be reduced.
[0081] (4) By adding more positioning points or different motion patterns, the robustness and accuracy of the calibration result can be improved.
[0082] (5) By taking the distance between the original coordinates of the positioning points in the rigid - body coordinate system and the coordinates of the positioning points in the rigid - body coordinate system obtained through coordinate transformation (i.e., the spatial registration error) as the optimization target, the error accumulation in the intermediate steps is reduced, the calibration accuracy can be improved, the calculation efficiency can be increased, and it is applicable to real - time application scenarios.
[0083] (6) This application does not need to use the camera internal parameter matrix during the optimization process, reduces the dependence of the system on image quality and feature - point detection, and enhances the robustness and applicability of the system.
[0084] (7) The present application obtains the coordinates of the positioning points on the rigid body through an action capture system, without relying on specific calibration targets (such as checkerboards), so it has better adaptability and flexibility and can be applied to various scenarios.
[0085] As Figure 5 shown, the present application also provides a hand-eye calibration transpose for calibrating the hand-eye conversion relationship between a camera and a rigid body fixedly connected to the camera. A plurality of positioning points are distributed around the rigid body; the transpose includes:
[0086] A first acquisition module 302 for acquiring a plurality of samples. Any one sample includes a camera pose conversion matrix and a rigid body pose conversion matrix. The camera pose conversion matrix is used to represent the relative pose relationship of the camera between two moments, and the rigid body pose conversion matrix is used to represent the relative pose relationship of the rigid body between the two moments; the rigid body pose conversion matrix is determined based on the position information of the plurality of positioning points;
[0087] A first determination module 304 for determining a plurality of initial hand-eye conversion relationships and their qualities. Among them, any one initial hand-eye conversion relationship and its quality are determined based on the following method: randomly select some samples from the plurality of samples, determine the initial hand-eye conversion relationship based on each selected sample, and determine the quality of the initial hand-eye conversion relationship based on the matching degree between each selected sample and the initial hand-eye conversion relationship;
[0088] A second determination module 306 for determining the initial hand-eye conversion relationship with the highest quality and determining all the samples used to determine the initial hand-eye conversion relationship with the highest quality as target samples;
[0089] A second acquisition module 308 for acquiring the target images corresponding to each target sample respectively and determining the optimized hand-eye conversion relationship based on the position information of the positioning points in the acquired target images; wherein, the target image corresponding to the target sample includes the images of the plurality of positioning points collected at the target moment corresponding to the target sample.
[0090] The functions or modules included in the device provided by the embodiments of the present application can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0091] The present application also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the methods of any embodiment of the present application are implemented.
[0092] Figure 6FIG. 0 shows a more specific schematic diagram of the hardware structure of a computer device provided by an embodiment of the present application. The device may include: a processor 402, a memory 404, an input / output interface 406, a communication interface 408, and a bus 410. Among them, the processor 402, the memory 404, the input / output interface 406, and the communication interface 408 are communicatively connected to each other inside the device through the bus 410.
[0093] The processor 402 may be implemented in a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application. The processor 402 may further include a graphics card, and the graphics card may be an Nvidia titan X graphics card or a 1080Ti graphics card, etc.
[0094] The memory 404 may be implemented in the form of a read-only memory (ROM), a random access memory (RAM), a static storage device, a dynamic storage device, etc. The memory 404 may store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of the present application through software or firmware, the relevant program codes are stored in the memory 404 and are called and executed by the processor 402.
[0095] The input / output interface 406 is used to connect to an input / output module to implement information input and output. The input / output module may be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.
[0096] The communication interface 408 is used to connect to a communication module (not shown in the figure) to implement communication interaction between this device and other devices. Among them, the communication module may implement communication in a wired manner (such as USB, network cable, etc.) or in a wireless manner (such as mobile network, WIFI, Bluetooth, etc.).
[0097] The bus 410 includes a path for transmitting information between various components of the device (such as the processor 402, the memory 404, the input / output interface 406, and the communication interface 408).
[0098] It should be noted that although the above device only shows the processor 402, the memory 404, the input / output interface 406, the communication interface 408, and the bus 410, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of the present application, and does not necessarily include all the components shown in the figure.
[0099] An embodiment of the present application provides a computer program product, including a computer program, which when executed by a processor implements the method described in any embodiment of the present application.
[0100] The present application also provides a computer-readable storage medium, on which computer instructions are stored, and when the computer instructions are executed by a processor, the addressing method described in any embodiment of the present application is implemented. Among them, the computer-readable storage medium may be a phase change memory (PRAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), other types of random access memory (RAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory or other memory technologies, a CD-ROM, a digital versatile disc (DVD) or other optical storage, a magnetic cassette tape, a magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0101] From the description of the above embodiments, those skilled in the art can clearly understand that the embodiments of the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments of the present application.
[0102] The systems, devices, modules, or units described in the above embodiments can be specifically implemented by a computer device or an entity, or by a product with a certain function. A typical implementation device is a computer, and the specific form of the computer can be a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email transceiver device, a game console, a tablet computer, a wearable device, or a combination of any several of these devices.
[0103] Each embodiment in this application is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the corresponding descriptions in the method embodiments. The device embodiments described above are only illustrative. The modules described as separate components may or may not be physically separated. When implementing the solutions of the embodiments of this application, the functions of each module can be realized in one or more software and / or hardware. It is also possible to select some or all of the modules according to actual needs to achieve the purpose of the solutions of this embodiment. A person of ordinary skill in the art can understand and implement it without creative efforts.
[0104] The above are only specific implementation manners of the embodiments of this application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the embodiments of this application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the embodiments of this application.
Claims
1. A hand-eye calibration method for calibrating the hand-eye conversion relationship between a camera and a rigid body fixedly connected to the camera, wherein a plurality of positioning points are distributed around the rigid body; the method comprises: Acquire multiple samples, any one of which includes a camera pose conversion matrix and a rigid body pose conversion matrix, wherein the camera pose conversion matrix is used to represent the relative pose relationship of the camera between two moments, and the rigid body pose conversion matrix is used to represent the relative pose relationship of the rigid body between the two moments; the rigid body pose conversion matrix is determined based on the position information of the multiple positioning points; Determine a plurality of initial hand-eye conversion relationships and their qualities, wherein any one of the initial hand-eye conversion relationships and its quality is determined based on the following method: randomly selecting a portion of samples from the plurality of samples, determining an initial hand-eye conversion relationship based on each of the selected samples, and determining the quality of the initial hand-eye conversion relationship based on a matching degree between each of the selected samples and the initial hand-eye conversion relationship; Determine an initial hand-eye conversion relationship with the highest quality, and determine each sample used to determine the initial hand-eye conversion relationship with the highest quality as a target sample; The target images corresponding to the target samples are obtained, and the optimized hand-eye conversion relationship is determined based on the position information of the positioning points in the obtained target images; wherein the target images corresponding to the target samples include images of the multiple positioning points collected at the target time corresponding to the target samples.
2. The method according to claim 1, further comprising: Determining the residual of each selected sample based on the initial hand-eye conversion relationship; Determine the ratio between the number of samples whose residuals are greater than a preset threshold and the total number of selected samples; The matching degree between each selected sample and the initial hand-eye conversion relationship is determined based on the ratio.
3. The method according to claim 2, wherein determining the residual of each selected sample based on the initial hand-eye conversion relationship comprises: For any selected sample, determine a first product of a camera pose conversion matrix in the sample and the initial hand-eye conversion relationship, and determine a second product of the initial hand-eye conversion relationship and a rigid body pose conversion matrix in the sample; A residual for the sample is determined based on a difference between the first product and the second product.
4. The method according to claim 1, wherein the plurality of positioning points comprises at least one of the following: The positioning points provided on the rigid body; Bind to the anchor point on the rigid body; Anchor points on other rigid bodies to which the rigid body is connected; and / or The target image includes images collected when the rigid body is in multiple motion modes.
5. The method according to claim 1, wherein the step of determining and optimizing the hand-eye conversion relationship based on the position information of the positioning point in the acquired target image comprises: The position information of the positioning point in the acquired target image is passed as an input parameter to the hand-eye calibration function in the robot vision tool library, so that the hand-eye calibration function solves the mathematical model of the hand-eye conversion relationship based on the input parameter to obtain an initial optimized hand-eye conversion relationship; The initial optimized hand-eye conversion relationship is iteratively optimized based on a preset optimization step size to obtain an optimized hand-eye conversion relationship.
6. The method according to claim 5, wherein the iterative optimization of the initial optimized hand-eye conversion relationship based on a preset optimization step length to obtain the optimized hand-eye conversion relationship comprises: Establishing an optimization function based on a first transformation matrix between a world coordinate system and a rigid body coordinate system, a second transformation matrix between a calibration object coordinate system and a camera coordinate system, a third transformation matrix between a world coordinate system and a calibration object coordinate system, and the initial optimized hand-eye conversion relationship; wherein the calibration object is used to determine the camera pose conversion matrix in the plurality of samples; The initial optimized hand-eye conversion relationship and the third transformation matrix in the optimization function are iteratively optimized based on a preset optimization step size and a preset optimization target to obtain an optimized hand-eye conversion relationship.
7. According to the method of claim 5, the optimization step size used in any iterative optimization of the initial optimized hand-eye conversion relationship is inversely correlated with the accuracy of the initial optimized hand-eye conversion relationship at that iteration.
8. According to the method of claim 1, the camera pose transformation matrix in any sample is determined based on the following method: Acquire a first calibration object image obtained by the camera capturing an image of the calibration object at a first moment and a second calibration object image obtained by the camera capturing an image of the calibration object at a second moment different from the first moment; Detecting position information of a first feature point on the calibration object from the first calibration object image, and detecting position information of a second feature point on the calibration object corresponding to the first feature point from the second calibration object image; A camera pose conversion matrix between the first moment and the second moment is determined based on the position information of the first feature point and the position information of the second feature point.
9. The method according to claim 1, wherein the rigid body posture transformation matrix in any sample is determined based on the following method: Acquire a first positioning point image obtained by acquiring images of the plurality of positioning points at a first moment and a second positioning point image obtained by acquiring images of the plurality of positioning points at a second position different from the first position; Detecting first position information of the plurality of positioning points from the first positioning point image, and detecting second position information of the plurality of positioning points from the second positioning point image; A rigid body posture conversion matrix between the first moment and the second moment is determined based on the first position information and the second position information.
10. According to the method of claim 1, the multiple positioning points include points on the rigid body that can emit or reflect infrared light, and the target images corresponding to the respective target samples are acquired by an infrared camera in a motion capture system.
11. A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1 to 10.
12. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 10 when executing the computer program.
13. A computer program product, comprising a computer program, which implements the method according to any one of claims 1 to 10 when executed by a processor.
Citation Information
Patent Citations
Improved hand-eye calibration algorithm based on screening and least square method
CN107053177A
Robot hand-eye calibrating posture selection method and device, robot system and medium
CN112743546A
Sensor pose conversion relation determination method, device and equipment and medium
CN113759384A
Mechanical arm hand-eye calibration method and system based on improved SVD algorithm
CN114872039A
Robot hand-eye calibration method based on matrix solution
CN115091456A