A Calibration Method for a Lightweight ToF Sensor and an RGB Image Sensor
By using EM algorithms and nonlinear optimization methods in a calibration environment with structural information, the combined calibration of lightweight ToF sensors and RGB image sensors is realized, solving the problem of sensor data alignment and improving perception accuracy and resolution.
Patent Information
- Application Number
- CN202211067550.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-01
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-09-01
AI Technical Summary
Due to spatial uncertainty in measuring depth, the lightweight ToF sensor is difficult to coordinate with the RGB image sensor, affecting the perception accuracy and resolution.
In a calibration environment with structural information, using the continuity of the planar structure, the pixel position and scene plane equation of the depth point are iteratively optimized through the EM algorithm, which is converted into the alignment problem between the point cloud map and the plan map. The nonlinear optimization method is used to solve it to realize the joint calibration of the RGB image sensor and the lightweight ToF sensor.
It effectively solves the joint calibration problem of lightweight ToF sensors and RGB image sensors, realizes data alignment of the two sensors, improves perception accuracy and resolution, and is suitable for mixed reality and robots fields.
Smart Images

Figure CN115482293B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a calibration method for a lightweight ToF sensor and an RGB image sensor. Background Art
[0002] Fields such as driverless, intelligent robots, and mixed reality need to achieve precise perception of the environment. Generally, multiple sensors are used in combination during perception to improve the perception accuracy. However, the primary prerequisite for using multiple sensors in combination is the joint calibration of each sensor to ensure the coordinate unity of each sensor in space.
[0003] The lightweight ToF sensor is a relatively new type of depth sensor, such as VL53L0, VL53L1, VL53L3, and VL53L5 under the FlightSense series of STMicroelectronics. These sensors are designed for applications such as simple gesture recognition, autofocus, and simple obstacle detection, and have the characteristics of low power consumption and low cost. Compared with general ToF (Time-of-Flight) sensors, the main difference of these lightweight ToF sensors is that due to the lightweight design concept, their resolution is extremely low (generally not exceeding 10x10), while general ToF sensors usually have a resolution of tens of thousands or more. At the same time, each pixel of the lightweight ToF sensor measures the depth distribution of a region of the scene, so there is a large uncertainty, while each pixel of a general ToF sensor measures the depth value corresponding to an accurate three-dimensional position.
[0004] Although such lightweight ToF sensors do not have the ability to measure high-resolution and dense scene depth information, compared with structured light-based depth sensors or general ToF (Time-of-Flight) sensors, their price and power consumption are at least one to two orders of magnitude lower. Therefore, they are widely used in some mobile scenarios with relatively strict requirements for power consumption and cost. Additionally, if this sensor is combined with an image sensor with a higher resolution to form a multi-sensor system, the perception accuracy and resolution can be further improved. At the same time, since the image sensor also has a similarly low level of power consumption and price, the power consumption and cost of this multi-sensor system will not increase significantly. Of course, to form this multi-sensor system, the joint calibration problem of this lightweight ToF sensor and the image sensor needs to be solved first. In view of the characteristics of the lightweight ToF sensor, the present invention proposes a calibration method for a lightweight ToF sensor and an RGB image sensor, which solves the problem that it is inconvenient to jointly calibrate the lightweight ToF with the RGB image sensor due to the spatial uncertainty of the measured depth. Summary of the Invention
[0005] The present invention proposes a calibration method for a lightweight ToF sensor and an RGB image sensor. When the lightweight ToF sensor measures, it cannot obtain the depth information of a precise point, but instead returns the average value and variance of the depth within a region, which makes the traditional calibration method inapplicable during the joint calibration of the RGB image sensor and the lightweight ToF sensor. In view of this characteristic of the lightweight ToF sensor, the present invention appropriately selects a calibration environment with structural information and solves this calibration problem by utilizing the continuity of the planar structure. The method pastes calibration paper on more than three non-parallel planar structures, and uses the RGB image sensor and the lightweight ToF sensor to collect two-dimensional image information and depth information respectively. The SLAM system is used to generate the pose of each image and the corresponding map point cloud from the two-dimensional images collected by the RGB image sensor, and the scene plane is estimated from the depth distribution data measured by the lightweight ToF sensor through the EM algorithm. In this way, the calibration problem is transformed into the alignment problem between the point cloud map and the plane map, and then this problem is solved by using the method of nonlinear optimization, thus realizing the joint calibration of the RGB image sensor and the lightweight ToF sensor.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] The present invention provides a calibration method for a lightweight ToF sensor and an RGB image sensor, which includes the following steps:
[0008] Step 1: Keep the relative fixed state of the RGB image sensor and the lightweight ToF sensor, and call the fixed whole the device to be calibrated;
[0009] Step 2: Select a suitable calibration scene, requiring that there are at least three non-parallel planes in the scene; collect the RGB image data and depth distribution data of each plane of the scene through the RGB image sensor and the lightweight ToF sensor of the device to be calibrated respectively;
[0010] Step 3: For the data recorded by the RGB image sensor, use the visual SLAM system to process the pictures taken by the RGB image sensor, generate the pose of each picture in the world coordinate system, and perform point cloud reconstruction on the scene; since the calibration environment is at least three non-parallel planes, extract these planes from the reconstructed point cloud, and classify all the valid points in the point cloud according to the plane to which they belong;
[0011] Step 4: For the depth distribution information collected by the lightweight ToF sensor, first, according to its field of view range and the distribution of the sensing area, back-project the depth values based on the assumption that the measured depth mean is located at the center of the measurement area to generate the corresponding point cloud and fit the scene plane equation in the current frame coordinate system; then use the EM algorithm to iteratively optimize the accurate position of the measured depth mean in the measurement area and the scene plane equation in the current frame coordinate system; finally, save the scene plane information corresponding to each frame.
[0012] Step 5: Construct an optimization problem to solve a set of optimal rotation and translation matrices R and t, such that after aligning each frame of the plane generated by the lightweight ToF sensor with the point cloud generated by the RGB image sensor using R and t, the sum of the residuals of the points in the point cloud to the corresponding plane is minimized. The obtained rotation and translation matrices R and t are the calibration results to be calculated.
[0013] Existing calibration techniques need to assume the pixel positions of known depth values, so they are not applicable to the calibration of lightweight ToF sensors and RGB image sensors. In view of the fact that the lightweight ToF sensor measures the average depth and variance within the area rather than the depth information of a certain precise point, the present invention appropriately selects a calibration environment with structural information, and uses the continuity of the plane structure to iteratively optimize the pixel positions of the depth points and the scene plane equation through the EM algorithm, thus solving this calibration problem. The calibration method of the present invention can be used for hardware settings equipped with both lightweight ToF sensors and RGB image sensors to align the depth distribution data measured by the lightweight ToF sensor with the RGB image data collected by the RGB image sensor. This device has a relatively low price and power consumption. After aligning the measurement data of the two sensors, it combines the high resolution of the RGB image sensor and the depth measurement information of the lightweight ToF sensor, and can be applied to fields such as mixed reality and robotics. Description of the Drawings
[0014] Figure 1 is the device diagram used in the embodiment of the present invention;
[0015] Figure 2 is the flow chart of the joint calibration method of the present invention;
[0016] Figure 3 is the schematic diagram of the preparation state of the joint calibration of the RGB image sensor and the lightweight ToF sensor in the present invention;
[0017] Figure 4 is the flow chart of using the EM algorithm to extract the corresponding main plane in the observation area of the lightweight ToF sensor in the present invention. Detailed Embodiments
[0018] A calibration method for an RGB image sensor and a lightweight ToF sensor of the present invention. First, the RGB image sensor and the lightweight ToF sensor are used to collect two-dimensional image information and depth distribution information respectively. Then, the visual SLAM system is used to generate the pose of each image and the corresponding map point cloud from the two-dimensional images collected by the RGB image sensor, and the scene plane is estimated from the depth distribution data measured by the lightweight ToF sensor through the EM algorithm. Finally, the rotation and translation matrix between the lightweight ToF sensor and the RGB image sensor is calculated by aligning the point cloud generated by the RGB image sensor and the scene plane generated by the lightweight ToF sensor, so as to realize the joint calibration of the data of the RGB image sensor and the lightweight ToF sensor.
[0019] The calibration method proposed by the present invention specifically includes the following steps:
[0020] Step 1: Keep the relative fixed state of the RGB image sensor and the lightweight ToF sensor, and the fixed whole is called the device to be calibrated.
[0021] Step 2: Select a suitable calibration scene, and it is required that there are at least three non-parallel planes in the scene; collect the depth distribution data and RGB image data of each plane of the scene through the RGB image sensor and the lightweight ToF sensor of the device to be calibrated respectively.
[0022] Step 3: For the data recorded by the RGB image sensor, use the visual SLAM system to process the pictures taken by the RGB image sensor, generate the pose of each picture in the world coordinate system, and perform point cloud reconstruction on the scene; since the calibration environment is at least three non-parallel planes, extract these planes from the reconstructed point cloud, and classify all valid points in the point cloud according to the plane to which they belong.
[0023] Step 4: For the depth distribution information collected by the lightweight ToF sensor, first, according to its field of view range and the distribution of the sensing area, perform back-projection on the depth value according to the assumption that the mean value of the measured depth is located at the center of the measurement area, generate the corresponding point cloud and fit the scene plane equation in the current frame coordinate system; then use the EM algorithm to iteratively optimize the accurate position of the mean value of the measured depth in the measurement area and the scene plane equation in the current frame coordinate system; finally, save the scene plane information corresponding to each frame.
[0024] Step 5: Construct an optimization problem, and solve a set of optimal rotation and translation matrices R and t, so that after aligning each frame of the plane generated by the lightweight ToF sensor with the point cloud generated by the RGB image sensor using R and t, the sum of the residuals of the points in the point cloud to the corresponding plane is minimized, and the obtained rotation and translation matrices R and t are the calibration results to be calculated.
[0025] As a preferred solution of the present invention, in step 2, three or more non-parallel planes can be natural scenes or artificial scenes.
[0026] As a preferred solution of the present invention, in step 3, first use the RANSAC algorithm to sequentially extract the planes in the point cloud. The specific method is as follows: Use RANSAC to fit a plane in the scene, numbered 0, record the plane equation, and then remove all inliers of the fitted plane from the point cloud. Repeat the steps of fitting the plane for the remaining points, gradually fitting the other remaining planes in the scene, and obtaining the plane equations of each plane. Assume that there are 3 non-parallel planes in the scene. Then, the present invention first fits a plane 0 in the scene, and then gradually fits to obtain plane 1 and plane 2.
[0027] As a preferred solution of the present invention, in step 4, first, according to the assumption that the measured depth mean is located at the center of the measurement area, recover the plane from the depth distribution measured by the lightweight ToF sensor; since this assumption does not always hold, the plane recovered in this step is not very accurate; however, due to the continuity of the plane, it can be known that there is a point in the measured area whose depth is equal to the depth mean measured by the lightweight ToF sensor; therefore, after obtaining the plane by solving, use the EM algorithm to alternately optimize the pixel position of the depth mean in each area and the scene plane equation of the current frame to obtain a more accurate plane equation.
[0028] As a preferred solution of the present invention, in step 4, use the EM algorithm to alternately optimize the pixel position of the depth mean in each area and the scene plane equation, specifically as follows:
[0029] 4.1) Assume that the scene plane parameters of the current frame to be optimized are {n, d}, and the pixel position of the point where the depth in the k-th area is equal to the depth mean measured by ToF is (x k , y k ), then the loss term to be jointly optimized is written as:
[0030]
[0031]
[0032] where Z represents the set of measurement areas of the lightweight ToF sensor, is the boundary range of the k-th measurement area, m k is the depth mean of the k-th measurement area, and K is the internal parameter attribute of the lightweight ToF sensor, which can be calculated from the field of view range and resolution of the sensor;
[0033] Initialize the pixel positions of the depth means of all measurement areas to the center points of the measurement areas;
[0034] 4.2) Back-project these two-dimensional positions according to the depth mean to obtain a number of three-dimensional points, and then fit the corresponding plane;
[0035] 4.3) Adjust the pixel positions of the depth mean in each region to minimize the distance from the back-projected three-dimensional points to the plane fitted in the previous step;
[0036] 4.4) Iteratively run steps 4.2) and 4.3) until convergence; during the iteration, discard the outliers that are too far from the plane.
[0037] As a preferred solution of the present invention, in step 5, first determine the correspondence between the plane captured by each frame of ToF and the plane reconstructed by the RGB image sensor; subsequently, for each frame of data, first transform the point cloud reconstructed by the RGB image sensor from the world coordinate system to the RGB image sensor coordinate system through the frame-by-frame pose obtained in step 3; at this time, the plane fitted from the ToF observation data and the point cloud reconstructed from the RGB image sensor data are in their respective coordinate systems, and the two only differ by the rotation and translation matrix from the RGB image sensor coordinate system to the lightweight ToF sensor coordinate system, and this matrix is the calibration result to be obtained; to obtain this matrix, use the point-to-plane alignment method to construct an optimization problem, that is, find a set of R, t to minimize the following formula
[0038]
[0039] where, R is the rotation matrix from the RGB image sensor coordinate system to the lightweight ToF sensor coordinate system, t is the translation vector from the RGB image sensor coordinate system to the lightweight ToF sensor coordinate system: F represents the set of captured frames, P i represents all the points in the main plane captured by the i-th frame RGB image sensor in the point cloud reconstructed by the visual SLAM system; p is the three-dimensional space coordinate of the point in P i inside, n i and d i are the plane equation coefficients corresponding to the main plane of the measurement area of the i-th frame lightweight ToF sensor, which are the plane normal direction and the offset distance from the plane to the origin respectively; the above formula is iteratively solved by a non-linear optimization method.
[0040] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. The described embodiments are some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0041] Embodiment
[0042] As Figure 1 shown, the present invention constructs a device equipped with both an RGB image sensor and a lightweight ToF sensor to verify the effect of the present invention. In this device, the RGB image sensor uses the color camera in Intel RealSense D435i, and the lightweight ToF sensor uses the VL53L5 model of STMicroelectronics. As Figure 2 shown, the present embodiment realizes the joint calibration of the RGB image sensor and the lightweight ToF sensor, mainly including the following steps:
[0043] Step 1: Keep the RGB image sensor and the lightweight ToF sensor installed relatively fixed. The lightweight ToF sensor adopts the multi-region measurement version with a resolution of 8*8.
[0044] Step 2: Select a suitable calibration environment, which requires that there are three non-parallel planes in the environment. In this embodiment, a corner formed by two walls and the ground in a room without sundries is selected, and calibration paper is pasted on the walls and the ground to increase texture information, as Figure 3 shown. Denote these three planes as Plane A, Plane B, and Plane C respectively. First, hold the device to be calibrated and align it with Plane A, and slowly move the device to be calibrated up, down, left, and right in turn. During the movement, there is also a small change in the pitch angle. When recording the data of Plane A, make the main object in the field of view be Plane A. After the data of Plane A is recorded, smoothly rotate the lens to switch to Plane B and record according to the same requirements as Plane A. After the recording of Plane B is completed, smoothly rotate the lens to switch to Plane C and record according to the same requirements. After the data of Plane A, Plane B, and Plane C are all recorded, the data acquisition work is completed.
[0045] Step 3: For the data recorded by the RGB image sensor, use the visual SLAM system to process the pictures taken by the RGB image sensor, generate the pose of each picture in the world coordinate system, and perform point cloud reconstruction on the scene. In this example, the open-source visual SLAM algorithm ORB-SLAM2 is used to process the two-dimensional images taken by the RGB image sensor. After processing, the pose of each photo and the point cloud map of the overall scene are saved. Since the calibration environment is three non-parallel planes, these three planes can be extracted from the reconstructed point cloud, and all the valid points in the point cloud are divided into three categories according to the plane to which they belong. Here, the RANSAC algorithm is mainly used to extract the planes from the point cloud, and three planes are extracted from the point cloud in turn. The specific steps of the operation are as follows:
[0046] 3.1 Plane extraction uses the segment_plane method in the open source algorithm library open3d. When extracting the plane, the RANSAC parameters are set to: n = 3, distance_threshold = 0.02, and the number of iterations is 1000 times.
[0047] 3.2 Use the algorithm to extract the three planes in the point cloud in sequence. The specific steps are as follows: use RANSAC to fit the first plane in the scene, numbered 0, record the equation of the plane, and then remove all the inner points of the plane fitted in this step in the point cloud map. Then, repeat this step for the remaining points to continue to obtain the equations of plane 1 and plane 2. At this point, the equations of the three planes are extracted.
[0048] 3.3 Mark all the internal points in the point cloud map that belong to a certain plane, and record the correspondence between the point and the plane to which it belongs.
[0049] Step 4: For the depth distribution information collected by the lightweight ToF sensor, first, according to its field of view and the distribution of the sensing area, the depth value is back-projected based on the assumption that the measured depth mean is located at the center of the measurement area, the corresponding point cloud is generated, and the scene plane equation in the current frame coordinate system is fitted; then the EM algorithm is used to iteratively optimize the accurate position of the measured depth mean in the measurement area and the scene plane equation in the current frame coordinate system, and the scene plane equation obtained for each frame is saved. Figure 4 As shown, the specific steps of the operation are:
[0050] 4.1 Generate a projection direction according to the field of view angle range and the number of regional distributions of the lightweight ToF sensor. The projection direction is the direction of a ray from the center of the lightweight ToF sensor to the center of each of its measurement areas, which is the same as the number of receiving areas included in the ToF.
[0051] 4.2 Read and analyze the depth distribution data measured by the lightweight ToF sensor. For each distribution data, project it in the ray direction of the corresponding area, that is, starting from the optical center of the lightweight ToF sensor, move forward along the direction to measure the depth mean distance and then reach the point as the corresponding 3D measurement point of the area. Perform the same processing on all areas in turn, and the depth distribution data of multiple areas measured by the lightweight ToF sensor will be converted into corresponding point clouds.
[0052] 4.3 Use the RANSAC plane fitting algorithm used in step 3 to extract the plane from the point cloud data generated by the lightweight ToF sensor. According to the requirements of the shooting data in step 2, the main body of the environmental content observed by the lightweight ToF sensor at the same time is one of the three non-parallel surfaces in the environment. Therefore, the corresponding plane equation will be extracted after running the plane fitting algorithm.
[0053] 4.4 The depth measured by the lightweight ToF sensor is not the depth of a specific point, but the mean and variance of the depth within a regional range. However, due to the continuity of the plane, it can be known that there is a point within the measured area whose depth is equal to the depth measured by the lightweight ToF sensor. Therefore, after obtaining the initial estimated scene plane in step 4.3, the EM algorithm is continued to alternately optimize the pixel positions of the points whose depth within each area is equal to the mean depth measured by the lightweight ToF sensor and the scene plane equation.
[0054] 4.5 Perform the above operations on each frame of data recorded by the lightweight ToF sensor and store the corresponding plane equations.
[0055] Step 5: Construct an optimization problem to solve a set of optimal rotation and translation matrices R and t, such that after aligning each frame of plane generated by the lightweight ToF sensor with the map generated by the RGB image sensor using R and t, the sum of the residuals from the map points generated by the RGB image sensor to the corresponding planes generated by the lightweight ToF sensor is minimized. The obtained rotation and translation matrices R and t are the calibration results to be calculated. The specific steps of the operation are as follows:
[0056] 5.1 View the calibration data recorded in step 2 and record the timestamps when the RGB image sensor captures the three planes respectively to determine the correspondence between each frame of the plane captured by the lightweight ToF sensor and the plane reconstructed by the RGB image sensor. Here, the timestamp of capturing the plane is recorded together with the serial number of the plane in a text file.
[0057] 5.2 For each frame of data, first transform the point cloud reconstructed by the RGB image sensor from the world coordinate system to the RGB image sensor coordinate system through the per-frame pose obtained in step 3. The implementation method is to left-multiply the homogeneous coordinates of each point in the point cloud by the transformation matrix from the world coordinate system to the RGB image sensor coordinate system, as shown in the following formula: P c = T cw P w . Where P c represents the point in the RGB image sensor coordinates after transformation, P w represents the point in the world coordinate system before transformation, and T cw is the transformation matrix obtained in step 3 from the world coordinate system to the RGB image sensor coordinate system.
[0058] 5.3 After the transformation, the plane fitted by the observation data of the lightweight ToF sensor and the point cloud reconstructed by the RGB image sensor data are in their respective coordinate systems. The only difference between the two is the rotation and translation matrix from the RGB image sensor coordinate system to the lightweight ToF sensor coordinate system. This matrix is the required calibration result. In order to obtain this matrix, an optimization problem is constructed using the point-to-plane alignment method, that is, to obtain a set of R, t so that the following formula obtains the minimum value, where R is the rotation matrix from the RGB image sensor coordinate system to the lightweight ToF sensor coordinate system, and t is the translation vector from the RGB image sensor coordinate system to the lightweight ToF sensor coordinate system:
[0059]
[0060] Where F represents the set of captured frames, P i Represents all points in the main plane captured by the RGB image sensor in the i-th frame in the point cloud reconstructed by the visual SLAM system. i and d i The plane equation coefficients of the main plane corresponding to the measurement area of the lightweight ToF sensor in the i-th frame are the plane normal direction and the offset distance from the plane to the origin. That is, the requirement is:
[0061]
[0062] The above formula can be solved iteratively through nonlinear optimization method to obtain the final calibration result. Figure 1 The device shown can calibrate a rotation matrix with a difference of nearly 90 degrees and a translation matrix of nearly 5 cm between the two sensor coordinate systems through the joint calibration method of the present invention. The average residual error from the map point generated by the RGB image sensor to the plane generated by the corresponding lightweight ToF sensor before calibration is 7.5 cm. After the rotation and translation matrices obtained by calibration are aligned, the residual error is reduced to 1.5 cm.
[0063] The above examples are only specific embodiments of the present invention. Obviously, the present invention is not limited to the above examples, and many variations are possible. All variations that can be directly derived or associated with the contents disclosed by a person skilled in the art should be considered as the protection scope of the present invention.
Claims
1. A calibration method for a lightweight ToF sensor and an RGB image sensor, characterized in that It includes the following steps: Step 1: Keep the RGB image sensor and the lightweight ToF sensor in a relatively fixed state, and call the fixed whole the device to be calibrated; Step 2: Select a suitable calibration scene, requiring that there are at least three non-parallel planes in the scene; respectively collect the RGB image data and depth distribution data of each plane of the scene through the RGB image sensor and the lightweight ToF sensor of the device to be calibrated; Step 3: For the data recorded by the RGB image sensor, use a visual SLAM system to process the pictures taken by the RGB image sensor, generate the pose of each picture in the world coordinate system, and perform point cloud reconstruction on the scene; since the calibration environment is at least three non-parallel planes, extract these planes from the reconstructed point cloud, and classify all valid points in the point cloud according to the planes to which they belong; Step 4: For the depth distribution information collected by the lightweight ToF sensor, first, according to its field of view range and sensing area distribution, perform backprojection on the depth value according to the assumption that the measured depth mean is located at the center of the measurement area, generate the corresponding point cloud and fit the scene plane equation in the current frame coordinate system; then use the EM algorithm to iteratively optimize the accurate position of the measured depth mean in the measurement area and the scene plane equation in the current frame coordinate system; Finally, save the scene plane information corresponding to each frame; Step 5: Construct an optimization problem, solve a set of optimal rotation and translation matrices R and t, so that after aligning each frame of the plane generated by the lightweight ToF sensor with the point cloud generated by the RGB image sensor using R and t, the sum of the residuals of the points in the point cloud to the corresponding plane is minimized, and the obtained rotation and translation matrices R and t are the calibration results to be calculated.
2. The calibration method of the lightweight ToF sensor and the RGB image sensor according to claim 1, characterized in that, The three or more non-parallel planes in Step 2 are natural scenes or artificial scenes.
3. The calibration method of the lightweight ToF sensor and the RGB image sensor according to claim 1, characterized in that, In Step 3, first use the RANSAC algorithm to sequentially extract the planes in the point cloud. The specific method is as follows: use RANSAC to fit a plane in the scene, numbered 0, record the plane equation, and then remove all inliers of the plane fitted in this step from the point cloud. Repeat the step of fitting the plane for the remaining points, gradually fit the other remaining planes in the scene, and obtain the plane equations of each plane.
4. The calibration method of the lightweight ToF sensor and the RGB image sensor according to claim 1, wherein In Step 4, first, according to the assumption that the measured depth mean is located at the center of the measurement area, restore the plane from the depth distribution measured by the lightweight ToF sensor; since this assumption does not always hold, the plane restored in this step is not very accurate; however, due to the continuity of the plane, it can be known that there is a point in the measured area whose depth is equal to the depth mean measured by the lightweight ToF sensor; therefore, after obtaining the plane, use the EM algorithm to alternately optimize the pixel position of the depth mean in each area and the scene plane equation in the current frame to obtain a more accurate plane equation.
5. The calibration method of the lightweight ToF sensor and the RGB image sensor according to claim 4, characterized in that, In Step 4, use the EM algorithm to alternately optimize the pixel position of the depth mean in each area and the scene plane equation, specifically: 4.1) Assume that the scene plane parameters of the current frame to be optimized are {n, d}, and the position of the pixel point with a depth equal to the mean value of the ToF measurement in the k-th region is (x k , y k ). Then, the loss term to be jointly optimized is written as follows: where $Z$ represents the set of measurement regions of the lightweight ToF sensor, is the boundary range of the $k$-th measurement region, $m$ k is the average depth of the $k$-th measurement region, and $K$ is the internal parameter property of the lightweight ToF sensor, which is calculated from the field of view and resolution of the sensor; Initialize the pixel positions of the depth means of all measurement areas to the center points of the measurement areas; 4.2) Perform backprojection on these two-dimensional positions according to the depth mean to obtain a number of three-dimensional points, and then fit the corresponding plane; 4.3) Adjust the pixel positions of the depth means in each region to minimize the distance from the back-projected 3D points to the plane fitted in the previous step. 4.4) Iteratively run steps 4.2) and 4.3) until convergence; during the iteration, discard the outliers that are too far from the plane.
6. The calibration method of the lightweight ToF sensor and the RGB image sensor according to claim 1, characterized in that In step 5, first determine the correspondence between the plane captured by each frame of ToF and the plane reconstructed by the RGB image sensor; subsequently, for each frame of data, first transform the point cloud reconstructed by the RGB image sensor from the world coordinate system to the RGB image sensor coordinate system through the frame-by-frame pose obtained in step 3; at this time, the plane fitted from the ToF observation data and the point cloud reconstructed from the RGB image sensor data are in their respective coordinate systems, and the two only differ by the rotation and translation matrix from the RGB image sensor coordinate system to the lightweight ToF sensor coordinate system, and this matrix is the required calibration result. To obtain this matrix, use the point-to-plane alignment method to construct an optimization problem, that is, find a set of R, t to minimize the following formula Among them, R is the rotation matrix from the RGB image sensor coordinate system to the lightweight ToF sensor coordinate system, and t is the translation vector from the RGB image sensor coordinate system to the lightweight ToF sensor coordinate system: F represents the set of frames captured, and P i represents all the points in the principal plane captured by the RGB image sensor in the i-th frame of the point cloud reconstructed by the visual SLAM system; p is the three-dimensional spatial coordinate of the point in P i , and n i and d i are the plane equation coefficients corresponding to the principal plane of the measurement area of the i-th frame lightweight ToF sensor, which are the plane normal direction and the offset distance from the plane to the origin, respectively; the above formula is iteratively solved by a non-linear optimization method.
Citation Information
Patent Citations
Mobile robot image visual positioning method in dynamic environment
CN108460779A
A method for accurately registering RGB and depth information
CN109816731A