Method of performing self-calibration and system of performing self-calibration

By combining 2D and 3D difference data optimization methods with a pseudo-target flow estimator and a learned neural network, the gradient optimization problem of the camera during monotonic motion is solved, realizing automatic calibration and stable self-calibration of camera parameters, which is suitable for autonomous driving and robot intelligence.

CN122176060APending Publication Date: 2026-06-09TOYOTA JIDOSHA KK

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TOYOTA JIDOSHA KK
Filing Date
2025-12-03
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Existing technologies cannot perform gradient optimization calculations when the camera is moving monotonically, and require pre-acquiring video of the assumed environment, which prevents the pose estimator from converging to the global optimum. Furthermore, the depth estimator and pose estimator need to share in-domain data of the same scale for self-supervised learning, making it difficult to adapt to rapid motion.

Method used

By using a combined optimization method with 2D and 3D differential data, a self-calibration system is optimized using camera parameters. This is combined with a pseudo-target flow estimator and a learned neural network for optical flow separation, thereby achieving automatic correction of camera parameters.

Benefits of technology

It enables automatic camera recalibration in a short time, adapts to rapid movements, reduces sensitivity to initial values, and improves the stability and efficiency of self-calibration, making it suitable for autonomous driving and intelligent robotics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122176060A_ABST
    Figure CN122176060A_ABST
Patent Text Reader

Abstract

The present invention provides a method of performing self-calibration using common optimization of 2D difference data and 3D difference data. The present invention provides a method of performing self-calibration, wherein both 2D difference data, which is data composed of difference components of at least two images included in a video, and 3D difference data, which is data composed of difference components of information in which depth components are added to at least two images included in the video, are used as variables to perform optimization of imaging parameters to perform self-calibration in such a manner that a difference between a target value and a current value of a function of variables including the 2D difference data and the 3D difference data becomes smaller.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method and system for self-calibration. Background Technology

[0002] Patent document 1 discloses a self-calibration method for at least one camera.

[0003] Patent Document 1: U.S. Patent Application Publication No. 2024 / 0070917 Summary of the Invention

[0004] However, as discussed later, there is a problem where the gradient used for optimization computation becomes 0 during monotonic camera motion, making gradient descent-based optimization impossible and hindering learning. Furthermore, to handle rapid camera motion, a pose estimator (PoseNet) is required through self-supervised learning, necessitating "pre-acquiring video assuming the usage environment." Since it is a non-convex optimization, convergence to the global optimum cannot be guaranteed if the initial values ​​are inappropriate. On the other hand, the solution proposed by Kanai et al. requires the pose estimator to share the same scale as the depth estimator used in conjunction with it [Reference 1 below]. Consequently, self-supervised learning using in-domain data to align the scales of the depth and pose estimators is required [Reference 2 below].

[0005] Therefore, one of the objectives of this invention is to provide a self-calibration method that utilizes both 2D and 3D differential data for optimization.

[0006] The self-calibration method of the present invention is as follows:

[0007] To reduce the difference between the target value and the current value of a function that includes 2D and 3D difference data as variables, self-calibration is performed by optimizing camera parameters using both 2D and 3D difference data as variables. The 2D difference data consists of the difference components of at least two images contained in the video, and the 3D difference data consists of the difference components of at least two images contained in the video with added depth information.

[0008] The above structure provides a self-calibration method that utilizes both 2D and 3D differential data for optimization.

[0009] The self-calibration system of the present invention comprises:

[0010] Camera device, which captures video; and

[0011] The optimization unit optimizes camera parameters by using both 2D and 3D difference data as variables in a way that minimizes the difference between the target value and the current value of a function containing 2D and 3D difference data as variables. The 2D difference data consists of difference components from at least two images contained in the video, and the 3D difference data consists of difference components from at least two images contained in the video to which depth information has been added.

[0012] The above structure provides a self-calibration system that utilizes both 2D and 3D differential data for optimization.

[0013] Invention Effects

[0014] This invention provides a self-calibration method that utilizes both 2D and 3D differential data for optimization. Attached Figure Description

[0015] Figure 1 This is a block diagram illustrating the structure of the self-calibrating system involved in the implementation method.

[0016] Figure 2 This is a flowchart of the self-calibration method involved in the implementation.

[0017] Figure 3 This is a schematic diagram illustrating the detailed process of the self-calibration method involved in the implementation. Detailed Implementation

[0018] Implementation

[0019] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, this is not intended to limit the invention to the following embodiments. Furthermore, not all structures described in the embodiments are necessary means to solve the problem. For clarity, the following description and drawings have been appropriately omitted and simplified. In the drawings, the same elements are labeled with the same symbols, and repeated descriptions are omitted as necessary.

[0020] (Description of the self-calibration system involved in the implementation)

[0021] Figure 1 This is a block diagram illustrating the structure of the self-calibrating system according to the implementation method. (Reference) Figure 1 The self-calibration system involved in the implementation method will be described. Self-calibration refers to automatically correcting the focal length, angle of view, magnification, etc., when the camera moves while shooting the surrounding area.

[0022] like Figure 1As shown, the self-calibration system 100 includes a camera device 101 and an optimization unit 102.

[0023] The video recording device 101 is a camera for capturing video. The video recording device 101 can be any camera, such as an RGB camera, an RGBD camera, an infrared camera, or a fisheye camera. The video recording device 101 is capable of acquiring two consecutive images.

[0024] The optimization unit 102 acquires 2D difference data and 3D difference data based on two images captured by the camera device, and optimizes the camera parameters based on the 2D difference data and 3D difference data. The 2D difference data consists of the difference components of at least two images contained in the video. The 3D difference data consists of the difference components of at least two images contained in the video with depth component information added.

[0025] The optimization unit 102 optimizes the camera parameters by treating both the 2D and 3D difference data as variables in a way that reduces the difference between the target value and the current value of the function containing the 2D difference data and the 3D difference data.

[0026] If camera parameters can be obtained, self-calibration of a single camera, such as focal length, angle of view, and magnification, can be performed.

[0027] 2D and 3D difference data can be optimized separately or in combination. For example, camera parameters can be adjusted to minimize the difference in 2D difference data, and then the camera parameters can be adjusted to minimize the difference in 3D difference data. Furthermore, camera parameters can be adjusted to minimize the result of addition, subtraction, multiplication, and division of the differences in 2D and 3D difference data.

[0028] The optimization unit 102 is executed by an information processing device. The information processing device includes a processor for executing programs and a memory for storing programs. The information processing device can be a single device or multiple devices. The information processing device can be a cloud server that provides some or all of the distributed processing capabilities.

[0029] Part or all of the processing in an information processing device can be implemented as a computer program. Such a program can be stored and supplied to a computer using various types of non-transitory computer-readable media. Non-transitory computer-readable media include various types of tangible recording media. Examples of non-transitory computer-readable media include magnetic recording media (e.g., floppy disks, magnetic tapes, hard disk drives), magneto-optical recording media (e.g., magneto-optical disks), read-only memory (CD-ROM), CD-R, CD-R / W, and semiconductor memories (e.g., mask ROM, programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), flash memory ROM, random access memory (RAM)). Furthermore, programs can also be supplied to a computer via various types of temporary computer-readable media. Examples of temporary computer-readable media include electrical signals, optical signals, and electromagnetic waves. Temporary computer-readable media can supply programs to a computer via wired or wireless communication paths such as wires and optical fibers.

[0030] The above structure provides a self-calibration system that utilizes both 2D and 3D differential data for optimization.

[0031] (Description of the self-calibration method involved in the implementation)

[0032] Figure 2 This is a flowchart of the self-calibration method involved in the implementation. Figure 3 This is a block diagram illustrating the detailed flow of the self-calibration method according to the implementation method. (See reference) Figure 2 and Figure 3 The self-calibration method involved in the implementation method will be described.

[0033] like Figure 2 and Figure 3 As shown, a trained Artificial Neural Network (ANN) 302 is prepared (step S201). The trained ANN 302 estimates depth based on at least one image. Furthermore, the trained ANN 302 estimates optical flow based on at least two images. The trained ANN 302 consists of two inference models. The trained ANN 302 includes an optical flow separator 303 at the end.

[0034] The optical flow separator 303 separates the optical flow into camera pose and camera internal parameters.

[0035] Next, video is recorded while the camera device 101 is moved (step S202). At this time, as... Figure 3 As shown, take at least 2 images.

[0036] Finally, the self-calibration algorithm is executed (step S203). Thus, self-calibration is performed.

[0037] Figure 3 The details of the self-calibration algorithm are shown below. For example... Figure 3 As shown, two images captured by camera device 101 are input into pseudo target flow estimator 304 and ANNs 302.

[0038] S-1 (Obtaining the machine learning model)

[0039] The pseudo-target flow estimator 304 and ANNs 302, for example, use existing OSS (Operating System) based on the proposal of Hagemann et al. The pseudo-target flow estimator 304 is a model that estimates the update amount and its weights for obtaining more accurate optical flow based on the image feature vector and optical flow.

[0040] S-2 (the initial value estimate of the value of the optimization variable)

[0041] ANNs302 estimates at least one (inverse) depth image z in a pair of images. 0 and optical flow f ij That is, the machine learning-processed information processing device takes at least one image as input and outputs depth, and takes at least two images, including one image, as input and outputs optical flow. The optical flow separator 303 separates the optical flow f... ij Separate the camera pose G into the optimization object. 0 and camera internal parameters θ 0 .

[0042] That is, the optical flow separator separates the optical flow into camera pose and camera intrinsic parameters. The optical flow separator 303 may, for example, be suitable for Procrustes Analysis (Equation (Eqn.) 1), which is not required in the linear intrinsic parameter representation.

[0043] [Formula 1]

[0044]

[0045] [Formula 2]

[0046]

[0047] Here, G ij To obtain from image I i To I j The relative camera movement, x i To make Ii The observation point group is projected onto the world coordinate system to obtain (Equation 2). W 1 / 2 These are weights relative to the optimized coordinates that can be used arbitrarily. STN b (x) j ,u ij () is based on optical flow estimation u ij Bilinear interpolation is applied to the coordinates (x, y) of the four nearest neighbors (top-left, tl, bottom-right, br, etc.) on the image grid. j tl ,x j tr The calculation of coordinate values ​​is obtained by (…) [see reference 3 below].

[0048] In addition, π -1 (·) is the reprojection function for the point group, which projects the grid points u on the image coordinates as shown in the following formula. i 1 / z estimated by depth i Projected onto three-dimensional space.

[0049] [Formula 3]

[0050]

[0051] u in equation (1) ij The estimation results of ANNs can be directly used to set u. i +f ij Furthermore, it can also be based on the optimized variable set {G} separated in one step. 0 ,z 0 ,θ 0 Reconstruct (Equation 2).

[0052] Unlike the method of Hagemann et al. in non-patent literature 1, ANNs assign an initial estimate {G} 0 ,z 0 ,θ 0 This reduces the sensitivity to initial values ​​for problems based on gradient descent. This is how initial values ​​for optimization calculations are obtained.

[0053] S-3 (Temporary Optical Flow Estimation)

[0054] Based on the variable group {G} that becomes the optimization target 0 ,z 0 ,θ 0 The optical flow is temporarily estimated by the pseudo-target flow estimator 304 using image information and image data to obtain the target value. The required update amount Δu ij and update weight w ij(Equation 3). For example, the same mechanism as DROID-SLAM can be considered, i.e., using UpdateModule [see reference 4 below].

[0055] [Formula 4]

[0056]

[0057] S-4 (Collaborative Optimization)

[0058] The optimization unit 102 performs nonlinear optimization that makes it difficult to achieve a gradient of zero during the linear motion of the camera. For example, in addition to the objective function L used by Hagemann et al. Opt.1 In addition to Equation 4, the objective function L proposed by Smith et al. can also be considered. Opt.2 The use of (Equation 5) together. The former is the residual in the two-dimensional representation of optical flow, while the latter is the objective function based on the residual in the three-dimensional representation of the scene flow that is proposed and used together.

[0059] [Formula 5]

[0060]

[0061] [Formula 6]

[0062]

[0063] [Formula 7]

[0064]

[0065] Here, π(·) is a function projected onto the image plane, (i,j)∈ε and (i,j)∈ε′ are pairs of keyframes, and ||·|| Σ The Mahalanobis distance is given by Σ. ij =diagw ij constitute.

[0066] [Formula 8]

[0067]

[0068] [Formula 9]

[0069]

[0070] Unlike the S-2, u ij Reliably use u by directly utilizing the estimation results of ANNs i +f ij Therefore, during the generation of the first-order differential, a differential term that might lead to gradient vanishing, i.e., du, will not be generated. ij / dθ. And, using W 1 / 2 or w ij Assist L Opt.2 The optimization methods are also included in this proposal. In particular, it is believed that w ij Applicable to W 1 / 2 This approach plays an important role in more reliably performing the optimization of equation (5), which is assumed to be susceptible to the adverse effects of estimation noise.

[0071] Unlike the method of Hagemann et al. in non-patent literature 1, the learning incompleteness during monotonous camera movement can be expected to be mitigated through the optimization described in equation (5).

[0072] S-5 ({G) K ,z K ,θ K} (308)

[0073] The variable set {G} is obtained by performing the optimization described in equations (4) and (5) a specified number of times. 1 ,z 1 ,θ 1} (306). When the error is sufficiently small ("yes" in 306), the estimated value {G} is obtained. K ,z K ,θ K} (308). If the error does not decrease sufficiently (No in 306), the count is incremented by 1 (307), and {G 1 ,z 1 ,θ 1 The input is fed into the pseudo-target flow estimator 304. That is, the process from S-2 to S-4 is repeated between multiple keyframes until the objective function error is minimized, thereby obtaining the final estimated value {G}. K ,z K ,θ K According to the implementation by Teed et al., this is repeated at each time point in the specified pair of keyframes [see reference 4 below].

[0074] Other implementation methods

[0075] In S-1, to improve the performance of the pseudo-target flow estimator 304, for example, learning data generated by simulating various camera parameters such as fisheye, and the application during learning, can be considered. That is, data captured by one camera can be applied to data captured by another camera.

[0076] In S-1, given the availability of video data in the hypothetical application environment, ANNs can be obtained through self-supervised learning. For example, the method of Fang et al. [see reference 5 below] can be considered.

[0077] In S-2, when information such as depth or self-position is available and does not require estimation, these sensor values ​​can be used instead of ANNs. For example, the use of LiDAR, GPS, and RADAR can be considered.

[0078] In S-4, the optimization function {L} Opt.1 ,L Opt.2 The usage and performance variations of} are diverse. For example, 2D difference data optimization and 3D difference data optimization can be performed separately. Furthermore, 2D difference data optimization and 3D difference data optimization can be performed through combinations of arithmetic operations.

[0079] This invention enables automatic camera recalibration in a short time. In the event of a malfunction in a (autonomous) driving vehicle with at least one camera, after confirming that it is safe to proceed, the minimum camera parameters can be restored simply by moving the vehicle forward or backward.

[0080] This invention enables robots to scale learning data. Even when the learning data used to make robots intelligent requires 3D reconstruction as a preprocessing step from video, it can be collected and integrated without considering individual differences between cameras.

[0081] (Proof of gradient vanishing)

[0082] If we assume that the objective function used by Hagemann et al. is L... Opt.1 For (1) a pinhole camera and (2) the camera movement is monotonic, then the residual r represents... ij The Jacobian matrix relative to θ becomes 0. At this point, the Jacobian matrix is ​​given by the following equation (6).

[0083] [Formula 10]

[0084]

[0085] [Formula 11]

[0086]

[0087] Here, the flow x of the three-dimensional point group ij The slope of θ is as follows.

[0088] [Formula 12]

[0089]

[0090] In addition to equation (6), if we assume the camera model is θ = [f x f y c x c y ] T(Equation 8) then each item obtains the following structure.

[0091] [Formula 13]

[0092]

[0093] [Formula 14]

[0094]

[0095] [Formula 15]

[0096]

[0097] When the camera movement becomes monotonic, that is, when it consists only of "translational motion in the z-axis direction without rotation", the following approximation holds.

[0098] [Formula 16]

[0099]

[0100] From equations (8) and (9), we can see that the following approximations hold true.

[0101] [Formula 17]

[0102]

[0103] As can be seen from the above, if equation (10) is substituted into equation (6), dr can be obtained immediately. ij / dθ=0. This makes it difficult to find the value of θ in nonlinear optimization based on gradient descent. Furthermore, equation (6) is...

[0104] [Formula 18]

[0105]

[0106] This holds true under the same conditions. Therefore, the same gradient vanishing problem may also occur in the method of Fang et al., which achieves learning through differentiable optical flow representation [see reference 5 below].

[0107] On the other hand, in Smith et al.'s method (Equation 5), the optical flow is estimated as a constant (directly using the output of the learned model), therefore it does not have du... ij The component of / dθ. At this point, the deviation e relative to θ. ij The gradient uses the coefficient γ of the constant assigned by the optical flow. k As given in equation (11).

[0108] [Formula 19]

[0109]

[0110] [Formula 20]

[0111]

[0112] [Formula 21]

[0113]

[0114] [Formula 22]

[0115]

[0116] For simplicity, we only need to focus on f x If the approximation of equation (9) is applied to the components, the following approximation is obtained.

[0117] [Formula 23]

[0118]

[0119] [Formula 24]

[0120]

[0121] Thus, even if the camera movement is monotonous, it is possible to derive L. Opt.2 The gradient can thus facilitate optimization based on gradient descent.

[0122] In addition, for detailed formula variations, see supplementary information from the paper by Hagemann et al. [reference 6 below].

[0123] Furthermore, the present invention is not limited to the above-described embodiments, and appropriate modifications can be made without departing from the spirit of the invention.

[0124] References

[0125] [Document 1] T. Kanai et al. “Self-supervised geometry-guidedinitialization for robust monocular visual odometry,” arXiu, 2024.

[0126] [Document 2] T. Zou et al. “Unsupervised learning of depth and ego-motion from video,” in CVPR, 2017.

[0127] [Literature 3] M. Jaderberg et al. “Spatial transformer networks,” inAdvances in Neural Information Processing Systems, C. Cortes, N. Lawrence, D.Lee, M. Sugiyama, and R. Garnett, Eds., vol. 28. Curran Associates, Inc., 2015.

[0128] [Literature 4] Z. Teed et al. "DROID-SLAM: Deep Visual SLAM for monocular, Stereo, and RGB-D Cameras," Advances in neural information processing systems, 2021.

[0129] [Document 5] J. Fang et al. "Self-supervised camera self-calibration from video," in ICRA, 2022, pp.8468-8475.

[0130] [Document 6] P. Hagemann et al. "Deep geometru-aware camera pose learning (supplemental)," 2023, ccessed:2023-10-05.[Online]. Available: https: / / openaccess.thecvf.com / comtent / ICCV2023 / supplemental / Hagemann_Deep_Geometry_Aware_Camera_ICCV_2023_supplemental.pdf

[0131] Symbol Explanation

[0132] 100-System, 101-Camera device, 102-Optimization unit, 302-ANNs, 303-Optical flow separator, 304-Pseudo-target flow estimator.

Claims

1. A method for self-calibration, characterized in that, To reduce the difference between the target value and the current value of a function that includes 2D and 3D difference data as variables, self-calibration is performed by optimizing camera parameters using both 2D and 3D difference data as variables. The 2D difference data consists of the difference components of at least two images contained in the video, and the 3D difference data consists of the difference components of at least two images contained in the video with added depth information.

2. The self-calibration method according to claim 1, characterized in that, The machine learning-based information processing device takes at least one image as input and outputs depth, and takes at least two images including the one image as input and outputs optical flow. The optical flow separator obtains the initial values ​​for the optimized calculation from the optical flow separation of the camera pose and camera internal parameters.

3. The self-calibration method according to claim 1, characterized in that, The optimization of the 2D difference data and the optimization of the 3D difference data are performed separately.

4. The self-calibration method according to claim 1, characterized in that, The optimization of the 2D difference data and the 3D difference data are performed by combining four arithmetic operations.

5. A self-calibration system, characterized in that, have: Camera device, which captures video; and The optimization unit optimizes camera parameters by using both 2D and 3D difference data as variables in a way that minimizes the difference between the target value and the current value of a function containing 2D and 3D difference data as variables. The 2D difference data consists of difference components from at least two images contained in the video, and the 3D difference data consists of difference components from at least two images contained in the video to which depth information has been added.