Calibration method of monocular camera

Through the monocular camera calibration method, combined with the Zhang Zhengyou calibration method and the fixed-speed motor, the problem of poor robustness and large calculation volume of stereo vision in autonomous driving and intelligent navigation of ships is solved, and low-cost, high-precision depth perception and simplified operation are achieved.

CN120431184APending Publication Date: 2025-08-05HARBIN ENG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510428019.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

Traditional stereo vision methods have problems such as poor robustness and high computational volume in fields such as autonomous driving and intelligent ship navigation.

Method used

The single-eye camera calibration method is used, and the Zhang Zhengyou calibration method and the speed motor are used to take pictures of calibration boards at different angles and distances, combined with supporting programs in C# and Python languages, calibration calculations are automatically completed, camera internal references are recorded, and camera movement is performed through experimental devices and guides to simplify depth calculations.

Benefits of technology

It realizes low-cost and high-precision depth perception, reduces calculation amount, improves robustness, simplifies the operation process, and facilitates understanding of the calculation principles of stereo vision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120431184A_ABST
    Figure CN120431184A_ABST
Patent Text Reader

Abstract

The invention provides a monocular camera calibration method, and belongs to the technical field of visual estimation. The problems that a traditional stereoscopic vision method can achieve high-precision depth calculation, but is poor in robustness, large in calculation amount and the like are solved. Calibration is completed by using a Zhang Zhengyou calibration method through a program. In the calibration process, all angular points can be accurately extracted, the problems of missing detection, position relation errors and the like do not exist, the movement displacement of the camera can be calculated by means of a constant-speed motor, so that a translation matrix is accurately expressed, the pixel coordinates of a target object are manually marked, coordinate information can be substituted into an equation, depth information of the object can be accurately solved, and the accuracy of calibration is improved. The matched program can visually display all functions, operation is simple, and calculation is accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of visual estimation technology, and in particular to a monocular camera calibration method. Background Art

[0002] As accuracy improves, the application of visual depth estimation in areas such as autonomous driving and intelligent ship navigation will continue to expand, enabling low-cost, high-precision distance perception. Current environmental perception technologies primarily use lidar for depth perception, which is expensive and very accurate. While traditional stereo vision methods can also achieve high-precision depth calculations, they also suffer from poor robustness and high computational complexity. Summary of the Invention

[0003] In view of this, the present invention aims to propose a monocular camera calibration method to solve the problems that traditional stereo vision methods can also achieve high-precision depth calculation, but also have poor robustness and large computational complexity.

[0004] To achieve the above objectives, the present invention adopts the following technical solutions. According to one aspect of the present invention, a method for calibrating a monocular camera is provided, comprising:

[0005] S1. Monocular camera calibration: Adjust the calibration plate and take at least a predetermined number of pictures containing the complete calibration plate in different states relative to the camera. Import the pictures into the designed supporting program, which automatically completes the calibration calculation and records the camera's internal parameters.

[0006] S2. Installation of the experimental device: installing the monocular camera on a horizontal guide rail, the camera base and the constant speed motor;

[0007] S3, experimental picture shooting, connecting the monocular camera to the computer, starting the supporting program, connecting the fixed speed motor, and shooting the object video;

[0008] S4, experimental data processing, read the picture, record the coordinates of 5 groups of pixel coordinate systems, and calculate K -1 x converts pixel coordinates to camera coordinates, normalizes them, and solves the depth through the equation.

[0009] Furthermore, the different states in S1 are different angles and different distances.

[0010] Furthermore, the not less than predetermined number in S1 is not less than 50 sheets.

[0011] Furthermore, the adjustment of the calibration plate in S1 is performed by hand-holding the calibration plate.

[0012] Furthermore, the program in S1 is matlabr.

[0013] Furthermore, the camera base and the constant speed motor in S2 are connected by a thin rope.

[0014] Furthermore, the supporting program described in S3 is designed based on C# and Python languages.

[0015] Furthermore, the recording of the multiple groups of pixels in S4 is recording 5 groups of pixels.

[0016] Furthermore, the Matrix K is the intrinsic parameter matrix of the camera, fx and fy are the focal lengths of the camera along the x-axis and y-axis of the image, and Cx and Cy are the coordinates of the principal point of the image on the x-axis and y-axis.

[0017] Furthermore, the equation in S4 is s2x2=s1Rx1+t, which describes the process of transforming a point X1 in one coordinate system to a point X2 in another coordinate system through the rotation matrix R and the translation vector t, as well as the scale factors S1 and S2.

[0018] Beneficial effects:

[0019] 1. A supporting program for the experiment was designed based on C# and Python languages. The program includes calibration and depth calculation functions, which can easily complete these steps and avoid the impact of large calculations on the experiment.

[0020] 2. Using a fixed-speed motor, the guide rails, and the slider prevent rotation during camera movement, making translation easy to calculate. This simplified design allows for a simple simulation of autonomous driving scenarios where the speed is known. Once the speed is known, depth can be calculated from the captured image sequence.

[0021] 3. The traditional stereo vision alignment and matching algorithms are not used. Instead, the target position is manually marked, which reduces the amount of program calculation and facilitates intuitive understanding of the calculation principles of stereo vision. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The accompanying drawings, which constitute part of the present invention, are provided to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are provided to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings:

[0023] Figure 1 This is a flow chart of a monocular camera calibration method of the present invention;

[0024] Figure 2 A schematic diagram of a homogeneous coordinate system for a monocular camera calibration method according to the present invention;

[0025] Figure 3Schematic diagram of epipolar constraints for a monocular camera calibration method according to the present invention;

[0026] Figure 4 A schematic diagram of a triangulation solution for a monocular camera calibration method according to the present invention. DETAILED DESCRIPTION

[0027] The following will be combined with the accompanying drawings in the embodiments of the present invention to clearly and completely explain the technical solutions in the embodiments of the present invention. It should be noted that the embodiments of the present invention and the features therein can be combined with each other in the absence of conflict, and the embodiments described are only part of the embodiments of the present invention, not all of the embodiments.

[0028] It should be noted that the descriptions of the present invention regarding directions such as "left", "right", "left side", "right side", "upper", "lower", "top", and "bottom" are all defined based on the relationship between the orientations or positions shown in the accompanying drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the structure must be constructed and operated in a specific orientation. Therefore, they cannot be understood as limitations on the present invention. In the description of the present invention, the meaning of "multiple" is more than two, unless otherwise clearly and specifically defined.

[0029] In the description of the present invention, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood broadly. For example, they may refer to fixed connections, detachable connections, or integral connections; they may refer to direct connections, indirect connections through an intermediary, or internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.

[0030] Referring to the accompanying drawings, this embodiment is described. According to one aspect of the present invention, a method for calibrating a monocular camera is provided, comprising:

[0031] S1. Monocular camera calibration: Hold the camera in place and hold the calibration plate in hand, taking at least 50 images of the complete calibration plate at different angles and distances relative to the camera. Import the images into the designed supporting program, set the size and number of checkerboard grids, and let the program automatically perform the calibration calculations and record the camera's internal parameters.

[0032] In the camera imaging process, the conversion from the world coordinate system to the camera coordinate system and from the camera coordinate system to the pixel coordinate system is involved. The parameters such as focal length and translation amount in the conversion process are fixed. They are written in matrix form and the intrinsic parameters of the camera. The conversion form from the world coordinate system to the pixel coordinate system is as follows

[0033]

[0034] The first two terms on the right side of the equation are generally denoted as K, which is called the camera intrinsic parameter;

[0035]

[0036] The process of solving the intrinsic parameters of a monocular camera is called monocular camera calibration. The calibration method used is Zhang Zhengyou calibration method. The obtained intrinsic parameters will be used for conversion between coordinate systems in subsequent steps.

[0037] In the common Cartesian coordinate system, two parallel lines cannot intersect. However, in perspective space, two lines at infinity intersect at one point. Homogeneous coordinates are more convenient when dealing with camera projection and imaging issues. This experiment will utilize the important property of homogeneous coordinates—equal up to scale. In a homogeneous coordinate system, we can represent a spatial point P as follows:

[0038]

[0039] Epipolar constraint, according to the pinhole camera model, the projection point p of the space point P in the image coordinate system is expressed as:

[0040] s1p1=KP,s2p2=K(RP+t)

[0041] Since they are equal in scale, the same projection point can be represented by a homogeneous coordinate multiplied by a non-zero constant, and the following homotopy relation exists:

[0042]

[0043] So the above projection relationship can be written as:

[0044]

[0045] Get coordinates;

[0046] x1=K -1 p1,x2=K -1 p2

[0047] Observe that the result on the left side of the equation is 0, then the equation can be written as:

[0048]

[0049] This equation is the epipolar constraint. Conventionally, the value of the matching points in space can be solved. When solving for this value, it is fully expanded into a 1×9 vector. Eight pairs of matching points can be used to construct an 8×9 vector, thus solving this rank-8 linear system of equations. This is the classic eight-point method.

[0050] To determine the relationship between the coordinates in the 2D image captured by the camera and the coordinates in the 3D world, it is necessary to first establish a camera imaging model. Camera imaging models are generally divided into linear pinhole imaging models and nonlinear distortion models.

[0051] The pinhole imaging model is an idealized imaging model. It assumes that the optical center of the lens is a very small aperture, through which all light enters, and that any point on the object forms an image point on the image plane. While this ideal model offers high accuracy, it ignores factors such as the magnification effect of the camera lens and light scattering.

[0052] In the process of lens imaging, the object distance, focal length and image distance satisfy the relationship

[0053]

[0054] In theory, the imaging of the lens is the ideal pinhole model mentioned above. However, in actual experiments, due to problems such as the manufacturing process and installation accuracy of the lens, the image collected by the lens will be distorted, which will cause errors in the measurement of the target.

[0055] There are two main types of image distortion: radial distortion and tangential distortion. Radial distortion is minimal at the exact center of the image and increases with radius. Radial distortion can be categorized as pincushion distortion and barrel distortion. Tangential distortion occurs when the lens is not parallel to the imaging plane, similar to perspective transformation.

[0056] The camera needs to be calibrated to eliminate the resulting errors and obtain more accurate camera parameters. The nonlinear model of the camera is:

[0057]

[0058] Among them, represents the real coordinates of the point, (x, y) is the position of the ideal point obtained according to the pinhole imaging model, and is the nonlinear distortion value, that is, the radial distortion value caused by lens processing and the tangential distortion value caused by installation deviation. The mathematical expression of radial distortion is:

[0059]

[0060] Where p1 and p2 are tangential distortion coefficients. Substituting into the formula, we can get the relationship between the actual imaging point and the ideal imaging point of a point in space on the image:

[0061]

[0062] Therefore, for nonlinear imaging cameras, it is composed of linear parameters f / dx, f / dy, u0, v0 and distortion parameters k1, k2, k3, p1, p2.

[0063] In this experiment, the camera is fixed on a guide rail and moved at a certain speed. The speed and movement time are measured, and R and t can be simply expressed, paving the way for subsequent steps.

[0064] Triangulation, observing the same object from different angles, the corresponding normalized plane coordinates are normalized and still meet

[0065] s2x2=s1Rx1+t

[0066] Multiply both sides of the formula by At this time, the left side is 0, so:

[0067]

[0068] Only s1 is unknown, from which the depth of the object can be calculated.

[0069] The calibration can be completed through the program using Zhang Zhengyou calibration method. During the calibration process, all corner points can be accurately extracted without problems such as missed detection and incorrect positional relationships.

[0070] Zhang Zhengyou calibration method:

[0071] The image is divided into four quadrants, with the calibration plates evenly distributed in each quadrant. At least two pictures with different tilt angles are taken in each quadrant. The calibration plate pictures need to cover the entire measurement field of view, and the number of calibration pictures is usually between 15 and 25. The imaging area of the calibration plate should roughly occupy 1 / 3 to 1 / 4 of the entire picture. If the calibration plate image is too dark, an auxiliary light source needs to be used to fill in the light. If it is too bright, the exposure time needs to be adjusted to ensure that the brightness of the calibration plate is sufficient and uniform. During the calibration process, the aperture and focal length of the camera cannot be changed. If they are changed, recalibration is required.

[0072] System parameter calibration is crucial for the entire 3D reconstruction system. Only accurate calibration can provide precise data for subsequent reconstruction calculations. Camera calibration involves the transformation between the 3D geometric position of a point on the surface of a spatial object and its corresponding point in the image. Determining this transformation requires establishing a geometric model of the camera's imaging system. These geometric model parameters are the camera parameters. This is essentially the process of calculating the parameters of the camera's geometric model based on the geometric principles of camera imaging. Once the camera parameters are determined, light plane calibration is performed. The light plane equation, combined with the camera's internal and external parameters, determines the relationship between a point in space and the image point.

[0073] Currently, cameras are categorized by their calibration methods. The most common methods are based on whether a calibration object is required: traditional calibration, self-calibration, and active vision-based calibration. Traditional calibration methods use a high-precision calibration plate as a reference to construct a calibration model. Using the known reference object information, the camera parameters are derived by determining the correspondence between spatial midpoints and image pixels. These parameters are then optimized. Traditional calibration methods include the Zhang Zhengyou calibration method, the Tsai two-step method, and 3D target calibration, all of which offer high accuracy.

[0074] The Zhang Zhengyou calibration method, proposed by Professor Zhang Zhengyou in the late 20th century, overcomes the strict calibration object requirements of traditional calibration methods and optimizes self-calibration. Combining traditional and self-calibration methods, it eliminates the need for a high-precision calibration object and instead uses a black and white checkerboard calibration plate. The plate's angular position is rotated and translated, and an image of the plate is captured. The camera's parameters are then determined by comparing the plate's physical information with the pixel information in the captured image, combining the advantages of both. This method boasts high calibration accuracy, simple operation, and robustness, making it widely used. This paper employs the Zhang Zhengyou calibration method to determine the projection matrix and then applies this to linear transformations to calculate the camera's parameters.

[0075] Assuming that the two-dimensional calibration plate is on the plane of ZW=0 in the world coordinate system, the world coordinates of the feature point M on the calibration plate are (Xw, Yw, 0, 1), and its corresponding point m in the pixel coordinates is

[0076] (u,v,1), we can get the following relationship;

[0077]

[0078] Where λ is a constant, M1 is the camera intrinsic parameter, ri is the i-column vector of the rotation matrix, and t is the translation vector. Let H = [h1,h2,h3] = λM1[r1,r2,t], also known as the homography matrix. Then we have:

[0079]

[0080] According to the properties of the rotation matrix r1 T r2=0,||r1||=||r2||=1, and the constraint conditions on the internal parameters can be obtained.

[0081]

[0082] A homography matrix can provide two equations, and the camera parameter matrix contains 5 unknown parameters. If you want to solve it, you need at least 3 homography matrices. In order to obtain 3 homography matrices, you need to shoot a chessboard with a number greater than or equal to 3 for calibration, and then you can uniquely solve M1.

[0083]

[0084] where a x =f / d x , a y =f / d y , γ is the parameter of the CCD photosensitive grid, representing the d in the pixel coordinate system x ,d y The degree of skewness. Then

[0085]

[0086] When the solution of b is found, the internal parameter matrix M1 can be obtained

[0087]

[0088] The λ can be calculated from the characteristics of the rotation matrix. In this way, the intrinsic and extrinsic parameters of the camera are obtained.

[0089] Copy the camera calibration photos into the calibration program folder. The steps to start camera calibration are as follows:

[0090] 1. Scroll down in MATLAB and select Camera Calibrator; click Add Images to import the photos.

[0091] 2. For a standard camera, select Third-Order Radial Distortion and Skew in the Options menu. Click Calibrate to perform camera calibration. Calculate the average reconstruction error. As long as the average error is less than 0.5, the camera calibration result is considered reliable.

[0092] 3. Export the camera parameters by clicking Export Camera Parameters. Click OK to save the Camera Parameters.

[0093] 4. You can see the camera parameters appear in the MATLAB workspace. Click on this parameter to get the various parameters of the camera.

[0094] S2. Install the experimental device. Install the monocular camera on a horizontal guide rail. Use a thin rope to connect the camera base and the fixed-speed motor. Make the thin rope horizontal to ensure that the camera can move horizontally at a uniform speed; ensure that the camera can fully capture the target object in both left and right positions.

[0095] S3. Experimental picture shooting: Connect the monocular camera to the computer, start the supporting program, turn on the fixed-speed motor, and shoot the object video. During the movement, click the capture button to record the images captured on both sides. At the same time, the time interval between the two shots is automatically calculated after two shots, that is, the camera's movement time. Combined with the known motor speed, the lateral displacement of the camera can be calculated. The supporting program is designed based on C# and Python languages. The C# code of the supporting visualization interface is as follows:

[0096] private voidbutton5_Click(object sender,EventArgs 3)

[0097] Process p = new Process();

[0098] p.StartInfo.FileName="cmd.exe";

[0099] p.StartInfo.CreteNoWindow=true;

[0100]

[0101] The Python code for the backend depth calculation is as follows:

[0102] import numpy as np

[0103] defdepth(args):

[0104] k = args.k

[0105] x1=args.coor1

[0106] x2 = args.coor2

[0107] defstr2arr(arr):

[0108] arr=arr.replace("[","").replace("]","").replace(",","").split()print(arr)

[0109] iflen(arr)>2:

[0110] arr=np.array([float(i)for iin arr])

[0111] arr = np.reshape((3,3))

[0112] else:

[0113] arr=np.array([float(i)for iin arr])

[0114] arr = np.append(arr, 1)

[0115] return arr

[0116] k=str2aerr(k)

[0117] print(k)

[0118] S4, experimental data processing, read the picture, record the coordinates of 5 groups of pixel coordinate systems, and calculate K -1 x converts pixel coordinates to camera coordinates, normalizes them, and solves the depth through the equation.

[0119] In the above description, the sensors, controllers, control programs, etc. that may be involved are all existing technologies and will not be described in detail.

[0120] The embodiments of the present invention disclosed above are intended only to illustrate the present invention. The embodiments do not describe all details in detail, nor do they limit the present invention to the specific embodiments described. Numerous modifications and variations are possible based on the contents of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the present invention, thereby enabling those skilled in the art to better understand and utilize the present invention.

Claims

1. A method for calibrating a monocular camera, characterized in that: The following steps are involved: S1. Monocular camera calibration: Adjust the calibration plate and take at least a predetermined number of pictures containing the complete calibration plate in different states relative to the camera. Import the pictures into the designed supporting program, which automatically completes the calibration calculation and records the camera's internal parameters. S2. Install the experimental device: install the monocular camera on the horizontal guide rail, and connect the camera base to the fixed speed motor; S3, experimental picture shooting, connecting the monocular camera to the computer, starting the supporting program, connecting the fixed speed motor, and shooting the object video; S4, experimental data processing, read the picture, record the coordinates of multiple groups of pixel coordinate systems, and calculate K -1 x converts pixel coordinates to camera coordinates, normalizes them, and solves the depth through the equation.

2. The method for calibrating a monocular camera according to claim 1, wherein: The different states in S1 are different angles and different distances.

3. The method for calibrating a monocular camera according to claim 1, wherein: The predetermined number of sheets or more in S1 is not less than 50 sheets.

4. The method for calibrating a monocular camera according to claim 1, wherein: The adjustment of the calibration plate in S1 is performed by hand-holding the calibration plate.

5. The method for calibrating a monocular camera according to claim 1, wherein: The program in S1 is matlabr.

6. The monocular camera calibration method according to claim 1, wherein: The camera base and the constant speed motor in S2 are connected by a thin rope.

7. The monocular camera calibration method according to claim 1, wherein: The supporting program described in S3 is designed based on C# and Python languages.

8. The method for calibrating a monocular camera according to claim 1, wherein: The recording of the plurality of groups of pixels in S4 is recording 5 groups of pixels.

9. The method for calibrating a monocular camera according to claim 1, wherein: The S4 Matrix K is the intrinsic parameter matrix of the camera, f x 、f y is the camera focal length along the x-axis and y-axis of the image, C x 、C y are the coordinates of the principal point of the image on the x and y axes.

10. The monocular camera calibration method according to claim 1, wherein: The equation in S4 is s2x2=s1Rx1+t, which describes the process of transforming a point X1 in one coordinate system to a point X2 in another coordinate system through the rotation matrix R and the translation vector t, as well as the scale factors S1 and S2.