Method for determining at least one camera parameter for calibrating a camera, and camera
Patent Information
- Application Number
- EP2023730828
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-05
- Filing Date
- 2023-06-12
- Publication Date
- 2025-05-14
AI Technical Summary
Existing camera calibration methods for intelligent cameras, such as traffic surveillance cameras, are complex and impractical, especially in complicated environments, as they require manual input of object dimensions for accurate geometric imaging behavior, which is not feasible in real-time monitoring scenarios.
An automatic calibration method using machine learning algorithms, specifically deep neural networks, to determine camera parameters by analyzing images of objects in the environment, such as vehicles, by projecting 3D bounding boxes onto the camera plane and optimizing extrinsic and intrinsic camera parameters based on object distances and orientations.
Enables fully automatic and precise calibration of camera parameters, reducing the need for manual input and improving the accuracy of geometric imaging behavior in various environments, including those with multiple planes or curved surfaces, with fewer length references required.
Smart Images

Figure 1.1
Abstract
Description
[0001] Description
[0002] title
[0003] Method for determining at least one camera parameter for calibrating a camera and camera
[0004] The present invention relates to a method for determining at least one camera parameter for calibrating a camera, as well as a computing unit, a camera and a computer program for carrying out the method.
[0005] Background of the invention
[0006] Cameras can be used to monitor an environment such as highways, buildings, fences, or offices. So-called intelligent cameras or camera systems are increasingly being used for this purpose. For these to function properly, camera calibration is usually necessary. For example, EP 2 044 573 B1, EP 2 798 611 B1, and US 2019 / 0311494 A1 describe ways to calibrate cameras.
[0007] Disclosure of the invention
[0008] According to the invention, a method for determining at least one camera parameter for calibrating a camera, as well as a computing unit, a camera, and a computer program for carrying out the method are proposed, having the features of the independent patent claims. Advantageous embodiments are the subject of the subclaims and the following description. The invention relates to cameras, in particular traffic surveillance cameras, and in particular to their calibration or the determination of one or more camera parameters for calibrating such a camera, i.e., camera parameters based on which the geometric imaging behavior of the camera is described. Calibration can comprise both extrinsic and intrinsic calibration. Extrinsic camera parameters are, for example, roll angle, pitch angle, and camera height (above the ground); intrinsic camera parameters are, for example, a focal length, or principal point and distortion.
[0009] Such camera parameters can be obtained, for example, based on manually determined and provided information. For example, dimensions of objects in images captured by the camera can be manually specified and provided, allowing calibration to be performed. However, this is very time-consuming and often impractical, especially in complex environments to be monitored.
[0010] Within the scope of the present invention, a procedure is proposed that, in particular, allows for automatic calibration of the camera. For this purpose, one or more images of an environment, e.g., a highway, captured by the camera (to be calibrated) are provided or received, e.g., in an executing processing unit. In the preferred case of multiple images, these can be, for example, individual frames of a video, but also separately captured images.
[0011] For a 3D bounding box of one or each of several objects in the one or more images, coordinates of at least two corner points are then determined in a camera coordinate system. The objects considered here are primarily vehicles, but people or trailers, for example, can also be considered.
[0012] A 3D bounding box is a cuboid into which the object is fitted, i.e., an enveloping cuboid. This 3D bounding box is then projected onto the image or camera plane, or the plane of the camera sensor, i.e., into the image. A projected 3D bounding box (as used here in particular) is thus defined by eight vertices, each of which has coordinates in the image plane (image coordinates). By projecting the 3D bounding box onto the camera plane, each vertex has two degrees of freedom, e.g., an x- and a y-coordinate.
[0013] However, the proposed method does not require all eight corner points or their coordinates; two corner points may be sufficient. However, it is expedient to determine the coordinates of at least three, in particular at least four, and in particular up to eight further corner points for the 3D bounding box of one or each of the multiple objects.
[0014] It is particularly preferred if the determination of the coordinates of the at least two corner points is carried out by means of a machine learning algorithm, in particular an artificial neural network (e.g. a deep neural network, DNN), which receives the one or more images as input data.
[0015] The machine learning algorithm then predicts (or estimates), for example, a certain number of vertices (or keypoints) for each of the objects in the image. The input data for the machine learning algorithm can be either individual images (i.e., multiple images individually) or an image series comprising several related or temporally separated images that are processed simultaneously as a batch. The architecture of a DNN can vary and can, for example, be an anchor-based approach or an anchor-free approach, or it can function based on a 'convolutional neural network' or a 'transformer network' (such as a Swin transformer). The prediction result of the DNN can then be, for example, all eight vertices of the enveloping cuboid (projected 3D bounding box) of the object, projected onto the camera plane.In this case, the DNN would have a full 16 degrees of freedom (each of the eight projected vertices has an x- and y-coordinate). Alternatively, a subset of the eight vertices can be predicted (determined), e.g., only four vertices, from which the remaining four vertices result from the linear combination of the predicted four vertices. In the latter case, the system has eight degrees of freedom.
[0016] Such a machine learning algorithm or neural network can be obtained, for example, by training the projected 3D bounding boxes and / or the corner points of interest or all of the projected 3D bounding boxes for the objects as training data for a large number of images with objects. It should be noted that the 3D bounding boxes and projected 3D bounding boxes themselves do not necessarily have to be completely determined; rather, the coordinates of the corner points of interest are sufficient. If, for example, the camera optics produce a distorted image, i.e., lines in the world / camera system are mapped onto curves, the eight corner points and not the curves are of interest.
[0017] Furthermore, the at least one camera parameter is then determined, specifically based on at least one distance in the image plane (i.e., in an image coordinate system, which is typically 2D), defined by the coordinates of the at least two corner points of the respective projected 3D bounding box, and at least one corresponding, predetermined distance in an environmental coordinate system or camera coordinate system. The at least one camera parameter is then provided. The predetermined distance is, for example, a width of a vehicle and / or a length of a vehicle and / or a height of a vehicle and / or a size of a vehicle that can be determined therefrom, for example, a length of a diagonal of a vehicle.
[0018] As mentioned, a 3D bounding box is a cuboid that encloses an object such as a vehicle. Accordingly, certain distances are defined by the vertices, namely the distances between two vertices, in particular the edges of the cuboid. Examples of such distances defined by the vertices are, for example, a length, a width, and a height. For such distances, there are corresponding distances in the environmental coordinate system (i.e., the real environment), for example, the length, a width, and a height of the real object, e.g., the vehicle. A real width or length of a vehicle is known for different vehicle types, for example.
[0019] The at least one camera parameter can then be determined, since the distance between two vertices in the camera coordinate system must correspond to the corresponding real distance; this requires, for example, certain values for the extrinsic camera parameters such as pitch angle, roll angle, and altitude.
[0020] It is particularly expedient if the at least one camera parameter is determined based on distances for multiple 3D bounding boxes, or in other words, based on the at least one distance defined by the coordinates of the at least two corner points of multiple 3D bounding boxes. This can include multiple images, each with one or more objects, with each object being assigned a 3D bounding box. The more distances (or objects and images) used, the more accurately the at least one camera parameter can be determined, for example, within the framework of an optimization process.
[0021] If distances between several (projected) 3D bounding boxes (in the image plane) are determined, in the simplest case one and the same dimensions, i.e. distances between the corner points or keypoints, of the 3D bounding boxes can be assumed, for example that all vehicles in the surrounding coordinate system have the same width. If, for example, the average width is assumed for the width of vehicles, the deviations between the actual width of individual vehicles and the average width are statistically averaged out (e.g. assuming the law of large numbers), so that an exact result is achieved for a sufficiently large number of detected vehicles. Within the framework of, for example, an optimization process, different distances between the 3D bounding boxes then balance each other out.
[0022] Preferably, at least one object parameter of the one or each of the multiple objects in the one or more images is also determined. Possible object parameters include, for example, the position in the plane and orientation (rotation about the vertical axis) of the objects. To determine the at least one camera parameter, the at least one camera parameter and the at least one object parameter are then optimized simultaneously within the framework of an optimization process.
[0023] This ingenious approach or calibration takes particular advantage of the fact that the axes resulting from the corner points of the 3D bounding boxes are at right angles to one another. By assuming at least one length (or, for example, length and width), a hypothesis for the extrinsics of the camera can be created from a single vehicle observation. The extrinsics are described in particular by the three parameters roll angle, pitch angle and height above ground (or floor, i.e. in particular the road) and are therefore different for each plane in the scene or environment. Here, it can initially be assumed that there is only one plane in the scene. In addition to the three extrinsic parameters, three object parameters can also be determined for each object or vehicle, for example: the position in the plane (two parameters) and the rotation or rotation around the vertical axis or plane normal, so that a top view can be created at any time.All object or vehicle observations (or the four ground vertices of each 3D bounding box) can serve as measurements here. In an optimization process, the extrinsic parameters and object parameters are then simultaneously estimated or optimized, minimizing, for example, a backprojection error. Specifically, the four ground vertices of the 3D bounding box are projected into the image, assuming an average vehicle size and the extrinsic camera and object parameters, and compared with the detected vertices.
[0024] If, for example, average values are always assumed for the object or vehicle size, it is expedient to use only cars (passenger vehicles). This can be achieved, for example, by the machine learning algorithm mentioned above only considering cars as objects. It should be noted that the estimated parameters are correct if the observed vehicles (or general objects) have the assumed average sizes (distances) on average. In a preferred embodiment, observations (i.e. images and the coordinates and parameters obtained from them) are not immediately discarded, so that more and more object parameters have to be estimated. It is also possible to always use a fixed, i.e. constant, number of vehicle observations (objects).
[0025] Observations or objects can be discarded according to various schemes. For example, the age of the observation can be taken into account. Older observations can be discarded to allow for continuous re-estimation of the parameters. In general, the multiple images considered can only include recent images determined according to a predefined criterion, e.g., only images from the last ten minutes or the last 100 images. This is crucial, for example, if the orientation of the camera or the mount (of the camera) can change over time (e.g., due to thermal influences).
[0026] Likewise, for example, the coverage of the image can be taken into account. For calibration, it is advantageous if measurements or observations (i.e., objects present in the image) are available in the entire relevant area of the captured environment. This is especially true if a non-curved plane (on which the objects are located, e.g., a flat road) is assumed, but the actual road is slightly curved (e.g., due to an incline). In this case, an estimate of an average plane is often desirable.
[0027] Likewise, different vehicles or generally different objects, i.e., types of objects, can be taken into account. Since, for example, vehicle sizes can deviate from the assumed average values, different vehicles / observations should be used wherever possible. This can be ensured, for example, by tracking vehicles based on the generated top view. Furthermore, different vehicle (or person) classes can be used if an average size is available for them. In this case, for example, there is not just one specified distance overall, but one for each vehicle class (e.g., cars and trucks). In addition to the extrinsic camera parameters, a focal length (e.g., expressed by the aperture angle) can also be determined (or estimated) as an intrinsic parameter. For example, perspective effects of the individual 3D bounding boxes can be considered, e.g.the position of the vanishing points resulting from the connections of the vertices.
[0028] In general, other intrinsic parameters (e.g., principal point, distortion) can also be estimated. Since each individual projected 3D bounding box (here, especially assuming at least four corner points that do not lie in the same plane) already allows the determination of the extrinsic parameters (angle and height), the intrinsic parameters can be determined from the observation of multiple vehicles and world assumptions, such as a common ground plane, since in this case, the same extrinsic calibration parameters would have to be determined for all vehicles. A (systematic) deviation can therefore be used to determine (calibrate) the intrinsic parameters.
[0029] Above, particular attention was paid to four corner points at the bottom of the 3D bounding box and an average length and width measurement was assumed. Generally, however, only a single measurement (distance), e.g., only width or height, is required. The four upper corner points can also be used. These are located in a plane parallel to the road. Preferably, there should be at least three (connected) corner points and two lengths, or four points and one length. However, it is also possible to include observations without a length reference, e.g., for vehicles with a higher variance in length / width / height, such as vans or commercial vehicles. In this case, only the perpendicularity of the legs is exploited.
[0030] In the above explanations, one (main) plane (the roadway) was assumed. In general, however, it is also possible to assume multiple planes or curved surfaces. For curved surfaces, parametric surfaces (e.g. described by polynomials) can be used. At the same time, each image can be decomposed into grid cells, for example, and at least one camera parameter (in particular an extrinsic calibration) can be determined for each grid cell. Neighborhood properties between grid cells can also be exploited (e.g. smoothing, for example via a Markov random field). For these approaches, the approach described above can be modified so that each object or vehicle generates a hypothesis. These would then be clustered locally, and corresponding parameters derived for each cluster.This would provide each cluster or grid cell with its own, distinct extrinsic calibration (with corresponding calibration parameters). These can then be used in various ways, e.g., by interpolation with polynomials, like a changing surface (change in height and orientation).
[0031] One property utilized in the proposed method is the perpendicularity of the 3D bounding boxes and their position on the ground (a vehicle resting on the roadway). Therefore, 3D bounding boxes of people, trailers, or other objects, not just vehicles, can generally be used.
[0032] A preferred extension of the approach presented here is calibration via temporal tracking of objects. Even if no assumed length reference is present, the fact that an object's dimensions do not change over time can be exploited. For example, object dimensions can be determined based on the currently assumed calibration or the camera parameters. These can then be used in subsequent images. In particular, this can be embedded in a holistic process or estimator, where the object dimensions are also estimated but assumed to be constant over time.
[0033] A further application or extension is, for example, the automatic, simultaneous construction of maps that contain, for example, lane information. Simple methods (e.g., grid-based) based on the generated top view can be used for this. In summary, simple and versatile 3D bounding box corner point detectors are used for the fully automatic calibration of (particularly stationary) cameras. In particular, the projections of 3D bounding box coordinates into the image are used to derive an extrinsic and (in extension) also intrinsic calibration. The advantages here are in particular the direct estimation of the corner points of the enveloping cuboid (as keypoints) in contrast to the estimation of, for example, special prominent keypoints on vehicles (e.g., tires, license plates, or lights). This means that no adaptation to specific vehicle types (if necessary) is necessary.This makes it easier to extrapolate to other vehicle types (beyond a single truck / car / van class). Furthermore, fewer length references are required (only at least one distance must be known) than with other approaches. Classes with high variability can even be used without a length reference.
[0034] The approach generally allows for the estimation of extrinsics with respect to multiple planes or even curved surfaces and curves. The number of required predicted vertices (eight for a 3D bounding box) can be further reduced to three or even two. These can, for example, all be located on the floor. Another special feature is the proposed option of a single-stage detector (compared to multi-stage methods).
[0035] A computing unit according to the invention, e.g. a processor of a camera, is configured, in particular in terms of programming, to carry out a method according to the invention.
[0036] A camera according to the invention, in particular a traffic surveillance camera, has a computing unit according to the invention. However, it should be noted that a computing unit can also be provided separately from the camera, e.g., by outsourcing the process to a server as the computing unit (e.g., in the so-called cloud).
[0037] The implementation of a method according to the invention in the form of a computer program or computer program product with program code for carrying out all method steps is also advantageous, as this entails particularly low costs, in particular if an executing control unit is also used for other tasks and is therefore already present. Finally, a machine-readable storage medium is provided with a computer program stored thereon, as described above. Suitable storage media or data carriers for providing the computer program are, in particular, magnetic, optical, and electrical memories, such as hard disks, flash memories, EEPROMs, DVDs, and others. Downloading a program via computer networks (Internet, intranet, etc.) is also possible. Such a download can be wired or cable-based or wireless (e.g., via a WLAN network, a 3G, 4G, 5G, or 6G connection, etc.).
[0038] Further advantages and embodiments of the invention will become apparent from the description and the accompanying drawings.
[0039] The invention is illustrated schematically in the drawing using an embodiment and is described below with reference to the drawing.
[0040] Short description of the drawings
[0041] Figure 1 shows schematically an environment with a camera to explain an embodiment.
[0042] Figure 2 shows an image of an environment to explain another embodiment.
[0043] Figure 3 shows schematically a sequence of an embodiment.
[0044] Figure 4 shows diagrams of camera parameters that can be determined within the scope of one embodiment.
[0045] Embodiment(s) of the invention Figure 1 schematically shows an environment 110 with a camera 100 to explain an embodiment of the invention. In the environment 110, a roadway 112 is shown, on which an object 120, designed as a vehicle, is located, for example. The camera 100 has a computing unit 102 and is used, for example, as a traffic monitoring camera, ie the environment 100 and objects present or appearing therein, in particular vehicles such as the vehicle 120, are to be observed or monitored. For this purpose, calibration of the camera 100 may be necessary, in particular.
[0046] The vehicle 120 has, for example, a length 122 and a height 124, which are defined in an environmental coordinate system x, y, z. A width of the vehicle 120 is not shown here.
[0047] Figure 2 shows an image 200 of an environment to illustrate another embodiment of the invention. Image 200 may be an image of an environment with objects, particularly vehicles, captured by a camera such as camera 100 in Figure 1.
[0048] For example, image 200 shows a roadway 212, as well as a multitude of vehicles as objects. For example, a vehicle approaching the camera used to capture image 200 is designated 220, in this case a car. Another vehicle approaching the camera is designated 250, in this case a truck. Furthermore, a vehicle moving away from the camera is designated 252, in this case a car.
[0049] Using vehicle 200 as an example, a so-called 3D bounding box 230 will now be explained, as used in the context of the present invention. The 3D bounding box 230 is a cuboid that encloses the vehicle 220 (or an object in general). The 3D bounding box or cuboid thus comprises eight vertices or is defined by these vertices. By way of example, four of these vertices are designated 231, 232, 233, and 234. Two of the vertices are connected to one another by lines or edges (there are a total of 12 edges for a cuboid), with three edges being designated 241, 242, and 243, by way of example. These edges are perpendicular to one another at the vertices, and the cuboid or 3D bounding box lies on a plane defined by the roadway 212.
[0050] Because the 3D bounding box 230 encloses the vehicle 220, the lengths of certain edges correspond to corresponding maximum dimensions of the vehicle 220. Thus, the length of edge 241 (the distance between the vertices 231 and 232) corresponds to the width of the vehicle 220, the length of edge 242 (the distance between the vertices 232 and 233) corresponds to the length of the vehicle 220 (see also length 122 in Figure 1), and the length of edge 243 (the distance between the vertices 231 and 234) corresponds to the height of the vehicle 220 (see also height 124 in Figure 1).
[0051] By projecting the 3D bounding box 230, the corner points have coordinates in the image plane y', z', i.e., a 2D coordinate system in the camera plane and the sensor plane, respectively. Accordingly, these coordinates are in 2D, i.e., two-dimensional, while the 3D bounding box 230 itself is three-dimensional. The 3D bounding box 230 is thus projected into the 2D plane.
[0052] In the same way, 3D bounding boxes with corner points can be defined for the other vehicles, as shown in Figure 2. 3D bounding boxes can also be defined for other objects such as people or trailers. An estimated horizon is also shown at 260. The position of the horizon in the image is shown here for clarity, as it is easy to determine whether such a horizon is plausible. However, the position and shape of the horizon curve also directly depend on the estimated angles of the ground plane and the intrinsic camera parameters.
[0053] Figure 3 schematically shows a sequence of a method in one embodiment, which will be explained in more detail below, particularly with reference to Figure 2.
[0054] In a step 300, first one, but preferably several images, such as image 200, are provided. For this purpose, the images can be recorded with the camera. In a step 310, for example, for each of the objects such as vehicle 230 in image 200, the coordinates 312 of corner points of a respective projected 3D bounding box are determined, specifically in the image plane y', z'. For this purpose, an artificial neural network 314 (or another machine learning algorithm) can be used, for example, which receives images 200 as input data. As already mentioned, the coordinates of all eight corner points of each 3D bounding box can be determined, but two or three corner points (e.g., 231, 232, 233) are also sufficient. The corner points or their coordinates then also define distances between the corner points, e.g., the edges 241, 242.
[0055] Optionally, in a step 320, object parameters 322 of the vehicles can also be determined, e.g. their position on the roadway 212 and their orientation or rotation about a perpendicular to the roadway 212.
[0056] In a step 330, one, but preferably several camera parameters are then determined, e.g., the extrinsic camera parameters roll angle 332, pitch angle 334 and height of the camera above the ground (roadway) 336, and an intrinsic camera parameter 338, e.g., the focal length. This is done based on the distances between corner points (in the camera coordinate system x', y') and at least one corresponding, predetermined distance such as the length, height, or width of a vehicle (in the environmental coordinate system x, y, z). This can also be done, in particular, within the framework of an optimization process. The aforementioned object parameters 322 can also be optimized within the framework of this optimization process.
[0057] In step 340, the camera parameters obtained during this calibration are then provided. They can then be applied, if necessary, in step 350, to adjust the camera settings.
[0058] In Figure 4, diagrams show, among other things, the camera parameters roll angle
[0059] 332 (upper line) and pitch angle 332 (lower line), each in degrees, for example, in diagram (A); altitude 336, e.g., in meters, in diagram (B); and the number of detected objects (400) in diagram (C), each plotted against a number N of processed images or measurements. It can be assumed that the estimation or determination of the camera parameters becomes more accurate with the increasing number of images or measurements.
[0060] Diagram (D) also shows a top view of the environment in the environmental coordinate system (x, y), with vehicles or objects, i.e., their positions and orientations (object parameters), being shown in a field of view 410 of the camera (located at x=0, y=0). These object parameters, and thus such a view, can, as mentioned, be obtained, for example, during optimization. Repeatedly determined top views can thus be used to determine, for example, maps of the lanes, since a vehicle typically does not change lanes frequently.
Claims
Claims 1. A method for determining at least one camera parameter (332, 334, 336, 338) for calibrating a camera (100), in particular a traffic monitoring camera, comprising: Providing (300) one or more images (200) of an environment (110) taken by the camera (100), Determining (310), for a 3D bounding box (230) of one or each of a plurality of objects (220, 250, 252), in particular vehicles, in the one or more images (200), coordinates (312) of at least two corner points (231-234) in an image plane (y', z'), Determining (330) the at least one camera parameter (332, 334, 336, 338) based on at least one distance (241-243) in the image plane defined by the coordinates (312) of the at least two corner points (231-234) of the respective 3D bounding box (230) and at least one corresponding predetermined distance (122, 144) in an environmental coordinate system (x, y, z), and Providing the at least one camera parameter (332, 334, 336, 338).
2. The method according to claim 1, wherein the coordinates of at least three, in particular at least four, further in particular eight corner points for the 3D bounding box (230) of one or each of the plurality of objects are determined.
3. The method according to claim 1 or 2, wherein the at least one camera parameter (332, 334, 336, 338) is determined based on the at least one distance (241-243) defined by the coordinates (312) of the at least two corner points (231-234) of a plurality of 3D bounding boxes (230), and in particular within the framework of an optimization method, wherein a corresponding predetermined distance in the environmental coordinate system is used for the distances for the plurality of 3D bounding boxes, in particular for an average value of the distances of the plurality of 3D bounding boxes.
4. Method according to one of the preceding claims, further comprising: Determining (320) at least one object parameter (322) of the one or each of the plurality of objects (220, 250, 252) in the one or more images (200), wherein for determining (330) the at least one camera parameter (332, 334, 336, 338) the at least one camera parameter (332, 334, 336, 338) and the at least one object parameter (322) are optimized simultaneously within the framework of an optimization method.
5. The method according to any one of the preceding claims, wherein the at least one camera parameter (332, 334, 336, 338) is determined based on a plurality of images, wherein the plurality of images comprise only temporally current images determined according to a predetermined criterion.
6. Method according to one of the preceding claims, wherein at least one camera parameter is determined for different objects, in particular for each of several objects.
7. Method according to one of the preceding claims, wherein the determination of the coordinates of the at least two corner points (231-234) is carried out by means of a machine learning algorithm, in particular an artificial neural network, which receives the one or more images as input data.
8. The method according to any one of the preceding claims, wherein the at least one camera parameter (332, 334, 336, 338) comprises at least one extrinsic camera parameter, in particular at least one of: roll angle, pitch angle and camera height, and / or wherein the at least one camera parameter comprises at least one intrinsic camera parameter, in particular a focal length.
9. A computing unit (102) which is designed to carry out all the process steps of a To carry out the method according to one of the preceding claims.
10. Camera (100), in particular a traffic monitoring camera, with a computing unit (102) according to claim 9.
11. A computer program which causes a computing unit (102) to carry out all method steps of a method according to one of claims 1 to 8 when it is executed on the computing unit (102).
12. A machine-readable storage medium having stored thereon a computer program according to claim 11.