Camera coordinate system calibration

The method uses a 3D model and nested transformations to efficiently calibrate the camera's local coordinate system within a vehicle, addressing inefficiencies and robustness issues in existing methods, enhancing accuracy and reducing complexity.

JP2025541957APending Publication Date: 2025-12-24SMART EYE AB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025525328
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-02
Filing Date
2023-10-27
Publication Date
2025-12-24

AI Technical Summary

Technical Problem

Existing camera calibration methods for vehicle interiors are inefficient, require extensive manual effort, and are not robust, especially when dealing with multiple degrees of freedom, leading to false positives and increased complexity.

Method used

A method and system that utilize a 3D model of the vehicle interior to identify fixed physical structures, apply nested coordinate transformations, and calibrate the camera's local coordinate system efficiently, using a series of transformations to model mechanical mounting and minimize errors through least-squares minimization.

Benefits of technology

Provides a computationally efficient and robust calibration of the camera's local coordinate system, reducing reliance on specific vehicle interiors and minimizing errors, thus improving accuracy and reducing complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025541957000001_ABST
    Figure 2025541957000001_ABST
Patent Text Reader

Abstract

A method and system for calibrating a coordinate transformation CT between a reference coordinate system and a camera coordinate system, the reference coordinate system relating to a vehicle interior to which a camera is mounted, the method including the steps of obtaining a 3D model of the vehicle interior, the 3D model including at least one physical structure identifiable in the image, acquiring an image of the vehicle interior using a camera, identifying at least one feature in the image corresponding to one of the physical structures and selecting a set of pixel coordinates for the feature, each pixel coordinate associated with a particular point in the 3D model, forming a 2D projection of the particular points onto the camera image plane by applying the coordinate transformation and a model of the camera optics, and calibrating the CT based on the relationship between the selected set of pixel coordinates and the 2D projection.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the calibration of the local coordinate system of a camera mounted inside a vehicle. [Background technology]

[0002] In vehicle imaging applications such as driver monitoring systems (DMS), cameras are mounted inside the vehicle. The exact location of the camera is not necessarily fixed or known. For example, if the camera is mounted on the steering column, it is often moved in at least two degrees of freedom (axial and elevation). If the camera is mounted on the rearview mirror, it moves whenever the mirror is adjusted. And even in a stationary position, such as on the dashboard, the camera may be unintentionally displaced.

[0003] For this reason, it is important to frequently update or calibrate the camera's position and orientation (i.e., the camera's local coordinate system) with respect to a known reference coordinate system (i.e., the coordinate system inside the vehicle). The camera's local coordinate system is often called the camera coordinate system (CCS). Calibrating the CCS is fundamental for the correct functioning of features such as head tracking, eye tracking, and visual targets. An example of an existing solution is provided by document WO 2018 / 000037.

[0004] Existing solutions are designed to track features in an image using template matching. Using the 2D displacement of the tracked feature as input, linear interpolation of the CCS is performed for the four ends of the steering column. The interpolated result is considered the calibrated CCS. Unfortunately, finding suitable features within a vehicle interior can be difficult or even impossible. Therefore, this approach requires extensive experience and manual effort to configure the system, e.g., to find suitable features / templates. In addition, satisfactory configurations are often application-specific and may not work well in different vehicle interiors. Even with a carefully designed configuration, template matching is not very robust, resulting in false positives, such as on the back of the steering wheel. To increase robustness, complex control logic and filtering are required, which increases complexity and cost.

[0005] Furthermore, in conventional systems, camera calibration is typically limited to two degrees of freedom, e.g., position in / out and tilt up / down. As mentioned above, calibration is often required in more complex situations and with more degrees of freedom. Summary of the Invention [Problem to be solved by the invention]

[0006] The object of the present invention is to alleviate the above problems and provide a more computationally efficient calibration of CCS. [Means for solving the problem]

[0007] According to a first aspect of the present invention, this and other objects are achieved by a method for calibrating a coordinate transformation (CT) between a reference coordinate system and a camera coordinate system, the reference coordinate system being related to the interior of a vehicle to which a camera is attached. The method includes the steps of obtaining a 3D model of the vehicle interior, the 3D model including at least one physical structure identifiable in the image; acquiring an image of the vehicle interior using a camera; identifying at least one feature in the image corresponding to one of the physical structures and selecting a set of pixel coordinates for the feature, each pixel coordinate associated with a specific point in the 3D model; forming a 2D projection of the specific point onto the camera image plane by applying the coordinate transformation and a model of the camera optics; and calibrating the CT based on the relationship between the selected set of pixel coordinates and the 2D projection. The coordinate transformation is represented as a series of nested transformations, each representing a possible way for the camera to move, such that the coordinate transformation models the mechanical mounting of the camera.

[0008] Typically, the first transform represents the fixed position of the mount. Such a transform can have up to six degrees of freedom (DOF) and represent mount tolerances or displacement from the expected position. Other transforms may have limited degrees of freedom due to the mechanical mount. Typical motion limitations include rotation around only one angle, translation along one axis, etc.

[0009] In one example, the camera is mounted on a steering column, in which case the nested transforms may include a first transform representing the (typically known, or at least approximately known) pivot point of the column, a second transform representing the rotation (tilt) of the column, and a third transform representing the translation (extension) of the column.

[0010] In another example, the camera is mounted at a 3D pivot point, e.g., a rearview mirror. In this case, the nested transformations may include a first transformation representing the pivot point (typically known, or at least approximately known), and a second transformation with three rotational degrees of freedom.

[0011] By expressing the coordinate transformations as a set of nested transformations, CT calibration becomes computationally more efficient; for example, the numerical solution may converge faster. It is important to note that it is not necessarily a limitation on the number of DOFs (the first transformation may often have 6 DOFs); rather, motions with a larger range of expected motions are modeled by more constrained transformations (e.g., a single DOF).

[0012] The 3D model can include geometric data (e.g., CAD data) of the vehicle, providing an accurate description of all mechanical structures. The features to be identified can then be defined (constructed) by identifying the 3D descriptions of the corresponding structures in the CAD data. This ensures more robust identification and tracking that is less dependent on specific vehicle interior conditions such as color, surface characteristics, etc.

[0013] It should be noted that the 2D projection is not necessarily explicitly determined, but may occur implicitly during calibration. For example, the calibration step may include creating a set of equations, each of which defines one element of an error vector as the difference between one of the pixel coordinates and the associated 2D projection, and minimizing the error vector, for example, by a least-squares method. This set of equations may be solved numerically. The equations may be nonlinear.

[0014] The number of equations in this set is determined by the number of features identified in the image and localized to corresponding features in the projected 3D model. The features selected should be stationary (i.e., non-moving) physical structures within the vehicle that can be reliably identified in the image and easily extracted from the projected 3D model. Examples of such features include the shape and size of B- or C-pillars, door handles, center console, parts of the rear seats, etc.

[0015] Calibration is not limited to cameras mounted on the steering column: in fact, CCSs (and transform CTs) can generally be described and calibrated in six degrees of freedom (orientation and position).

[0016] According to a second aspect of the present invention, the above object is achieved by a system for calibrating a coordinate transformation CT between a reference coordinate system and a camera coordinate system, the reference coordinate system relating to an interior of a vehicle, the system comprising: a camera mounted inside the vehicle; and a controller for controlling the camera to acquire images of the vehicle interior. The system further comprises a processing circuit configured to obtain a 3D model of the vehicle interior, the 3D model including at least one physical structure identifiable in the image, the processing circuit configured to identify at least one feature in the image corresponding to one of the physical structures and select a set of pixel coordinates of the feature, each pixel coordinate associated with a specific point in the 3D model, the processing circuit configured to form a 2D projection of the specific point onto the camera image plane by applying the coordinate transformation CT and a model of the camera optics, and the processing circuit configured to calibrate the coordinate transformation CT based on the relationship between the selected set of pixel coordinates and the 2D projection.

[0017] The present invention will now be described in more detail with reference to the accompanying drawings, which show presently preferred embodiments of the invention. [Brief explanation of the drawings]

[0018] [Figure 1] Schematic diagram of the eye-tracking system mounted on the dashboard of a vehicle. [Figure 2] Figure 1. A more detailed diagram of the eye-tracking system. [Figure 3] 1 is a flowchart of a method according to an embodiment of the present invention. [Figure 4] FIG. 1 illustrates three nested transformations representing a model of a steering column. [Figure 5a] FIG. 1 illustrates a calibration process according to an embodiment of the present invention. [Figure 5b]FIG. 1 illustrates a calibration process according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0019] Embodiments of the present invention will now be discussed with reference to an eye tracking system, however the principles of the present invention are equally useful in any application in which a camera is mounted inside a vehicle, for example, any driver monitoring system (DMS) or cabin monitoring system (CMS).

[0020] FIG. 1 shows a driver 1 of a vehicle. A camera 2 is mounted in front of the driver 1, here on the steering column 3. Alternatively, the camera 2 may be mounted on the dashboard, fixed to the ceiling, or in any other location suitable for a particular application. The camera forms part of an imaging system 4 (see FIG. 2) used to capture images of the driver and / or the interior of the vehicle. For example, the system 4 may be a driver monitoring system (DMS).

[0021] Referring to FIG. 2, the components of the imaging system 4 are shown in more detail. The camera 2 includes an image sensor 5, e.g., a CMOS image sensor, and appropriate optics 6. The optics are configured to project incident light onto an image plane of the sensor 5. In the illustrated example, the system further includes at least one light source 7 having a known geometric relationship to the sensor 5. The light source 7 is typically configured to emit light outside the visible range, such as infrared (IR) or near infrared (NIR). The light source may also be a solid-state light source, such as an LED. In the illustrated example, the light source 7 is an LED configured to emit light having a light spectrum centered in a 50 nm band centered around 850 or 940 nm (NIR). Additionally, an optical bandpass filter, e.g., an interference filter, may be positioned between the user and the sensor 5. The filter (not shown) is configured to have a pass band substantially corresponding to the emission spectrum of the light source 7. Thus, in the above example, the filter 6 should have a pass band centred around 850 nm or 940 nm, for example a pass band of 825-875 nm or 915-965 nm.

[0022] The controller 8 is connected to the camera 2 and the LED 7 and is programmed to control the sensor 5 to capture successive images under illumination by the LED 7. Typically, the LED 7 is driven at a given duty cycle, and the controller 8 then controls the sensor 5 to capture images in synchronization with the light pulses from the LED 7.

[0023] The system further comprises processing circuitry 9 (also referred to as a processor) and memory 10. The memory stores program code executable by the processor 9, enabling the processor to receive and process images acquired by the sensor 5. The processor 9 may be configured to determine and track eye pose to determine a user's gaze direction, i.e., where the user is looking. The system of FIG. 1 has many different applications, including, for example, automotive applications where the driver's eyes are tracked for safety reasons, and various human-machine interfaces.

[0024] During eye tracking execution, the user 1 is illuminated by a light source 7, and light reflected from the object (the user's face) is received by the imaging sensor 5 through the camera optics 6. Note that most of the ambient light is blocked by a filter, thereby reducing the power required for the light source.

[0025] Gaze direction can be determined by determining the head pose (the position and orientation of the head in a reference coordinate system (RCS)) and then the eye pose (the position and orientation of the eyes relative to the head). In simple cases, sometimes called estimated eye tracking, the eye pose is determined based on the position of the iris relative to the head. However, in many applications, more accurate gaze detection can be obtained by using a light source. Under light source illumination, the acquired image contains the reflection (glint) of the eye's cornea, which can be used to more accurately determine gaze.

[0026] In the illustrated embodiment, a 3D model 11, for example a CAD model, of the vehicle interior is stored in memory 10. The 3D model 11 includes geometric data of the vehicle interior expressed in a reference coordinate system RCS.

[0027] The position of camera 2 can also be expressed in a reference coordinate system, called the camera coordinate system (CCS). Knowing the coordinate transformation CT between RCS and CCS, any point in the 3D model can be located in the camera coordinate system by applying this coordinate transformation. Using knowledge of the camera optics, any point in the camera coordinate system can be projected onto the camera's image plane.

[0028] In FIG. 1, the reference coordinate system RCS is indicated at 12 while the camera coordinate system CCS is indicated at 13.

[0029] The operation of the system will now be discussed with reference to Figures 3 to 5. The steps of Figure 3 may be performed by processing circuitry 9.

[0030] In a first initialization step S10, a 3D model 11 of the vehicle interior is obtained and stored in memory 10. In embodiments where the camera is mounted on a pivotable steering column 3, this model may include the location of the pivot point P.

[0031] The calibration process begins in step S11 by acquiring an image 20 of the vehicle interior using camera 1. In step S12, at least one geometric feature 21 is identified in image 20 using appropriate image processing. Each feature corresponds to a fixed physical structure within the vehicle interior that is present in the 3D model. Suitable structures may be, for example, a B-pillar 14 (see FIG. 1) or C-pillar, a door handle, or any other non-moving physical structure that can be easily identified in an image of the vehicle interior. Note that movable structures, such as front seat backs, are less suitable for use in the calibration process. After the features 21 are identified, a set of pixel coordinates 22 for each identified feature 21 is selected. Each selected pixel coordinate 22 corresponds to a specific point within the 3D model.

[0032] Identifying features in an image can be achieved using traditional template tracking, but this requires creating templates using actual images from inside the vehicle. A more sophisticated approach is to use a neural network system trained to identify a set of features in an image (landmark detection). Neural network systems can be trained on many different vehicle interiors, making them robust to vehicle variations (e.g., color, shape, etc.). Feature detectors using such trained neural networks can also provide an internal confidence signal that represents how similar the detected feature is to the one they were trained to identify. A low confidence signal can indicate that a particular feature was not correctly identified, for example, because it was not visible (the feature was not in the field of view). As an alternative approach, neural networks can be trained to perform pixel segmentation and extract stable inter-segment boundaries.

[0033] In step S13, a current (or initial) approximation of the transformation CT is applied to specific points of the 3D model to locate these points in the camera coordinate system. Then, in step S14, a mathematical model of the camera optics 6 is used to form 2D projections 23 of these specific points onto the image plane of the camera 2. Thus, each 2D projection 23 of a specific point is a mathematical representation obtained by applying the transformation CT and the model of the camera optics to selected points of the 3D model data.

[0034] In step S15, the CT is calibrated (updated) based on the relationship between the selected pixel coordinates 22 and the corresponding 2D projections 23.

[0035] Step S15 may be performed as a single calculation by a processing circuit, but for the sake of understanding the principle, step S15 is divided into sub-steps S16 and S17 here.

[0036] In step S16, a set of equations is formed, each of which defines one element of an error vector as the difference between 1) one of the pixel coordinates identified in the image and 2) the 2D projection of the point in the 3D point model data associated with that pixel coordinate. The equations further include a calibration vector containing one variable for each degree of freedom.

[0037] In step S17, the error vector is minimized, for example, by the least squares method, thereby calibrating the coordinate transformation.

[0038] As explained earlier, the solution to the minimization problem has six variables, one for each degree of freedom (three positions and three rotations). However, this process can be made more effective and robust by modeling the mechanical mounting of the camera using a series of nested transformations. Each transformation then corresponds to one possible way for the camera to move, possibly with limited degrees of freedom.

[0039] In this example, where camera 2 is mounted on steering column 3, the coordinate transformation can be modeled as three nested transformations: 1) position P, 2) rotation in the vertical plane (tilt, Φ), and 3) scaling (z). This is shown in Figure 4. As mentioned previously, the pivot point P of the steering column can be assumed to be a known point (known coordinate) in the 3D model data. However, mounting tolerances (or unexpected displacements) can be included as an additional six degrees of freedom (see below).

[0040] Using diagrams, the process of determining the CCS can be shown as a two-step process, as shown schematically in Figures 5a-5b. Figure 5a shows an edge 21 of the B-pillar 14 identified in the image 20 and a set of pixel coordinates 22 along the edge 21. These pixel coordinates therefore correspond to specific points in the 3D model 11. Figure 5b shows 2D projections 23 of two of these specific points in the image plane. Because the coordinate transformation CT at this time is not calibrated, the 2D projections 23 are not properly aligned with the corresponding pixel coordinates 22.

[0041] For purposes of explanation, CT calibration can be viewed as an adjustment of the camera coordinate system CCS, as shown in Figure 4. During calibration, the CCS is rotated in the vertical plane to minimize errors as much as possible, thereby aligning pixel coordinates 22 and projections 23 in the plane. The CCS is also translated in the z direction to further minimize errors.

[0042] In practice, the process of calibrating a CT is typically performed in a single computational process. For a camera mounted on a steering column, this process can be expressed as a minimization problem with two variables: minf(Φ,z), where the function f involves a first (known) transformation of the pivot point P of the steering column and two additional geometric transformations. The second transformation describes the rotation about the pivot point, while the third transformation describes the translation along the z-axis. The solution of this problem gives the tilt and translation z.

[0043] Limited calibration may be allowed in other degrees of freedom to accommodate camera mounting tolerances. By introducing these additional DOF into the quadratic transformation, the minimization problem can be expressed as minf(tilt angle, 6DOF).

[0044] The calibration process uses the most recent calibration as the starting point for each successive iteration. In the example shown above, the CT contains three nested transformations (pivot, rotation, and translation). The pivot transformation remains unchanged, but the second and third transformations are calibrated with each iteration.

[0045] In another example, where the camera is mounted on a rearview mirror, the CT includes a first transform representing the pivot point of the mirror and at least one transform representing the rotation about the pivot point. If appropriate, there may be a separate transform, one for each rotational degree of freedom.

[0046] Those skilled in the art will appreciate that the present invention is in no way limited to the above-described embodiments. On the contrary, many modifications and variations are possible within the scope of the appended claims. For example, the system may include multiple cameras that are calibrated.

Claims

1. A method for calibrating a current coordinate transformation CT between a reference coordinate system (12) and a camera coordinate system (13), said reference coordinate system being relative to the interior of a vehicle in which a camera (2) is mounted, said method comprising: obtaining a 3D model of the vehicle interior, the 3D model including geometric data of the vehicle interior expressed in the reference coordinate system (12), the geometric data including at least one physical structure (14) identifiable in the image; acquiring an image (20) of the interior of the vehicle using the camera (2); identifying at least one feature (21) in the image (20) that corresponds to one of the physical structures (14) and selecting a set of pixel coordinates (22) of the feature, each pixel coordinate being associated with a particular point in the 3D model; forming a 2D projection (23) of said particular point onto the camera image plane by applying said coordinate transformation and a model of the camera optics (6); calibrating the current CT based on a relationship between the set of selected pixel coordinates (22) and the 2D projection (23); The method, wherein the coordinate transformation is expressed as a series of nested transformations such that the coordinate transformation models the mechanical mounting of the camera.

2. said step of calibrating further comprising: creating a set of equations, each equation defining one element of an error vector as the difference between one of said pixel coordinates and an associated 2D projection; minimizing the error vector, for example by the least squares method; The method of claim 1 , comprising:

3. 3. The method of claim 1, wherein the camera is mounted on a steering column, and the series of nested transformations includes a first transformation representing the position of a pivot point P of the column, a second transformation representing a rotation of the column, and a third transformation representing a translation along the column.

4. 3. The method of claim 1, wherein the camera is mounted on a rearview mirror, and the series of nested transformations includes a first transformation representing a pivot point of the mirror and at least one second transformation representing a rotation about the pivot point.

5. The method of claim 4 , wherein the at least one second transformation comprises a separate transformation for each rotational degree of freedom.

6. 6. The method of claim 1, wherein the 3D model is a CAD model of the vehicle, and each identified feature is defined by identifying a 3D description of a corresponding structure in the CAD data.

7. The method of any one of claims 1 to 6, wherein the at least one physical structure includes at least one of a B-pillar, a C-pillar, a center console, a portion of a rear seat, and a door handle.

8. 1. A system for calibrating a current coordinate transformation CT between a reference coordinate system (12) and a camera coordinate system (13), said reference coordinate system being related to the interior of a vehicle, said system comprising: a camera (2) mounted inside the vehicle; a controller (8) for controlling the camera (2) to acquire an image (20) of the interior of the vehicle; a processing circuit (9); the processing circuitry (9) is configured to obtain a 3D model of the vehicle interior, the 3D model including geometric data of the vehicle interior expressed in the reference coordinate system (12), the geometric data including at least one physical structure (14) identifiable in the image (20); the processing circuitry (9) is further configured to identify at least one feature (21) in the image (20) corresponding to one of the identifiable physical structures and to select a set of pixel coordinates (22) of the feature, each pixel coordinate relating to a particular point in the 3D model; the processing circuitry (9) is further configured to form a 2D projection (23) of the particular point onto a camera image plane by applying the coordinate transformation CT and a model of the camera optics (6); the processing circuitry (9) is further configured to calibrate the current coordinate transformation CT based on a relationship between the set of selected pixel coordinates (22) and the 2D projection (23); The system, wherein the coordinate transformation is represented as a series of nested transformations such that the coordinate transformation models the mechanical mounting of the camera.

9. the processing circuitry creating a set of equations, each equation defining one element of an error vector as the difference between one of said pixel coordinates and an associated 2D projection; minimizing the error vector, for example by the least squares method; The system of claim 8 , configured to calibrate the coordinate transformation by:

10. 10. The system of claim 8 or 9, wherein the camera is mounted on a steering column, and the series of nested transformations includes a first transformation representing the position of a pivot point P of the column, a second transformation representing the rotation of the column, and a third transformation representing the translation of the column.

11. 10. The system of claim 8 or 9, wherein the camera is mounted on a rearview mirror, and the series of nested transformations includes a first transformation representing a pivot point of the mirror and at least one second transformation representing a rotation about the pivot point.

12. The system of claim 11 , wherein the at least one second transformation comprises a separate transformation for each rotational degree of freedom.

13. 13. The system of claim 8, wherein the 3D model is a CAD model of the vehicle, and each identified feature is defined by identifying a 3D description of a corresponding structure in the CAD data.

14. 14. The system of any one of claims 8 to 13, wherein the at least one physical structure includes at least one of a B-pillar, a C-pillar, a center console, a portion of a rear seat, and a door handle.