Method for determining a line of sight of a driver in a vehicle
The method leverages existing vehicle sensors and an augmented reality projector to project a coded pattern for 3D reconstruction and triangulation, addressing the limitations of existing eye-tracking methods by providing accurate gaze detection without additional hardware or calibration, suitable for various ADAS applications.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-08-25
- Publication Date
- 2026-03-26
AI Technical Summary
Existing eye-tracking methods in vehicles require additional hardware components, calibration, and large amounts of training data, and are prone to false positives, especially during head rotations and varying environmental conditions.
A method utilizing existing vehicle sensors and an augmented reality projector to project a coded pattern onto the windshield, enabling 3D reconstruction and triangulation to determine the driver's line of sight without additional hardware or calibration, using a cornea reflective image and triangulation vectors.
Stable against head rotations and environmental variations, requiring no additional hardware or calibration, and providing accurate gaze detection without training data, suitable for various ADAS applications.
Smart Images

Figure EP2025074096_26032026_PF_FP_ABST
Abstract
Description
[0001] Mercedes-Benz Group AG
[0002] Method for determining a driver's line of sight in a vehicle
[0003] The invention relates to a method for determining a driver's line of sight in a vehicle according to the preamble of claim 1.
[0004] Eye-tracking systems for vehicles often rely on the use of optical sources (most commonly infrared diodes). Based on these projections, the position of the eye is determined, and the driver's line of sight is derived. One of the most common methods is Purkinje-based eye tracking, which analyzes characteristic reflections and their relative positions to estimate the eye's position relative to the projection. Disadvantages include the need for additional IR diodes and / or other hardware components, and the requirement to know all Purkinje reflections. This can be particularly challenging when the driver's head is rotating. Another approach involves detecting the pupil using conventional image processing algorithms (such as Canny-Edge detectors).In this method, the pupil's silhouette is extracted, and its relative position is inferred from the deformation of this silhouette. However, this requires initial calibration for individual adjustment. Furthermore, data-driven systems exist that, for example, can extract the relative position of the eye from a 2D image using convolutional neural networks. However, these require a relatively large amount of data, and false positives are possible if the network is not sufficiently generalized.
[0005] In Purkinje-based eye tracking, a light source (for example, an IR diode) creates a reflection on the cornea. These reflections are also called Purkinje reflections (a total of four reflections resulting from the refraction of light within and outside the cornea and lens). These reflections and their relative positions can be used to determine the eye's position relative to the recording system. A disadvantage of this method is the additional requirement for a light source and the need to detect the four reflections, which is particularly difficult during large head rotations.
[0006] Furthermore, data-driven approaches are known. These involve training neural networks to derive gaze from 2D image captures. In many cases, convolutional neural networks are used, with, for example, the results of Purkinje detection serving as ground truth for the training phase. A disadvantage of this approach is the large amount of training data required and the need to account for significant variance in that data. Otherwise, a high number of false positives and incorrect gaze determinations will result.
[0007] Methods for carrying out eye tracking in a vehicle are known from WO 2016 / 12458 A1 and DE 10 2014 009638 A1.
[0008] MORANO, Raymond A. [et al.]: Structured light using pseudorandom codes. In: IEEE transactions on pattern analysis and machine intelligence, Vol. 20, 1998, No. 3, pp. 322-327. ISSN 1939-3539. https: / / doi.org / 10.1109 / 34.667888 [accessed on 2025-07-09] describes the solution to the correspondence problem in active stereo vision using pseudorandomly coded structured light.
[0009] DE 102024 002 430 B3 describes a method for operating a
[0010] A field-of-view display device in a vehicle, wherein an in-vehicle computing unit generates an augmented reality display on a projection surface of the field-of-view display device for a driver. This display provides a contact-analog representation of an object marker with an object detected by the vehicle based on an evaluation of environmental sensor data, and wherein a positional difference perceptible to the driver between the object marker and the object within the projection surface is determined. To determine this positional difference, the computing unit considers the reflection of the object marker from the projection surface, as detectable in a corneal reflection image of the driver.
[0011] The invention is based on the objective of providing a novel method for determining a driver's line of sight in a vehicle, as well as a novel vehicle. This objective is achieved according to the invention by a method for determining a driver's line of sight in a vehicle with the features of claim 1 and by a vehicle with the features of claim 9.
[0012] Advantageous embodiments of the invention are the subject of the dependent claims.
[0013] A method for determining a driver's line of sight in a vehicle is proposed, wherein the vehicle's imaging sensors, including at least one sensor for capturing an environment, an augmented reality projector for a head-up display, and a driver observation camera are extrinsically calibrated with respect to a coordinate system of the vehicle, and the environment is captured by the at least one sensor and reconstructed three-dimensionally. According to the invention, a cornea reflective image (CRI) is derived from 2D image data captured by the driver observation camera, wherein the pupil center of one of the driver's eyes is detected in the CRI image. A coded pattern is cyclically projected onto the vehicle's windshield by the augmented reality projector to establish a correspondence with the CRI image, and characteristic features are detected in the CRI image at a first time point.wherein a mapping of the characteristic features to the coded patterns is performed, wherein triangulation vectors are calculated based on the 3D point projection by the augmented reality projector and the focal point features on the 2D image plane of the driver observation camera, wherein a 3D detection is performed using the triangulation vectors for all characteristic features to determine a depth to the characteristic features, wherein an ellipse is fitted into the determined depths, having its main vertex at the pupil center, to determine the 3D position of the eye in the vehicle's coordinate system and thus the driver's line of sight.
[0014] In one embodiment, the CRI image is generated at least substantially as an image and / or reflection of the environment on the cornea of the driver's eye.
[0015] In one embodiment, post-processing of the CRI image is performed, including rectification of a spherical projection. In another embodiment, a camera, a lidar sensor, and / or a radar sensor is used as a sensor for environmental detection.
[0016] In one embodiment, a coordinate system is used whose coordinate origin is located at the center of a front axle of the vehicle.
[0017] In one embodiment, a 3D vector is formed by deriving the CRI image from 2D image data captured by the driver observation camera, which is determined by a focal point of the driver observation camera and a feature point on a 2D camera image plane of the driver observation camera, from which the pupil center of the driver's eye is detected.
[0018] In one embodiment, a Morano code is used as the coded pattern.
[0019] In one embodiment, edges are detected as characteristic features.
[0020] According to one aspect of the present invention, a vehicle is proposed comprising imaging sensors, including at least one sensor for sensing an environment, an augmented reality projector for a head-up display, a driver observation camera, and a windshield. According to the invention, the vehicle is configured to carry out the method described above.
[0021] In one embodiment, the sensor for detecting an environment is designed as a camera, a lidar sensor and / or a radar sensor.
[0022] The present invention disclosure provides an approach to enable eye-tracking in vehicles using existing hardware components, without requiring training data and eye-tracking-specific calibration.
[0023] The solution according to the invention allows the detection of a driver's gaze direction based on 3D information, is stable against head rotations, requires no eye-tracking-specific calibration, no additional IR diodes, no additional human intervention, no further hardware components than those already in series production, and is independent of any existing
[0024] Scenery and lighting and weather conditions. Furthermore, it can be repeated cyclically within a control loop and allows input data for a variety of other ADAS applications (e.g., Attention Assist, etc.).
[0025] Exemplary embodiments of the invention are explained in more detail below with reference to drawings.
[0026] This shows:
[0027] Fig. 1 shows a schematic flowchart of a method for performing eye-tracking in a vehicle,
[0028] Fig. 2 shows a schematic view of a projection of a coded pattern by an augmented reality projector onto a vehicle windshield.
[0029] Fig. 3 shows a schematic view of a CRI image with detected characteristic features,
[0030] Fig. 4 shows a schematic view of the CRI image and the projection of the coded pattern by the augmented reality projector onto the vehicle's windshield, with certain characteristic features assigned to specific points of the coded pattern.
[0031] Fig. 5 shows a schematic view of a driver's eye with a pupil center in the driver's CRI image, and
[0032] Fig. 6 is a schematic view to illustrate an ellipse fitting.
[0033] Corresponding parts are marked with the same reference symbols in all figures.
[0034] Figure 1 is a schematic flowchart of a procedure for performing eye-tracking in a vehicle.
[0035] The invention describes a method for performing eye-tracking in a vehicle based on existing hardware components. The method utilizes coded passive triangulation with the following steps VS1 to VS10: The present invention provides an approach to extract a driver's line of sight (eye-tracking) based on 3D sensor data and a cornea reflective image (CRI).
[0036] In step VS1, the hardware components involved, in particular at least one sensor (e.g., a camera, lidar sensor, and / or radar sensor), an augmented reality (AR) projector, especially for a head-up display, and / or a driver observation camera, are extrinsically calibrated (calibration to determine the orientation and translation of the components) so that the orientation and position of the hardware components are defined within a specific coordinate system. For example, a coordinate system is used whose origin is located at the center of the vehicle's front axle (e.g., MB CAD). This calibration can be performed during the vehicle's manufacturing process.
[0037] In step VS2, during field operation of the vehicle, a 3D reconstruction of the vehicle's surroundings is performed using the vehicle's environmental sensors. The scene around the vehicle is reconstructed three-dimensionally (based on data from at least one sensor, such as the camera, lidar sensor, and / or radar sensor).
[0038] If a resulting 3D image is available, a CRI image CRI-B (Cornea Reflective Image) is generated in a further step, VS3. This image essentially represents a projection and / or reflection of the environment onto the driver's cornea. Post-processing of the CRI image CRI-B can be performed in step VS3, for example, rectification of a spherical projection. By deriving the CRI image CRI-B from 2D image data captured by the driver observation camera, a 3D vector can already be created (focal point of the driver observation camera and feature point on a 2D camera image plane of the driver observation camera). From this vector, the pupillary center PZ of the driver's eye A can be detected in step VS4. However, two vectors may be relevant for triangulation.
[0039] Therefore, in one step, VS5 projects a coded pattern KM using the augmented reality projector to establish the correspondence between the 3D point and the reflection on the CRI image CRI-B. Different coded patterns can be used. For example, the Morano code, which is essentially a spatial color coding, can be used.
[0040] Figure 2 is a schematic view of the projection of the coded pattern KM by the augmented reality projector onto a windshield WSS of the vehicle.
[0041] The projection with the augmented reality projector is repeated at a certain cycle time, which is not visible to the driver. However, the projection must not be continuously displayed in order not to disrupt and / or block the 3D image on the CRI-B image.
[0042] The coded projection now reveals how each 3D object hits the windshield WSS (correspondence to the CRI image CRI-B), allowing a vector to be constructed and triangulation to be performed based on the 2 x 3D points (i.e., two 3D points, e.g., CM1 and CM6 in Figure 5).
[0043] In a further step VS6, feature detection takes place in the CRI image CRI-B at time T1. Characteristic features in the CRI image CRI-B (e.g., edges, etc.) are detected. For this purpose, the CRI image CRI-B is searched for characteristic features CM1 to CM6. For example, edges are sought. Figure 3 is a schematic view of the CRI image CRI-B at time T1 with detected characteristic features CM1 to CM6, in particular edges.
[0044] In step VS7, a correspondence mapping is performed in the CRI image CRI-B (at time T1+1). This involves mapping the characteristic features CM1 to CM6 to the coded patterns KM in order to establish triangulation vectors. Specifically, the characteristic features CM1 to CM6 from the CRI image CRI-B are mapped to the coded patterns KM projected by the augmented reality projector. Figure 4 is a schematic view of the CRI image CRI-B and the projection of the coded pattern KM by the augmented reality projector onto the vehicle's windshield WSS, where certain characteristic features CM3 and CM4 are assigned to specific points of the coded pattern KM. In step VS8, the triangulation vectors are calculated based on the 3D point projection by the augmented reality projector and the focal point features on the 2D image plane of the driver observation camera.
[0045] In step VS9, a 3D detection is performed using the triangulation vectors calculated in step VS8 for all characteristic features CM1 to CM6 generated in step VS6. Based on these characteristic features, 3D triangulation is carried out, thus determining the depth to these features. Figure 5 is a schematic view of the driver's eye A with the pupil center PZ in the driver's CRI image CRI-B. Furthermore, two three-dimensional characteristic features CM1, CM6, which were detected by the sensors, the characteristic features CMT, CM6' detected by means of the coded pattern KM on the windshield WSS, the optical center OZ of the driver observation camera, the characteristic features CM1“, CM6“ represented in the CRI image CRI-B on an image plane BE and the characteristic features CMT“, CM6'“ reconstructed three-dimensionally on the driver CRI image CRI-B are shown.
[0046] In step VS10, an ellipse fitting and gauze determination (driver's view) are performed. The ellipse fitting is carried out based on the depth points generated in step VS9.
[0047] In step VS10, an ellipse E1,E2 is fitted into the 3D space reconstructed in step VS9. This ellipse has its main vertex at the pupil center PZ. This pupil center PZ was previously determined from the CRI image CRI-B. This establishes the 3D position of the eye A in the vehicle reference system (the vehicle's coordinate system), allowing the line of sight to be determined and passed on to the subsequent driving functions.
[0048] Figure 6 is a schematic view of the three-dimensionally reconstructed environment 3DU with three-dimensional characteristic features CM1, CM2, CM5, CM6 detected by the sensors, the characteristic features CMT, CM2' detected by means of the coded pattern KM on the windshield WSS, the optical center OZ of the driver observation camera, the characteristic features CM1", CM5" represented in the CRI image CRI-B on an image plane BE, and the three-dimensionally reconstructed characteristic features CMT", CM2'", CM5'", CM6'", with a first one defined by the reconstructed characteristic features
[0049] ellipse E1 fitted by CM2'" and a second ellipse E2 fitted by the reconstructed characteristic features CM5'", CM6'".
[0050] This approach has the advantage of determining the eye position in 3D space. It is also relatively stable against head rotation and requires no eye-tracking-specific calibration. Furthermore, it requires no additional hardware components and does not use IR diodes. It is also relatively independent of weather conditions and the surrounding environment.
[0051] Mercedes-Benz Group AG
[0052] Reference symbol list
[0053] 3DU three-dimensionally reconstructed environment
[0054] A eye
[0055] BE Image plane
[0056] CM1 to CM6 characteristic feature
[0057] CMT to CM6' characteristic feature
[0058] CM1" to CM6" characteristic feature
[0059] CM1'“ to CM6'“ characteristic feature
[0060] CRI-B CRI image
[0061] E1, E2 Ellipse
[0062] KM coded pattern
[0063] OZ Optical Center
[0064] PZ Pupillary Center
[0065] VS1 to VS10 step
[0066] WSS windshield
Claims
Mercedes-Benz Group AG Patent claims 1. A method for determining a driver's line of sight in a vehicle, wherein the vehicle's imaging sensors, including at least one sensor for capturing an environment, an augmented reality projector for a head-up display, and a driver observation camera are extrinsically calibrated with respect to a coordinate system of the vehicle, wherein the environment is captured by means of the at least one sensor and reconstructed three-dimensionally, characterized in that a CRI image (CRI-B) is derived from 2D image data captured by the driver observation camera, wherein a pupil center (PC) of an eye (A) of the driver is detected in the CRI image (CRI-B), wherein a coded pattern (CM) is cyclically projected onto a windshield (WSS) of the vehicle by means of the augmented reality projector to establish a correspondence with the CRI image (CRI-B).wherein at a first time point (T1) characteristic features (CM1 to CM6) are detected in the CRI image (CRI-B), wherein a mapping of the characteristic features (CM1 to CM6) to the coded patterns (KM) is performed, wherein triangulation vectors are calculated based on the 3D point projection by the augmented reality projector and the focal point features on the 2D image plane of the driver observation camera, wherein a 3D detection is performed using the triangulation vectors for all characteristic features (CM1 to CM6) to determine a depth to the characteristic features (CM1 to CM6), wherein an ellipse (E1, E2) is fitted into the determined depths, having its principal vertex at the pupil center (PZ) to determine the 3D position of the eye (A) in the coordinate system of the vehicle and thus the driver's line of sight.
2. Method according to claim 1, characterized in that the CRI image (CRI-B) is generated at least substantially as an image and / or reflection of the environment on the cornea of the driver's eye (A).
3. Method according to claim 1 or 2, characterized in that post-processing of the CRI image (CRI-B), including rectification of a spherical projection, is carried out.
4. Method according to one of the preceding claims, characterized in that a camera, a lidar sensor and / or a radar sensor is used as a sensor for detecting an environment.
5. Method according to one of the preceding claims, characterized in that a coordinate system is used whose coordinate origin is located at the center of a front axle of the vehicle.
6. Method according to one of the preceding claims, characterized in that a 3D vector is formed by deriving the CRI image (CRI-B) from 2D image data captured by the driver observation camera, which is determined by a focal point of the driver observation camera and a feature point on a 2D camera image plane of the driver observation camera, from which the pupil center (PC) of the eye (A) of the driver is detected.
7. Method according to one of the preceding claims, characterized in that a Morano code is used as the coded pattern (KM).
8. Method according to one of the preceding claims, characterized in that Edges are detected as characteristic features (CM1 to CM6).
9. Vehicle comprising imaging sensors, including at least one sensor for sensing an environment, an augmented reality projector for a head-up display, and a driver observation camera, characterized in that the vehicle is configured to perform the method according to one of the preceding claims.
10. Vehicle according to claim 9, characterized in that the sensor for detecting an environment is designed as a camera, as a lidar sensor and / or as a radar sensor.
Citation Information
Patent Citations
Method for detecting the gaze direction of a driver in a motor vehicle and gaze direction detection device for a motor vehicle
DE102014009638A1
Method for operating a field of vision display device and vehicle
DE102024002430B3
Method and apparatus for detecting and following an eye and / or the gaze direction thereof
WO2016012458A1