Method for determining a driver's line of sight in a vehicle

A method using existing vehicle sensors and an augmented reality projector for 3D reconstruction and triangulation addresses the limitations of existing eye-tracking methods by determining the driver's gaze direction accurately and stably, without additional hardware or calibration, suitable for applications like Attention Assist.

DE102024003084B3Active Publication Date: 2025-12-31MERCEDES BENZ GROUP AG
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
DE102024003084
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-09-23
Publication Date
2025-12-31
Estimated Expiration
2044-09-23

AI Technical Summary

Technical Problem

Existing eye-tracking methods in vehicles require additional hardware components, calibration, and large amounts of training data, and are prone to false positives and instability due to head rotations and varying lighting conditions.

Method used

A method utilizing existing vehicle sensors and an augmented reality projector to project a coded pattern onto the windshield, enabling 3D reconstruction and triangulation to determine the driver's line of sight without additional hardware or calibration, using a cornea reflective image and triangulation vectors.

Benefits of technology

Stable against head rotations, independent of lighting and weather, and requiring no additional hardware or calibration, this method accurately determines the driver's gaze direction for various applications like Attention Assist.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to a method for determining a driver's line of sight in a vehicle, wherein at least one sensor for detecting an environment, a projector for a head-up display, and a driver observation camera are extrinsically calibrated with respect to a coordinate system of the vehicle, wherein the environment is detected by means of the sensor and reconstructed three-dimensionally, wherein a CRI image (CRI-B) is derived from image data acquired by the driver observation camera, wherein a pupil center (PC) of an eye (A) of the driver is detected in the CRI image (CRI-B), wherein a coded pattern (CM) is cyclically projected by the projector, wherein characteristic features (CM1 to CM6) are detected in the CRI image (CRI-B), and wherein a mapping of the characteristic features (CM1 to CM6) to the coded patterns (CM) is performed.where triangulation vectors are calculated based on the projection and the focal point features on the 2D image plane of the driver observation camera, where 3D detection is performed using the triangulation vectors for all characteristic features (CM1 to CM6) to determine a depth to the characteristic features (CM1 to CM6), where an ellipse (E1, E2) is fitted into the determined depths, having its main vertex at the pupil center (PZ), to determine the 3D position of the eye (A) in the coordinate system of the vehicle and thus the driver's line of sight.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a method for determining a driver's line of sight in a vehicle according to the preamble of claim 1.

[0002] Eye-tracking systems for vehicles often rely on the use of optical sources (most commonly infrared diodes). Based on these projections, the position of the eye is determined, and the driver's line of sight is derived. One of the most common methods is Purkinje-based eye tracking, which analyzes characteristic reflections and their relative positions to estimate the eye's position relative to the projection. Disadvantages include the need for additional IR diodes and / or other hardware components, and the requirement to know all Purkinje reflections. This can be particularly challenging when the driver's head is rotating. Another approach involves pupil detection using classic image processing algorithms (such as Canny-Edge detectors).In this method, the pupil's silhouette is extracted, and its relative position is inferred from the deformation of this silhouette. However, this requires initial calibration for individual adjustment. Furthermore, data-driven systems exist that, for example, can extract the relative position of the eye from a 2D image using convolutional neural networks. However, these require a relatively large amount of data, and false positives are possible if the network is not sufficiently generalized.

[0003] In Purkinje-based eye tracking, a light source (for example, an IR diode) creates a reflection on the cornea. These reflections are also called Purkinje reflections (a total of four reflections due to the refraction of light inside and outside the cornea and lens). These reflections and their relative positions can be used to determine the position of the eye relative to the imaging system.

[0004] The disadvantage here is the additional need for a light source and the necessity of detecting the four reflections, which is particularly difficult with larger rotations of the head.

[0005] Furthermore, data-driven approaches are known. These involve training neural networks to derive gaze from 2D image captures. In many cases, convolutional neural networks are used, with, for example, the results of Purkinje detection serving as ground truth for the training phase. A disadvantage of this approach is the large amount of training data required and the need to account for significant variance in that data. Otherwise, a high number of false positives and incorrect gaze determinations will result.

[0006] Methods for carrying out eye tracking in a vehicle are known from WO 2016 / 012 458 A1 and DE 10 2014 009 638 A1.

[0007] MORANO, Raymond A. [et al.]: Structured light using pseudorandom codes. In: IEEE transactions on pattern analysis and machine intelligence, Vol. 20, 1998, No. 3, pp. 322-327. ISSN 1939-3539. https: / / doi.org / 10.1109 / 34.667888 [accessed on 2025-07-09] describes the solution to the correspondence problem in active stereo vision using pseudorandomly coded structured light.

[0008] DE 10 2024 002 430 B3 describes a method for operating a field-of-view display device in a vehicle, wherein an in-vehicle computing unit generates an augmented reality display on a projection surface of the field-of-view display device for the driver. This display is contact-analogous to an object marker and corresponds to an object detected by the vehicle based on an evaluation of environmental sensor data. Furthermore, the computing unit determines a positional difference perceptible to the driver between the object marker and the object within the projection surface. To determine this positional difference, the computing unit considers the reflection of the object marker from the projection surface, which is detectable in a corneal reflection image of the driver.

[0009] The invention is based on the objective of providing a novel method for determining a driver's line of sight in a vehicle, as well as a novel vehicle.

[0010] The problem is solved according to the invention by a method for determining a line of sight of a driver in a vehicle with the features of claim 1 and by a vehicle with the features of claim 9.

[0011] Advantageous embodiments of the invention are the subject of the dependent claims.

[0012] A method for determining a driver's line of sight in a vehicle is proposed, wherein the vehicle's imaging sensors, including at least one sensor for capturing an environment, an augmented reality projector for a head-up display, and a driver observation camera are extrinsically calibrated with respect to a coordinate system of the vehicle, and the environment is captured by the at least one sensor and reconstructed three-dimensionally. According to the invention, a cornea reflective image (CRI) is derived from 2D image data captured by the driver observation camera, wherein the pupil center of one of the driver's eyes is detected in the CRI. A coded pattern is cyclically projected onto the vehicle's windshield by the augmented reality projector to establish a correspondence with the CRI image, and characteristic features are detected in the CRI image at a first time point.wherein a mapping of the characteristic features to the coded patterns is performed, wherein triangulation vectors are calculated based on the 3D point projection by the augmented reality projector and the focal point features on the 2D image plane of the driver observation camera, wherein a 3D detection is performed using the triangulation vectors for all characteristic features to determine a depth to the characteristic features, wherein an ellipse is fitted into the determined depths, having its main vertex at the pupil center, to determine the 3D position of the eye in the vehicle's coordinate system and thus the driver's line of sight.

[0013] In one embodiment, the CRI image is generated at least substantially as an image and / or reflection of the environment on the cornea of ​​the driver's eye.

[0014] In one embodiment, post-processing of the CRI image is performed, including rectification of a spherical projection.

[0015] In one embodiment, a camera, a lidar sensor and / or a radar sensor is used as a sensor to detect an environment.

[0016] In one embodiment, a coordinate system is used whose coordinate origin is located at the center of a front axle of the vehicle.

[0017] In one embodiment, a 3D vector is formed by deriving the CRI image from 2D image data captured by the driver observation camera, which is determined by a focal point of the driver observation camera and a feature point on a 2D camera image plane of the driver observation camera, from which the pupil center of the driver's eye is detected.

[0018] In one embodiment, a Morano code is used as the coded pattern.

[0019] In one embodiment, edges are detected as characteristic features.

[0020] According to one aspect of the present invention, a vehicle is proposed comprising imaging sensors, including at least one sensor for sensing an environment, an augmented reality projector for a head-up display, a driver observation camera, and a windshield. According to the invention, the vehicle is configured to carry out the method described above.

[0021] In one embodiment, the sensor for detecting an environment is designed as a camera, a lidar sensor and / or a radar sensor.

[0022] The present invention disclosure provides an approach to enable eye-tracking in vehicles using existing hardware components, without requiring training data and eye-tracking-specific calibration.

[0023] The solution according to the invention allows the detection of a driver's gaze direction based on 3D information, is stable against head rotations, requires no eye-tracking-specific calibration, no additional IR diodes, no additional human intervention, no further hardware components beyond those already in series production, and is independent of the surrounding scenery and of lighting and weather conditions. Furthermore, it can be repeated cyclically within a control loop and allows input data for a variety of other ADAS applications (e.g., Attention Assist, etc.).

[0024] Exemplary embodiments of the invention are explained in more detail below with reference to drawings.

[0025] This shows: Fig. 1. A schematic flowchart of a procedure for performing eye-tracking in a vehicle, Fig. 2. A schematic view of a projection of a coded pattern by an augmented reality projector onto a vehicle windshield. Fig. 3 a schematic view of a CRI image with detected characteristic features, Fig. 4 a schematic view of the CRI image and the projection of the coded pattern by the augmented reality projector onto the windshield of the vehicle, wherein certain characteristic features are assigned to certain points of the coded pattern, Fig. 5 a schematic view of a driver's eye with a pupil center in the CRI image of the driver, and Fig. 6 A schematic view to illustrate an ellipse fitting.

[0026] Corresponding parts are marked with the same reference symbols in all figures.

[0027] Fig. Figure 1 is a schematic flowchart of a procedure for performing eye-tracking in a vehicle.

[0028] The invention describes a method for performing eye-tracking in a vehicle based on existing hardware components. The method utilizes coded passive triangulation with the following steps VS1 to VS10: The present invention provides an approach to extract a line of sight of a vehicle driver (eye-tracking) based on consideration of 3D sensor data and a CRI (Cornea Reflective Image).

[0029] In step VS1, the hardware components involved, in particular at least one sensor (e.g., a camera, lidar sensor, and / or radar sensor), an augmented reality (AR) projector, especially for a head-up display, and / or a driver observation camera, are extrinsically calibrated (calibration to determine the orientation and translation of the components) so that the orientation and position of the hardware components are defined within a specific coordinate system. For example, a coordinate system is used whose origin is located at the center of the vehicle's front axle (e.g., MB CAD). This calibration can be performed during the vehicle's manufacturing process.

[0030] In step VS2, during field operation of the vehicle, a 3D reconstruction of the vehicle's surroundings is performed using the vehicle's environmental sensors. The scene around the vehicle is reconstructed three-dimensionally (based on data from at least one sensor, such as the camera, lidar sensor, and / or radar sensor).

[0031] If a resulting 3D image is available, a CRI image CRI-B (Cornea Reflective Image) is generated in a further step, VS3. This image essentially represents an image and / or reflection of the environment onto the driver's cornea. Post-processing of the CRI image CRI-B can be performed in step VS3, for example, rectification of a spherical projection. By deriving the CRI image CRI-B from 2D image data captured by the driver observation camera, a 3D vector can already be created (focal point of the driver observation camera and feature point on a 2D camera image plane of the driver observation camera). From this vector, the pupillary center PZ of the driver's eye A can be detected in step VS4. However, two vectors may be relevant for triangulation.

[0032] Therefore, in one step, VS5 projects a coded pattern KM using the augmented reality projector to establish the correspondence between the 3D point and the reflection on the CRI image CRI-B. Different coded patterns can be used. For example, the Morano code, which is essentially a spatial color coding, can be used.

[0033] Fig. Figure 2 is a schematic view of the projection of the coded pattern KM by the augmented reality projector onto a windshield WSS of the vehicle.

[0034] The projection with the augmented reality projector is repeated at a certain cycle time, which is not visible to the driver. However, the projection must not be continuously displayed in order not to disrupt and / or block the 3D image on the CRI-B image.

[0035] The coded projection now reveals which 3D object hits the windshield (WSS) and how (correspondence to the CRI image CRI-B), which allows the determination of the 3D object's position based on the two 3D points (e.g., CM1 and CM6). Fig. 5) a vector can be set up and the triangulation can be performed.

[0036] In a further step VS6, feature detection takes place in the CRI image CRI-B at time T1. Characteristic features in the CRI image CRI-B (e.g., edges, etc.) are detected. For this purpose, the CRI image CRI-B is searched for characteristic features CM1 to CM6. For example, edges are sought. Fig. Figure 3 is a schematic view of the CRI image CRI-B at time T1 with detected characteristic features CM1 to CM6, especially edges.

[0037] In step VS7, a correspondence mapping is performed in the CRI image CRI-B (at time T1+1), that is, a mapping of the characteristic features CM1 to CM6 to the coded patterns KM in order to establish triangulation vectors. This involves mapping the characteristic features CM1 to CM6 from the CRI image CRI-B onto the coded patterns KM, which were projected by the augmented reality projector. Fig. Figure 4 is a schematic view of the CRI image CRI-B and the projection of the coded pattern KM by the augmented reality projector onto the windscreen WSS of the vehicle, with certain characteristic features CM3, CM4 assigned to certain points of the coded pattern KM.

[0038] In step VS8, the triangulation vectors are calculated based on the 3D point projection by the augmented reality projector and the focal point features on the 2D image plane of the driver observation camera.

[0039] In step VS9, a 3D detection is performed using the triangulation vectors calculated in step VS8 for all characteristic features CM1 to CM6 generated in step VS6. Based on these characteristic features CM1 to CM6, the 3D triangulation is carried out, thus determining the depth to these features. Fig. Figure 5 is a schematic view of the driver's eye A with the pupil center PZ in the driver's CRI image CRI-B. Also shown are two three-dimensional characteristic features CM1, CM6 detected by the sensors; the characteristic features CM1', CM6' detected on the windshield WSS using the coded pattern KM; the optical center OZ of the driver monitoring camera; the characteristic features CM1", CM6" represented on an image plane BE in the CRI image CRI-B; and the characteristic features CM1''', CM6''' reconstructed three-dimensionally on the driver's CRI image CRI-B.

[0040] In step VS10, an ellipse fitting and gauze determination (driver's view) are performed. The ellipse fitting is carried out based on the depth points generated in step VS9.

[0041] In step VS10, an ellipse E1,E2 is fitted into the 3D space reconstructed in step VS9. This ellipse has its main vertex at the pupil center PZ. This pupil center PZ was previously determined from the CRI image CRI-B. This establishes the 3D position of the eye A in the vehicle reference system (the vehicle's coordinate system), allowing the line of sight to be determined and passed on to the subsequent driving functions.

[0042] Fig.Figure 6 is a schematic view of the three-dimensionally reconstructed environment 3DU with three-dimensional characteristic features CM1, CM2, CM5, CM6 detected by the sensors, the characteristic features CM1', CM2' detected by means of the coded pattern KM on the windshield WSS, the optical center OZ of the driver observation camera, the characteristic features CM1'', CM5'' represented in the CRI image CRI-B on an image plane BE and the three-dimensionally reconstructed characteristic features CM1''', CM2''', CM5''', CM6''', with a first ellipse E1 fitted by the reconstructed characteristic features CM1''', CM2''' and a second ellipse E2 fitted by the reconstructed characteristic features CM5''', CM6'''.

[0043] This approach has the advantage of determining the eye position in 3D space. It is also relatively stable against head rotation and requires no eye-tracking-specific calibration. Furthermore, it requires no additional hardware components and does not use IR diodes. It is also relatively independent of weather conditions and the surrounding environment. Reference symbol list 3DU three-dimensionally reconstructed environment A eye BE Image plane CM1 to CM6 characteristic feature CM1' to CM6' characteristic feature CM1'' to CM6'' characteristic feature CM1''' to CM6''' characteristic feature CRI-B CRI image E1, E2 Ellipse KM coded pattern OZ Optical Center PZ Pupillary Center VS1 to VS10 step WSS windshield

Claims

[1] Method for determining a driver's line of sight in a vehicle, wherein imaging sensors of the vehicle, including at least one sensor for sensing an environment, an augmented reality projector for a head-up display, and a driver observation camera are extrinsically calibrated with respect to a coordinate system of the vehicle, wherein the environment is sensed by means of the at least one sensor and reconstructed three-dimensionally, characterized bythat a CRI image (CRI-B) is derived from 2D image data captured by the driver observation camera, wherein a pupil center (PC) of one eye (A) of the driver is detected in the CRI image (CRI-B), wherein a coded pattern (CM) is cyclically projected onto a windshield (WSS) of the vehicle using the augmented reality projector to establish a correspondence with the CRI image (CRI-B), wherein characteristic features (CM1 to CM6) are detected in the CRI image (CRI-B) at a first time point (T1), wherein a mapping of the characteristic features (CM1 to CM6) to the coded patterns (CM) is performed, wherein triangulation vectors are calculated based on the 3D point projection by the augmented reality projector and the focal point features on the 2D image plane of the driver observation camera, wherein 3D detection is performed using the triangulation vectors for all characteristic features (CM1 to CM6) are performed,to determine a depth to the characteristic features (CM1 to CM6), whereby an ellipse (E1, E2) is fitted into the determined depths, which has its main vertex in the pupil center (PZ), in order to determine the 3D position of the eye (A) in the coordinate system of the vehicle and thus the driver's line of sight. [2] Method according to claim 1, characterized by , that the CRI image (CRI-B) is generated at least substantially as an image and / or reflection of the surroundings on the cornea of ​​the driver's eye (A). [3] Method according to claim 1 or 2, characterized by , that post-processing of the CRI image (CRI-B), including rectification of a spherical projection, is carried out. [4] Method according to any one of the preceding claims, characterized by that a camera, a lidar sensor and / or a radar sensor is used as a sensor to detect an environment. [5] Method according to any one of the preceding claims, characterized by , that a coordinate system is used whose origin is located at the center of a front axle of the vehicle. [6] Method according to any one of the preceding claims, characterized by , that by deriving the CRI image (CRI-B) from 2D image data captured by the driver observation camera, a 3D vector is formed, which is determined by a focal point of the driver observation camera and a feature point on a 2D camera image plane of the driver observation camera, from which the pupil center (PC) of the eye (A) of the driver is detected. [7] Method according to any one of the preceding claims, characterized by , that a Morano code is used as the coded pattern (CM). [8] Method according to any one of the preceding claims, characterized by , that edges are detected as characteristic features (CM1 to CM6). [9] Vehicle comprising imaging sensors, including at least one sensor for sensing an environment, an augmented reality projector for a head-up display, and a driver observation camera, characterized by that the vehicle is configured to carry out the procedure according to one of the preceding claims. [10] Vehicle according to claim 9, characterized by that the sensor for detecting an environment is designed as a camera, a lidar sensor and / or a radar sensor.

Citation Information

Patent Citations

  • Method for operating a field of vision display device and vehicle

    DE102024002430B3