System and method for estimating gaze direction of a user with or without glasses

The system uses a CNN-based gaze estimation method to accurately determine a driver's gaze direction, even with glasses, by detecting keypoints on obscured eyes, thereby preventing distractions and accidents.

JP2026501804APending Publication Date: 2026-01-16MERCEDES BENZ GROUP AG
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025540287
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-01-10
Filing Date
2023-12-21
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

Existing gaze detection techniques struggle to accurately estimate a driver's gaze direction when wearing glasses, as glasses obscure eye landmarks and cause glare, requiring excessive computational power and degrading detection accuracy.

Method used

A system utilizing a convolutional neural network (CNN) architecture with a case-based reasoning (CBR) block and a learning module, such as a random forest model, to detect keypoints on the eyes obscured by glasses, calculating yaw and pitch to determine gaze direction, and generating alerts or switching the vehicle to automatic mode if gaze deviates from the road.

Benefits of technology

Accurately identifies the driver's gaze direction with or without glasses, preventing distractions and traffic accidents by alerting the driver or switching to automatic mode when gaze is off the road for a few seconds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026501804000001_ABST
    Figure 2026501804000001_ABST
Patent Text Reader

Abstract

A system (100) for estimating gaze direction of a user with or without eyeglasses includes a processing device (102) incorporating a model (212) having a first branch configured using a convolutional neural network (CNN) architecture and a learning module (214), where the model (212) is pre-trained to detect one or more keypoints from one or more images associated with one or both eyes of the user obscured by the eyeglasses. The processing device (102) is configured to use the CNN architecture to detect one or more keypoints in real time from one or more images associated with one or both eyes of a user of a vehicle, and to determine the user's gaze direction based on the detected keypoints using the learning module.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to the field of gaze estimation, and in particular to a system and method for estimating the gaze direction of a user with or without glasses in a vehicle. [Background technology]

[0002] Vehicle drivers are commonly distracted while driving by performing other tasks, such as looking at their phones, caring for children or pets in the vehicle, or manually operating the vehicle's infotainment system. This can potentially result in traffic accidents. Therefore, it would be advantageous from a safety perspective to provide a simple, automated, efficient, and feasible solution for monitoring where a vehicle driver is looking. This would also provide a user with a reminder or alert, or switch the vehicle into an automatic mode and / or stop the vehicle for safety purposes, if the driver's vision or gaze direction deviates from the road for more than a predetermined period of time.

[0003] EP 3789848 A1 discloses a method for estimating a user's gaze direction, which includes the steps of acquiring an image of a user's face, determining an approximate gaze direction based on a current head pose and a relationship between the head pose and the gaze direction, determining an estimated gaze direction based on detected eye features, determining a fine gaze direction based on a glint / reflection position and the eye features, and combining the approximate gaze direction with at least one of the estimated gaze direction and the fine gaze direction to provide a corrected gaze direction.

[0004] The above-cited documents estimate the user's current gaze direction based on reflective landmarks on the cornea of ​​the eye along with the driver's head pose, which may require more computational power. Furthermore, the driver may wear glasses or sunglasses while driving a vehicle, which may obscure the eyes. Therefore, the gaze detection techniques of the above-cited documents may not be able to properly estimate the gaze direction of the hair dryer because the glasses or obscuration make it difficult to easily estimate important landmarks on the eyes. Furthermore, glasses that cover the eyes may reflect the surrounding environment and glare, which may also degrade the gaze detection process of the above-cited documents.

[0005] Therefore, there is a need in the art to overcome the above-mentioned drawbacks, limitations, and shortcomings associated with conventional gaze detection techniques and the above-cited documents by accurately estimating the gaze direction of a user / driver with or without glasses while operating a vehicle. Summary of the Invention [Problem to be solved by the invention]

[0006] [Purpose of this disclosure] A general objective of the present disclosure is to monitor the gaze direction of drivers with or without eyeglasses.

[0007] One object of the present disclosure is to provide an efficient, reliable and faster system and method for estimating the gaze direction of a driver wearing eyeglasses or when the driver's eyes are occluded.

[0008] Another object of the present disclosure is to accurately identify the direction in which a driver wearing glasses is looking while operating a vehicle, to help avoid distractions and prevent traffic accidents.

[0009] Yet another object of the present disclosure is to alert the driver or switch the vehicle into automatic mode to stop the vehicle for the driver's safety if the driver's gaze direction is off the road for a few seconds.

[0010] One objective of the present disclosure is to provide an efficient and reliable system and method for estimating the gaze direction of a driver with or without glasses, which will either call the driver's attention or alert them, or switch the vehicle into an automatic mode to stop the vehicle, if the gaze direction deviates from the road for a few seconds. [Means for solving the problem]

[0011] Aspects of the present disclosure relate to the field of gaze estimation. In particular, the present disclosure provides systems and methods for estimating the gaze direction of a user with or without glasses in a vehicle.

[0012] One aspect of the present disclosure relates to a system for estimating the gaze direction of a user with or without eyeglasses. The system includes a processing device incorporating a model having a first branch configured using a convolutional neural network (CNN) architecture. The CNN architecture may include a case-based reasoning (CBR) block, and the CNN architecture may be fed with a heat map to detect one or more keypoints. The model is pre-trained to detect one or more keypoints from one or more images associated with one or both eyes of a user covered by eyeglasses, and a learning module. The processing device includes a processor coupled to a memory, the memory storing one or more instructions executable by the processor to use the CNN architecture to detect one or more keypoints in real time from one or both eyes of a user of a vehicle, and to use the learning module to determine the user's gaze direction based on the detected keypoints.

[0013] In one aspect, the system may include an image capture unit comprising a camera that captures one or more images of a user's face, and a processing device configured to crop one or more eye regions from the captured one or more images of the user's face, detect one or more keypoints from the cropped one or more images using a CNN architecture, and regress the detected keypoints using a learning module to calculate yaw and pitch indicative of the user's gaze direction.

[0014] In one embodiment, the pre-trained model of the processing device may have a second branch including a classification network, which may be fed with the output of at least one of the CBR blocks and may be trained to reduce entropy loss so that the model can determine one or more actual keypoints of the eye based on training with synthetic data.

[0015] In one aspect, the processing device may be configured to train the learning module using synthetic training data and / or real-time data including a first set of images associated with an eye region and a second set of images associated with an eye region obscured by glasses, thereby enabling the learning module to determine the user's gaze direction. One or more key points may be associated with the sclera, the iris, and the center of the eye.

[0016] In one aspect, the learning module can be a random forest model. The random forest model can be fed with four vectors derived from any one or combination of the center of the iris and the extreme junction of the sclera, and the angular distance between the extreme sclera point and the center of the iris. The random forest model can be trained using one or more noise, keypoint removal, and one or more additional features to identify yaw and gaze.

[0017] In one aspect, the processing device may be configured to generate a set of alert signals when the user's gaze direction is determined to be outside a region of interest (ROI) for a predetermined time period, and to stop the vehicle or switch the vehicle into an autonomous driving mode when the user's gaze direction is determined to be outside the ROI for a predetermined time period, the ROI being a road on which the vehicle is traveling.

[0018] In one aspect, when one of the user's eyes is at least partially occluded, the processing device may be configured to determine the user's gaze direction using one or more key points associated with the user's other, unoccluded eye.

[0019] Another aspect of the present disclosure relates to a method for estimating a user's gaze direction, the method including detecting, in real time, by a processing device configured with a CNN architecture from one or more images associated with one or both eyes of a user of a vehicle, the CNN architecture being pre-trained to detect one or more keypoints from one or more images associated with the user's eye or eyes obscured by glasses, and determining, using a learning module associated with the processing device, the user's gaze direction based on the detected keypoints.

[0020] In one aspect, the method may include capturing, by a camera, one or more images of a user's face; cropping, by a processing device, one or more eye regions from the captured one or more images of the user's face; detecting one or more keypoints from the cropped one or more images using a CNN architecture; and regressing, using a learning module, the detected keypoints to calculate yaw and pitch indicative of the user's gaze direction.

[0021] Various objects, features, aspects and advantages of the present subject matter will become more apparent from the following detailed description of preferred embodiments, taken in conjunction with the accompanying drawings in which like numerals represent like elements.

[0022] The accompanying drawings are included to provide a further understanding of the present disclosure, and are incorporated in and constitute a part of this specification. The drawings illustrate exemplary embodiments of the present disclosure and, together with the description, serve to explain the principles of the disclosure. [Brief explanation of the drawings]

[0023] [Figure 1] 1 is an exemplary block diagram of a proposed system for estimating the gaze direction of a user wearing glasses in a vehicle, according to an embodiment of the present invention; FIG. [Figure 2] 1 is an exemplary block diagram illustrating functional units of a processing device associated with the proposed system according to an embodiment of the present invention; [Figure 3A] FIG. 1 is a flow diagram illustrating steps of a proposed method for estimating gaze direction of a user wearing glasses in a vehicle, in accordance with one or more embodiments of the present disclosure. [Figure 3B] 1 is an exemplary flowchart illustrating the operation of a proposed system for estimating gaze direction of a user wearing glasses in a vehicle, in accordance with one or more embodiments of the present disclosure. [Figure 3C] FIG. 1 is an exemplary logic diagram of the proposed system to detail the operation of the system, in accordance with one or more embodiments of the present disclosure. [Figure 4] FIG. 10 illustrates exemplary key points associated with a user's eyes that are useful for detecting the user's gaze direction, in accordance with one or more embodiments of the present disclosure. [Figure 5] FIG. 1 illustrates an exemplary architecture of a random forest model implemented in the proposed system, in accordance with one or more embodiments of the present disclosure. [Figure 6] FIG. 1 illustrates an exemplary representation of a KeyPoint architecture implemented in the proposed system, in accordance with one or more embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0024] The following is a detailed description of embodiments of the present disclosure, as illustrated in the accompanying drawings. The embodiments are in such detail as to clearly communicate the disclosure. However, the gist of the details provided is not intended to limit the possible variations of the embodiments. Rather, the intent is to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the disclosure as defined by the appended claims.

[0025] FIELD OF THE INVENTION The embodiments described herein relate to the field of gaze estimation. In particular, the present disclosure provides a system and method for estimating the gaze direction of a user with or without glasses in a vehicle.

[0026] 1, a proposed system 100 for estimating the gaze direction of a user (also referred to herein as a driver, passenger, or crew member) wearing glasses is disclosed. In particular, the disclosed system 100 can accurately identify the gaze direction (the direction in which the user / driver wearing the glasses is looking) while driving a vehicle, and if the user / driver's gaze direction deviates from the road for a few seconds, it can alert the user / driver or switch the vehicle into automatic mode to stop the vehicle for the driver's safety, avoiding distractions and helping to prevent traffic accidents.

[0027] In one embodiment, the system 100 includes a processing device 102 capable of communicating with a vehicle control unit (VCU) 106 of a vehicle. In another embodiment, the processing device 102 may be a server that may reside in one or more vehicle units associated with one or more vehicles.

[0028] In another embodiment, system 100 may further include an image acquisition unit 104 comprising one or more image sensors or cameras (collectively referred to herein as cameras or image sensors) installed within the vehicle to monitor one or more images or videos of a user (also referred to herein as a driver, passenger, or occupant) seated / moving within the vehicle. Image acquisition unit 104 may also capture images or videos of the road or an image or video of a scene outside the vehicle. Image acquisition unit 104 may communicate with the vehicle's processing device 102 and / or VCU 106.

[0029] In an exemplary embodiment, an image acquisition unit 104 including camera(s) or image sensor(s) 104 may be installed inside the vehicle interior to capture or monitor images or videos of the user / driver's face in real time. One of the cameras 104 may be directed toward the front seat of the vehicle where the user / driver may be seated. The other camera 104 may be directed toward the road or toward the exterior of the vehicle. The camera or image sensor 104 may be located on the dashboard or ceiling of the vehicle to fully cover the interior of the vehicle and capture images of the user / driver's face, as well as capture the road or exterior view. The camera 104 may be an infrared (IR) camera, such as a near-infrared camera, a mid-wave infrared camera, and a long-wave infrared camera.

[0030] In one embodiment, the processing device 102 is communicatively or operatively connectable with the vehicle's image capture unit 104 and VCU 106 via a network. Furthermore, the network may be a wireless network, a wired network, or a combination thereof, which may be implemented as one of different types of networks, such as an intranet, a local area network (LAN), a wide area network (WAN), or the Internet. Furthermore, the network may be either a dedicated network or a shared network. A shared network may represent a federation of different types of networks that may use various protocols, such as, for example, Hypertext Transfer Protocol (HTTP), Transmission Control Protocol / Internet Protocol (TCP / IP), Wireless Application Protocol (WAP), etc. Furthermore,

[0031] In one embodiment, the processing device 102 may be implemented using any one or a combination of hardware and software components, such as a cloud, a server, a computing system, a computing device, a network device, etc. Furthermore, the processing device 102 may interact with the image acquisition unit 104 and the VCU 106 via a wired or wireless network.

[0032] The image acquisition unit 104 may be configured to capture images / videos of the user / driver's face. The processing device 102 may be configured with a model and learning module having a first branch configured with a convolutional neural network (CNN) architecture, where the model may be pre-trained to detect one or more keypoints from one or more images associated with one or both of the user's eyes that are covered by eyeglasses. In one embodiment, the one or more keypoints may be associated with the user / driver's sclera, iris, and eye center. The processing device 102 may be configured to use the pre-trained CNN architecture to detect one or more keypoints associated with one or both of the user's eyes in real time from images captured by the image acquisition unit 104. Furthermore, the processing device 102 may be configured with a learning module, such as, but not limited to, a random forest model, that enables the processing device 102 to determine the user's gaze direction based on the detected keypoints. The learning module enables the processing device 102 to regress the detected keypoints to calculate yaw and pitch, which indicate the gaze direction of the user / driver with or without eyeglasses. Thus, the processing device 102 can accurately identify the direction in which the user / driver wearing the glasses is looking while driving a vehicle, which can help avoid distractions and prevent traffic accidents.

[0033] In one embodiment, the CNN architecture may further include a case-based reasoning (CBR) block. Furthermore, the heatmap may be fed to the CNN architecture to detect one or more keypoints. Furthermore, the pre-trained model of the processing device 102 may also have a second branch including a classification network, and the output of at least one of the CBR blocks may be fed to the second branch. The second branch may be trained to reduce entropy loss so that the model can determine one or more actual keypoints of the eye based on training with synthetic data. In one embodiment, a random forest model or learning module may be fed with four vectors derived from the iris center and the sclera's extreme junction and / or the angular distance between the sclera's extreme point and the iris center. Furthermore, for better accuracy and reliability, the random forest model may be trained using one or more noise, keypoint removal, and one or more additional features to identify yaw and gaze.

[0034] In another embodiment, the processing device 102 may be configured to pre-train or train a learning module (a random forest model) in real time using synthetic training data and / or real-time data including a first set of images associated with eye regions and a second set of images associated with eye regions obscured by glasses, thereby enabling the learning module to determine a user's gaze direction. During training, the processing device 102 may receive the first set of images associated with the eye regions and one or more corresponding keypoints present in the first set of images. Additionally, the processing device 102 may render glasses with a predetermined transparency over at least one of a predetermined percentage of the eye regions in the first set of images to provide the second set of images. Thus, the processing device 102 may train and test the CNN and random forest model using synthetic data including the first set of images, the second set of images, and the corresponding one or more keypoints. This enables the learning module to accurately detect one or more keypoints in captured images of the user in real time.

[0035] In one example, when one of the user's eyes is detected to be at least partially occluded, the processing device 102 may be configured to determine the user's gaze direction using key points associated with the user's other, unoccluded eye.

[0036] In one embodiment, the processing device 102 can be configured to generate a set of alert signals using the vehicle's VCU 106 when it is determined that the user's gaze direction is outside a region of interest (ROI), where the ROI is the road on which the vehicle is traveling. In another embodiment, the processing device 102 can be configured to activate the VCU 106 to stop the vehicle or switch the vehicle to an autonomous driving mode when it is determined that the user's gaze direction is outside the ROI (road) for a predetermined time.

[0037] Referring to FIG. 2 , the block diagram depicted therein illustrates exemplary functional units of the processing device 102, which may include one or more processor(s) 202, memory 204, interface(s) 206, processing engine(s) 208, and database 210. The one or more processor(s) 202 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing devices 102s, logic circuits, and / or any device that manipulates data based on operational instructions. Among other capabilities, the one or more processor(s) 202 are configured to fetch and execute computer-readable instructions stored in memory 204 of the server 102. The memory 204 may store one or more computer-readable instructions or routines that may be fetched and executed to create or share data units via a network service. The memory 204 may include any non-transitory storage device, including, for example, volatile memory such as RAM, or non-volatile memory such as EPROM, flash memory, etc.

[0038] In one embodiment, processing device 102 may also include interface(s) 206. Interface(s) 206 may include various interfaces, such as data input and output devices, also referred to as I / O devices, storage devices, etc. Interface(s) 206 may facilitate communication between processing device 102 and various devices connected to server 102, such as the vehicle's infotainment system, image acquisition unit 104, VCU 106, and power supply. Interface(s) 206 may also provide a communication path to one or more components of processing device 102. Examples of such components include, but are not limited to, processing engine(s) 208 and database 210.

[0039] In one embodiment, the processing engine(s) 208 may be implemented as a combination of hardware and programming (e.g., programmable instructions) for implementing one or more functionality of the processing engine(s) 208. In the examples described herein, such a combination of hardware and programming may be implemented in several different ways. For example, the programming of the processing engine(s) 208 may be processor-executable instructions stored on a non-transitory machine-readable storage medium, and the hardware of the processing engine(s) 208 may include processing resources (e.g., one or more processors) for executing such instructions. In this example, the machine-readable storage medium may store instructions that, when executed by the processing resources, implement the processing engine(s) 208. In such an example, the processing device 102 may include a machine-readable storage medium that stores instructions and the processing resources for executing the instructions, or the machine-readable storage medium may be separate from but accessible to the processing device 102 and the processing resources. In other examples, the processing engine(s) 208 may be implemented by electronic circuitry. Database 210 may include data stored or generated as a result of functionality provided by any of the components of processing engine(s) 208.

[0040] In one embodiment, the processing engine(s) 208 may include a pre-trained model 212 having a CNN architecture and a classification network, a learning module 214, an actuation and control unit 216, an alert unit 218, and other unit(s) 220. The other unit(s) 220 may implement functionality complementary to the applications or functions performed by the processing device 102 or the processing engine(s) 208.

[0041] According to one embodiment, the processing device 102 may enable the image acquisition unit 104 to capture images / videos of the user / driver's face wearing the glasses, as well as images of the road or images of the vehicle's exterior. As shown in detail in FIG. 5, the pre-trained model 212 may cause the processing device 102 to crop one or more eye regions from the captured image of the user / driver's face obscured by the glasses and correspondingly detect one or more key points from the captured image. In one embodiment, the one or more key points may be associated with the user / driver's sclera, iris, and eye center, as shown in detail in FIG. 4.

[0042] In one embodiment, the processing device 102 can be further configured with a learning module 214, such as a random forest model, as shown in FIG. 6, which allows the processing device 102 to determine the user's gaze direction based on the detected keypoints. The learning module 214 enables the processing device 102 to regress the detected keypoints to calculate yaw and pitch, which indicate the gaze direction of a user / driver with or without glasses. Thus, the pre-trained model 212 and the learning module 214 enable the processing device 102 to accurately identify the direction in which a user / driver with glasses is looking while operating a vehicle. This can help avoid distractions and prevent traffic accidents.

[0043] In an exemplary embodiment, as shown in FIG. 5, the pre-trained model 212 may include a first branch configured using a convolutional neural network (CNN) architecture and a second branch including a classification network. The model is pre-trained to detect one or more keypoints from one or more images associated with one or both of a user's eyes that are covered by glasses. The CNN architecture may include a case-based reasoning (CBR) block. Additionally, a heat map may be fed to the CNN architecture to detect one or more keypoints. The second branch may be fed with the output of at least one of the CBR blocks and may be trained to reduce entropy loss so that the model can determine one or more actual keypoints of the eyes based on training with synthetic data.

[0044] In an exemplary embodiment, the random forest model 214 can be fed four vectors derived from the center of the iris and the extreme junction of the sclera and / or the angular distance between the extreme sclera point and the center of the iris, as shown in Figure 6. Furthermore, for better accuracy and reliability, the random forest model 214 can be trained with one or more noise, keypoint removal, and one or more additional features to identify yaw and gaze.

[0045] In another exemplary embodiment, the processor 202 can cause the processing device 102 to pre-train or train the learning module 214 (random forest model) using synthetic training data and / or real-time data including a first set of images associated with eye regions and a second set of images associated with eye regions obscured by glasses. This enables the learning module 214 to determine the user's gaze direction. During training, the processing device 102 can receive the first set of images associated with the eye regions and one or more corresponding keypoints present in the first set of images. Further, as shown in FIG. 3B , the processing device 102 can provide the second set of images by rendering glasses with a predetermined transparency over at least one of the eye regions in a predetermined percentage of the first set of images. Thus, the processing device 102 can train and test the model 212 and the learning module (random forest model) 214 using synthetic data including the first set of images, the second set of images, and the corresponding one or more keypoints. This enables the learning module 214 to accurately detect one or more keypoints in captured images of the user in real time.

[0046] In one embodiment, the actuation and control unit 216 can cause the processing device 102 to generate and send a set of alert signals to the vehicle's VCU 106 to alert the user / driver when the user's gaze direction is determined to be outside a region of interest (ROI) for a predetermined period of time, where the ROI is the road the vehicle is traveling on.

[0047] In another embodiment, the actuation and control unit 216 can cause the processing device 102 to actuate the VCU 106 to stop the vehicle or switch the vehicle to autonomous driving mode if it is determined that the user's line of sight is outside the ROI (road) for a predetermined period of time.

[0048] In another embodiment, the ROI may be a road on which the vehicle is traveling. The alert unit 218 may cause the processing device 102 to generate an alert when the estimated gaze direction of the user / driver is estimated to be away from the road (ROI) for a first predetermined time period, indicating that the user / driver's attention is diverted. Furthermore, the alert unit 218 may also cause the processing device 102 to generate an alert when the user / driver's eyes are found to be closed for a second predetermined time period, indicating that the user / driver is asleep or unconscious.

[0049] 3A, a proposed method 300 for estimating the gaze direction of a user with or without glasses in a vehicle involves an image acquisition unit and a processing device connected to the vehicle's VCU. The method 300 includes step 302 of capturing one or more images of the user's face by the image acquisition unit, and then step 304 of cropping one or more eye regions from the captured one or more images of the user's face by the processing device.

[0050] The method 300 further includes detecting 306, in real time, one or more keypoints from one or more images associated with one or both eyes of a user of the vehicle by a processing device configured with a CNN architecture, which may be pre-trained to detect one or more keypoints from one or more images associated with one or both eyes of the user that are obscured by glasses.

[0051] Further, method 300 includes determining 308 a user's gaze direction based on the detected keypoints using a learning module associated with the processing device, wherein the processing device can regress the detected keypoints to calculate yaw and pitch indicative of the user's gaze direction.

[0052] 3B and 3C, the image acquisition unit can capture an image of a user / driver seated in a vehicle, and the face detection module can detect a face in the captured image. The processing device can then crop the face from the captured image and further crop the eye region. The CNN model can then detect key points in the cropped eye image. If no key points are detected, the gaze situation is not identified by the processing device. Furthermore, the random forest model can then select the unoccluded eye and estimate the gaze direction of the user / driver.

[0053] Therefore, the present invention (system and method) accurately estimates the gaze direction of a user / driver with or without glasses while driving a vehicle, and if the driver's gaze direction deviates from the road for a few seconds, the present invention will alert the driver or switch the vehicle into automatic mode to stop the vehicle for the driver's safety, thereby avoiding distraction and preventing traffic accidents.

[0054] While various embodiments of the present invention have been described above, other and further embodiments of the invention may be devised without departing from the basic scope thereof, the scope of which is determined by the claims that follow. The invention is not limited to the described embodiments, variations, or examples, which, when combined with information and knowledge available to those skilled in the art, are included to enable one to make and use the invention.

[0055] Benefits of this disclosure The present disclosure monitors the gaze direction of drivers with or without glasses.

[0056] The present disclosure provides an efficient, reliable and faster system and method for estimating gaze direction of a driver wearing glasses or when the driver's eyes are occluded.

[0057] The present invention accurately identifies the direction in which a driver wearing glasses is looking while operating a vehicle in order to avoid distractions and prevent traffic accidents.

[0058] The present invention will alert the driver or switch the vehicle into automatic mode to stop the vehicle for the driver's safety if the driver's gaze direction is off the road for a few seconds.

[0059] The present disclosure provides an efficient and reliable system and method for estimating the gaze direction of a driver with or without glasses, which will either prompt or alert the driver or switch the vehicle into an automatic mode to stop the vehicle if the driver's gaze direction deviates from the road for a few seconds. [Prior art documents] [Patent documents]

[0060] [Patent Document 1] European Patent Application Publication No. 3789848

Claims

1. A system (100) for estimation of gaze direction of a user with or without glasses, comprising: a processing device (102) incorporating a model (212) having a first branch constructed using a convolutional neural network (CNN) architecture and a learning module (214), the CNN architecture including a case-based reasoning (CBR) block, the CNN architecture being supplied with a heat map to detect the one or more keypoints, the model (212) being pre-trained to detect one or more keypoints from one or more images associated with one or both eyes of a user obscured by eyeglasses, the processing device (102) comprising a processor (202) coupled to a memory (204), the memory (204) comprising: using the CNN architecture of the model (212) to detect one or more keypoints in real time from one or more images associated with the one or both eyes of the user of the vehicle; determining the gaze direction of the user based on the detected keypoints using the learning module (214); storing one or more instructions executable by said processor (202); System (100).

2. an image capture unit (104) comprising a camera for capturing one or more images of the user's face; The processing device (102) cropping one or more eye regions from the captured one or more images of the face of the user; Detecting the one or more keypoints from the one or more cropped images using the CNN architecture of the model (212); and using the learning module (214) to regress the detected keypoints to calculate yaw and pitch indicative of the gaze direction of the user. The system (100) of claim 1.

3. the pre-trained model (212) of the processing device (102) also has a second branch including a classification network, the second branch being fed with an output of at least one of the CBR blocks, the second branch being trained to reduce entropy loss so as to enable the model (212) to determine the actual keypoint(s) of the eye based on training with synthetic data; The system (100) of claim 1.

4. the processing device (102) is configured to train the learning module (214) using synthetic training data and / or real-time data including a first set of images associated with an eye region and a second set of images associated with the eye region obscured by glasses, thereby enabling the learning module (214) to determine the gaze direction of the user; the one or more key points are associated with the sclera, the iris, and the center of the eye; The system (100) of claim 1.

5. the learning module (214) is a random forest model; the random forest model (214) is fed with four vectors derived from any one or combination of the iris center and the sclera extreme junction, and the angular distance between the sclera extreme point and the iris center; the random forest model (214) is trained using one or more of noise, keypoint removal, and one or more additional features to identify the yaw and gaze; The system (100) of claim 4.

6. The processing device (102) generating a set of alert signals when the gaze direction of the user is determined to be outside a region of interest (ROI) for a predetermined time; and stopping the vehicle or switching the vehicle to an autonomous driving mode when it is determined that the line of sight of the user is outside the ROI for the predetermined time, the ROI being a road on which the vehicle is traveling. The system (100) of claim 1.

7. and when one of the eyes of the user is at least partially occluded, the processing device (102) is configured to determine the gaze direction of the user using the one or more key points associated with the other, unoccluded eye of the user. The system (100) of claim 1.

8. A method (300) for estimation of a user's gaze direction, said method (300) comprising: detecting (306) in real time one or more keypoints from one or more images associated with one or both eyes of a user of the vehicle by a processing device (102) incorporating a model (212) having a first branch constructed using a convolutional neural network (CNN) architecture, the CNN architecture being pre-trained to detect one or more keypoints from one or more images associated with one or both eyes of the user obscured by eyeglasses; determining (308) the gaze direction of the user based on the detected keypoints using a learning module (214) associated with the processing device (102); Method (300).

9. The method (300) comprises the steps of: capturing (302) one or more images of the user's face by an image acquisition unit (104); Cropping (304) one or more eye regions from the captured one or more images of the face of the user by the processing device (102); Detecting (306) the one or more keypoints from the one or more cropped images using the CNN architecture; and regressing the detected keypoints using the learning module (214) to calculate yaw and pitch indicative of the user's gaze direction (308).

9. The method (300) of claim 8.

Citation Information

Patent Citations

  • Image processing device and method therefor

    JP2009041972A

  • Driver state detecting device

    JP2018116428A

  • Information processing apparatus for estimating person's line of sight and estimation method, and learning device and learning method

    JP2019028843A

  • Occupant monitoring device of vehicle

    JP2022129368A

  • Determination of gaze direction

    EP3789848A1