System and method for estimating eye gaze direction of user wearing or not wearing glasses
Through the combination of convolutional neural network and random forest model, the key points of eye contact drivers wearing glasses are detected in real time, solving the problem of inaccurate estimation of gaze direction in the prior art, and achieving accurate identification of driver gaze direction and accident prevention.
Patent Information
- Application Number
- CN202380091129.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-10
- Filing Date
- 2023-12-21
- Publication Date
- 2025-08-15
AI Technical Summary
The prior art is difficult to accurately estimate the eye gaze direction of drivers wearing or not wearing glasses, especially when glasses are blocked, which leads to an increased risk of distraction for drivers and is unable to effectively prevent road accidents.
The convolutional neural network (CNN) architecture and random forest model are adopted, combined with image acquisition units and processing devices, and the user's eye key points are detected in real time, the gaze direction is accurately estimated through the pre-trained model and learning module, and alarms are generated or the vehicle is switched to the autonomous driving mode if necessary.
It realizes accurate gaze direction identification for the driver wearing glasses, avoids distraction, promptly warns or switches vehicle modes, and reduces the risk of accidents.
Smart Images

Figure CN120500709A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of eye gaze estimation. Specifically, the present disclosure provides a system and method for estimating the eye gaze direction of a user wearing or not wearing glasses in a vehicle. Background Art
[0002] Drivers of vehicles are often distracted while driving by engaging in other tasks, such as checking a mobile device, caring for children or pets in the vehicle, or manually operating the vehicle's infotainment system, which can lead to road accidents. Therefore, it would be advantageous from a safety perspective if a simple, automated, and efficient operational solution could be provided to monitor where a vehicle driver is looking, and if the driver's vision or gaze direction is off the road for more than a defined period of time, the solution should also alert or warn the user or switch the vehicle to an automated mode and / or bring the vehicle to a stop for safety.
[0003] Patent document No. EP3789848A1 discloses a method for estimating a user's gaze direction, the method comprising the following steps: acquiring an image of the user's face, determining an approximate gaze direction based on a current head posture and a relationship between the head posture and the gaze direction, determining an estimated gaze direction based on detected eye features, determining a precise gaze direction based on a glint / reflection position and eye features, and combining the approximate gaze direction with at least one of the estimated gaze direction and the precise gaze direction to provide a corrected gaze direction.
[0004] The above-mentioned reference estimates the user's current gaze direction based on reflective feature points in the cornea of the eye together with the driver's head posture, which may require more computing power. In addition, the driver may wear glasses or sunglasses while driving the vehicle, so their eyes may be occluded. Therefore, the eye gaze detection technology of the above-mentioned reference may not be able to correctly estimate the dry driver's eye gaze direction because the key feature points of the eye cannot be easily inferred due to the glasses or occlusion. In addition, the glasses on the eyes may also reflect the surrounding environment and glare, which may again degrade the eye gaze detection process of the above-mentioned reference.
[0005] Therefore, there is a need in the art to overcome the aforementioned shortcomings, limitations and deficiencies associated with conventional eye gaze detection techniques and the above-cited references by accurately estimating the eye gaze direction of a user / driver while driving a vehicle with or without glasses.
[0006] Purpose of this disclosure
[0007] The general purpose of the present disclosure is to monitor the eye gaze direction of a driver with or without glasses.
[0008] One object of the present disclosure is to provide an efficient, reliable, and faster system and method for estimating the eye gaze direction of a driver wearing glasses or when the driver's eyes are occluded.
[0009] Another object of the present disclosure is to accurately identify the direction a driver wearing glasses is looking while driving a vehicle to avoid distraction and help prevent road accidents.
[0010] Yet another object of the present disclosure is to warn the driver or switch the vehicle to an automatic mode and stop the vehicle for the driver's safety if the driver's gaze direction is off the road for a few seconds.
[0011] An object of the present disclosure is to provide an efficient and reliable system and method for estimating the eye gaze direction of a driver with or without glasses, which alerts or warns the driver or switches the vehicle to automatic mode and stops the vehicle if the gaze direction is off the road for a few seconds. Summary of the Invention
[0012] Aspects of the present disclosure relate to the field of eye gaze estimation. In particular, the present disclosure provides a system and method for estimating the eye gaze direction of a user in a vehicle, wearing or not wearing glasses.
[0013] One aspect of the present disclosure relates to a system for estimating the eye gaze direction of a user wearing or not wearing glasses. The system includes a processing device incorporating a model having a first branch configured with a convolutional neural network (CNN) architecture. The CNN architecture may include a case-based reasoning (CBR) block, and wherein a heat map may be fed to the CNN architecture to detect one or more key points. The model is pre-trained to detect one or more key points from one or more images associated with one or more eyes of the user that are occluded by the glasses. The processing device includes a processor coupled to a memory, wherein the memory stores one or more instructions executable by the processor to: detect one or more key points in real time from one or more images associated with the one or more eyes of the user of a vehicle using the CNN architecture; and determine the eye gaze direction of the user based on the detected key points using the learning module.
[0014] In one aspect, the system may include an image acquisition unit including a camera for capturing one or more images of a user's face. The processing device may be configured to crop one or more eye regions from the one or more captured images of the user's face, detect one or more key points from the cropped one or more images using a CNN architecture, and regress the detected key points using a learning module to calculate yaw and pitch indicating a direction in which the user's eyes are looking.
[0015] In one aspect, a pre-trained model of a processing device may include a second branch comprising a classification network, wherein the second branch may be fed with an output of at least one of the CBR blocks, wherein the second branch may be trained to reduce entropy loss so as to enable the model to determine one or more true keypoints of the eye based on training with synthetic data.
[0016] In one aspect, the processing device can be configured to train the learning module using synthetic training data and / or real-time data, which can enable the learning module to determine the user's gaze direction, the synthetic training data and / or real-time data including a first set of images associated with an eye region and a second set of images associated with the eye region occluded by glasses. The one or more key points can be associated with the sclera, the iris, and the center of the eyeball.
[0017] In one aspect, the learning module can be a random forest model. The random forest model can be fed with four vectors derived from any one or a combination of: the distal junction of the iris center and the sclera, and the angular distance between the distal sclera point and the iris center. The random forest model can be trained using one or more of noise, keypoint removal, and one or more additional features to recognize the yaw and gaze.
[0018] In one aspect, the processing device can be configured to generate a set of alarm signals when it is determined that the user's gaze direction is outside the region of interest (ROI) and lasts for a predefined time, and to stop the vehicle or switch the vehicle to an autonomous driving mode when it is determined that the user's gaze direction is outside the ROI and lasts for a predefined time, where the ROI is the road on which the vehicle is traveling.
[0019] In one aspect, when one of the user's eyes is at least partially occluded, the processing device may be configured to determine the user's gaze direction using one or more keypoints associated with the user's other unoccluded eye.
[0020] Another aspect of the present disclosure relates to a method for estimating a user's eye gaze direction. The method includes the following steps: detecting, in real time, one or more key points from one or more images associated with one or more eyes of a user of a vehicle by a processing device configured with a CNN architecture, wherein the CNN architecture has been pre-trained to detect one or more key points from one or more images associated with one or more eyes of the user that are occluded by glasses; and determining, using a learning module associated with the processing device, the user's eye gaze direction based on the detected key points.
[0021] In one aspect, the method may include the following steps: capturing one or more images of a user's face by a camera; cropping one or more eye regions from the captured one or more images of the user's face by a processing device; detecting the one or more key points from the cropped one or more images using the CNN architecture; and regressing the detected key points using the learning module to calculate yaw and pitch indicating the direction of the user's eye gaze.
[0022] Various objects, features, aspects, and advantages of the present subject matter will become more apparent from the following detailed description of preferred embodiments, taken in conjunction with the accompanying drawings, wherein like numerals represent like parts, wherein: BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The accompanying drawings are included to provide a further understanding of the disclosure and are incorporated in and constitute a part of this specification. The drawings illustrate exemplary embodiments of the disclosure and together with the description serve to explain the principles of the disclosure.
[0024] Figure 1 An exemplary block diagram of a proposed system for estimating the eye gaze direction of a user wearing glasses in a vehicle according to an embodiment of the present invention is illustrated.
[0025] Figure 2 An exemplary block diagram representing functional units of a processing device associated with the proposed system according to an embodiment of the present invention is illustrated.
[0026] Figure 3A A flow chart representing the steps of a proposed method for estimating the eye gaze direction of a user wearing glasses in a vehicle is illustrated, according to one or more embodiments of the present disclosure.
[0027] Figure 3B An exemplary flow chart depicting the operation of the proposed system for estimating the eye gaze direction of a user wearing glasses in a vehicle is illustrated, in accordance with one or more embodiments of the present disclosure.
[0028] Figure 3C An exemplary logic diagram of the proposed system according to one or more embodiments of the present disclosure is illustrated to explain the working of the system in detail.
[0029] Figure 4 Exemplary key points associated with a user's eyes that facilitate detecting the user's eye gaze direction according to one or more embodiments of the present disclosure are illustrated.
[0030] Figure 5 An exemplary architecture of a random forest model implemented in the proposed system according to an embodiment of the present disclosure is illustrated.
[0031] Figure 6 An exemplary representation of a key point architecture implemented in the proposed system according to one or more embodiments of the present disclosure is illustrated. DETAILED DESCRIPTION
[0032] The following is a detailed description of the embodiments of the present disclosure as depicted in the accompanying drawings. The details of the embodiments are provided for clarity of the disclosure. However, the amount of detail provided is not intended to limit the intended variations of the embodiments; on the contrary, the invention is intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the present disclosure as defined by the appended claims.
[0033] The embodiments explained herein relate to the field of eye gaze estimation. Specifically, the present disclosure provides a system and method for estimating the eye gaze direction of a user in a vehicle, wearing or not wearing glasses.
[0034] See also Figure 1 , discloses a proposed system 100 for estimating the eye gaze direction of a user wearing glasses (also referred to herein as a driver, passenger, or occupant). Specifically, the disclosed system 100 accurately identifies the gaze direction (the direction in which the user / driver wearing glasses is looking) while driving a vehicle to avoid distraction and help prevent road accidents. If the user / driver's gaze direction is off the road for several seconds, the system warns the user / driver or switches the vehicle to automatic mode and stops the vehicle for the driver's safety.
[0035] In one embodiment, the system 100 may include a processing device 102 that may communicate with a vehicle control unit (VCU) 106 of a vehicle. In another embodiment, the processing device 102 may be a server that may reside in one or more vehicle units associated with one or more vehicles.
[0036] In another embodiment, the system 100 may further include an image acquisition unit 104, which includes one or more image sensors or cameras (collectively referred to herein as cameras or image sensors) installed within the vehicle to monitor one or more images or videos of a user (also referred to herein as a driver or passenger or occupant) sitting in / riding in the vehicle. The image acquisition unit 104 may also capture images or videos of the road or the external field of view of the vehicle. The image acquisition unit 104 may communicate with the processing device 102 and / or the vehicle's VCU 106.
[0037] In one exemplary embodiment, an image acquisition unit 104 including a camera and / or an image sensor 104 may be installed inside a vehicle to capture or monitor images or videos of a user / driver's face in real time. One of the cameras 104 may face the front seat of the vehicle, where the user / driver may be seated. Other cameras 104 may face the road or the exterior of the vehicle. The camera or image sensor 104 may be positioned on the dashboard or ceiling of the vehicle to provide complete coverage of the interior of the vehicle and capture images of the user / driver's face and also capture the exterior view or road. The camera 104 may be an infrared (IR) camera, such as a near-infrared camera, a medium-wave infrared camera, a long-wave infrared camera, or the like.
[0038] In one embodiment, the processing device 102 can communicate or be operably connected to the image acquisition unit 104 and the VCU 106 of the vehicle via a network. In addition, the network can be a wireless network, a wired network, or a combination thereof, which can be implemented as one of different types of networks, such as an intranet, a local area network (LAN), a wide area network (WAN), the Internet, etc. In addition, the network can be a dedicated network or a shared network. A shared network can represent an association of different types of networks that can use various protocols (e.g., Hypertext Transfer Protocol (HTTP), Transmission Control Protocol / Internet Protocol (TCP / IP), Wireless Application Protocol (WAP), etc.). In addition,
[0039] In one embodiment, the processing device 102 may be implemented using any one or a combination of hardware components and software components, such as a cloud, a server, a computing system, a computing device, a network device, etc. In addition, the processing device 102 may interact with the image acquisition unit 104 and the VCU 106 via a wired network or a wireless network.
[0040] Image acquisition unit 104 may be configured to capture images / videos of the user / driver's face. Processing device 102 may be configured with a model and a learning module, the model having a first branch configured with a convolutional neural network (CNN) architecture, wherein the model may be pre-trained to detect one or more key points from one or more images associated with one or more eyes of the user that are occluded by glasses. In one embodiment, the one or more key points may be associated with the sclera, iris, and eye center of the user / driver. Processing device 102 may be configured to detect one or more key points associated with one or more eyes of the user in real time from images captured by image acquisition unit 104 using the pre-trained CNN architecture. Furthermore, processing device 102 may be configured with a learning module, such as, but not limited to, a random forest model, that enables processing device 102 to determine the user's eye gaze direction based on the detected key points. The learning module enables processing device 102 to regress the detected key points to calculate yaw and pitch that indicate the eye gaze direction of the user / driver with or without glasses. Thus, the processing device 102 can accurately identify the direction in which a user / driver wearing the glasses is looking while driving a vehicle, which can help avoid distractions and help prevent road accidents.
[0041] In one embodiment, the CNN architecture may further include a case-based reasoning (CBR) block. In addition, the heat map may be fed to the CNN architecture to detect one or more key points. In addition, the pre-trained model of the processing device 102 may further include a second branch comprising a classification network, wherein the second branch may be fed with the output of at least one of the CBR blocks. The second branch may be trained to reduce entropy loss so that the model can determine one or more true key points of the eye based on training with synthetic data. In one embodiment, the random forest model or learning module may be fed with four vectors derived from the terminal junction of the iris center and the sclera, and / or the angular distance between the terminal sclera point and the iris center. In addition, the random forest model may be trained with one or more noise, key point removal, and one or more additional features to identify the yaw and gaze to obtain better accuracy and reliability.
[0042] In another embodiment, processing device 102 may be configured to pre-train or train a learning module (a random forest model) in real time using synthetic training data and / or real-time data, which includes a first set of images associated with the eye region and a second set of images associated with the eye region occluded by the glasses, thereby enabling the learning module to determine the user's gaze direction. During training, processing device 102 may receive the first set of images associated with the eye region and one or more corresponding key points present in the first set of images. Furthermore, processing device 102 may render the glasses with a predefined transparency over at least one of the eye regions in a predefined percentage of the first set of images to provide the second set of images. Thus, processing device 102 may train and test a CNN and random forest model using synthetic data including the first set of images, the second set of images, and the corresponding one or more key points, thereby enabling the learning module to accurately detect one or more key points in captured images of the user in real time.
[0043] In one example, when it is detected that one of the user's eyes is at least partially occluded, the processing device 102 may be configured to determine the user's gaze direction using key points associated with the user's other unoccluded eye.
[0044] In one embodiment, the processing device 102 may be configured to generate a set of warning signals using the vehicle's VCU 106 when it is determined that the user's gaze direction is outside a region of interest (ROI) for a predefined time, where the ROI is the road the vehicle is traveling on. In another embodiment, the processing device 102 may be configured to actuate the VCU 106 to stop the vehicle or switch the vehicle to an autonomous driving mode when it is determined that the user's gaze direction is outside the ROI (road) for a predefined time.
[0045] See also Figure 2, the block diagram shown herein depicts an exemplary functional unit of the processing device 102, which may include one or more processors 202, a memory 204, an interface 206, a processing engine 208, and a database 210. The one or more processors 202 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units 102, logic circuits, and / or any device that manipulates data based on operational instructions. Among other capabilities, the one or more processors 202 may be configured to retrieve and execute computer-readable instructions stored in the memory 204 of the server 102. The memory 204 may store one or more computer-readable instructions or routines that may be retrieved and executed to create or share data units through a network service. The memory 204 may include any non-transitory storage device, including, for example, volatile memory (such as RAM) or non-volatile memory (such as EPROM, flash memory, etc.).
[0046] In one embodiment, processing device 102 may also include interface 206. Interface 206 may include various interfaces, such as interfaces for data input and output devices (referred to as I / O devices, storage devices, etc.). Interface 206 may facilitate communication between processing device 102 and various devices coupled to server 102, such as an infotainment system, image acquisition unit 104, VCU 106, and the vehicle's power supply. Interface 206 may also provide a communication path for one or more components of processing device 102. Examples of such components include, but are not limited to, processing engine 208 and database 210.
[0047] In one embodiment, processing engine 208 can be implemented as a combination of hardware and programming (e.g., programmable instructions) to implement one or more functions of processing engine 208. In the examples described herein, such a combination of hardware and programming can be implemented in several different ways. For example, the programming for processing engine 208 can be processor-executable instructions stored on a non-transitory machine-readable storage medium, and the hardware for processing engine 208 can include processing resources (e.g., one or more processors) to execute such instructions. In this example, the machine-readable storage medium can store instructions that, when executed by the processing resources, implement processing engine 208. In such examples, processing device 102 can include a machine-readable storage medium that stores instructions and a processing resource that executes the instructions, or the machine-readable storage medium can be separate but accessible to both processing device 102 and the processing resource. In other examples, processing engine 208 can be implemented by electronic circuitry. Database 210 can include data stored or generated as a result of the functions implemented by any of the components of processing engine 208.
[0048] In one embodiment, the processing engine 208 may include a pre-trained model 212 having a CNN architecture and a classification network, a learning module 214, an actuation and control unit 216, an alarm unit 218, and other units 220. The other units 220 may implement functions that supplement the applications or functions executed by the processing device 102 or the processing engine 208.
[0049] According to one embodiment, the processing device 102 may enable the image acquisition unit 104 to capture images / videos of the face of the user / driver wearing the glasses and images of the road or the exterior of the vehicle. Figure 5 The pre-trained model 212 shown in detail may enable the processing device 102 to crop one or more eye regions from a captured image of the user / driver's face obscured by glasses and correspondingly detect one or more key points from the captured image. In one embodiment, the one or more key points may be associated with the user / driver's sclera, iris, and eye center, such as Figure 4 Shown in detail.
[0050] In one embodiment, the processing device 102 may also be configured with a learning module 214, such as Figure 6 The random forest model shown in FIG. 1 may enable processing device 102 to determine the user's eye gaze direction based on the detected keypoints. Learning module 214 enables processing device 102 to regress the detected keypoints to calculate yaw and pitch, which indicate the eye gaze direction of a user / driver wearing or not wearing glasses. Thus, pre-trained model 212 and learning module 214 may enable processing device 102 to accurately identify the direction a user / driver wearing glasses is looking while driving a vehicle, which may help avoid distractions and help prevent road accidents.
[0051] In an exemplary embodiment, Figure 5 As shown, the pre-trained model 212 may include a first branch configured with a convolutional neural network (CNN) architecture and a second branch including a classification network. The model is pre-trained to detect one or more key points from one or more images associated with one or more eyes of the user that are occluded by glasses. The CNN architecture may include a case-based reasoning (CBR) block. In addition, a heat map may be fed to the CNN architecture to detect one or more key points. The second branch may be fed with the output of at least one of the CBR blocks, wherein the second branch may be trained to reduce entropy loss so that the model can determine one or more true key points of the eye based on training using synthetic data.
[0052] In an exemplary embodiment, Figure 6As shown, the random forest model 214 can be fed with four vectors derived from the distal junction of the iris center and the sclera, and / or the angular distance between the distal sclera point and the iris center. In addition, the random forest model 214 can be trained using one or more noise, keypoint removal, and one or more additional features to identify the yaw and gaze, thereby achieving better accuracy and reliability.
[0053] In another exemplary embodiment, the processor 202 may cause the processing device 102 to pre-train or train the learning module 214 (random forest model) in real time using synthetic training data and / or real-time data, which may enable the learning module 214 to determine the user's gaze direction. The synthetic training data and / or real-time data include a first set of images associated with the eye region and a second set of images associated with the eye region occluded by the glasses. During training, the processing device 102 may receive the first set of images associated with the eye region and one or more corresponding key points present in the first set of images. In addition, the processing device 102 may render the glasses with a predefined transparency on at least one eye region in the eye region in a predefined percentage of the first set of images to provide the second set of images, such as Figure 3B Therefore, the processing device 102 may utilize synthetic data including the first set of images, the second set of images, and the corresponding one or more key points to train and test the model 212 and the learning module (random forest model) 214, which may enable the learning module 214 to accurately detect the one or more key points in the captured image of the user in real time.
[0054] In one embodiment, when it is determined that the user's gaze direction is outside the region of interest (ROI) for a predefined time, the actuation and control unit 216 can cause the processing device 102 to generate and transmit a set of alarm signals to the vehicle's VCU 106 to warn the user / driver, where the ROI is the road on which the vehicle is traveling.
[0055] In another embodiment, when it is determined that the user's gaze direction is outside the ROI (road) for a predefined time, the actuation and control unit 216 can cause the processing device 102 to actuate the VCU 106 to stop the vehicle or switch the vehicle to autonomous driving mode.
[0056] In another embodiment, the ROI may be a road on which the vehicle is traveling. When the user / driver's estimated gaze direction is estimated to be away from the road (ROI) for a first predefined time, the alarm unit 218 may cause the processing device 102 to generate an alarm indicating that the user / driver is distracted. Furthermore, when the user / driver's eyes are found to be closed for a second predefined time, the alarm unit 218 may cause the processing device 102 to generate an alarm indicating that the user / driver is sleeping or unconscious.
[0057] See also Figure 3A The proposed method 300 for estimating the eye gaze direction of a user wearing or not wearing glasses in a vehicle involves an image acquisition unit and a processing device connected to a vehicle control unit (VCU) of the vehicle. The method 300 includes a step 302 of capturing one or more images of the user's face by the image acquisition unit, followed by a step 304 of cropping one or more eye regions from the one or more captured images of the user's face by the processing device.
[0058] Method 300 also includes detecting, in real time, by a processing device configured with a CNN architecture, one or more key points from one or more images associated with one or more eyes of a user of the vehicle, step 306. The CNN architecture may be pre-trained to detect the one or more key points from one or more images associated with one or more eyes of a user that are occluded by glasses.
[0059] Furthermore, method 300 includes determining, using a learning module associated with the processing device, a user's eye gaze direction based on the detected key points at step 308. At step 308, the processing device may regress the detected key points to calculate yaw and pitch indicating the user's eye gaze direction.
[0060] See also Figure 3B and Figure 3C The image acquisition unit may capture an image of a user / driver sitting in a vehicle, and the face detector module may detect a face in the captured image. The processing device may then crop the face and further crop the eye region from the captured image. The CNN model may then detect key points in the cropped eye image. If no key points are detected, the processing device does not identify a gaze condition. Furthermore, the random forest model may then select unobstructed eyes and estimate the user / driver's gaze direction.
[0061] Thus, the present invention (system and method) accurately estimates the eye gaze direction of a user / driver wearing or not wearing glasses while driving a vehicle. In addition, if the driver's gaze direction is off the road for a few seconds, the present invention warns the driver or switches the vehicle to automatic mode and stops the vehicle for the driver's safety, thereby avoiding distraction and preventing road accidents.
[0062] Although various embodiments of the present invention have been described above, other and further embodiments of the invention may be devised without departing from the basic scope of the invention. The scope of the invention is determined by the appended claims. The invention is not limited to the described embodiments, versions or examples, which are included to enable one of ordinary skill in the art to make and use the invention when combined with the information and knowledge available to one of ordinary skill in the art.
[0063] Advantages of the present disclosure
[0064] The present disclosure monitors the eye gaze direction of a driver with or without glasses.
[0065] The present disclosure provides an efficient, reliable, and faster system and method for estimating the eye gaze direction of a driver wearing glasses or when the driver's eyes are occluded.
[0066] The present disclosure accurately identifies the direction a driver wearing glasses is looking while driving a vehicle to avoid distraction and help prevent road accidents.
[0067] If the driver's gaze direction is off the road for several seconds, the present disclosure warns the driver or switches the vehicle to automatic mode and stops the vehicle for the driver's safety.
[0068] The present disclosure provides an efficient and reliable system and method for estimating the eye gaze direction of a driver wearing or not wearing glasses, which alerts or warns the driver or switches the vehicle to automatic mode and stops the vehicle if the driver's gaze direction is off the road for several seconds.
Claims
1. A system (100) for estimating the eye gaze direction of a user wearing or not wearing glasses, the system (100) comprising: A processing device (102) incorporating a model (212) and a learning module (214), the model having a first branch configured with a convolutional neural network (CNN) architecture, wherein the CNN architecture includes a case-based reasoning (CBR) block, and wherein a heat map is fed to the CNN architecture to detect one or more key points, the model (212) being pre-trained to detect the one or more key points from one or more images associated with one or more eyes of a user that are occluded by glasses; the processing device (102) comprising a processor (202) coupled to a memory (204), wherein the memory (204) stores one or more instructions executable by the processor (202) to: detecting, in real time, one or more keypoints from one or more images associated with the one or more eyes of the user of the vehicle using the CNN architecture of the model (212); as well as The learning module (214) is used to determine the eye gaze direction of the user based on the detected key points.
2. The system (100) of claim 1, wherein the system (100) comprises an image acquisition unit (104) comprising a camera for capturing one or more images of the user's face, The processing device (102) is configured to: cropping one or more eye regions from the one or more captured images of the face of the user; detecting the one or more keypoints from the cropped one or more images using the CNN architecture of the model (212); as well as The detected keypoints are regressed using the learning module (214) to calculate yaw and pitch indicating the eye gaze direction of the user.
3. The system (100) of claim 1 , wherein the pre-trained model (212) of the processing device (102) further comprises a second branch comprising a classification network, the second branch being fed with an output of at least one of the CBR blocks, wherein the second branch is trained to reduce entropy loss so that the model (212) is able to determine one or more true key points of the eye based on training with synthetic data.
4. The system (100) of claim 1 , wherein the processing device (102) is configured to train the learning module (214) using synthetic training data and / or real-time data, such that the learning module (214) is able to determine the gaze direction of the user, the synthetic training data and / or real-time data comprising a first set of images associated with an eye region and a second set of images associated with the eye region occluded by glasses; and The one or more key points are associated with the sclera, the iris, and the center of the eyeball.
5. The system (100) of claim 4, wherein the learning module (214) is a random forest model; wherein the random forest model (214) is fed with four vectors derived from any one or a combination of: the distal junction of the iris center and the sclera, and the angular distance between the distal sclera point and the iris center; and in, The random forest model (214) is trained using one or more of noise, keypoint removal, and one or more additional features to identify the yaw and gaze.
6. The system (100) according to claim 1, wherein the processing device (102) is configured to: generating a set of alarm signals when it is determined that the gaze direction of the user is outside a region of interest (ROI) for a predefined time; and When it is determined that the gaze direction of the user is outside the ROI for the predefined time, the vehicle is stopped or switched to an automatic driving mode, wherein the ROI is a road on which the vehicle is traveling.
7. A system (100) according to claim 1, wherein when one of the eyes of the user is at least partially occluded, the processing device (102) is configured to use the one or more key points associated with the other unoccluded eye of the user to determine the gaze direction of the user.
8. A method (300) for estimating a user's eye gaze direction, the method (300) comprising the following steps: detecting (306) in real time, by a processing device (102) incorporating a model (212) having a first branch configured with a convolutional neural network (CNN) architecture, one or more keypoints from one or more images associated with one or more eyes of a user of a vehicle, wherein the CNN architecture has been pre-trained to detect the one or more keypoints from one or more images associated with one or more eyes of a user that are occluded by glasses; as well as The eye gaze direction of the user is determined (308) based on the detected keypoints using a learning module (214) associated with the processing device (102).
9. The method (300) according to claim 8, wherein the method (300) comprises the following steps: capturing (302) one or more images of the user's face by an image acquisition unit (104); cropping (304), by the processing device (102), one or more eye regions from the one or more captured images of the face of the user; detecting (306) the one or more keypoints from the cropped one or more images using the CNN architecture; The detected keypoints are regressed (308) using the learning module (214) to calculate yaw and pitch indicative of the eye gaze direction of the user.
Citation Information
Patent Citations
Determination of gaze direction
EP3789848A1