Method and system for training a model to predict an eye-opening of a driver of a vehicle
The method enhances driver gaze estimation in vehicles by training a predictive model to handle low-quality images, ensuring accurate and safe driving conditions.
Patent Information
- Application Number
- DE102024132867
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-04
- Filing Date
- 2024-11-11
- Publication Date
- 2025-06-05
AI Technical Summary
Existing image processing techniques for monitoring driver gaze direction in vehicles face challenges due to degraded image quality, leading to unreliable results and potential safety hazards.
A method and system for training a predictive model to estimate driver gaze by identifying low-quality images, generating updated images using pixel intensities, and training the model with reconstruction and gaze losses to enhance accuracy.
The system provides reliable and accurate gaze prediction, enabling safer driving by improving model performance on degraded images.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
PREAMBLE TO THE DESCRIPTION:
[0001] The following description explains the invention and the manner in which it is carried out in more detail: DESCRIPTION OF THE INVENTION: Technical field
[0002] The present disclosure relates generally to monitoring driving safety in motor vehicles. More particularly, but not exclusively, the disclosure relates to methods and systems for training models to predict the eye movement of vehicle drivers. Background of the disclosure
[0003] Vehicles are capable of monitoring drivers and detecting whether a driver is paying attention to the road while driving. Drivers may be distracted by other tasks such as looking at mobile phones, caring for children or pets in the vehicle, manually operating the vehicle's infotainment system, and the like. Therefore, it is important to monitor where the driver is looking while driving. This monitoring is done so that appropriate measures can be taken to prevent a collision with a vehicle if the driver is not paying attention. For example, warning systems can be activated to alert the driver that they are inattentive. In addition, automatic braking and steering systems can be activated to bring the vehicle to a stop if it is determined that the driver has not paid attention even after a warning.
[0004] To monitor alertness, the gaze direction of the eyes is estimated using image processing techniques. In image processing techniques, images are provided as input for processing. The input images may also have degraded image quality. The quality of the images can degrade due to various external factors, including but not limited to the quality of the camera, different lighting factors, environmental reflections, different accessories (e.g., different types of glasses), etc. Therefore, training a model to estimate the gaze direction of drivers using images with degraded quality is difficult and reduces the performance of the model because the results obtained based on the images with degraded quality may be subject to errors.The conclusions drawn from the input images can therefore lead to unreliable results and endanger the safety of drivers and passengers in the vehicle.
[0005] The information disclosed in this Background of the Disclosure section is provided merely to facilitate an understanding of the general background of the invention and is not to be construed as an acknowledgement or indication that such information constitutes prior art already known to a person skilled in the art. SUMMARY OF DISCLOSURE
[0006] In one non-limiting embodiment of the present disclosure, a method for training a model to predict the eye blink of a driver of a vehicle is disclosed. The method includes identifying one or more low-quality images from a plurality of images based on imaging features associated with the plurality of images. The plurality of images includes the one or more low-quality images and one or more acceptable-quality images. The method includes generating updated low-quality images for the one or more low-quality images based on pixel intensities associated with the one or more low-quality images and the one or more acceptable-quality images using a predefined image comparison technique.The method then includes training a predictive model to predict the gaze of a driver of a vehicle using the updated low-quality images and the one or more acceptable-quality images based on loss parameters associated with a reconstruction loss and a gaze loss. The loss parameters are determined by simultaneously performing the reconstruction of the updated low-quality images and the prediction of the gaze of the driver of the vehicle on the updated low-quality images.
[0007] In one embodiment of the disclosure, the imaging features include contrast levels associated with the plurality of images.
[0008] In one embodiment of the disclosure, identifying the one or more low-quality images includes identifying whether the contrast values associated with the plurality of images are within predefined thresholds.
[0009] In one embodiment of the disclosure, generating the updated low-quality images comprises transferring the pixel intensities associated with the one or more low-quality images to the one or more acceptable-quality images.
[0010] In one embodiment of the disclosure, training the predictive model includes determining the gaze loss for the updated low-quality images based on an angular error and a mean absolute error (MAE) associated with corresponding updated high-quality images. Training the predictive model includes determining the reconstruction loss for the updated low-quality images based on a predefined function between the one or more acceptable-quality images and reconstructed images obtained from the reconstruction of the updated low-quality images.
[0011] In another non-limiting embodiment of the disclosure, an eye-blink training system is disclosed for training a model to predict the eye-blink of a driver of a vehicle. The eye-blink training system includes a processor and a memory communicatively coupled to the processor, the memory storing processor-executable instructions that, when executed, can cause the eye-blink training system to identify one or more low-quality images from a plurality of images based on imaging features associated with the plurality of images. The plurality of images consists of one or more low-quality images and one or more acceptable-quality images.The eye-gaze training system generates updated low-quality images for the one or more low-quality images based on pixel intensities associated with the one or more low-quality images and the one or more acceptable-quality images using a predefined image matching technique. The eye-gaze training system then trains a predictive model to predict the gaze of a driver of a vehicle using the updated low-quality images and the one or more acceptable-quality images based on loss parameters associated with a reconstruction loss and a gaze loss. The loss parameters are determined by simultaneously reconstructing the updated low-quality images and predicting the gaze of the driver of the vehicle.
[0012] The foregoing summary is for illustrative purposes only and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments, and features described above, other aspects, embodiments, and features will become apparent by reference to the drawings and the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The accompanying drawings, which are incorporated in this disclosure, illustrate exemplary embodiments and, together with the description, serve to explain the disclosed principles. In the figures, the leftmost digit(s) of a reference number indicates the figure in which the reference number first appears. Throughout the figures, the same numerals are used to refer to like features and components. Some embodiments of systems and / or methods according to embodiments of the present subject matter will now be described, by way of example only, and with reference to the accompanying figures, in which: Fig. 1 shows an example environment for training a model to predict the eye blink of a driver of a vehicle according to some embodiments of the present disclosure; Fig. 2 shows a detailed block diagram of an eye-opening training system in accordance with some embodiments of the present disclosure; Fig. 3 shows an exemplary illustration of low quality image detection in accordance with some embodiments of the present disclosure; Fig. 4a-4b show exemplary illustrations of generating updated low-quality images according to some embodiments of the present disclosure; Fig. 5 shows an exemplary representation of a network architecture for training a predictive model in accordance with some embodiments of the present disclosure; Fig. 6 shows a flowchart illustrating a method for training a model to predict the gaze of a driver of a vehicle in accordance with some embodiments of the present disclosure; and Fig. 7 shows a block diagram of an exemplary computer system for implementing embodiments consistent with the present disclosure.
[0014] Those skilled in the art should appreciate that all block diagrams contained herein represent conceptual views of systems embodying the principles of the present subject matter. Likewise, it will be understood that all flowcharts, sequence diagrams, state transition diagrams, pseudocode, and the like represent various processes substantially capable of being represented in a computer-readable medium and executed by a computer or processor, whether or not such a computer or processor is explicitly depicted. DETAILED DESCRIPTION
[0015] As used herein, the word "exemplary" means "serving as an example, instance, or illustration." Any embodiment or implementation of the present subject matter described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments.
[0016] While the disclosure is susceptible to various modifications and alternative forms, a specific embodiment thereof has been shown by way of example in the drawings and will be described in detail below. It should be understood, however, that the disclosure is not intended to limit the disclosure to the forms disclosed; on the contrary, the disclosure is intended to include all modifications, equivalents, and alternatives falling within the spirit and scope of the disclosure.
[0017] The terms "comprise," "including," or other variations thereof are intended to cover non-exclusive inclusion, such that an assembly, device, or method comprising a list of components or steps not only includes those components or steps, but may also include other components or steps not expressly listed or included in such assembly, device, or method. In other words, one or more elements in a system or device introduced with "comprises" does not exclude, without further limitation, the presence of other elements or additional elements in the system or method.
[0018] The terms "comprises," "including," or other variations thereof are intended to cover non-exclusive inclusion, such that an assembly, device, or method comprising a list of components or steps not only includes those components or steps, but may also include other components or steps not expressly listed or included in such assembly, device, or method. In other words, one or more elements in a system or device introduced with "comprises" does not exclude, without further limitation, the presence of other elements or additional elements in the system or method.
[0019] In the following detailed description of the embodiments of the disclosure, reference is made to the accompanying drawings, which form a part of this document, and in which is shown by way of illustration specific embodiments in which the disclosure may be practiced. These embodiments are described in sufficient detail to enable one skilled in the art to practice the disclosure, and it is to be understood that other embodiments may be utilized and that changes may be made without departing from the scope of the present disclosure. The following description, therefore, is not to be taken in a limiting sense.
[0020] Fig. 1 shows an example environment for training a model to predict the eye blink of a driver of a vehicle according to some embodiments of the present disclosure.
[0021] As in Fig. 1, an environment 100 comprises an eye-opening training system 101 connected via a communication network 105 to one or more sources 103 (a source 103 1 , a source 103 2 , ................... and a source 103 N , collectively referred to as one or more sources 103) associated with a vehicle. The one or more sources 103 may include, among other things, image acquisition sensors such as cameras, a database with an image dataset corresponding to a plurality of images, and the like. In one embodiment, the eye-opening training system 101 may be configured in a control unit (CU) of the vehicle. The eye-opening training system 101 may include an input / output (I / O) interface 107, a memory 109, and a processor 111, as shown in Fig. 1. The I / O interface 107 may receive the image data set comprising the plurality of images from the one or more sources 103 connected to the vehicle. In one embodiment, the gaze training system 101 may be connected to an eye gaze prediction application, which may be implemented in a personal device of a driver or passenger of the vehicle. The eye gaze prediction system may enable the driver and passengers to take action to prevent vehicle-to-vehicle collisions when the driver is inattentive while operating the vehicle. Thus, the eye gaze training system 101 of the present disclosure is used to train a predictive model in a manner that produces accurate prediction results.
[0022] The eye-open training system 101 receives the image dataset from one or more sources 103. The image dataset may include a plurality of images. In one embodiment, the plurality of images may correspond to two-dimensional (2-D) images of the driver's eyes. The 2-D images captured in the vehicle may be appropriately cropped to extract the plurality of images including the driver's eyes. In one embodiment, the cropping process may be performed prior to sending the plurality of images to the eye-open training system 101. Additionally, the plurality of images includes one or more low-quality images and one or more acceptable-quality images. The one or more low-quality images may include images of reduced quality due to clarity issues.The one or more acceptable quality images may consist of images that can be classified as either high or average quality. Furthermore, the eye-lift training system 101 may identify one or more low-quality images from the plurality of images based on imaging features associated with the plurality of images. In one embodiment, the imaging features may refer to pixel intensities associated with the plurality of images in the image dataset. In one embodiment, a pixel intensity is defined as a value that comprises primary information stored in pixels of the corresponding images. In one embodiment, the eye-lift training system 101 may calculate a root mean square (RMS) contrast value for each of the plurality of images based on the pixel intensities.A person skilled in the art will understand that in addition to the RMS contrast level, any other technique can be used to filter the multiple images.
[0023] Specifically, the eye-lift training system 101 identifies the one or more low-quality images by determining whether the contrast value associated with the pixel intensities is within a predefined threshold range. For example, if the eye-lift training system 101 determines that one of the plurality of images is not within the predefined threshold range, the image is identified as a low-quality image. In one embodiment, the predefined threshold range may be defined based on the image dataset. Therefore, the predefined threshold range may vary for different image datasets.
[0024] Furthermore, the eye-lift training system 101 generates updated low-quality images for the one or more low-quality images based on the pixel intensities associated with the one or more low-quality images and the one or more acceptable-quality images. The eye-lift training system 101 uses a predefined image adjustment technique to generate the updated low-quality images. In one embodiment, the predefined image matching technique may correspond to the histogram matching technique. The histogram matching technique refers to the transformation of an image such that a histogram associated with the corresponding image matches a specific histogram associated with another image.The eye-lift training system 101 generates the updated low-quality images by transferring the pixel intensities associated with the one or more low-quality images into the one or more acceptable-quality images, resulting in the updated low-quality images. Specifically, the acceptable-quality images are downgraded by transferring the pixel intensities associated with the one or more low-quality images into the one or more acceptable-quality images. In one embodiment, prior to transferring the pixel intensities, the eye-lift training system 101 performs an appropriate selection of an acceptable-quality image based on the imaging features that approximate the corresponding low-quality image.Thus, the updated low-quality images generated based on the transfer of pixel intensities include the pixel intensities associated with both categories of images, namely the low-quality images and the acceptable-quality images. Therefore, an updated low-quality image may be generated based on an image pair comprising the corresponding low-quality image and the selected acceptable-quality image. In this way, the imaging characteristics of the low-quality images, which may be related to gaze information, may be preserved because the updated low-quality images include pixel intensities associated with both image categories. In one embodiment, the eye-lift training system 101 may add a predefined amount of Gaussian noise to the updated low-quality images.
[0025] Next, the eye-open training system 101 trains a predictive model to predict the eye-open of the vehicle's driver. Training is performed using the updated low-quality images and the one or more images of acceptable quality based on loss parameters associated with reconstruction loss and gaze loss. Training predictive models based on machine learning (ML) approaches is associated with a loss function that evaluates the performance of the predictive model. If the predictive model's prediction is accurate, the loss may be close to zero. However, if the prediction deviates, the loss may be higher.The loss parameters are determined by the eye-open training system 101 by performing two tasks simultaneously: reconstructing the updated low-quality images and predicting the eye-open of the vehicle's driver. In one embodiment, predefined techniques may be used to perform the reconstruction of the updated low-quality images and the prediction of the driver's eye-open. The predefined techniques include, among others, a variational autoencoder method for reconstruction and an EyeNet neural network method for predicting the eye-open.
[0026] Therefore, the eye-open training system 101 is configured to learn noise data associated with the updated low-quality images based on the loss parameters based on the predefined techniques. The eye-open training system 101 is further configured to generate a latent spatial embedding that can be used by the two tasks. The latent spatial embedding includes the updated low-quality images in an embedded or compressed format. Thus, the eye-open training system 101 uses the generated latent spatial embedding to perform reconstruction of the updated low-quality images, thereby obtaining reconstructed images for each of the one or more updated low-quality images. In doing so, the updated low-quality images are enhanced to match the acceptable-quality images.Furthermore, the eye-blink training system 101 uses the generated latent spatial embedding to perform the prediction of the driver's eye-blink.
[0027] The eye-lift training system 101 determines the reconstruction loss and gaze loss based on the two tasks. In one embodiment, the reconstruction loss may be determined based on a predefined function between the one or more acceptable quality images and the reconstructed images. The predefined function corresponds to a difference between the pixel intensities associated with the one or more acceptable quality images and the reconstructed images. Further, the gaze loss may be determined based on an angular error and a mean absolute error (MAE) associated with the corresponding low-quality images. In one embodiment, the MAE may be associated with yaw and pitch of the updated low-quality images. Yaw and pitch define the orientation of an object in an image.Therefore, the MAEs associated with yaw and pitch are used by the eye-lift training system 101 to determine the gaze loss. Thus, the eye-lift training system 101 uses the determined total loss to evaluate the performance of the predictive model and further improve the accuracy of the results obtained based on the total loss.
[0028] Once the model is trained, the predictive model is used to estimate the driver's gaze in real time based on images captured by one or more sensors in the driver's vehicle.
[0029] Fig. 2 shows a detailed block diagram of an eye-opening training system in accordance with some embodiments of the present disclosure.
[0030] The eye-opening training system 101 may include data 200 and one or more modules 209, which are described in detail herein. In one embodiment, the data 200 may be stored in the memory 109. The data 200 may include, for example, image data 201, generated data 203, training data 205, and other data 207.
[0031] The image data 201 may include data received from the one or more sources 103, which may comprise an image data set that may include the plurality of images. The plurality of images may include the one or more acceptable quality images and the one or more low quality images. The image data 201 may further include information about imaging features associated with the plurality of images and a predefined threshold range for performing the identification of the one or more low quality images.
[0032] The generated data 203 may include updated low-quality images that may be generated based on an image pair formed from the image data 201. The image pair includes a low-quality image and an acceptable-quality image.
[0033] The training data 205 may include a prediction model, reconstructed images, results related to the predicted gaze of a driver, loss parameters including a reconstruction loss and a gaze loss, a predefined function, and errors related to the image data 201.
[0034] The other data 207 may store data, including temporary data and temporary files, generated by the modules 209 to perform the various functions of the eye-opening training system 101.
[0035] In one embodiment, the data 200 in memory 109 is processed by the one or more modules 209 present in memory 109 of the eye-opening training system 101. In one embodiment, the one or more modules 209 may be implemented as dedicated units. As used herein, the term module refers to an application-specific integrated circuit (ASIC), an electronic circuit, a field-programmable gate array (FPGA), a programmable system-on-chip (PSoC), a combinational logic circuit, and / or other suitable components that provide the described functionality. In some embodiments, the one or more modules 209 may be communicatively coupled to the processor 111 to perform one or more functions of the eye-opening training system 101.When said modules 209 are configured with the functionality defined in the present disclosure, a novel hardware results.
[0036] In one implementation, the one or more modules 209 may include, among other things, an image quality detection module 211, a low-quality updated image generation module 213, and a training module 215. The one or more modules 209 may also include other modules 217 to perform various other functions of the eye-opening training system 101.
[0037] The image quality detection module 211 may receive the image data 201 from one or more sources 103. Since the image data 201 includes a plurality of images corresponding to one or more acceptable quality images and one or more low quality images, the image quality detection module 211 is configured to identify the one or more low quality images. Furthermore, the image quality detection module 211 identifies the remaining images (i.e., the images adjacent to the low quality images) as acceptable quality images. Therefore, the image quality detection module 211 is configured to detect the one or more low quality images and the one or more acceptable quality images.Specifically, the image quality detection module 211 performs the identification of the one or more low-quality images based on imaging features associated with the plurality of images. In one embodiment, the imaging features may refer to pixel intensities associated with the plurality of images in the image dataset. In one embodiment, the image quality detection module 211 calculates a root mean square (RMS) contrast value for each of the plurality of images based on the pixel intensities. The RMS contrast value may be determined as shown in equation (1). ContrastRMS:1MN∑i=0N−1∑j=0M−1(Iij−I¯)2 where I = average pixel intensity of an image; I ij = pixel intensity at the i - j th element of the image; and MN = image size.
[0038] Therefore, the image quality detection module 211 identifies one or more low-quality images by determining whether the contrast value associated with the pixel intensities is within the predefined threshold range.
[0039] Fig. Figure 3 shows an example of the identification of low-quality images. As can be seen from Fig. 3, the contrast values assigned to a plurality of images 301 are calculated, after which it is checked whether the contrast values lie within the predefined threshold range or not. Let us assume that the predefined threshold range corresponds to 65 to 75 (65 ≤ RMS Contrast≤ 75). Then, any contrast outside the specified range can be considered a low-quality image. For example, if the contrast of an image A is 70, the image quality detection module 211 identifies the image A as an acceptable-quality image, displayed within an acceptable-quality image 303. If the contrast of an image B is 50, the image quality detection module 211 identifies the image B as a low-quality image, displayed as 305. Thus, based on the pixel intensities associated with the plurality of images 301, one or more low-quality images 305 are identified.
[0040] Back to Fig. 2: The updated low-quality image generation module 213 may receive the one or more low-quality images from the image quality detection module 211. Further, the updated low-quality image generation module 213 generates an updated low-quality image for each of the one or more low-quality images based on the pixel intensities associated with the corresponding low-quality images and the one or more acceptable-quality images. A predefined image matching technique may be used to generate the updated low-quality images. In one embodiment, the predefined image matching technique may correspond to the histogram matching technique.The image adjustment technique is performed by the updated low-quality image generation module 213 by transferring the pixel intensities associated with the one or more low-quality images into the one or more acceptable-quality images, resulting in the updated low-quality images. In one embodiment, prior to transferring the pixel intensities, the updated low-quality image generation module 213 performs an appropriate selection of an acceptable-quality image based on imaging features that are close to the corresponding low-quality image. Therefore, an updated low-quality image may be generated based on an image pair comprising the corresponding low-quality image and the selected acceptable-quality image. Fig. Figure 4a shows the generation of updated low-quality images. As can be seen from Fig. 4a, the pixel intensities associated with the low quality images 305 and the acceptable quality images 303, as shown in Fig. 3, are merged to produce an updated low-quality image 401. In particular, the acceptable-quality images are downgraded by transferring the pixel intensities associated with the one or more low-quality images into the one or more acceptable-quality images. As in Fig. As can be seen in Figure 4a, a predefined amount of Gaussian noise 403 is added to the updated low-quality image 401. Fig. Figure 4b shows an image pair 405 formed on the basis of the acceptable quality images 303 and the updated low quality images 401.
[0041] Back to Fig. 2: The training module 215 may receive the updated low-quality images from the low-quality image generation module 213 and also receives the acceptable-quality images from the image quality detection module 211. The training module 215 is configured to train the predictive model for predicting the gaze of the vehicle driver based on the loss parameters associated with the reconstruction loss and the gaze loss. The loss parameters are determined by the training module 215 by simultaneously performing two tasks: reconstructing the updated low-quality images and predicting the gaze of the vehicle driver.
[0042] Equation (2) shows the total loss associated with the prediction model. Furthermore, equations (3) and (4) correspond to the reconstruction loss and gaze loss, respectively. Furthermore, equations (5) and (6) correspond to the MAEs and cosine similarities associated with gaze loss. Lmodel=α Lrecon+β Lgaze where, Lrecon=‖Yi−Yi~‖; where, Y i = Images of acceptable quality; and Y i ~ = reconstructed images; Lgaze=MAEyaw(5)+MAEpitch(5)+Cosine similarity(6) where, MAE=∑i=1n|yi−xi|n where MAE = Mean Absolute Error; y i = predicted value; x i = true value; and n = total number of data points. Cosine similarity = SC(A, B): = cos(θ) = AB‖A‖‖B‖ = ∑i = 1n Ai Bi ∑i = 1n Ai2 ∑i = 1n Bi 2 where A and B are yaw and pitch in radians.
[0043] Thus, by substituting the values into equations (3) and (4), the total loss can be determined, which is used by the training module 215 to evaluate the performance of the prediction model and further improve the accuracy based on the total loss results.
[0044] Thus, the training module 215 may consist of two sub-modules to perform the two tasks, namely the reconstruction of the updated low-quality images and the prediction of the eye-opening of the vehicle driver. The first sub-module may correspond to an image reconstruction sub-module, and the second sub-module may correspond to an eye-opening prediction sub-module. In one embodiment, predefined techniques may be used by the sub-modules to perform the reconstruction of the updated low-quality images and the prediction of the driver's gaze. The predefined technique for performing the reconstruction may correspond to a variational auto-encoder technique. Thus, the first sub-module may be connected to a variational auto-encoder (VAE) configured to reconstruct the updated low-quality images.Additionally, a neural network (EyeNet) may be used for eye gaze prediction. Thus, the gaze prediction sub-module may correspond to a neural network module configured to predict the gaze of the vehicle driver. In one embodiment, an output generated by the VAE and connected to the image reconstruction sub-module may be used by the neural network module, thus performing the two tasks concurrently. In one embodiment, the output may be a latent spatial embedding, which may include the updated low-quality images in an embedded or compressed format.
[0045] Fig. Figure 5 shows an exemplary representation of a network architecture for training a prediction model. The network architecture includes updated low-quality images 401, reconstructed images 505, and a Variational Auto Encoder (VAE) 507, which receives the updated low-quality images 401 as input 501 and generates an output 503. The output 503 corresponds to the reconstructed images 505. As shown in Fig. As can be seen in Figure 5, the prediction model is trained using the VAE 507, which is used for the two tasks (image reconstruction and gaze prediction). First, the updated low-quality images 401 are fed into the prediction model as input 501. Then, the VAE 507 generates the latent space embedding for the given input 501, which is used to reconstruct the updated low-quality image and predict the driver's gaze based on the updated low-quality images. When performing the reconstruction, the reconstructed images 505 may be generated as output 503 for each of the one or more updated low-quality images 401. In doing so, the updated low-quality images are enhanced to match the acceptable-quality images.Furthermore, the predictive model performs the eye-open prediction based on the updated low-quality images 401. The predictive model is trained using a neural network comprising multiple hidden layers 509. By performing the two tasks simultaneously, the loss parameters are determined. The loss parameters associated with the reconstruction loss and the gaze loss can be determined. Equations (2)-(4) above show the determination of the loss parameters. Thus, the training module 215 can measure the performance of the predictive model based on the loss parameters. When measuring the performance of the predictive model, the training module 215 can generate appropriate feedback and continue learning to produce improved results based on the feedback.Therefore, the predictive model includes a self-learning feature based on the loss parameters determined from improved-quality images (i.e., the updated low-quality images) rather than from degraded-quality images (the low-quality images 305). Once trained, the predictive model is used to estimate the driver's gaze in real time based on images captured by one or more sensors in the driver's vehicle.
[0046] Therefore, the present disclosure provides a reliable eye-blink training system 101 that trains the prediction model to produce accurate results.
[0047] Fig. 6 shows a flowchart illustrating a method for training a model to predict the gaze of a driver of a vehicle in accordance with some embodiments of the present disclosure.
[0048] As in Fig. 6, the method 600 includes one or more blocks illustrating a method for training a model to predict the eye movement of a driver of a vehicle. The method 600 may be described in the general context of computer-executable instructions. In general, computer-executable instructions may include routines, programs, objects, components, data structures, procedures, modules, and functions that perform particular functions or implement particular abstract data types.
[0049] The order in which the method 600 is described is not limiting, and any number of the described method blocks may be combined in any order to perform the method. Furthermore, individual blocks may be deleted from the methods without affecting the scope of the subject matter described herein. Furthermore, the method may be implemented in any suitable hardware, software, firmware, or a combination thereof.
[0050] At block 601, method 600 may include an eye-opening training system 101 identifying one or more low-quality images from a plurality of images received from one or more sources based on imaging characteristics associated with the plurality of images. The plurality of images consists of one or more low-quality images and one or more acceptable-quality images. Identifying the one or more low-quality images includes identifying whether the pixel intensities associated with the plurality of images are within a predefined threshold range.
[0051] At block 603, method 600 may include generating, by the eye-open training system 101, updated low-quality images for the one or more low-quality images based on pixel intensities associated with the one or more low-quality images and the one or more acceptable-quality images using a predefined image matching technique. Generating the updated low-quality images includes transferring the pixel intensities associated with the one or more low-quality images to the one or more acceptable-quality images.
[0052] At block 605, the method 600 may include training a predictive model for predicting the gaze of a driver of a vehicle by the eye-gaze training system 101 using the updated low-quality images and the one or more acceptable-quality images, based on loss parameters associated with a reconstruction loss and a gaze loss. The loss parameters are determined by simultaneously reconstructing the updated low-quality images and predicting the gaze of the driver of the vehicle. The reconstruction loss is determined based on a predefined function between the one or more acceptable-quality images and the reconstructed images obtained from reconstructing the updated low-quality images.Gaze loss is determined based on an angular error and a mean absolute error (MAE) associated with the corresponding updated low-quality images.
[0053] At block 607, the method 600 may include an eye-opening training system 101 determining a driver safety score based on a function of a maximum driver safety score and a risk factor associated with the one or more violations. The risk factor is a function of the violation score and a distance traveled by the vehicle in the predefined time period. The method further includes subtracting the risk factor from the maximum safe driving score.
[0054] An embodiment of the present disclosure trains the predictive model based on the updated low-quality images, which are an improved version of the degraded quality images. Therefore, the eye-gaze training system 101 achieves accurate results by efficiently analyzing the driver's attention based on their gaze, so that the driver and passengers are warned when the driver's attention deviates. Accordingly, the driver and passengers can take the necessary precautions.
[0055] Fig. 7 is a block diagram of an exemplary computer system for implementing embodiments consistent with the present disclosure.
[0056] In some embodiments, Fig. 7 is a block diagram of an exemplary computer system 700 for implementing embodiments consistent with the present disclosure. In some embodiments, the computer system 700 may be the eye-opening training system 101, which includes a processor (in this Fig. 7 (also referred to as processor 702) used to predict an eye blink of a driver of a vehicle. The processor 702 may include at least one data processor for executing program components for performing user- or system-generated business processes. The processor 702 may include specialized processing units such as integrated system (bus) controllers, memory management controllers, floating-point units, graphics processing units, digital signal processing units, etc.
[0057] The processor 702 may communicate with the input devices 710 and the output devices 711 via the I / O interface 701. The I / O interface 701 may use communication protocols / methods such as, without limitation, audio, analog, digital, stereo, IEEE-1394, serial bus, Universal Serial Bus (USB), infrared, PS / 2, BNC, coaxial, component, composite, Digital Visual Interface (DVI), High-Definition Multimedia Interface (HDMI), radio frequency (RF) antennas, S-Video, Video Graphics Array (VGA), IEEE 802.n / b / g / n / x, Bluetooth, cellular (e.g., Code-Division Multiple Access (CDMA), High-Speed Packet Access (HSPA+), Global System For Mobile Communications (GSM), Long-Term Evolution (LTE), WiMax, or the like), microphones, etc.
[0058] The computer system 700 can communicate with input devices 710 and output devices 711 via the I / O interface 701.
[0059] In some embodiments, the processor 702 may be connected to a communications network 709 via a network interface 703. The network interface 703 may communicate with the communications network 709. The network interface 703 may use connection protocols including, without limitation, direct connect, Ethernet (e.g., twisted pair 10 / 100 / 1000 Base T), Transmission Control Protocol / Internet Protocol (TCP / IP), Token Ring, IEEE 802.1 1a / b / g / n / x, OBD-II (On-Board Diagnostics) port, etc. The computer system 700 may communicate with one or more sources 103 via the network interface 703 and the communications network 709.
[0060] The communication network 709 may be implemented as one of various types of networks, such as an intranet or local area network (LAN), and the like within the organization. The communication network 709 may be either a dedicated network or a shared network, which is an aggregation of various network types that use a variety of protocols, such as Hypertext Transfer Protocol (HTTP), Transmission Control Protocol / Internet Protocol (TCP / IP), Wireless Application Protocol (WAP), etc., to communicate with each other.
[0061] In addition, the communication network 709 may include a variety of network devices, including routers, bridges, servers, computing devices, storage devices, etc. In some embodiments, the processor 702 may be connected to a memory 705 (e.g., RAM, ROM, etc., in Fig.7 not shown). The storage interface 704 may be connected to storage 705, including, without limitation, storage drives, removable disk drives, etc., using connection protocols such as Serial Advanced Technology Attachment (SATA), Integrated Drive Electronics (IDE), IEEE-1394, Universal Serial Bus (USB), Fibre Channel, Small Computer Systems Interface (SCSI), etc. The storage drives may further include a drum, a magnetic disk drive, a magneto-optical drive, an optical drive, a RAID (Redundant Array of Independent Discs), solid-state storage devices, solid-state drives, etc.
[0062] Memory 705 may store a collection of program or database components, including, without limitation, a user interface 706, an operating system 707, a web browser 708, etc. In some embodiments, computer system 700 may store user / application data, such as the data, variables, records, etc. described in this invention. Such databases may be implemented as fault-tolerant, relational, scalable, secure databases such as Oracle or Sybase.
[0063] The operating system 707 may facilitate resource management and operation of the computer system 700. Examples of operating systems include, without limitation, APPLE® MACINTOSH® OS X®, UNIX®, UNIX-like system distributions (e.g., BERKELEY SOFTWARE DISTRIBUTION® (BSD), FREEBSD®, NETBSD®, OPENBSD, etc.), LINUX® DISTRIBUTIONS (e.g., RED HAT®, UBUNTU®, KUBUNTU®, etc.), IBM® OS / 2®, MICROSOFT® WINDOWS® (XP®, VISTA® / 7 / 8, 10, etc.), APPLE® IOS®, GOOGLE™ ANDROID™, BLACKBERRY® OS, or the like. The user interface 706 may facilitate the display, execution, interaction, manipulation, or operation of program components through textual or graphical means. For example, user interfaces may provide computer interaction interface elements on a display system operatively connected to computer system 700, such as cursors, icons, check boxes, menus, scroll bars, windows, widgets, etc.Graphical user interfaces (GUIs) may be used, including, without limitation, Aqua® of the Apple® Macintosh® operating system, IBM® OS / 2®, Microsoft® Windows® (e.g., Aero, Metro, etc.), web interface libraries (e.g., ActiveX®, Java®, Javascript®, AJAX, HTML, Adobe® Flash®, etc.), or similar.
[0064] Computer system 700 may implement a web browser 708 with stored program components. Web browser 708 may be a hypertext viewing application such as MICROSOFT® INTERNET EXPLORER®, GOOGLE™ CHROME™, MOZILLA® FIREFOX®, APPLE® SAFARI®, etc. Secure web browsing may be provided using Secure Hypertext Transport Protocol (HTTPS), Secure Sockets Layer (SSL), Transport Layer Security (TLS), etc. Web browsers 708 may utilize facilities such as AJAX, DHTML, ADOBE® FLASH®, JAVASCRIPT®, JAVA®, application programming interfaces (APIs), etc. Computer system 700 may implement a stored program component of a mail server. The mail server may be an Internet mail server such as Microsoft Exchange or the like. The mail server can use facilities such as ASP, ACTIVEX®, ANSI® C++ / C#, MICROSOFT®, NET, CGI SCRIPTS, JAVA®, JAVASCRIPT®, PERL®, PHP, PYTHON®, WEBOBJECTS®, etc.The mail server may use communication protocols such as the Internet Message Access Protocol (IMAP), Messaging Application Programming Interface (MAPI), MICROSOFT® exchange, Post Office Protocol (POP), Simple Mail Transfer Protocol (SMTP), or the like. In some embodiments, computer system 700 may implement a stored mail client program component. The mail client may be a mail viewing application such as APPLE® MAIL, MICROSOFT® ENTOURAGE®, MICROSOFT® OUTLOOK®, MOZILLA® THUNDERBIRD®, etc.
[0065] Furthermore, one or more computer-readable storage media may be used in implementing embodiments according to the present invention. A computer-readable storage medium refers to any type of physical storage on which information or data readable by a processor can be stored. Thus, a computer-readable storage medium may store instructions for execution by one or more processors, including instructions that cause the processor(s) to perform steps or stages consistent with the embodiments described herein. The term "computer-readable medium" should be understood to include tangible objects and exclude carrier waves and transient signals, i.e., to be non-transitory.Examples include Random Access Memory (RAM), Read-Only Memory (ROM), volatile memory, non-volatile memory, hard disks, Compact Disc (CD) ROMs, Digital Video Disc (DVD), flash drives, floppy disks, and all other known physical storage media.
[0066] An embodiment of the present disclosure determines a total loss to evaluate the performance of the prediction model and to improve the accuracy of the results obtained based on the total loss.
[0067] An embodiment of the present disclosure facilitates the simultaneous performance of two tasks, namely, reconstructing the updated low-quality images and predicting the eye movement of the vehicle driver. Because the updated low-quality images are used instead of degraded images, the present disclosure provides reliable and accurate results. EQUIVALENTS
[0068] Regarding the use of terms in the plural and / or singular, the skilled person may translate from the plural to the singular and / or from the singular to the plural, depending on the context and / or application. The various singular / plural permutations may be explicitly listed here for the sake of clarity.
[0069] Those skilled in the art will understand that the terms used herein, and particularly in the appended claims (e.g., in the parts of the appended claims), are generally to be understood as "open-ended" terms (e.g., the term "including" should be interpreted as "including, but not limited to," the term "with" as "at least including," the term "comprises" as "comprises, but not limited to," etc.).
[0070] Those skilled in the art will further understand that when a particular number of introduced claims is intended, that intention will be expressly stated in the claim, and that in the absence of such a statement, no such intention exists. For example, in the following appended claims, the introductory phrases "at least one" and "one or more" may be used to introduce claim statements to facilitate understanding.
[0071] However, the use of such phrases should not be interpreted to mean that the introduction of a claim by the indefinite articles "a" or "an" limits a particular claim containing such an introductory claim statement to inventions containing only such a statement, even if the same claim contains the introductory phrases "one or more" or "at least one" and indefinite articles such as "a" or "an" (e.g., "a" and / or "an" should normally be interpreted to mean "at least one" or "one or more"); the same applies to the use of particular articles to introduce claim formulations. Even if a particular number of introductory claims is expressly mentioned, the skilled person will recognize that such a mention is usually to be understood as meaning at least the mentioned number (e.g.the mere mention of “two mentions” without other modifiers usually means at least two mentions or two or more mentions).
[0072] Furthermore, where a convention is used analogously to "at least one of A, B, and C, etc.", such a construction is generally meant in the sense in which a person skilled in the art would understand the convention (e.g., "a system having at least one of A, B, and C" would include, but is not limited to, systems having A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). Where a convention is used analogously to "at least one of A, B, or C, etc.", such a construction is generally meant in the sense in which a person skilled in the art would understand the convention (e.g., "a system having at least one of A, B, or C" would include, but is not limited to, systems having A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.).Those skilled in the art will appreciate that virtually any disjunctive word and / or disjunctive sentence containing two or more alternative terms in the description, claims, or drawings should be understood to encompass either of the terms, either of the terms, or both of the terms. For example, the phrase "A or B" should be understood to encompass the possibilities "A" or "B" or "A and B."
[0073] When features or aspects of the disclosure are described in terms of Markush groups, the skilled person will recognize that the disclosure is thereby also described in terms of an individual member or a subset of members of the Markush group.
[0074] Although various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are illustrative and not restrictive, with a true scope and spirit being indicated by the following claims. Reference symbols: 100 surroundings 101 Eye-Blink Training System 103 One or more sources 105 Communication network 107 I / O interface 109 storage 111 processor 200 data 201 image data 203 Created data 205 training data 207 Other data 209 modules 211 Image quality detection module 213 Updated module for generating low-quality images 215 Training Module 217 Other Modules 301 A variety of images 303 Acceptable image quality 305 Low image quality 401 Updated Low Quality Images 403 Gaussian noise 405 image pair 501 Input 503 Edition 505 Reconstructed Images 507 Variational Auto Encoder 509 Hidden Layers
Claims
[1] A method for training a model to predict an eye blink of a driver of a vehicle, the method comprising: Identifying, by an eye-opening training system, one or more low-quality images from a plurality of images received from one or more sources based on imaging features associated with the plurality of images, wherein the plurality of images comprises the one or more low-quality images and one or more acceptable-quality images; generating updated low-quality images for the one or more low-quality images by the eye-open training system based on pixel intensities associated with the one or more low-quality images and the one or more acceptable-quality images using a predefined image comparison technique; and Training a predictive model for predicting an eye-blink of a driver of a vehicle by the eye-blink training system using the updated low-quality images and the one or more acceptable-quality images based on loss parameters associated with a reconstruction loss and a gaze loss, wherein the loss parameters are determined by simultaneously performing a reconstruction of the updated low-quality images and a prediction of the eye-blink of the driver of the vehicle. [2] The method of claim 1, wherein the imaging features comprise pixel intensities associated with the plurality of images. [3] The method of claim 1, wherein identifying the one or more low quality images comprises identifying whether the pixel intensities associated with the plurality of images are within a predefined threshold range. [4] The method of claim 1, wherein generating the updated low quality images comprises transferring the pixel intensities associated with the one or more low quality images to the one or more acceptable quality images. [5] The method of claim 1, wherein training the predictive model comprises: Determining the reconstruction loss for the updated low-quality images based on a predefined function between the one or more images of acceptable quality and reconstructed images obtained from the reconstruction of the updated low-quality images; and Determine the gaze loss for the updated low-quality images based on an angular error and a mean absolute error (MAE) associated with the corresponding updated low-quality images. [6] An eye blink training system for training a model to predict an eye blink of a driver of a vehicle, comprising: a processor; and a memory communicatively coupled to the processor, the memory storing processor instructions that, when executed, cause the processor to: Identifying one or more low-quality images from a plurality of images received from one or more sources based on imaging characteristics associated with the plurality of images, wherein the plurality of images comprises the one or more low-quality images and one or more acceptable-quality images; generating updated low-quality images for the one or more low-quality images based on pixel intensities associated with the one or more low-quality images and the one or more acceptable-quality images using a predefined image comparison technique; and Training a predictive model to predict an eye-open of a driver of a vehicle using the updated low-quality images and the one or more acceptable-quality images based on loss parameters associated with a reconstruction loss and a gaze loss, wherein the loss parameters are determined by simultaneously performing a reconstruction of the updated low-quality images and a prediction of the eye-open of the driver of the vehicle. [7] The eye-opening training system of claim 6, wherein the imaging features comprise pixel intensities associated with the plurality of images. [8] The eye-opening training system of claim 6, wherein the processor identifies the one or more low-quality images by determining whether the pixel intensities associated with the plurality of images are within a predefined threshold range. [9] The eye-opening training system of claim 6, wherein the processor generates the updated low-quality images by transferring the pixel intensities associated with the one or more low-quality images to the one or more acceptable-quality images. [10] The eye-opening training system of claim 6, wherein the processor trains the prediction model by: Determining the reconstruction loss for the updated low-quality images based on a predefined function between the one or more images of acceptable quality and reconstructed images obtained from the reconstruction of the updated low-quality images; and Determine the gaze loss for the updated low-quality images based on an angular error and a mean absolute error (MAE) associated with the corresponding updated low-quality images.