Driver fatigue monitoring method based on eye movement tracking
By using eye tracking technology and BP neural network in the driver's fatigue monitoring system, we can identify the driver's attention changes and fatigue state, and solve the problem of expensive equipment and complex settings in the prior art, and achieve efficient monitoring and early warning in complex driving environments.
Patent Information
- Application Number
- CN202411870026.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art has the problem of expensive equipment and complex settings in driver fatigue monitoring, and it is difficult to effectively apply in complex and changeable real driving environments.
The driver's fatigue monitoring method based on eye tracking technology is adopted to obtain facial image frames, identify the central area of the iris, establish the world coordinate system and camera coordinate system, calculate the coordinates of the gaze point in three-dimensional space, and identify the driver's attention changes and fatigue state through the BP neural network.
It realizes efficient identification of driver facial features and accurate prediction of line of sight direction, and can effectively identify and monitor driver's attention and fatigue status in complex traffic scenarios, and conduct early warning operations in advance.
Smart Images

Figure CN120032347A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of assisted driving, and in particular to a driver fatigue monitoring method using eye tracking technology and environmental perception technology. Background Art
[0002] With the rapid progress in the field of transportation, the importance of traffic safety has become increasingly prominent, prompting the scientific research community and the industry to continue to explore and deepen the improvement of road safety. With the rapid development of science and technology, eye-tracking technology, as a cutting-edge innovation, is gradually being applied to the research and practice of driving safety. This technology can accurately capture and analyze the driver's eye movement trajectory, and map the driver's current focus in real time, providing unprecedented possibilities for predicting driving behavior and issuing immediate safety warnings, greatly enhancing the safety factor of the driving process.
[0003] At present, the research on gaze capture technology in driving scenarios mainly focuses on identifying abnormal driving behavior. Such as fatigue driving, making phone calls, etc. The researchers constructed an eye movement video dataset in real road scenes, and proposed a model (TWNet) based on the dual-path principle of the human visual cortex to simulate the visual mechanism in response to challenges such as camera movement, lighting changes, and line of sight occlusion. Experimental results show that the model can effectively identify the driver's eye movement behavior. However, existing research still needs to be deepened to fully support its practical application in complex and changeable real driving environments.
[0004] In order to apply gaze capture technology to fatigue determination, researchers have established a data collection platform for simulated driver fatigue parameters with an eye tracker as the core, and used a rough set attribute reduction method based on binary channels to screen out key attributes, and then developed a driver fatigue determination method based on BP neural network. Compared with the traditional PERCLOS (percentage of eyelid closure duration) method, this method has significantly improved the recognition accuracy. This study has opened up a new research direction in the field of driving safety, and at the same time shows that gaze capture technology has broad application potential and practical value in multiple high safety standard fields such as aviation and driving.
[0005] The goal of the present invention is to create a head-mounted driving warning system that can collect environmental information and driver's facial and eye information, and perform real-time analysis of the driver's attention and fatigue. This system can be implemented by using cameras installed on the device. The environmental perception camera will collect and analyze the environmental information in front of the vehicle in real time, and the facial camera will perform real-time analysis of the driver's facial information and eye feature information, and accurately estimate the driver's line of sight and gaze area. By analyzing the driver's line of sight and external environmental information, the system further expands the application prospects of line of sight capture technology in improving driving safety, optimizing driving training, etc.
[0006] Most of the existing technologies collect and analyze the driver's eye feature information by using various eye trackers, such as eye trackers. The equipment required for the overall system is usually expensive and complicated to set up. The present invention will play a positive role in reducing the cost of eye movement data collection, improving the capabilities of existing technologies, etc. Summary of the invention
[0007] The purpose of the present invention is to provide a cost-effective and easy-to-use driver fatigue monitoring method based on eye tracking technology in order to solve the problems mentioned in the above background technology.
[0008] The above-mentioned purpose of the present application is achieved through the following technical solutions:
[0009] S1: Obtain facial image frame;
[0010] S2: Collect facial image frames and Mediapipe detection through the front-end camera, record head posture, and identify the iris center area; extract the driver's continuous eye closure time through facial information and mouth status information;
[0011] S3: Use the PNP algorithm to establish the world coordinate system and the camera coordinate system so that the sight vector is consistent with the external The environment is located in the same space;
[0012] S4: Model the human eye and gaze area, and calculate the three-dimensional space through the geometric model Gaze point coordinates;
[0013] S5: Conduct real-time attention analysis through fixation points and determine the analysis results;
[0014] S6: Through the BP neural network, combined with the gaze point analysis results, the driver's attention changes and fatigue status can be effectively identified.
[0015] Optionally, step S1 includes:
[0016] S11: Obtaining real-time video of the driver;
[0017] S12: Training the BP neural network using the eye feature dataset and the mouth feature dataset
[0018] S13: Use the trained BP neural network to perform real-time recognition on the collected video data and extract facial feature information; locate the eye area and mouth area; detect the facial area, eye area and mouth area through the BP neural network to obtain a facial image frame.
[0019] Optional step S13 includes:
[0020] S13a: After the facial region recognition is completed, the eye region and the mouth region are located, and the eye feature information and the mouth feature information are extracted; the facial image is divided, and the eye and mouth regions are marked;
[0021] S13b: cutting out the eye area and the mouth area according to the marked image;
[0022] S13c: Detect the eye area image and the mouth area image respectively through the BP neural network.
[0023] Optional step S2 includes:
[0024] S21: Convert the facial image frame color space and convert the image into RGB mode;
[0025] S22: Detect key points of the face in the image through the Mediapipe face grid detector;
[0026] S23: Draw a face mesh through the detected key points and locate the pupil position.
[0027] Optional, Figure 2 The schematic diagram of the PNP algorithm is shown in FIG. 1 , and step S3 includes:
[0028] S31: Obtain the conversion relationship between the coordinate system and the camera coordinate system through the PNP algorithm;
[0029] S32: Establish camera coordinate system and world coordinate system for eye tracking camera and external environment camera Standard system;
[0030] S34: By using the obtained transformation relationship, the position relationship between cameras, and the internal parameters, the sight vector is transformed into the environment perception camera coordinate system through the rotation and translation relationship, such as Figure 2 shown.
[0031] Optionally, step S4 includes:
[0032] S41: After coordinate conversion, determine the unit vector of the optical axis in the world coordinate system;
[0033] S42: The optical axis is defined as a three-dimensional vector connecting the center of the pupil and the center of the cornea. The visual axis is a three-dimensional vector connecting the center of the cornea and the fovea. The angle between the two is called the kappa angle. The positional relationship between the visual axis and the optical axis is modeled through the kappa angle. Figure 4 Schematic diagram of the angle between the visual axis and the optical axis.
[0034] Define a simple visual axis-optical axis angle model, the unit vector of the optical axis is v o =[0 0 1]T , the angle between the visual axis and the optical axis is expressed as κ = (α, β), and the visual axis vectors of the left and right eyes can be expressed as:
[0035]
[0036] Decomposing the optical axis into horizontal and vertical angles is expressed as but:
[0037]
[0038] Calculate the intersection of the left and right eye visual axis vectors in space to estimate the subject's current gaze point. Define the three-dimensional coordinates of the gaze point as g(o,κ)=(g x ,g y ,g z ), the three-dimensional coordinate of the eyeball is o=(o x ,o y ,o z ), the fixation point g(o,κ) is determined by the following formula:
[0039]
[0040] Different subjects have different kappa angles. During the personal calibration process, the subjects are asked to look at N specific calibration points on the screen: g i , i = 1, 2, ..., N, and then optimize the model parameters by minimizing the distance between the estimated predicted fixation point and the actual fixation point, where g i are the actual coordinates of the calibration points, is the subject's current gaze point, and the diagram of the angle between the visual axis and the optical axis is as follows Figure 3 .
[0041]
[0042] Optionally, step S5 includes:
[0043] S51: integrating the driver's gaze point information and external environment perception information;
[0044] S52: Analyzing the driver's gaze area in the external environment by fusing the information;
[0045] S53: Determine the real-time attention status by analyzing the results.
[0046] Optional step S6 includes:
[0047] S61: Recognize the driver's mouth, eyes, and head through the BP neural network;
[0048] S62: (The average time a person's mouth is wide open when yawning is at least 5 seconds.) The camera monitoring frequency is 25 frames per second. If the driver's mouth state time state series in 150 frames shows that the mouth is open If the number of consecutive yawns exceeds 125, the driver is considered to be in a yawning state;
[0049] The PERCLOS method uses the P80 standard to count the number of times the eyes are closed within one minute. The proportion of time taken up;
[0050] S63: By fusing the driver's gaze area, the external driving environment image and the BP neural network output structure The results are used to analyze the current driver's attention and driving fatigue status;
[0051] S64: Set an extreme threshold, track the line of sight based on the camera, and determine the driver's eye movement direction and gaze area based on the algorithm. If the driver's line of sight continues to deviate from the front within a preset time, If the driver's sight is still in the deviation state after the time of deviation exceeds the preset value;
[0052] The driver yawned multiple times and the driver's eye features met the preset standards of the PERCLOS method. The system determined that the driver was driving fatigued and took early warning measures.
[0053] An electronic device, characterized by a data acquisition unit, a wireless communication unit, a central processing unit, and a storage unit, wherein the data acquisition unit is used to collect the driver's forward-looking image information and the driver's facial feature information, the storage unit is used to store the collected image data information, the wireless communication unit is used to communicate with other devices, and the central processing unit is used to execute pre-set instruction operations so that the electronic device executes the method described in any one of claims 1 to 8.
[0054] The technical solution provided by this application brings the following beneficial effects:
[0055] Use Mediapipe detection to locate and identify the iris center area of the acquired facial image frame; Monitor and analyze the driver's eye movement status in real time to determine changes in the driver's line of sight; Record the current head posture and facial feature information; by collecting facial information, establish a facial grid, record the driver's continuous eye closure time and mouth state changes; combine external environment information to determine the driver's gaze area in the external environment. Greatly improve the ability to recognize the driver's facial features and the level of prediction of the driver's line of sight and gaze area. In complex traffic scenarios, through real-time collection and analysis of the driver's eye feature information, it can effectively identify and monitor the driver's current attention status and current fatigue status characteristics, and provide early warning operations for the driver's fatigue state or abnormal driving behavior.
[0056] Use the PNP algorithm; establish the world coordinate system and the camera coordinate system so that the sight vector is in the same space as the external environment. The "gold standard" solution of PNP is to estimate the six parameters of the transformation by minimizing the norm of the reprojection error using a nonlinear minimization method.
[0057] The human eye and gaze area are modeled, and the coordinates of the gaze point in three-dimensional space are calculated through geometric model. The facial camera can observe in real time whether the driver is paying attention to the road traffic conditions through line of sight estimation, and give corresponding output results. It only needs to perform information processing calculations according to the budget algorithm, simplifying the system operation, complete packaging, and easy to operate.
[0058] BP neural network imitates the response of animal neurons to external stimulus signals, establishes a multi-layer perceptron model, and uses the learning mechanism of signal forward propagation and error reverse regulation. Through multiple iterative learning, it successfully builds an intelligent network model for processing nonlinear information, providing strong support for data analysis and processing.
[0059] Through the driver's gaze point, combined with external environment information, real-time attention analysis is performed to determine the analysis results; combined with the driver's facial state information, the attention mechanism is used to perform data fusion analysis to determine whether it meets the warning standards. The lightweight feature of the attention module ensures that the system overhead will not be excessively increased, and it can accept multiple forms of data input; it enhances its robustness while ensuring the system's operating speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 System flow chart
[0061] Figure 2 PNP algorithm schematic diagram
[0062] Figure 3 Schematic diagram of the angle between the visual axis and the optical axis Figure 4 Schematic diagram of double sphere model projection Specific implementation methods
[0063] In order to have a clearer understanding of the technical features, purposes and effects of the present application, the specific implementation methods of the present application are now described in detail with reference to the accompanying drawings.
[0064] An embodiment of the present application provides a driver fatigue monitoring method based on eye tracking.
[0065] Please refer to Figure 1 , Figure 1 This is a step diagram of a driver attention analysis method using eye tracking technology in an embodiment of the present application, including:
[0066] S1: Obtain facial image frame;
[0067] S2: Collect facial image frames and Mediapipe detection through the front-end camera, record head posture, and identify the iris center area; extract the driver's continuous eye closure time and mouth status information through facial information;
[0068] S3: Using the PNP algorithm; establishing the world coordinate system and the camera coordinate system so that the sight vector and the external environment are located in the same space;
[0069] S4: Modeling the human eye and the gaze area, and calculating the gaze point coordinates in the three-dimensional space through the geometric model;
[0070] S5: Conduct real-time attention analysis through fixation points and determine the analysis results;
[0071] S6: Through the BP neural network, combined with the gaze point analysis results, the driver's attention changes and fatigue status can be effectively identified.
[0072] Step S1 includes:
[0073] S11: Obtaining real-time video of the driver;
[0074] S12: Training the BP neural network using the eye feature dataset and the mouth feature dataset
[0075] S13: Use the trained BP neural network to perform real-time recognition on the collected video data and extract facial feature information; locate the eye area and mouth area; detect the facial area, eye area and mouth area through the BP neural network to obtain a facial image frame.
[0076] Specifically, in the BP neural network training stage, it is divided into three subtasks. The recognition of the mouth and eyes is completed by a single neural network respectively. Each neural network is learned separately, and then the recognition results of each neural network are combined for the final judgment.
[0077] Specifically, the BP neural network structure is a three-layer structure. The input layer has 3 neurons, which are the characteristic points of the driver's mouth, including the maximum width of the mouth area, the maximum height of the mouth area, and the height between the upper and lower lips; the driver's eye feature vector is input into the BP network, and the output of the network can be used to judge the open and closed state of the driver's eyes. The BP neural network has a three-layer structure. The input layer has 4 neurons, which represent the driver's eye blinking frequency, continuous eye closure time, pupil diameter, and PERCLOS respectively.
[0078] Specifically, the system will collect video information in real time during driving, mainly from facial cameras and environmental perception cameras. The data information collected in this part is mainly used to detect the driver's facial information and external environment information.
[0079] Step S13 includes:
[0080] S13a: After the facial region recognition is completed, the eye region and the mouth region are located, and the eye feature information and the mouth feature information are extracted; the facial image is divided, and the eye and mouth regions are marked;
[0081] S13b: cutting out the eye area and the mouth area according to the marked image;
[0082] S13c: Detect the eye area image and the mouth area image respectively through the BP neural network.
[0083] Specifically, the captured video is preprocessed through the BP neural network to detect and identify the eye area and the mouth area, extract the complete image frame information, filter out blurry, missing and other unusable image information, and perform preparatory operations for subsequent image processing.
[0084] Step S2 includes:
[0085] S21: Convert the facial image frame color space and convert the image into RGB mode;
[0086] S22: Detect key points of the face in the image through the Mediapipe face grid detector;
[0087] S23: Draw a face mesh through the detected key points and locate the pupil position.
[0088] Specifically, the facial camera collects real-time facial information, and the facial key point network is established through Mediapipe to extract eye feature point information and pupil location. In order to further improve the accuracy of facial recognition and eye feature information, the system screens the extracted facial images and cuts the eye area and mouth area maps through mouth and eye features. By filtering useless facial feature information, unnecessary computing power waste is reduced, detection efficiency is improved, and detection results are optimized.
[0089] Specifically, according to the physiological structure of the eyeball, a double-sphere mathematical model is established for the real eyeball structure. The model consists of two spheres, namely the pupil sphere and the eyeball sphere. The light path generated by the line of sight can be regarded as the line connecting the centers of the two spheres. By calculating their relative position and size in the three-dimensional Cartesian coordinate system, the optical axis vector of the eyeball is determined.
[0090] In the two-sphere model, the optical axis is represented as the line connecting the centers of the two spheres. For the pupil fitting ellipse in the two-dimensional image, the optical axis ||v 3d =(v x ,v y ,v z )|, we need to first obtain the optical axis in the two-dimensional plane
[0091] The pupil image is used as input data. The schematic diagram of the double-sphere model projected onto the two-dimensional pupil image is as follows Figure 4 shown.
[0092] When projected onto a two-dimensional image, the three-dimensional optical axis vector passes through the center of the eyeball and the center of the pupil in the two-dimensional image, and the direction of the vector is the same as the direction of the minor axis of the pupil fitting ellipse. According to the rotational invariance of the sphere, the above two features always hold true when the pupil moves, and the eyeball coordinates and radius remain unchanged.
[0093] By calculating the coordinates of the center point (x e ,y e ) and radius R, let the center of the eyeball be located on the XOY plane of the three-dimensional coordinate system, and determine the three-dimensional coordinate equation of the eyeball (xx e ) 2 +(yy e ) 2 +z 2 =R 2 . Define the three-dimensional coordinate equation of the pupil sphere (xi) 2 +(yj) 2 +(zk) 2 =r 2 , combined eyeball and pupil sphere equations:
[0094]
[0095] Solve the equations to get the plane where the two spheres intersect, and substitute it into any equation to get the projection equation of the intersection line on the XOY plane. And the projection of the center of the pupil sphere on the XOY plane falls on the minor axis l:y=kx+b of the ellipse fitting equation. Two constraint equations can be obtained:
[0096]
[0097] By solving the above equation, we can find the center of the pupil sphere (x p ,y p ,z p ), and the center of the eyeball solved in the previous article, the optical axis of the double-ball model is calculated, and the modeling of the double-ball model of the eyeball is completed.
[0098]
[0099] Step S3 includes:
[0100] S31: Obtain the conversion relationship between the coordinate system and the camera coordinate system through the PNP algorithm;
[0101] S32: establishing a camera coordinate system and a world coordinate system for the eye tracking camera and the external environment camera;
[0102] S34: By using the obtained transformation relationship, position relationship between cameras, and internal parameters, the sight vector is transformed into the environment perception camera coordinate system through the rotation and translation relationship.
[0103] Specifically, the PNP function estimates the object pose given a set of object points, their corresponding image projections, and the camera intrinsic parameter matrix and distortion coefficients.
[0104] According to the PNP algorithm theory, we only need to know n three-dimensional points in the world coordinate system and their corresponding two-dimensional points, as well as the intrinsic parameters of the camera, to obtain the transformation relationship between the world coordinate system and the camera coordinate system. In OPENCV, using the CV..SOLVEPNP_ITERATIVE function, only four point pairs are needed to calibrate the rotation matrix R and translation vector T.
[0105] Specifically, given a 2D feature point, we can check the distance between the projected 3D point and the 2D facial feature. When the estimated pose is perfect, the 3D point projected onto the image plane will be almost perfectly aligned with the 2D feature. When the pose estimate is incorrect, we can calculate the reprojection error metric - the sum of the squared distances between the projected 3D point and the 2D facial feature point.
[0106] Specifically, after using the PNP algorithm to obtain the transformation relationship between the two coordinate systems, the camera coordinate system and the world coordinate system are established for the macro camera used for eye tracking. The camera coordinate system and the world coordinate system are also established for the external environment perception camera. The transformation relationship between the camera coordinate system and the world coordinate system of each camera can be obtained by the PNP algorithm.
[0107] Specifically, through the known position information between the two cameras, the three-dimensional coordinates of the sight vector in its camera coordinate system can be converted to the external environment perception camera coordinate system through a certain rotation and translation relationship, realizing the fusion of the sight vector and the vehicle's external environment in the same coordinate system, and converted to image coordinates through the camera's internal reference relationship, so that the sight line and environmental information are located in the same picture. The specific formula is as follows:
[0108] The PNP function estimates the object pose given a set of object points, their corresponding image projections, and the camera intrinsic matrix and distortion coefficients, expressed in the world frame. W is projected into the image plane [u,v] using the perspective projection model Π and the camera intrinsic parameter matrix A:
[0109]
[0110]
[0111] The estimated pose is therefore the rotation R and translation T vectors that transform a 3D point represented in the world frame to the camera frame:
[0112]
[0113]
[0114] Equations of the above form can be solved using some algebraic derivation by using a method called Direct Linear Transformation (DLT).
[0115] The "gold standard" solution to PNP is to estimate the six parameters of the transformation by minimizing the norm of the reprojection error using a nonlinear minimization method. Minimizing this reprojection error provides the maximum likelihood estimate when assuming Gaussian noise in the measurements. The problem can be stated as:
[0116]
[0117] Where d is the Euclidean distance between two points. Solving the above equation involves minimizing the cost function E(q) = ke(q), defined as follows:
[0118]
[0119] After using the PNP algorithm to obtain the transformation relationship between the two coordinate systems, we will then use these transformation relationships to transform between multiple coordinate systems. For the macro camera used for eye tracking, we establish its camera coordinate system and world coordinate system. For the external environment perception camera, we also establish its camera coordinate system and world coordinate system. The transformation relationship between the camera coordinate system and the world coordinate system of each camera can be obtained using the PNP algorithm.
[0120] Optionally, step S4 includes:
[0121] S41: After coordinate conversion, determine the unit vector of the optical axis in the world coordinate system;
[0122] S42: The optical axis is defined as a three-dimensional vector connecting the center of the pupil and the center of the cornea. The visual axis is a three-dimensional vector connecting the center of the cornea and the fovea. The angle between the two is called the kappa angle. The positional relationship between the visual axis and the optical axis is modeled through the kappa angle. The specific formula is as follows:
[0123] Define a simple visual axis-optical axis angle model, the unit vector of the optical axis is v o =[0 0 1] T , the angle between the visual axis and the optical axis is expressed as κ = (α, β), and the visual axis vectors of the left and right eyes can be expressed as:
[0124]
[0125] Decomposing the optical axis into horizontal and vertical angles is expressed as but:
[0126]
[0127] Calculate the intersection of the left and right eye visual axis vectors in space to estimate the subject's current gaze point. Define the three-dimensional coordinates of the gaze point as g(o,κ)=(g x ,g y ,g z ), the three-dimensional coordinate of the eyeball is o=(o x ,o y ,o z ), the fixation point g(o,κ) is determined by the following formula:
[0128]
[0129]
[0130] Different subjects have different kappa angles. During the personal calibration process, the subjects are asked to look at N specific calibration points on the screen: g i , i = 1, 2, ..., N, and then optimize the model parameters by minimizing the distance between the estimated predicted fixation point and the actual fixation point, where g i are the actual coordinates of the calibration points, is the subject's current gaze point.
[0131]
[0132] Specifically, the kappa angle of adults is fixed, with horizontal and vertical rotation angles of about 5° and 1.5° respectively. The visual axis is used to measure the current gaze direction of the subject, that is, the gaze point is defined as the intersection of the binocular visual axes, not the intersection of the optical axes, so the positional relationship between the optical axis and the visual axis needs to be further modeled.
[0133] Specifically, in the driver state perception module, the camera in the driver's eyes uses line of sight estimation to confirm whether the driver is paying attention to the road traffic conditions. The driving road environment is also collected through the external environment perception module, and the fusion results are displayed in a visual way, which requires a series of coordinate transformations and video overlays.
[0134] This system uses a pinhole camera model. The P of the projected scene is obtained by projecting the three-dimensional points of the scene. W , use perspective transformation to enter the image plane and form pixel p in the pixel plane, P w and p are expressed in homogeneous coordinates, i.e. as three-dimensional and two-dimensional homogeneous vectors respectively.
[0135] Step S5 includes:
[0136] S51: integrating the driver's gaze point information and external environment perception information;
[0137] S52: Analyzing the driver's gaze area in the external environment by fusing the information;
[0138] S53: Determine the real-time attention status by analyzing the results.
[0139] Specifically, we selected the lightweight version of YOLOv8 in the YOLO series, YOLOv 8s, as the model basis for experimental training. By collecting external environmental information in real time, we used the preset model to perform semantic recognition and instance segmentation of traffic elements in the external environment.
[0140] The driver's eye feature information is recognized and processed through the BP neural network, and the coordinate transformation is performed using the PNP algorithm to obtain the estimation of the driver's gaze point in the external environment picture.
[0141] By adopting the attention mechanism and the system data fusion module, it is judged whether the driver has noticed the traffic elements in the external environment based on the driver's eye movement status.
[0142] Step 6 includes:
[0143] S61: Recognize the driver's mouth, eyes, and head through the BP neural network;
[0144] S62: When a person yawns, the average time that the mouth is wide open is at least 5 seconds, and the camera monitoring frequency is 25 frames per second. If the number of consecutive occurrences of the driver's mouth open in the time state series of the mouth state in 150 frames exceeds 125 times, it is determined that the driver is in a yawning state;
[0145] The PERCLOS method uses the P80 standard to count the proportion of time the eyes are closed within one minute;
[0146] S63: Analyze the current driver's attention status and driving fatigue status by fusing the driver's gaze area, the external driving environment image and the BP neural network output result;
[0147] S64: setting an extreme threshold, tracking the line of sight according to the camera, and judging the driver's eye movement direction and gaze area according to the algorithm. If the driver's line of sight continues to deviate from the front within a preset time, and if the driver's line of sight deviates for a time exceeding the preset value, the driver's line of sight is still in a deviated state;
[0148] The driver yawned multiple times and the driver's eye features met the preset standards of the PERCLOS method. The system determined that the driver was driving fatigued and took early warning measures.
[0149] Specifically, P80 means that the eyes are considered closed when the area of the pupil covered by the eyelid exceeds 80%, and the proportion of time closed within a certain period of time is calculated.
[0150] This system uses the P80 standard for the PERCLOS method, which counts the proportion of time the eyes are closed within one minute.
[0151] The normal threshold of PERCLOS is set at 40%, and the slightly lower threshold is set at 35%. The pupil diameter can be measured by the geometric grayscale binary method of the eye tracker. After consulting the data, it is determined that the normal threshold is set at a change rate of 30% between its diameter and the normal state; the slightly lower threshold is set at 20%.
[0152] Specifically, the algorithm flow of this application is as follows:
[0153] 1. Start: Turn on the camera and the system starts running.
[0154] 2. Collect facial images: Obtain the driver’s real-time facial image frame information from the facial camera.
[0155] 3. Recognize face, eyes and mouth: Use BP neural network to recognize the face, eyes and mouth position information in the image.
[0156] 4. Locate the key points of the eyes and mouth: Use Mediapipe detection to establish a key point network to locate the eye area and mouth area.
[0157] 5. Extract eye and mouth image frames: Cut out the eye and mouth areas from the acquired video data.
[0158] 6. Extraction of driver’s eye movement: Use BP neural network to determine the driver’s eye movement status and obtain pupil and rotation angle information including line of sight direction.
[0159] 7. Identify external environment information: The external camera collects external environment information in real time
[0160] 8. Coordinate conversion: Use the PNP algorithm to convert the estimated driver's line of sight to the same coordinates of the external environment image.
[0161] 9. Gaze point determination: Determine the coordinates of the driver’s gaze point by combining external environment information and the driver’s line of sight estimation.
[0162] 10. Fusion analysis of driver attention: Determine the driver’s current attention status based on the external environment image and the driver’s gaze area.
[0163] 11. Analyze the driver’s facial status information: count the proportion of time when the eyes are closed within one minute.
[0164] 12. Driver attention fatigue status analysis: Integrate the driver's gaze area, external driving environment image and BP neural network output results to analyze the current driver's attention status and driving fatigue status.
[0165] 13. Determine the driver's line of sight: Determine the driver's eye movement direction and gaze area. If the line of sight continues to deviate from the front and exceeds the preset value, the driver's line of sight is still in a deviated state; and the driver yawns multiple times, and the driver's eye characteristics meet the preset standards in the PERCLOS method, the system will determine that the driver is driving fatigued and take early warning measures.
[0166] 14. End: The system is finished running
[0167] This system has achieved significant enhancements in the accuracy and response speed of attention assessment, especially in continuous driving monitoring of driver fatigue and identification of distracted driving behavior, showing higher stability and reliability. Furthermore, through the optimization of algorithms and calculation processes, the system not only greatly reduces the demand for computing resources and the overall system overhead while ensuring excellent performance, making it not only suitable for high-end automotive systems, but also widely integrated into ordinary civilian vehicles, greatly broadening its market applicability and potential user groups. In general, the present invention has created a novel, efficient and economical solution aimed at improving the safety and comfort of the driving process. Its unique technical advantages and broad application potential indicate that it will play a vital role in the field of road traffic safety in the future.
[0168] It should be made clear that the device shown in the aforementioned implementation case only divides and explains each functional module as an example when realizing its function. In actual application scenarios, it can be flexibly adjusted according to actual needs, and the functions can be assigned to different functional modules to complete, that is, the internal structure of the device is divided into different functional modules according to actual needs to realize all or part of the functions described above. In addition, the device in the aforementioned implementation case and the method case implemented by it are based on the same design concept. Please refer to the method implementation case for its specific execution details, which will not be repeated here.
Claims
1. A driver fatigue monitoring method based on eye tracking, characterized in that: The method comprises the following steps: S1: Obtain facial image frame; S2: Collect facial image frames and Mediapipe detection through the front-end camera, record head posture, and identify the iris center area; Through facial information, extract the driver's continuous eye closure time and mouth status information; S3: Using the PNP algorithm; establishing the world coordinate system and the camera coordinate system so that the sight vector and the external environment are located in the same space; S4: Modeling the human eye and the gaze area, and calculating the gaze point coordinates in the three-dimensional space through the geometric model; S5: Conduct real-time attention analysis through fixation points and determine the analysis results; S6: Through the BP neural network, combined with the gaze point analysis results, the driver's attention changes and fatigue status can be effectively identified.
2. The driver fatigue monitoring method based on eye tracking according to claim 1, characterized in that: Step S1 includes: S11: Obtaining real-time video of the driver; S12: Training the BP neural network using the eye feature dataset and the mouth feature dataset S13: Use the trained BP neural network to perform real-time recognition on the collected video data and extract facial feature information; locate the eye area and mouth area; detect the facial area, eye area and mouth area through the BP neural network to obtain a facial image frame.
3. A driver fatigue monitoring method based on eye tracking as claimed in claim 2, characterized in that: Step S13 includes: S13a: After the facial region recognition is completed, the eye region and the mouth region are located, and the eye feature information and the mouth feature information are extracted; the facial image is divided, and the eye and mouth regions are marked; S13b: cutting out the eye area and the mouth area according to the marked image; S13c: Detect the eye area image and the mouth area image respectively through the BP neural network.
4. The driver attention analysis method using eye tracking technology as claimed in claim 1, characterized in that: Step S2 includes: S21: Convert the facial image frame color space and convert the image into RGB mode; S22: Detect key points of the face in the image through the Mediapipe face grid detector; S23: Draw a face mesh through the detected key points and locate the pupil position.
5. The driver attention analysis method using eye tracking technology as claimed in claim 1, characterized in that: Step S3 includes: S31: Obtain the conversion relationship between the coordinate system and the camera coordinate system through the PNP algorithm; S32: establishing a camera coordinate system and a world coordinate system for the eye tracking camera and the external environment camera; S34: By using the obtained transformation relationship, position relationship between cameras, and internal parameters, the sight vector is transformed into the environment perception camera coordinate system through the rotation and translation relationship. Specifically, the PNP function estimates the object pose given a set of object points, their corresponding image projections, and the camera intrinsic parameter matrix and distortion coefficients, expressed in the world frame point X W is projected into the image plane [u,v] using the perspective projection model Π and the camera intrinsic parameter matrix A: The estimated pose is therefore the rotation R and translation T vectors that transform a 3D point represented in the world frame to the camera frame: Equations of the above form can be solved using some algebraic derivation by using a method called Direct Linear Transformation (DLT). The "gold standard" solution to PNP is to estimate the six parameters of the transformation by minimizing the norm of the reprojection error using a nonlinear minimization method. Minimizing this reprojection error provides the maximum likelihood estimate when assuming Gaussian noise in the measurements. The problem can be stated as: Where d is the Euclidean distance between two points. Solving the above equation involves minimizing the cost function E(q) = ke(q)k, which is defined as follows: E(q)=e(q) i e(q) e(q)=x(q)-x After using the PNP algorithm to obtain the transformation relationship between the two coordinate systems, we will then use these transformation relationships to transform between multiple coordinate systems. For the macro camera used for eye tracking, we establish its camera coordinate system and world coordinate system. For the external environment perception camera, we also establish its camera coordinate system and world coordinate system. The transformation relationship between the camera coordinate system and the world coordinate system of each camera can be obtained using the PNP algorithm.
6. The driver attention analysis method using eye tracking technology as claimed in claim 1, characterized in that: Step S4 includes: S41: After coordinate conversion, determine the unit vector of the optical axis in the world coordinate system; S42: The optical axis is defined as a three-dimensional vector connecting the center of the pupil and the center of the cornea. The visual axis is a three-dimensional vector connecting the center of the cornea and the fovea. The angle between the two is called the kappa angle. The positional relationship between the visual axis and the optical axis is modeled through the kappa angle. The specific formula is as follows: Define a simple visual axis-optical axis angle model, the unit vector of the optical axis is v o =[0 0 1] T , the angle between the visual axis and the optical axis is expressed as κ = (α, β), and the visual axis vectors of the left and right eyes can be expressed as: Decomposing the optical axis into horizontal and vertical angles is expressed as but: Calculate the intersection of the left and right eye visual axis vectors in space to estimate the subject's current gaze point. Define the three-dimensional coordinates of the gaze point as g(o,κ)=(g x ,g y ,g z ), the three-dimensional coordinate of the eyeball is o=(o x ,o y ,o z ), the fixation point g(o,k) is determined by the following formula: Different subjects have different kappa angles. During the personal calibration process, the subjects are asked to look at N specific calibration points on the screen: g i , i = 1, 2, ..., N, and then optimize the model parameters by minimizing the distance between the estimated predicted fixation point and the actual fixation point, where g i are the actual coordinates of the calibration points, is the subject's current gaze point.
7. The driver attention analysis method using eye tracking technology as claimed in claim 1, characterized in that: Step S5 includes: S51: integrating the driver's gaze point information and external environment perception information; S52: Analyzing the driver's gaze area in the external environment by fusing the information; S53: Determine the real-time attention status by analyzing the results.
8. The driver attention analysis method using eye tracking technology as claimed in claim 1, characterized in that: Step S6 includes: S61: Recognize the driver's mouth, eyes, and head through the BP neural network; S62: (The average time that a person's mouth is wide open when yawning is at least 5 seconds.) The camera monitoring frequency is 25 frames per second. If the number of consecutive occurrences of the driver's mouth open in the time state series of the mouth state in 150 frames exceeds 125 times, it is determined that the driver is in a yawning state; The PERCLOS (percentage of duration of eyelid closure) method uses the P80 standard to calculate the proportion of time the eyes are closed within one minute; (write the Chinese name and explanation of the special terms) S63: Analyze the current driver's attention status and driving fatigue status by fusing the driver's gaze area, the external driving environment image and the BP neural network output result; S64: Set extreme thresholds, track the driver's gaze based on the camera, and determine the driver's eye movement direction and gaze area based on the algorithm. If the driver's gaze direction continues to deviate from the front within a preset time, if the driver's gaze deviation time exceeds the preset value, the driver's gaze is still in a deviated state; if the driver yawns multiple times and the driver's eye features meet the preset standards in the PERCLOS method, the system will determine that it is fatigue driving and take early warning measures.
Citation Information
Patent Citations
Method and system for determining eyeball fixation point
CN115963931A
Distraction detection method, vehicle-mounted controller and computer storage medium
CN116052136A
Driving state monitoring method and device, electronic equipment and storage medium
CN118876988A
Systems and methods for anatomy-constrained gaze estimation
US20220050521A1
Cited By
User high-risk behavior identification method and system based on machine learning
CN120327531A
Intelligent zoom liquid crystal glasses system based on time eyeball tracking and positioning
CN122110542A
Driver staring target recognition method combining sight line estimation and scene understanding
CN122253907A
Driver gaze target recognition method combining line-of-sight estimation and scene understanding
CN122253907B
Eye movement tracking method and device based on slit lamp
CN122313555A