A visual-based fatigue driving detection method and device
By employing a fatigue detection method based on multi-dimensional facial feature extraction and environmental adaptive weight adjustment, the problems of illumination adaptability and individual differences in existing visual detection systems are solved, achieving highly accurate and safe driver fatigue monitoring and early warning.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 浪潮智慧科技有限公司
- Filing Date
- 2026-03-31
- Publication Date
- 2026-07-07
AI Technical Summary
Existing visual detection systems suffer from poor light adaptability, insufficient multi-feature fusion, weak handling of individual differences, and high misjudgment rate in driver fatigue detection. They are unable to adapt to the ever-changing driving environment, resulting in insufficient detection accuracy and safety.
By acquiring driver facial images and extracting multi-dimensional features such as eye, mouth, and head posture, and dynamically adjusting the weights in conjunction with vehicle environmental information, fatigue index prediction and fusion are performed to achieve multi-dimensional, adaptive real-time monitoring and graded early warning.
It achieves highly accurate, adaptive, and personalized real-time monitoring and early warning of driver fatigue, significantly improving driving safety and reliability, reducing misjudgments, and ensuring the safety of drivers and passengers.
Smart Images

Figure CN122347792A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of automotive technology, and in particular to a vision-based method and device for detecting driver fatigue. Background Technology
[0002] With the continuous increase in the number of motor vehicles, fatigued driving has become one of the major causes of road traffic accidents. Fatigue driving leads to decreased driver attention, slow reaction time, and impaired judgment; in severe cases, it can even cause drivers to fall asleep briefly, greatly increasing the risk of traffic accidents and threatening the lives and property of drivers and others.
[0003] In related technologies, methods for driver fatigue detection using visual detection systems suffer from problems such as poor light adaptability, insufficient multi-feature fusion, weak handling of individual differences, and high misjudgment rate, making it difficult to adapt to the ever-changing driving environment. Summary of the Invention
[0004] This application provides a vision-based fatigue driving detection method and device to solve the following technical problem: how to achieve multi-dimensional, adaptive, and highly accurate real-time monitoring and graded early warning of driver fatigue state, thereby improving driving safety.
[0005] In a first aspect, embodiments of this application provide a vision-based fatigue driving detection method, the method comprising: Acquire the driver's facial image at the first moment, and extract features from the facial image to obtain facial features in at least one dimension; Fatigue index prediction is performed on the facial features of each dimension to obtain the fatigue prediction index for each dimension. Based on the vehicle environment information at the first moment, determine the first weight corresponding to each dimension; Based on the first weight, the fatigue prediction index of each dimension is fused to obtain the comprehensive fatigue index of the driver at the first moment, and the fatigue level of the driver at the first moment is determined based on the comprehensive fatigue index of the driver at the first moment.
[0006] Secondly, embodiments of this application also provide a vision-based fatigue driving detection device, the device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a vision-based fatigue driving detection method as described above.
[0007] Thirdly, embodiments of this application also provide a computer storage medium storing computer-executable instructions, which, when executed, implement a vision-based fatigue driving detection method as described above.
[0008] The vision-based fatigue driving detection method and device provided in this application have the following beneficial effects: First, by acquiring the driver's facial image at the first moment and extracting facial features in at least one dimension, the system can comprehensively capture the driver's facial state information at the current moment, providing a rich and direct data foundation for subsequent fatigue analysis. Next, fatigue index prediction is performed on the facial features in each dimension, resulting in a corresponding fatigue prediction index. This transforms various facial features into quantifiable and comparable fatigue indicators, enabling accurate assessment of different facial behaviors. Then, based on the vehicle environment information at the first moment, the primary weight for each dimension is determined. This allows for dynamic adjustment of the importance of each dimension in fatigue judgment based on real-time environmental factors, making the system more closely aligned with actual driving scenarios and effectively reducing misjudgments caused by environmental interference. Finally, the fatigue prediction indices of each dimension are fused according to the primary weight to obtain a comprehensive fatigue index, which is then used to determine the driver's fatigue level. This comprehensive assessment of the driver's overall fatigue state, based on multi-dimensional information, enables highly accurate, adaptive, and personalized real-time monitoring and reasonable early warning of driver fatigue status at different times and under different driving environments, significantly improving the reliability, safety, and practicality of fatigue driving detection. Attached Figure Description
[0009] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart of a vision-based fatigue driving detection method provided in this application embodiment; Figure 2 This is a schematic diagram of the internal structure of a vision-based fatigue driving detection device provided in an embodiment of this application. Detailed Implementation
[0010] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0011] It is understood that in the embodiments of this application, data related to user information (such as user accounts) is involved. When the embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with relevant laws, regulations and standards.
[0012] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0013] In the following description, the terms “first, second, ...” are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that “first, second, ...” may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0014] With the continuous increase in the number of motor vehicles, fatigued driving has become one of the major causes of road traffic accidents. Fatigue driving leads to decreased driver attention, slow reaction time, and impaired judgment; in severe cases, it can even cause drivers to fall asleep briefly, greatly increasing the risk of traffic accidents and threatening the lives and property of drivers and others.
[0015] In related technologies, methods for driver fatigue detection using visual detection systems suffer from problems such as poor light adaptability, insufficient multi-feature fusion, weak handling of individual differences, and high misjudgment rate, making it difficult to adapt to the ever-changing driving environment.
[0016] Based on this, the embodiments of this application provide a vision-based fatigue driving detection method, which can realize multi-dimensional, adaptive, and highly accurate real-time monitoring and graded early warning of driver fatigue state, thereby effectively improving driving safety.
[0017] The technical solutions proposed in the embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0018] Figure 1This document provides a flowchart of a vision-based fatigue driving detection method for embodiments of this application. This method can be applied to various fatigue driving detection scenarios. For example, in passenger vehicle (private car, sedan, etc.) fatigue driving monitoring scenarios, an onboard camera monitors the driver's facial state in real time, detecting fatigue in daily commutes and long-distance driving scenarios, providing timely warnings, and ensuring the safety of the driver and passengers. In commercial vehicle driver status monitoring scenarios, a driving assistance system monitors the driver's fatigue state in real time to prevent traffic accidents caused by fatigue driving, improving transportation safety and efficiency. In ride-hailing and taxi driver status detection scenarios, detecting driver fatigue ensures passenger travel safety and also helps the platform manage and control driver behavior and risks.
[0019] This application provides a vision-based fatigue driving detection method. It should be noted that the execution entity in these embodiments can be a server or any terminal device with data processing capabilities. For example, the server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal device can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, in-vehicle terminal, etc., but is not limited to these.
[0020] like Figure 1 As shown in the figure, the vision-based fatigue driving detection method provided in this application embodiment specifically includes the following steps: Step 101: Obtain the driver's facial image at the first moment.
[0021] It should be noted that the first moment refers to the time point at which the system acquires the driver's facial image and begins fatigue detection at a specific moment. The driver's facial image can be acquired by an in-vehicle camera installed inside the vehicle, a dedicated camera in the driver monitoring system, or by means of infrared, mobile devices, etc. The driver's facial image acquired at the first moment can be one or multiple, and no specific limitation is made here.
[0022] As an example, in a scenario where a driver is traveling long distances in a private car, once the driver starts the vehicle and activates the fatigue driving monitoring function, the system immediately enters working mode. The system uses an infrared camera inside the vehicle to capture the driver's facial image at the current moment (the first moment).
[0023] Step 102: Extract features from the facial image to obtain facial features in at least one dimension.
[0024] It should be noted that the feature extraction mentioned above may include image preprocessing (grayscale conversion, denoising, image enhancement, etc.), face detection (face localization, face pose and angle detection, etc.), facial landmark localization, and feature extraction in various dimensions (eye dimension, mouth dimension, head pose dimension, etc.).
[0025] As an example, after obtaining the driver's facial image, the acquired color facial image of the driver is converted to grayscale to obtain a grayscale image of the driver. Then, denoising methods such as Gaussian filtering are used to remove noise interference in the image. In poor lighting conditions, enhancement operations such as brightness and contrast adjustment can be performed on the image. Subsequently, a face detection algorithm (such as the YOLO algorithm) is used to accurately locate the position and range of the face in the facial image, segment the face from the background, and preliminarily determine the face's posture, such as whether it is a frontal view or whether it is tilted. Then, a facial landmark detection algorithm is used to accurately find the key points of various key facial features (such as eyes, mouth, nose, facial contours, and other important facial feature areas) from the detected face region, and the coordinates of each key point are determined. Finally, features are extracted from at least one dimension (such as eyes, mouth, head posture, etc.) to obtain facial features of at least one dimension.
[0026] In some embodiments, the dimensions include eye dimension, mouth dimension, and head pose dimension; step 102 above can be implemented in the following ways: based on the eye key points of the facial image, the aspect ratio of the eyes is calculated to obtain the facial features of the eye dimension; based on the mouth key points of the facial image, the aspect ratio of the mouth is calculated to obtain the facial features of the mouth dimension; based on the facial key points of the facial image, head pose is detected to obtain the facial features of the head pose dimension.
[0027] Thus, by calculating the eye aspect ratio based on eye key points in facial images to obtain eye-dimensional facial features, calculating the mouth aspect ratio based on mouth key points to obtain mouth-dimensional facial features, and detecting head posture based on facial key points to obtain head posture-dimensional facial features, the system can comprehensively and accurately capture the driver's facial state information from multiple key dimensions. These multi-dimensional features complement each other and are comprehensively considered, effectively avoiding the limitations and misjudgment risks of single-feature detection. Furthermore, the calculation methods for eye aspect ratio, mouth aspect ratio, and head posture detection are highly objective and quantifiable, transforming the driver's fatigue state into specific feature data, providing a reliable basis for accurate prediction of the fatigue index. Moreover, analysis based on facial key points does not rely on the driver wearing additional equipment, offering non-contact, convenient, and efficient features applicable to various real-world driving scenarios. The extraction of facial features from different dimensions lays the foundation for building a more accurate and intelligent fatigue driving detection system, helping to promptly detect driver fatigue and provide early warnings, thereby significantly improving driving safety and reducing traffic accidents caused by fatigue driving.
[0028] It should be noted that the eye aspect ratio is a value calculated based on key points of the eyes, which can quantify the degree of eye opening; the mouth aspect ratio is a value calculated through analysis of key points of the mouth, used to measure the degree of mouth opening; head posture refers to the position and orientation of the head in space, used to describe tilting, turning, etc., relative to a normal upright or standard posture, mainly including pitch angle, yaw angle, and roll angle. Pitch angle is used to reflect the degree of head tilt in the forward and backward direction, yaw angle is used to indicate the head turning in the left and right direction, and roll angle is used to reflect the head's state in the left and right tilt direction.
[0029] As an example, in a long-haul freight scenario, after acquiring the driver's facial image, the system locates key points around the driver's eyes through facial landmark detection, and calculates the aspect ratio of the eyes based on the key points at the eye position, thereby obtaining facial features in the eye dimension. Simultaneously, regarding the mouth dimension, the system calculates the mouth's aspect ratio using key points at the mouth's location, thereby obtaining facial features in the mouth dimension. Simultaneously, regarding the head pose dimension, the system constructs a 3D head pose model based on facial key points, calculating the head's pitch, yaw, and roll angles to obtain facial features in the head pose dimension. .
[0030] In some embodiments, the above-mentioned calculation of the eye aspect ratio of the facial image based on the eye keypoints of the facial image to obtain the facial features of the eye dimension can be achieved in the following way: eye keypoint detection is performed on the facial image to obtain the eye keypoints of the facial image; the distance between the eye keypoint located at the upper edge of the eye and the eyeglass keypoint located at the lower edge of the eye in the facial image is taken as the vertical distance of the eye; the distance between the eye keypoint located at the outer corner of the eye and the eyeglass keypoint located at the inner corner of the eye in the facial image is taken as the horizontal distance of the eye; the ratio of the horizontal distance of the eye to the vertical distance of the eye is taken as the facial features of the eye dimension.
[0031] Thus, by detecting key points around the eyes and calculating the ratio of horizontal to vertical distance between the eyes as facial features in the eye dimension, the overall opening and closing state of the driver's eyes can be accurately quantified. The ratio of horizontal to vertical distance can intuitively reflect the degree of eye opening and closing; a large ratio indicates a high degree of eye opening, while a small ratio means the eyes tend to be closed. This quantification method allows for more objective and easier analysis of facial features in the eye dimension. Furthermore, by focusing on key areas of the eyes, subtle changes in the eyes in the early stages of driver fatigue can be accurately captured, allowing for timely detection of signs of fatigue. At the same time, the feature construction based on simple geometric distance relationships has a clear calculation logic and relatively small computational load, making it easier to process and analyze in real time. In addition, facial features in the eye dimension can complement other facial features detected subsequently (such as the mouth, head posture, etc.), comprehensively judging the driver's fatigue state from multiple dimensions, thereby improving the accuracy and reliability of fatigue detection and effectively ensuring driving safety.
[0032] As an example, in a nighttime highway driving scenario, when a truck driver starts the vehicle and activates the fatigue monitoring system, the infrared camera located above the rearview mirror captures the driver's facial image. The system then uses deep learning algorithms (such as the YOLO algorithm) to detect multiple key points in the driver's eye area, accurately locating the key points at the upper edge of the eyes (e.g., the point near the eyebrow on the upper eyelid), the lower edge of the eyes (e.g., the point near the cheek on the lower eyelid), the outer corner of the eye, and the inner corner of the eye. Subsequently, the system detects and measures the straight-line distance between the key points at the upper and lower edges of the eyes. (Vertical distance to the eye), assuming the vertical distance to the eye is calculated using keypoint coordinates. The system measures the straight-line distance between key points at the outer and inner corners of the eye, using 8-pixel units. (Horizontal distance of the eye), assuming the horizontal distance of the eye is calculated using the coordinates of key points. The unit is 12 pixels; then, the horizontal distance of the eye is calculated. vertical distance from the eyes The ratio of these ratios can be used to obtain the facial features in the eye dimension at the current moment. for .
[0033] In some embodiments, the above-mentioned calculation of the mouth aspect ratio of the facial image based on the mouth key points of the facial image to obtain the facial features of the mouth dimension can be achieved in the following way: detecting mouth key points in the facial image to obtain the mouth key points of the facial image; taking the distance between the mouth key points located on the upper lip and the mouth key points located on the lower lip in the facial image as the vertical distance of the mouth; taking the distance between the mouth key points located on the left side and the mouth key points located on the right side of the mouth in the facial image as the horizontal distance of the mouth; and taking the ratio of the horizontal distance of the mouth to the vertical distance of the mouth as the facial features of the mouth dimension.
[0034] Thus, by detecting key points of the mouth and calculating the ratio of horizontal to vertical distances as facial features in the mouth dimension, the overall degree of mouth opening can be accurately quantified. The ratio of horizontal to vertical distances can intuitively reflect the dimensional relationship of the mouth in the horizontal and vertical directions. This quantification method allows for more objective and easier analysis of facial features in the mouth dimension. Furthermore, by focusing on key areas of the mouth, changes in the mouth state in the early stages of driver fatigue can be accurately captured, allowing for timely detection of fatigue signs. At the same time, features are constructed based on simple geometric distance relationships, with clear calculation logic and relatively low computational load, making it easier to process and analyze in real time. In addition, facial features in the mouth dimension can complement other facial features (such as eyes, head posture, etc.) to comprehensively judge the driver's fatigue state from multiple dimensions, thereby improving the accuracy and reliability of fatigue detection and effectively ensuring driving safety.
[0035] As an example, in the scenario of driving an intercity long-distance bus, when the bus driver starts the vehicle and activates the fatigue monitoring system, after the camera located on the roof of the vehicle captures an image of the driver's face, the system uses a facial landmark detection algorithm (such as the YOLO algorithm) to detect multiple key points in the driver's mouth area, accurately locating the key points on the upper edge of the lips, the lower edge of the lips, and the left and right sides of the mouth; subsequently, the system detects and measures the straight-line distance between the key points on the upper lip and the key points on the lower lip. (Vertical distance to the mouth), assuming the vertical distance to the mouth is calculated using key point coordinates. The system measures the linear distance between key points on the left and right sides of the mouth, using units of 10 pixels. (Horizontal distance of the mouth), assuming the horizontal distance of the mouth is calculated using key point coordinates. The unit is 14 pixels; then, the horizontal distance of the mouth is calculated. Vertical distance from the mouth The ratio of the two values can be used to obtain the facial features (MAR) of the mouth at the current moment. .
[0036] In some embodiments, the above-mentioned method of performing head pose detection on the facial image based on facial key points to obtain facial features in the head pose dimension can be implemented in the following way: performing facial key point detection on the facial image to obtain facial key points; constructing a three-dimensional head pose model of the driver based on the relative positions of the facial key points; calculating the pitch angle, yaw angle, and roll angle of the driver's head based on the three-dimensional head pose model, and using the calculation results as facial features in the head pose dimension.
[0037] Thus, by detecting facial key points to construct a 3D head posture model of the driver and calculating pitch, yaw, and roll angles as facial features of the head posture dimension, the spatial position and orientation changes of the driver's head can be accurately captured as a whole. The 3D head posture model constructed based on the relative positions of key points can reflect the head's posture in 3D space. The calculated pitch angle (reflecting the head's forward and backward tilt), yaw angle (reflecting the head's left and right rotation), and roll angle (reflecting the head's left and right tilt) can quantify the driver's fatigue-related head movements. By accurately identifying the driver's fatigue state through angle changes, and compared to relying solely on 2D images, the quantitative analysis based on spatial geometry can be more objective and accurate. By comprehensively evaluating head posture from multiple angles, the driver's attention state and concentration can be fully reflected, making it suitable for real-time fatigue monitoring in in-vehicle environments. It can complement other facial features (such as eye and mouth features), thereby significantly improving the accuracy and reliability of fatigue detection and effectively preventing driving safety hazards caused by abnormal head posture.
[0038] As an example, in the scenario of driving a city bus, when the bus driver starts the bus during the morning rush hour and turns on the fatigue monitoring system, after the infrared camera next to the rearview mirror captures the driver's facial image, the system accurately identifies the positions of multiple key points on the driver's face, such as the corners of the eyes, the tip of the nose, the corners of the mouth, and the facial contours, using a high-precision facial key point detection algorithm. Then, the system uses the coordinates of the key points in the two-dimensional image and their known facial geometric constraints to construct a three-dimensional posture model of the driver's head to restore the true orientation of the head in real space. Subsequently, the system uses the three-dimensional posture model to calculate the three core rotation angles of the head using formulas (1)-(3), which gives the facial features of the head posture dimension at the current moment.
[0039] (1) (2) (3) in, The pitch angle, Here are the vertical coordinates of the key points at the tip of the nose. The vertical coordinates of the key points at the center of both eyes. This is the depth distance from the tip of the nose to the eye plane estimated based on the model. Yaw angle The horizontal coordinates of the key point of the nose tip. The horizontal coordinates of the key points at the center of the face. For roll angle, , These are the vertical coordinates of the left and right eyes, respectively. , These are the horizontal coordinates for the left and right eyes, respectively.
[0040] Step 103: Perform fatigue index prediction on the facial features of each dimension to obtain the fatigue prediction index for each dimension.
[0041] It should be noted that fatigue index prediction can be achieved through threshold comparison and normalized linear mapping, or through time series state duration analysis, or through frequency domain analysis of statistical regularities. No specific limitations are made here.
[0042] In some embodiments, step 103 described above can be implemented as follows: for each dimension of facial features, perform the following processing respectively: perform state analysis on the facial features of the dimension to obtain the facial state corresponding to the dimension; determine the first difference between the facial state corresponding to the dimension and the reference state corresponding to the dimension, and map the first difference to obtain the fatigue prediction index of the dimension.
[0043] Thus, by performing state analysis on facial features, continuous numerical features are transformed into discrete facial states with clear physical meaning, making the features more semantically readable and easier for the system to understand and make subsequent logical judgments. Furthermore, by determining the first difference between the current facial state and the preset reference state, the degree of deviation of the driver's facial behavior from the normal state can be accurately quantified, thereby directly reflecting the magnitude of behavioral changes caused by fatigue. In addition, by mapping the first difference to obtain a fatigue prediction index, individual differences and instantaneous interference can be smoothly handled, and the state difference can be transformed into a standardized quantitative indicator, making the fatigue level between different dimensions comparable. This provides more robust and reliable input data for subsequent multi-dimensional fusion, thereby significantly enhancing the system's robustness and adaptability in fatigue detection under complex driving environments.
[0044] As an example, in a nighttime highway driving scenario, the system performs state analysis on the currently collected eye features (eye aspect ratio EAR). Assuming the current EAR value is 0.18, the system compares it with preset state thresholds such as "open eyes," "half-closed eyes," and "completely closed eyes," determining that the driver's current eye state is "half-closed eyes." Next, the system determines the first difference between this state and the reference state corresponding to the dimension (i.e., the baseline state of the driver when awake, "open eyes," with a corresponding EAR reference value of 0.35). By calculating the relative change, the difference value is obtained (e.g., formula: first difference = (reference state value - current state value) / reference state value). The first difference is 0.486, meaning that the current eye opening / closing degree has decreased by 48.6% compared to the awake state. Finally, the system converts the first difference into a fatigue prediction index through a preset mapping function (e.g., a linear function or a sigmoid function), thus obtaining the fatigue prediction index for the eye dimension. For example, if the mapping rule is to directly amplify the first difference by 2 and limit it to the [0,1] interval, the fatigue prediction index for the eye dimension is 0.972.
[0045] Step 104: Based on the vehicle environment information at the first moment, determine the first weight corresponding to each dimension.
[0046] It should be noted that vehicle environmental information may include one or more of the following: lighting environment information (such as the intensity of light inside the vehicle, whether it is nighttime, whether it is in a tunnel, etc.), driving status information (such as current vehicle speed, acceleration, etc.), road and weather information (such as road type, current weather, etc.), driver status information (such as whether the driver is wearing sunglasses, a mask, etc.), and equipment operating status information (such as camera working status, image jitter level, etc.). No specific limitation is made here.
[0047] In some embodiments, step 104 described above can be implemented as follows: acquiring vehicle environment information at the first moment and determining a second weight corresponding to each preset dimension; performing an influence degree analysis on each dimension based on the vehicle environment information to determine the weight adjustment amount corresponding to each dimension; adjusting the second weight corresponding to each dimension based on the weight adjustment amount corresponding to the dimension to obtain the second weight after dimension adjustment; and normalizing the second weight after dimension adjustment to obtain the first weight corresponding to each dimension.
[0048] Thus, by establishing the baseline importance (second weight) of each dimension, and then deeply analyzing the specific impact of factors such as illumination, occlusion, and vehicle speed on the detection reliability of each dimension based on vehicle environmental information, the required weight adjustment amount for each dimension can be accurately determined, preserving expert experience while possessing environmental adaptability. Secondly, by normalizing the adjusted weights, the mathematical rationality and comparability of the weights of each dimension in the fusion calculation can be ensured, so that the final first weight can truly reflect the actual contribution of each dimension to fatigue judgment at the current specific moment, thereby effectively avoiding misjudgment or omission due to environmental interference, and providing a scientific, dynamic, and reliable weight basis for subsequent multi-dimensional fusion calculation of the comprehensive fatigue index.
[0049] As an example, in a long-distance highway driving scenario during a rainstorm at night, the system acquires the vehicle's environmental information at the first moment, including data such as "nighttime + heavy rain (low visibility)," "driver wearing sunglasses," and "vehicle speed 100km / h." The system's preset "secondary weights" (i.e., baseline weights) are: eye dimension 0.4, mouth dimension 0.3, and head posture dimension 0.3. Subsequently, based on the influence analysis of the environmental information, it is found that "nighttime + sunglasses" makes eye feature extraction extremely unreliable, determining the weight adjustment for the eye dimension to be -0.2. Furthermore, head posture during "high-speed driving" is more sensitive to fatigue. The weight adjustment for the head pose dimension is +0.1, while the mouth dimension is less affected and the adjustment is 0. Next, the system calculates the weight adjustment for each dimension, resulting in the following weights: eyes = 0.4 - 0.2 = 0.2, mouth = 0.3 + 0 = 0.3, and head = 0.3 + 0.1 = 0.4. Finally, the system normalizes all adjusted weights, i.e., normalized weight = adjusted weight of each dimension / sum of adjusted weights of all dimensions. Therefore, the first weight for the eyes dimension is 0.22, the first weight for the mouth dimension is 0.33, and the first weight for the head dimension is 0.44.
[0050] Step 105: Based on the first weight, the fatigue prediction index of each dimension is fused to obtain the driver's comprehensive fatigue index at the first moment.
[0051] As an example, suppose the first weight of the eye dimension is 0.2, and the corresponding fatigue prediction index is 0.8; the first weight of the mouth dimension is 0.3, and the corresponding fatigue prediction index is 0.6; and the first weight of the head dimension is 0.5, and the corresponding fatigue prediction index is 0.7. Then, based on the first weight, a weighted fusion is performed, and the comprehensive fatigue index can be obtained by weighted summation. That is, the total fatigue index of the driver at the first moment is 0.69.
[0052] Step 106: Determine the driver's fatigue level based on the driver's comprehensive fatigue index at the first moment.
[0053] In some embodiments, step 106 described above can be implemented as follows: in response to the comprehensive fatigue index being less than a first index threshold, the driver's fatigue level at the first moment is determined to be normal; in response to the comprehensive fatigue index being greater than or equal to the first index threshold and less than a second index threshold, the driver's fatigue level at the first moment is determined to be mild fatigue; in response to the comprehensive fatigue index being greater than or equal to the second index threshold and less than a third index threshold, the driver's fatigue level at the first moment is determined to be moderate fatigue; in response to the comprehensive fatigue index being greater than or equal to the third index threshold, the driver's fatigue level at the first moment is determined to be severe fatigue.
[0054] In this way, by adopting a hierarchical logic, different stages of fatigue development can be accurately distinguished, providing a clear and objective basis for the system to adopt differentiated and progressive early warning strategies. Furthermore, the logic of determining the level based on the interval division is clear and the rules are well-defined. This not only facilitates the system to make real-time and rapid judgments, but also allows for matching different intensities of reminders to different fatigue levels. This ensures driving safety while minimizing interference with normal driving, thereby significantly improving the practicality, reliability, and user experience of the fatigue driving monitoring system.
[0055] As an example, in a long-distance highway driving scenario, the system performs multi-dimensional feature fusion calculations and assumes that the driver's overall fatigue index at a certain moment is 0.75. The system's preset fatigue level judgment thresholds are: a first index threshold of 0.3 (distinguishing between normal and mild fatigue), a second index threshold of 0.6 (distinguishing between mild and moderate fatigue), and a third index threshold of 0.8 (distinguishing between moderate and severe fatigue). Subsequently, the system compares the overall fatigue index of 0.75 with each threshold and finds that 0.75 is greater than or equal to the second index threshold of 0.6, but less than the third index threshold of 0.8, thus satisfying the condition "responding to the overall fatigue index of 0.75". The system determines that the driver's fatigue level at this first moment is moderate fatigue, based on the condition that "the comprehensive fatigue index is greater than or equal to the second index threshold and less than the third index threshold". Similarly, if the comprehensive fatigue index of the driver at a certain moment is 0.25, then according to the above determination method, the driver's fatigue level at the current moment can be determined to be normal. If the comprehensive fatigue index of the driver at a certain moment is 0.4, then the current driver's fatigue level can be determined to be mild fatigue. If the comprehensive fatigue index of the driver at a certain moment is 0.85, then the current driver's fatigue level can be determined to be severe fatigue.
[0056] In some embodiments, after step 106, the following processing may also be performed: in response to the driver's fatigue level being mild fatigue at the first moment, a level one warning reminder is given to the driver, the level one warning reminder including a voice prompt and a flashing warning light on the instrument panel; in response to the driver's fatigue level being moderate fatigue at the first moment, a level two warning reminder is given to the driver, the level two warning reminder including a phoenix alarm and a fatigue driving warning pop-up window on the central control screen; in response to the driver's fatigue level being severe fatigue at the first moment, a level three warning reminder is given to the driver, the level three warning reminder including a seat vibration reminder and a high-decibel voice warning.
[0057] In this way, by matching differentiated alarm methods to different fatigue levels, mild fatigue uses low-interference methods such as voice prompts and dashboard flashing, moderate fatigue is upgraded to a phoenix alarm and screen pop-up, and severe fatigue is activated by strong stimulation methods such as seat vibration and high-decibel voice. This can avoid excessively disturbing the awake driver, and can also forcibly wake the driver's attention with high-intensity stimulation when fatigue worsens. Furthermore, by accurately matching the alarm intensity with the degree of fatigue, the system can take the most appropriate intervention measures at different times and under different fatigue states, so as to significantly improve the effectiveness of fatigue driving warning and user experience, thereby minimizing the risk of traffic accidents caused by fatigue driving.
[0058] As an example, in a cross-provincial long-haul freight scenario, the system determines that the driver is currently at the "moderate fatigue" level based on the comprehensive fatigue index. At this point, the system immediately triggers a level two warning. First, the in-vehicle speakers emit a rapid "phoenix-like" alarm sound, which is higher in pitch and has penetrating power than ordinary warning sounds. At the same time, a prominent "Fatigue Driving Warning" pop-up window appears on the central control screen, displaying the text prompt "You are in a state of moderate fatigue. Please drive to a service area to rest as soon as possible." This dual warning of sound and sight draws the driver's attention. If the driver ignores the warning and the fatigue continues to worsen, when the system determines that the driver has entered the "severe fatigue" level, it immediately upgrades to a level three warning. The vibration motors at the bottom and back of the driver's seat instantly start, generating strong vibrations. At the same time, the in-vehicle audio switches to a high-volume, stern voice command, "Danger! Please pull over immediately!" This high-intensity tactile and auditory stimulation forcibly wakes the driver, compelling them to take emergency evasive action, thereby maximizing driving safety.
[0059] The above are embodiments of the method proposed in this application. Based on the same inventive concept, embodiments of this application also provide a vision-based fatigue driving detection device, the structure of which is as follows: Figure 2 As shown.
[0060] Figure 2 This is a schematic diagram of the internal structure of a vision-based fatigue driving detection device provided in an embodiment of this application. Figure 2 As shown, the device includes: At least one processor 201; And a memory 202 that is communicatively connected to at least one processor; The memory 202 stores instructions that can be executed by at least one processor. The instructions are executed by at least one processor 201 to enable at least one processor 201 to perform the steps of the method corresponding to any of the above embodiments.
[0061] Some embodiments of this application provide corresponding to Figure 1 A non-volatile computer storage medium stores computer-executable instructions configured to perform the steps of the method corresponding to any of the above embodiments.
[0062] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments for IoT devices and media are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0063] The systems, media, and methods provided in this application are one-to-one correspondences. Therefore, the systems and media also have similar beneficial technical effects as their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the systems and media will not be repeated here.
[0064] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0065] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0066] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0067] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0068] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0069] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0070] Computer-readable media include both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0071] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0072] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A vision-based method for detecting driver fatigue, characterized in that, The method includes: Acquire the driver's facial image at the first moment, and extract features from the facial image to obtain facial features in at least one dimension; Fatigue index prediction is performed on the facial features of each dimension to obtain the fatigue prediction index for each dimension. Based on the vehicle environment information at the first moment, determine the first weight corresponding to each dimension; Based on the first weight, the fatigue prediction index of each dimension is fused to obtain the comprehensive fatigue index of the driver at the first moment, and the fatigue level of the driver at the first moment is determined based on the comprehensive fatigue index of the driver at the first moment.
2. The method according to claim 1, characterized in that, The dimensions include eye dimension, mouth dimension, and head posture dimension; The step of extracting features from the facial image to obtain facial features in at least one dimension includes: Based on the eye key points of the facial image, the aspect ratio of the eyes is calculated to obtain the facial features of the eye dimension; Based on the key points of the mouth in the facial image, the aspect ratio of the mouth is calculated to obtain the facial features of the mouth dimension. Based on the facial key points of the facial image, head pose detection is performed on the facial image to obtain facial features in the head pose dimension.
3. The method according to claim 2, characterized in that, The step of calculating the eye aspect ratio of the facial image based on the eye key points to obtain the facial features in the eye dimension includes: Eye key points are detected in the facial image to obtain the eye key points of the facial image; The distance between the eye key point located at the upper edge of the eye and the glasses key point located at the lower edge of the eye in the facial image is taken as the vertical distance of the eye; The distance between the key eye point located at the outer corner of the eye and the key eye point located at the inner corner of the eye in the facial image is taken as the horizontal distance of the eye; The ratio of the horizontal distance to the vertical distance of the eyes is used as the facial feature of the eye dimension.
4. The method according to claim 2, characterized in that, The step of calculating the aspect ratio of the mouth based on the key points of the mouth in the facial image to obtain the facial features of the mouth dimension includes: The mouth key points of the facial image are detected by performing mouth key point detection to obtain the mouth key points of the facial image; The distance between the key points of the mouth located on the upper lip and the key points of the mouth located on the lower lip in the facial image is taken as the vertical distance of the mouth. The distance between the key point of the mouth on the left side and the key point of the mouth on the right side in the facial image is taken as the horizontal distance of the mouth. The ratio of the horizontal distance of the mouth to the vertical distance of the mouth is used as the facial feature of the mouth dimension.
5. The method according to claim 2, characterized in that, The process of performing head pose detection on the facial image based on facial key points to obtain facial features in the head pose dimension includes: Facial landmark detection is performed on the facial image to obtain the facial landmarks of the facial image; Based on the relative positions of the facial key points, a three-dimensional head pose model of the driver is constructed. Based on the three-dimensional head posture model, the pitch angle, yaw angle and roll angle of the driver's head are calculated, and the calculation results are used as facial features of the head posture dimension.
6. The method according to claim 1, characterized in that, The step of predicting the fatigue index for each dimension of facial features to obtain the fatigue prediction index for each dimension includes: For each of the facial features in the aforementioned dimensions, the following processing is performed: Perform state analysis on the facial features of the stated dimension to obtain the facial state corresponding to the stated dimension; A first difference is determined between the facial state corresponding to the dimension and the reference state corresponding to the dimension, and the first difference is mapped to obtain the fatigue prediction index of the dimension.
7. The method according to claim 1, characterized in that, The determination of the first weight corresponding to each dimension based on the vehicle environment information at the first moment includes: Obtain the vehicle environment information at the first moment and determine the second weight corresponding to each preset dimension; Based on the vehicle environment information, an impact degree analysis is performed on each dimension to determine the weight adjustment amount corresponding to each dimension; For each dimension, the second weight corresponding to the dimension is adjusted based on the weight adjustment amount corresponding to the dimension, to obtain the second weight after dimension adjustment; The adjusted second weights for each dimension are normalized to obtain the first weights corresponding to each dimension.
8. The method according to claim 1, characterized in that, Determining the driver's fatigue level based on the driver's comprehensive fatigue index at the first moment includes: In response to the comprehensive fatigue index being less than the first index threshold, the driver's fatigue level at the first moment is determined to be normal; In response to the comprehensive fatigue index being greater than or equal to the first index threshold and less than the second index threshold, the driver's fatigue level at the first moment is determined to be mild fatigue. In response to the comprehensive fatigue index being greater than or equal to the second index threshold and less than the third index threshold, the driver's fatigue level at the first moment is determined to be moderate fatigue. In response to the comprehensive fatigue index being greater than or equal to the third index threshold, the driver's fatigue level at the first moment is determined to be severe fatigue.
9. The method according to claim 8, characterized in that, The method further includes: In response to the driver's fatigue level being mild at the first moment, a Level 1 warning reminder is issued to the driver, which includes a voice prompt and a flashing warning light on the dashboard; In response to the driver's fatigue level being moderate at the first moment, a level two warning reminder is issued to the driver, which includes a phoenix alarm and a fatigue driving warning pop-up window on the central control screen; In response to the driver's fatigue level being severe fatigue at the first moment, a three-level warning reminder is issued to the driver, which includes a seat vibration reminder and a high-decibel voice warning.
10. A vision-based fatigue driving detection device, characterized in that, The device includes: At least one processor; And, a memory communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method as described in any one of claims 1-9.