Fatigue driving detection method based on personalized double-threshold filtering and adaptive threshold mechanism

By adopting personalized double-threshold filtering and adaptive threshold mechanism in fatigue driving detection, the problem of insufficient detection accuracy and stability in the prior art is solved, and a more efficient and reliable fatigue driving detection effect is achieved.

CN120148012AActive Publication Date: 2025-06-13SOUTHEAST UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510249881.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-13
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

The prior art has problems such as cumbersome operation, weak anti-interference ability, single evaluation dimension, data dependence, large computing resource consumption, and limited model generalization ability in fatigue driving detection, resulting in insufficient detection accuracy and stability.

Method used

The fatigue driving detection method based on personalized double-threshold filtering and adaptive threshold mechanism is adopted, and accurate detection of face key points and efficient calculation of feature parameters are achieved through technologies such as image acquisition, preprocessing, FaRL model recognition and heat map fusion.

Benefits of technology

It significantly improves the accuracy, generalization ability and real-time performance of detection, reduces the risk of driver fraudulent systems, and provides a reliable and practical fatigue driving detection solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148012A_ABST
    Figure CN120148012A_ABST
Patent Text Reader

Abstract

The invention provides a fatigue driving detection method based on personalized double-threshold filtering and an adaptive threshold mechanism. A face key point geometrical relationship is embedded into a heat map through matrix operation, and complete feature representation is constructed in combination with specific key point information of eye and mouth areas. An improved FaRL model based on ViT-B / 16 is adopted, a novel CNN structure is designed to serve as a feature parameter head, and the aspect ratio of the eyes and the aspect ratio of the mouth are precisely fitted from a fusion heat map. A personalized double-threshold filtering method is introduced to replace traditional peak detection, personalized characteristic parameters are extracted by analyzing video key frames to serve as reference values, and meanwhile, a detection standard is dynamically adjusted in combination with an adaptive threshold mechanism so as to cope with environmental and individual differences. The method improves the accuracy, real-time performance and environmental adaptability of fatigue detection, provides an efficient and reliable solution for intelligent driving safety monitoring, and has important application value and prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of driver fatigue driving monitoring, and particularly relates to a fatigue driving detection method based on personalized double-threshold filtering and an adaptive threshold mechanism. Background Art

[0002] In recent years, the analysis of the causes of road traffic accidents in China has shown that fatigue driving has become one of the main risk factors leading to traffic accidents. According to the current road traffic regulations, continuous driving for 4 hours constitutes fatigue driving. Fatigue driving has become a universal problem that cannot be ignored. After driving for a long time, drivers experience a decline in driving skills, resulting in slower reactions, inability to concentrate, decreased judgment ability, and in severe cases, even loss of control of the vehicle, which is very likely to cause road traffic accidents. In the field of fatigue driving detection technology, the extraction of facial key points and the calculation of feature parameters such as the Eye Aspect Ratio (EAR), Mouth Aspect Ratio (MAR), blink frequency, and yawn frequency are crucial, and the prediction and early warning of fatigue driving states have become the core of effective prevention and control. Modern fatigue driving detection systems provide technical guarantees for preventing traffic accidents caused by fatigue driving by real-time monitoring of driver states and combining timely early warning mechanisms, which is of great significance for improving road traffic safety.

[0003] Currently, with the continuous development of fields such as artificial intelligence and biomedical engineering, fatigue driving detection technology is also continuously innovating and progressing. For example, fatigue driving detection systems based on devices such as in-vehicle cameras and physiological signal acquisition devices can accurately evaluate the fatigue state of drivers and issue warnings in a timely manner by monitoring information such as eye movements, facial expressions, and physiological indicators. In addition, there are also fatigue driving early warning and management platforms based on technologies such as intelligent vehicle systems and Internet big data, which predict potential fatigue driving risks by analyzing factors such as drivers' behavior patterns and traffic environments, and take corresponding measures to reduce the risk of accidents and improve driving safety.

[0004] However, traditional fatigue driving detection methods based on physiological signal acquisition, vehicle driving behavior, single facial features, etc. have some limitations. For example, these methods usually require drivers to wear additional devices, which are not only inconvenient to use, but also their detection accuracy is vulnerable to environmental factors. Under harsh weather or complex road conditions, the driving behavior of the vehicle itself will be disturbed due to environmental changes, thus significantly reducing the accuracy of fatigue detection based on driving behavior. At the same time, it is difficult for a single feature to comprehensively reflect the fatigue state of the driver and is vulnerable to external environmental interference. For example, the blink frequency may fluctuate due to light changes or road bumps, leading to misjudgment. To overcome the limitations of traditional detection methods, fatigue driving detection based on deep learning has gradually emerged. Compared with traditional methods, deep learning technology has many advantages, such as not relying on external sensors, high stability and not being easily affected by the environment. However, this technology also faces certain challenges. The performance of the deep learning model may be excellent on a specific dataset, but when generalized to other datasets or actual application scenarios, its performance may drop significantly. This lack of generalization ability may lead to false alarms or missed detections, especially when various scenarios and condition changes are not fully considered during the model training stage. In addition, deep learning models are usually regarded as "black box models", and their internal decision-making processes are difficult to explain. This non-explainability limits users' understanding and trust in the model output results and also poses certain obstacles to further optimization and improvement of the model.

[0005] Therefore, how to overcome the limitations of the existing technology, achieve accurate, stable and real-time face key point detection, and efficiently calculate feature parameters such as the eye aspect ratio EAR, mouth aspect ratio MAR, blink frequency, yawn frequency, etc. has become the key challenge to improve the monitoring performance of the fatigue driving system. Summary of the Invention

[0006] In view of the limitations of the existing technology, the present invention innovatively proposes a fatigue driving detection method based on personalized double-threshold filtering and adaptive threshold mechanism. This method has achieved significant breakthroughs in multiple key technical indicators. Compared with traditional detection methods, the present invention has successfully overcome the inherent defects such as cumbersome operation, weak anti-interference ability, and single evaluation dimension; compared with the current mainstream deep learning solutions, it effectively solves the technical bottlenecks such as high data dependence, large consumption of computing resources, and limited model generalization ability. This method has achieved significant improvements in detection accuracy, generalization ability and real-time performance, and effectively reduces the risk of drivers deceiving the system. By innovatively integrating the advantages of multiple technologies, this method not only comprehensively improves the technical indicators, but also provides a reliable and practical solution for the field of fatigue driving detection, showing broad application prospects and practical promotion value.

[0007] The present invention adopts the following technical solutions:

[0008] A fatigue driving detection method based on personalized double-threshold filtering and adaptive threshold mechanism, including the following:

[0009] Step S1, image acquisition, acquiring face image data of the driver to be detected;

[0010] Step S2, image preprocessing, performing standardization processing on the acquired face image data, including extraction processing such as size unification, cropping, color channel conversion, and normalization;

[0011] Step S3, inputting the preprocessed image into the trained FaRL (Facial Representation Learning) model (its basic model is Vision Transformer (ViT), specifically using the ViT-B / 16 basic structure);

[0012] Step S4, face recognition, using the face detection model MTCNN to extract the face region;

[0013] Step S5, key point recognition, using the visual module of the FaRL model to perform face key point recognition, and inputting the face image data into a preset face detection model for arithmetic processing. Innovatively, it avoids the traditional point-by-point regression and geometric calculation methods, and instead realizes the transformation from point-by-point coordinate regression to overall feature extraction through heat map fusion and coordinate embedding technology;

[0014] Step S6, based on the face key point coordinate data, calculating two important facial feature parameters: the eye aspect ratio EAR and the mouth aspect ratio MAR;

[0015] Step S7, extracting the driver's personalized facial feature parameters and setting the action threshold values and value threshold values of the eye aspect ratio EAR and the mouth aspect ratio MAR. When the system is used for the first time, it samples and analyzes the driver's facial video frames, and extracts the driver's personalized facial feature parameters from the key frames as the double-threshold filtering reference values, and respectively sets the action threshold values and value threshold values of the eye aspect ratio EAR and the mouth aspect ratio MAR according to this reference value;

[0016] Step S8, processing the eye aspect ratio EAR, comparing it with the adaptive eye aspect ratio EAR threshold (action threshold value) to determine whether a blinking action occurs, and comparing it with the eye aspect ratio EAR value threshold to determine whether this blinking action is completed; processing the mouth aspect ratio MAR, comparing it with the adaptive mouth aspect ratio MAR (action threshold value) threshold to determine whether a yawning action occurs, and comparing it with the mouth aspect ratio MAR value threshold to determine whether this yawning action is completed;

[0017] Step S9, Frequency Statistics: According to the blink action determination method in S8, count the number of blinks within a certain period of time to obtain the blink frequency; according to the yawn action determination method in S8, count the number of yawns within a certain period of time to obtain the yawn frequency.

[0018] Step S10, Fatigue Driving State Judgment: Compare the blink frequency per unit time and the yawn frequency per unit time in S9 with the preset blink frequency threshold and yawn frequency threshold respectively to determine whether the driver is in a fatigue driving state. Or determine whether the driver to be detected is in a fatigue driving state according to whether the time when the driver to be detected is in a closed-eye state exceeds the preset fatigue closed-eye time threshold.

[0019] Further, the specific steps of Step S1 are as follows:

[0020] Step S11, Image Acquisition: Use an independent image acquisition module, including a normal image sensor and an infrared image sensor. The infrared image sensor includes an infrared fill light and an infrared camera. At the beginning of acquisition, use the normal image sensor for acquisition. The image acquisition module judges whether the light is qualified. If it is not qualified, start the infrared fill light and then switch to the infrared camera for acquisition to obtain the face image data of the driver to be detected.

[0021] Further, the specific steps of Step S2 are as follows:

[0022] Step S21, Image Preprocessing: That is, perform extraction processing such as size unification, cropping, color channel conversion, and normalization on the face image data. First, generate a unified image size of 447×447.

[0023] Step S22, Color Channel Conversion of the Image: If the input image is in grayscale or BGR format, convert the image format to RGB format.

[0024] Step S23, Normalization Processing of the Image: Map the pixel values from the range [0, 255] to the range [0, 1].

[0025] Step S23, Standardization Processing of the Normalized Image: Further subtract the mean of the training set of the channel from each pixel value and divide the result by the standard deviation of the training set of the channel. Unify the numerical ranges of different features and adjust the distribution of the image to the standardized state during model training, thereby improving the performance and stability of the model.

[0026] Further, the specific steps of Step S4 are as follows:

[0027] Step S41, Use the face detection model MTCNN to detect the face position in the image; output the bounding box of each face and return the coordinates (x of the face bounding boxbox , y box , width, height);

[0028] Step S42: To ensure including the complete face region, expand the width and height by 10%-20% based on the bounding box; the calculation formula for the new face bounding box is as follows:

[0029]

[0030] where new_width and new_height are the width and height of the face bounding box after expanding by 10%-20% respectively, p is the expansion ratio, and its value range is [0.1, 0.2], x box_new and y box_new are the abscissa and ordinate of the upper left corner of the new face bounding box respectively.

[0031] Step S43: Face cropping, according to the updated face bounding box, crop the sub-region containing the face from the original image.

[0032] Furthermore, the specific steps of Step S5 are as follows:

[0033] Step S51: Adopt the trained FaRL visual encoder as the backbone;

[0034] Step S52: Select the feature maps of the 4th, 6th, 7th, and 12th layers for multi-level feature fusion, and use UperNet to integrate the multi-level feature maps;

[0035] Step S53: Use 1×1 convolution to generate the prediction heatmap (heatmap) of the key points as the output layer;

[0036] Step S54: Render the groundtruth key points into a Gaussian heatmap of size 127×127, use the Gaussian distribution to model the key point positions and set its standard deviation to 1 pixel, and the value range of the heatmap is [0, 1];

[0037] Step S55: According to the relationship between the detected key points and the pre-set standard template, adopt Affine Transformation, and apply geometric transformations such as translation, rotation, and scaling to adjust the face to the position of the target template.

[0038] Step S56: Heatmap fusion. Adopt the two-dimensional heatmap regression method soft-argmax, which is defined as follows:

[0039]

[0040] where d is the given component x or y, and P is the weight matrix corresponding to the coordinates (x, y) of W×H×2. The matrix P can be represented by its components P x and P y , both of which are two-dimensional discrete normalized linear mappings, defined as follows:

[0041]

[0042] Here, Φ(h i,j ) represents the softmax result of a single heatmap, defined as:

[0043]

[0044] This method can be regarded as a convolution with a kernel size of H×W, where the content of the convolution kernel contains the position information of x or y and is arranged according to the pixel position Φ(h i,j ). The simplified soft-argmax operation performs a convolution on the weight heatmap of the key points according to the variation characteristics of the x-axis and y-axis, and after normalization and other processes, obtains the x coordinate and y coordinate of the key points.

[0045] Step S57: Analyze the geometric relationship between facial key points. Skip the regression and geometric calculations of each key point and analyze the overall facial features. Taking the definition of facial key points in the 300-W dataset as an example, landmarks 36-47 are classified as eye feature key points, and the inner lip landmarks 61-67 are classified as mouth feature key points. Define the overall facial feature weight Φ(h' i,j ) as:

[0046]

[0047] where n and m represent the serial numbers of the eye feature key points and mouth feature key points respectively.

[0048] Define the overall heatmap M as:

[0049]

[0050] Step S58, Coordinate embedding. When processing the overall heatmaps extracted from multiple images, it is found that it is difficult to effectively extract the feature information for distinguishing fatigue states only relying on direct convolution operations. To optimize the relative geometric relationship between key points after fusion, more complex matrix operations need to be performed on the heatmaps. The present invention innovatively draws on the core idea of the soft-argmax algorithm and designs a dual coordinate embedding mechanism to enhance the feature expression in the heatmap matrix. In the specific implementation process, first, zero-padding preprocessing is performed on the overall heatmap matrix. When H ≤ W, (W - H) columns of zeros are filled on the right side of M, and when H > W, (H - W) rows of zeros are filled at the bottom of M to convert it into a standard square matrix structure, laying a foundation for subsequent feature enhancement operations.

[0051]

[0052] Next, construct the diagonal matrix diag(0, 1,... K - 1). This diagonal matrix is used to represent the coordinate information of each column or each row in the heatmap. Finally, by multiplying the square matrix M' on the left and right with the diagonal matrix respectively, linear weighting is performed row by row and column by column to obtain the coordinate information in the horizontal and vertical directions and embed the position information into the feature map, realizing coordinate embedding on the overall heatmap:

[0053]

[0054] where H and W respectively represent the height and width of the heatmap, and M x represents the result of the heatmap embedded in the x direction, and M y represents the result of the heatmap embedded in the y direction. In the formula, K = max(W, H) ensures that the dimension of the diagonal matrix matches that of the square matrix M'. In this way, the geometric relationship between the relative coordinates of key points in the entire overall heatmap is emphasized. After coordinate embedding, the facial feature parameters required for driver fatigue detection can be extracted.

[0055] Step S59, Extract the data containing face key points from the fused heatmap after coordinate embedding. Among them, the key points numbered 36 to 47 are classified as eye key points, and the key points numbered 61 to 67 are classified as mouth key points, and the corresponding face image data are extracted;

[0056] Furthermore, the specific steps of Step S6 are as follows:

[0057] Step S61, Dynamically extract the feature information of the image data of the eye opening and closing state to obtain the first dynamic feature information - the coordinates of the key points of both eyes. The coordinates of the key points of both eyes include the coordinates of the upper eyelid key points (2), the coordinates of the lower eyelid key points (2), and the coordinates of the corner of the eye key points (2);

[0058] The expression of the first feature information (taking the left eye as an example) is as follows:

[0059] Calculate the sum of the Euclidean distances between two pairs of eye key points in the vertical direction based on the coordinates of the upper eyelid key points and the lower eyelid key points. The eye height is the distance from the upper eyelid to the lower eyelid. At the same time, calculate the distance between the left and right eye corners in the horizontal direction based on the coordinates of the eye key points at the eye corners. The formula for calculating the eye aspect ratio EAR is:

[0060]

[0061] In the formula, p 1 to p 6 are the 6 key points of the eye. The numerator ||p 2 -p 6 ||, ||p 3 -p 5 || represent the Euclidean distances between two pairs of eye key points corresponding to the upper and lower eyelids in the vertical direction. ||p 1 -p 4 || represents the Euclidean distance between the eye key points at the left and right eye corners in the horizontal direction. ||p 1 -p 4 || is multiplied by 2 to ensure that the numerator and denominator have the same weight in the calculation.

[0062] The formula for calculating the Euclidean distance between two pairs of eye key points corresponding to the upper and lower eyelids is:

[0063]

[0064] x 38 and x 39 are the abscissa values of the coordinate points corresponding to the upper eyelid key points; x 42 and x 41 are the abscissa values of the coordinate points corresponding to the lower eyelid key points; y 38 and y 39 are the ordinate values of the coordinate points corresponding to the upper eyelid key points; y 42 and y 41 are the ordinate values of the coordinate points corresponding to the lower eyelid key points;

[0065] Calculate the eye pixel width. Calculate the eye pixel width corresponding to both eyes based on the coordinates of the eye corner key points. The calculation formula is:

[0066]

[0067] Among them, ||p 1 -p 4 || is the pixel width value between the left and right eye corners of the eye; x 37 and y 37 are respectively the abscissa and ordinate values of the coordinate point corresponding to the left eye corner key point; x 40 and y40 They are respectively the abscissa values of the coordinate points corresponding to the key points at the right eye corner;

[0068] Step S62: Extract dynamic features from the image data of the mouth opening and closing state to obtain the first dynamic feature information, the mouth key point coordinates, which include the coordinates of the key points inside the upper lip (3), the coordinates of the key points inside the lower lip (3), and the coordinates of the key points at the corners of the mouth (2);

[0069] The expression of the first feature information, the mouth key point coordinates, is as follows:

[0070] According to the mouth key point coordinates, including the coordinates of the key points inside the upper lip and the coordinates of the key points inside the lower lip, calculate the sum of the Euclidean distances of the key points of the upper and lower lips in the vertical direction through these coordinates. In addition, according to the coordinates of the corners of the mouth inside the lips, calculate the Euclidean distance of the left and right corners of the mouth in the horizontal direction. The calculation formula of the mouth aspect ratio MAR is:

[0071]

[0072] where p 61 to p 65 are 7 key points of the mouth. ||p 62 -p 68 ||, ||p 63 -p 67 ||, ||p 64 -p 66 || represent the Euclidean distances of the three pairs of key points inside the upper and lower lips; ||p 61 -p 65 || represents the Euclidean distance of the key points at the left and right corners of the mouth inside the lips; ||p 61 -p 65 || is multiplied by 3 to ensure that the numerator and denominator have the same weight in the calculation.

[0073] The formula for calculating the Euclidean distance of the three pairs of corresponding key points inside the upper and lower lips is:

[0074]

[0075] x 62 x 63 x 64 are the abscissa values of the coordinate points corresponding to the key points inside the upper lip; x 66 x 67 x 68 are the abscissa values of the coordinate points corresponding to the key points inside the lower lip; y 62 y 63 y 64 are the ordinate values of the coordinate points corresponding to the key points inside the upper lip; y 66 y67 y 68 is the vertical coordinate value of the key point corresponding to the inner side of the lower lip;

[0076] Calculate the width of the mouth corners. Calculate the width corresponding to the mouth corners according to the coordinates of the key points of the mouth corners. The calculation formula is:

[0077]

[0078] where, ||p 61 -p 65 || is the pixel width value of the left and right mouth corners; x 61 y 61 are respectively the horizontal coordinate value and the vertical coordinate value of the key point corresponding to the left mouth corner on the inner side of the lip; x 65 y 65 are respectively the horizontal coordinate value and the vertical coordinate value of the key point corresponding to the right mouth corner on the inner side of the lip.

[0079] Furthermore, the specific steps of step S7 are as follows:

[0080] Step S71: When the driver just gets in the car and starts the system, start video frame cutting and extract the sequence of facial feature parameters, denoted as P = {p 1 , p 2 ,..., p n}, where p i is the facial feature value of the i-th frame. Obtain the change curves of the key parameters EAR and MAR.

[0081] Step S72: Statistically calculate the average values of the change curves of the eye aspect ratio EAR and the mouth aspect ratio MAR, which are used to reflect the typical values of the parameters in the normal state, and use them as the reference benchmark values for the adaptive thresholds of the double-door filtering. The calculation formulas for the average values of the eye aspect ratio EAR and the mouth aspect ratio MAR are as follows:

[0082]

[0083] where u is the average value of the eye aspect ratio EAR and the mouth aspect ratio MAR, n is the total number of frames within a unit time, and p i is the facial feature value of the i-th frame.

[0084] Step S73: Set an action threshold value to determine whether an action is completed, and select a threshold setting method based on feature recovery. During the blink detection process, it is specified that only when the Eye Aspect Ratio (EAR) parameter of the eye drops below this threshold, can the newly generated peak be used as the basis for counting the blink action. During the yawn detection process, it is specified that only when the Mouth Aspect Ratio (MAR) parameter of the mouth drops below this threshold, can the newly generated peak be used as the basis for counting the yawn action. This is used to prevent the influence of multiple adjacent peaks with unclear rising and falling trends caused by lens jitter due to road bumps on the counting of blink and yawn actions.

[0085] Step S74: The Eye Aspect Ratio (EAR) threshold (action threshold value) in S7 is improved based on the traditional fixed empirical threshold. A dynamic adjustment method based on the average value of historical Eye Aspect Ratio (EAR) is adopted, that is, the threshold is dynamically adjusted by using the average value of EAR values within a period of time, and the historical record is updated according to the current EAR value in each frame. In addition, an adjustment coefficient α is introduced. By calculating the average Eye Aspect Ratio (EAR) value of the last 20 frames and multiplying it by the coefficient α (usually between 1.05 and 1.20, which can be fine-tuned according to the actual situation), the dynamic adjustment of the action threshold value is achieved. The adaptive threshold calculation formula is as follows:

[0086] dynamic_threshold = sum(historical_ears) / len(historical_ears)×α (16)

[0087] Among them, dynamic_threshold represents the adaptive action threshold value, historical_ears represents the historical Eye Aspect Ratio (EAR) value, sum(historical_ears) represents the sum of EAR values within a period of time, len(historical_ears) represents the total number of frames recorded within a period of time, and α represents the dynamic adjustment coefficient, and its value range is [1.05, 1.20].

[0088] Step S75: Select the Mouth Aspect Ratio (MAR) threshold in Step S7. From the Mouth Aspect Ratio (MAR) calculation formula in Step S62, it can be seen that the MAR value is positively correlated with the opening degree of the mouth: when the mouth is completely closed, the distance between the upper and lower lips is the smallest, corresponding to the minimum Mouth Aspect Ratio (MAR) value; as the opening degree increases, the Mouth Aspect Ratio (MAR) value increases accordingly. In the daily driving state, the Mouth Aspect Ratio (MAR) value when speaking is usually between the closed state and the yawn state. The Mouth Aspect Ratio (MAR) value when speaking is greater than the MAR value when closed and generally less than the MAR value when yawning.

[0089] Based on the driver's personalized MAR parameter reference value obtained in step S7, a critical threshold M, i.e., the action threshold value, can be determined to distinguish the yawning state from the normal speaking state. However, it should be noted that in actual applications, the MAR value during speaking may briefly exceed the threshold M at certain moments. This is due to the periodic opening and closing of the mouth during speaking. Even if it occasionally exceeds the threshold M, its duration is relatively short. Therefore, to improve the detection accuracy, this method adopts a dual judgment mechanism, which not only examines whether the MAR value exceeds the threshold M, but also takes the duration for which the MAR value exceeds the threshold M as an important reference index. Through this comprehensive evaluation method combining time and space, the precision of yawning detection is effectively improved.

[0090] Step S76. Calculate the standard deviation of the facial feature parameter change curve according to step S71 to reflect the fluctuation range of the parameters. The standard deviation calculation formulas for the eye aspect ratio EAR and the mouth aspect ratio MAR are as follows:

[0091]

[0092] Set a value threshold, which is used to exclude invalid peaks caused by picture jitter or minor facial movements (such as slight eyelid movement), ensuring that only peaks exceeding a certain amplitude (reflecting clear blinking or yawning actions) will be counted as valid actions. Set the value threshold as:

[0093] T υ = μ - k·σ (18)

[0094] where T υ represents the value threshold, which is used to filter out invalid peaks below T υ in the feature curve. μ is the average value of the eye aspect ratio EAR curve, σ is the standard deviation of the eye aspect ratio EAR curve, and k is a hyperparameter, usually adjusted between [0.5, 1.5].

[0095] Furthermore, the specific steps of step S8 are as follows:

[0096] Step S81. Blinking action determination and counting. According to the eye aspect ratio EAR obtained in step S6, compare it with the EAR adaptive threshold (action threshold value) and the preset value threshold to determine whether the detected person has generated and completed a blinking action. When the EAR values of both eyes are less than the preset EAR action threshold, it is determined that the detected driver has started to generate a blinking action. When the EAR values of consecutive image frames are detected to exceed the EAR value threshold again, it is determined that the detected driver has completed this blinking action, and the blink count is incremented by one;

[0097] Step S82: Yawning action determination and counting. According to the mouth aspect ratio MAR obtained in step S6, compare it with the preset MAR action threshold value and the value threshold value to determine whether the detected person generates and completes a yawning action. When the mouth MAR value is less than the preset MAR action threshold value, it is determined that the detected person starts to generate a yawning action. When it is detected that the mouth MAR value of consecutive image frames exceeds the MAR value threshold again, it is determined that the detected person completes this yawning action, and the yawning action count is incremented by one.

[0098] Further, the specific steps of step S9 are as follows:

[0099] Step S91: The blinking frequency calculation formula is as follows:

[0100]

[0101] where f wink is the blinking frequency, N wink is the number of blinks within the unit time, and T is the unit time.

[0102] Step S92: The yawning frequency calculation formula is as follows:

[0103]

[0104] In the above formula, f yawn is the yawning frequency, n represents the number of yawns within the unit time T, and the yawning frequency can also be calculated by the number of frames:

[0105]

[0106] where, N represents the total number of frames within the unit time period, f n represents whether the mouth of the nth frame is in the yawning state. If f n is equal to 1, it is in the yawning state, and if f n is equal to 0, it is in the normal state.

[0107] Further, the specific steps of step S10 are as follows:

[0108] Step S101: Set the blink frequency threshold in Step S10. Based on a large amount of driver behavior research data, this study found the regular changes in blink characteristics with the progression of fatigue. In the normal waking state, the blink frequency of drivers usually maintains at the level of 12 - 18 times per minute. As the fatigue level increases, the blink characteristics show obvious phased changes. In the mild fatigue stage, drivers will instinctively increase blinking to combat fatigue, resulting in a significant increase in blink frequency and an extension of eye closure time. Research data shows that the blink frequency in the fatigued state will increase by 50% to 60% compared to the normal level, reaching 18 - 27 times per minute. In the moderate to severe fatigue stage, although the driver may keep their eyes open, it is often accompanied by phenomena such as distraction and dull eyes, or there may be a persistent eye closure state, resulting in a sharp drop in blink frequency. In addition, in terms of the blink cycle, in the normal state, a single blink (from fully open to fully closed) takes 200 - 400 milliseconds, while in the fatigued state, this cycle will be significantly extended to 600 - 700 milliseconds. Considering individual differences, especially that some groups may have a relatively high baseline blink frequency, after in-depth analysis, this study determined the fatigue determination threshold of blink frequency to be 25 times per minute to balance the accuracy and universality of detection.

[0109] Step S102: Fatigue determination. When the blink frequency of the detected driver exceeds the preset upper blink frequency threshold or is lower than the preset lower blink frequency threshold, it is determined that the detected driver is in a fatigued driving state; or when the yawn frequency of the detected driver exceeds the preset yawn frequency threshold, it is determined that the detected driver is in a fatigued driving state.

[0110] The core innovation of this method lies in the organic combination of personalized double-threshold filtering and an adaptive threshold adjustment mechanism. Based on the driver's personalized facial feature parameters, the system can respond in real time to changes in the driver's state and dynamic environmental changes (including lighting conditions, driving environment, and individual differences), and automatically optimize and adjust the action threshold and value threshold of the eye aspect ratio (EAR) and mouth aspect ratio (MAR), thus maintaining excellent detection accuracy and system stability in various application scenarios. The introduction of the double-threshold filtering technology effectively solves the problem of lens jitter caused by road bumps, and successfully eliminates the interference of adjacent invalid peaks without significant change trends on the counting of blinking and yawning actions. This improvement significantly enhances the overall performance of the fatigue driving detection system, reaching a new level in terms of accuracy, stability, and environmental adaptability. In addition, in terms of the model architecture, by designing a new CNN (Convolutional Neural Network) structure as the facial feature parameter head, directly learning and fitting the eye aspect ratio (EAR) and mouth aspect ratio (MAR) from the fused heatmap, the efficient extraction of face key point features is achieved. This design of the fatigue driving detection method based on personalized double-threshold filtering and the adaptive threshold mechanism opens up a new direction for the research and application in the field of fatigue driving detection, promoting the practical and intelligent development of the system in complex scenarios.

[0111] The beneficial effects of this invention are as follows: A complete face recognition and feature extraction system is established. Through heatmap fusion and coordinate embedding, the geometric relationship of key points is directly embedded into the heatmap, and the heatmap information of specific key points in the eye and mouth regions is fused to construct a complete feature representation. It not only retains the advantages of traditional geometric features but also further integrates heatmap information to achieve a more comprehensive and in-depth state assessment. At the technical implementation level, this invention innovatively uses coordinate embedding and heatmap fusion technologies to extract the eye aspect ratio (EAR) and mouth aspect ratio (MAR) indicators, and breakthroughly adopts a personalized double-threshold filtering method to replace the traditional peak detection and counting mechanism, significantly improving the accuracy and stability of the counting process. In terms of system intelligence, this invention designs a unique threshold mechanism: when the driver first uses the system, it extracts personalized facial feature parameters by analyzing the key frames of the video and uses them as the reference values for double-threshold filtering. At the same time, combined with the innovative adaptive threshold adjustment mechanism, it can optimize the detection threshold in real time according to multi-dimensional factors such as lighting conditions, driving environment complexity, and individual differences, significantly enhancing the overall accuracy and environmental adaptability of the fatigue driving detection method.

[0112] Compared with the existing technologies, the present invention overcomes the deficiencies of traditional detection methods such as cumbersome operation, poor anti-interference ability, and single evaluation criteria. At the same time, it effectively solves the problems of data dependence, computational resource consumption, and limited generalization ability in deep learning methods, achieving an overall improvement in detection accuracy, real-time performance, and environmental adaptability. This comprehensive solution not only provides reliable technical support for fatigue driving detection but also makes an important contribution to improving road traffic safety levels, with broad practical application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0113] Figure 1 It is a flowchart of the fatigue driving detection method of the present invention.

[0114] Figure 2 It is a simplified soft-argmax operation process of the present invention.

[0115] Figure 3 It is the training and inference workflow of the fine-tuning model of the present invention.

[0116] Figure 4 It is a diagram of eye key points and inner lip key points of the mouth of the present invention.

[0117] Figure 5 It is a diagram of the heatmap coordinate embedding process of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0118] The present invention will be further clarified below in conjunction with the drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.

[0119] Figure 1 It is a flowchart of the fatigue driving detection system of the present invention. The present invention will be further described below in conjunction with the drawings.

[0120] Step 1: Image acquisition, acquiring real-time face image data of the driver to be detected;

[0121] Step 2: Image preprocessing, performing standardization processing on the acquired face image data, including size unification, region cropping, color channel conversion, and normalization, etc.;

[0122] Step 3: Input the preprocessed image into the trained FaRL model;

[0123] Step 4: Face region detection, using the face detection model MTCNN to extract the face region;

[0124] Step 5: Key point recognition. Use the visual module of the FaRL model to perform face key point recognition, and input the face image data into a preset face recognition model for computing and processing. Innovatively, it avoids the traditional point-by-point regression and geometric calculation methods, and instead realizes the transformation from point-by-point coordinate regression to overall feature extraction through heat map fusion and coordinate embedding techniques;

[0125] Step 6: Based on the face key point coordinate data, calculate two important facial feature parameters: the eye aspect ratio EAR and the mouth aspect ratio MAR;

[0126] Step 7: Extract the driver's personalized facial feature parameters and set the action threshold values and value threshold values of the eye aspect ratio EAR and the mouth aspect ratio MAR. When the system is used for the first time, sample and analyze the video frames of the driver's face, and extract the driver's personalized facial feature parameters from the key frames as the double-threshold filtering reference values, and set the action threshold values and value threshold values of the eye aspect ratio EAR and the mouth aspect ratio MAR respectively according to this reference value;

[0127] Step 8: Process the eye aspect ratio EAR, compare it with the adaptive eye aspect ratio EAR threshold (action threshold value) to determine whether a blinking action occurs, and compare it with the value threshold value of the eye aspect ratio EAR to determine whether this blinking action is completed; process the mouth aspect ratio MAR, compare it with the adaptive mouth aspect ratio MAR (action threshold value) threshold to determine whether a yawning action occurs, and compare it with the value threshold value of the mouth aspect ratio MAR to determine whether this yawning action is completed;

[0128] Step 9: Frequency statistics. Count the number of blinks within a period of time according to the blink action determination method in Step 8 to obtain the blink frequency; count the number of yawns within a period of time according to the yawn action determination method in Step 8 to obtain the yawn frequency;

[0129] Step 10: Fatigue driving state determination. Compare the blink frequency per unit time and the yawn frequency per unit time in Step 9 with the preset blink frequency threshold and yawn frequency threshold respectively to determine whether the driver is in a fatigue driving state. Or determine whether the driver to be detected is in a fatigue driving state according to whether the time when the driver to be detected is in a closed-eye state exceeds the preset fatigue closed-eye time threshold;

[0130] The specific implementation of Step 1 is:

[0131] First, an independent image acquisition module is used for image acquisition, which includes a normal image sensor and an infrared image sensor. The infrared image sensor includes an infrared fill light and an infrared camera. At the beginning of acquisition, the normal image sensor is used for acquisition. The image acquisition module determines whether the light is qualified. If it is not qualified, the infrared fill light is activated, and then the infrared camera is switched to for acquisition to obtain the face image data of the driver to be detected;

[0132] The specific implementation of step two is as follows:

[0133] First, image preprocessing is performed, that is, extraction processing such as size unification, cropping, color channel conversion, and normalization is performed on the face image data,

[0134] First, a unified image size of 447×447 is generated. Then, color channel conversion is performed on the image. If the input image is in grayscale or BGR format, the image format is converted to RGB format.

[0135] Then, normalization processing is performed on the image, mapping the value range of its pixels from [0,255] to [0,1].

[0136] Finally, standardization processing is performed on the normalized image, further subtracting the mean of the training set of the channel from each pixel value and dividing the result by the standard deviation of the training set of the channel. The numerical ranges of different features are unified, and the distribution of the image is adjusted to the standardized state during model training, thereby improving the performance and stability of the model.

[0137] The specific implementation of step four is as follows:

[0138] First, the face detection model MTCNN is used to detect the face positions in the image. The bounding box (BoundingBox) of each face is output, and the coordinates (x, y, width, height) of the face bounding box are returned.

[0139] Next, to ensure that the complete face area is included, a certain proportion is extended on the basis of the bounding box, for example, increasing the width and height by 10%-20%. The calculation formula for the new bounding box is as follows:

[0140]

[0141] Finally, face cropping is performed, and according to the updated face bounding box, the sub-region containing the face is cropped from the original image.

[0142] The specific implementation of step five is as follows:

[0143] In the first step, the trained FaRL visual encoder is used as the backbone;

[0144] Step 2: Select the feature maps of the 4th, 6th, 7th, and 12th layers for multi-level feature fusion, and use UperNet to integrate the multi-level feature maps;

[0145] Step 3: Use 1×1 convolution to generate the heat map prediction of the key points as the output layer;

[0146] Step 4: Render the ground truth key points into a Gaussian heat map of size 127×127, model the key point positions using a Gaussian distribution and set its standard deviation to 1 pixel, and the value range of the heat map is [0,1];

[0147] Step 5: According to the relationship between the detected key points and the pre-set standard template, use Affine Transformation to apply geometric transformations such as translation, rotation, and scaling to adjust the face to the position of the target template.

[0148] Step 6: Heat map fusion. Adopt the two-dimensional heat map regression method soft-argmax, which is defined as follows:

[0149]

[0150] In the formula, d is the given component x or y, and P is the weight matrix corresponding to the coordinates (x,y) of W×H×2. The matrix P can be represented by its components P x and P y These two components are both two-dimensional discrete normalized linear mappings, and are defined as follows:

[0151]

[0152] Here, Φ(h i,j ) represents the softmax result of a single heat map, which is defined as:

[0153]

[0154] This method can be regarded as a convolution with a kernel size of H×W, where the content of the convolution kernel contains the position information of x or y, and is arranged according to the pixel position Φ(h i,j ). The simplified soft-argmax operation is as shown in Figure 2 . According to the change characteristics of the x-axis and y-axis, perform a convolution on the weight heat map of the key points, and after normalization and other processes, obtain the x coordinate and y coordinate of the key points.

[0155] Step 7: Analyze the geometric relationship between the facial key points. Skip the regression and geometric calculations of each key point and analyze the overall facial features. Figure 4For the definition of facial key points in the 300-W dataset, landmarks 36 to 47 are classified as eye feature key points, and inner lip landmarks 61 to 67 are classified as mouth feature key points. Define the overall facial feature weight Φ(h′ i,j ) as:

[0156]

[0157] In the formula, n and m represent the serial numbers of eye feature key points and mouth feature key points respectively.

[0158] Define the overall heatmap M as:

[0159]

[0160] Eighth step, Figure 5 This is the process diagram of heatmap coordinate embedding of the present invention. The present invention will be further described below with reference to the accompanying drawings. Although heatmap fusion can extract the geometric information of key points, it is often difficult to effectively capture the spatial geometric relationship between key points by directly performing convolution operations on the fused heatmap. In addition, experiments show that when dealing with the overall heatmaps extracted from multiple images, it is difficult to extract effective features that can distinguish fatigue states simply by relying on convolution operations. This further illustrates that convolution operations alone are not sufficient to fully model the geometric relationship between key points. For this reason, the present invention innovatively draws on the core idea of the soft-argmax algorithm and designs a dual coordinate embedding mechanism to enhance the feature expression in the heatmap matrix. In the specific implementation process, first perform zero-padding preprocessing on the overall heatmap matrix. When H ≤ W, fill (W - H) columns of zeros on the right side of M. When H > W, fill (H - W) rows of zeros at the bottom of M to convert it into a standard square matrix structure, laying a foundation for subsequent feature enhancement operations.

[0161]

[0162] Then, construct the diagonal matrix diag(0, 1,... K - 1). This diagonal matrix is used to represent the coordinate information of each column or each row in the heatmap. Finally, by multiplying the square matrix M′ on the left and right with the diagonal matrix respectively, linear weighting is performed row by row and column by column to obtain the coordinate information in the horizontal and vertical directions respectively and embed the position information into the feature map, realizing coordinate embedding on the overall heatmap:

[0163]

[0164] Among them, H and W represent the height and width of the heatmap respectively, M x represents the result of embedding the heatmap in the x direction, M yIt represents the result after the heatmap is embedded in the y direction. In the formula, K = max(W, H) ensures that the dimension of the diagonal matrix matches that of the square matrix M'. In this way, the geometric relationship between the relative coordinates of the key points in the entire overall heatmap is emphasized. After coordinate embedding, the facial feature parameters required for driver fatigue detection can be extracted.

[0165] The training and inference workflow of the entire fine-tuning model is as Figure 3 shown. Initially, the driver's image undergoes a standard facial alignment procedure. After facial region detection, a facial alignment network is applied to generate a facial key point heatmap. Subsequently, the heatmap will be used in two different processes. One group will be regressed to determine the coordinates of the key points, and then geometric calculations will be performed and used as training data labels. The other group will be fused and filled to generate an overall heatmap <', and then coordinate embedding will be performed to generate a high-dimensional feature representation. These embedded features serve as the input to the CNN facial feature parameter extraction module, which first extracts significant features and reduces the dimension through a single max pooling layer, and then effectively captures the spatial feature relationship through a five-layer convolutional network. The extracted features then undergo a series of linear layers for feature transformation and fitting calculations, and finally, two highly generalized facial feature parameter scalars are output through a fully connected layer, providing an accurate quantification index for fatigue state judgment.

[0166] Step 9: Extract the data containing the face key points from the fused heatmap after coordinate embedding. Among them, the key points numbered 36 to 47 are classified as eye key points, and the key points numbered 61 to 67 are classified as mouth key points, and the corresponding face image data are extracted.

[0167] The specific implementation of Step 6 is as follows:

[0168] First, dynamic feature extraction is performed on the image data of the eye opening and closing state to obtain the first dynamic feature information - the coordinates of the two-eye key points. The coordinates of the two-eye key points include the coordinates of the upper eyelid key points (2), the coordinates of the lower eyelid key points (2), and the coordinates of the eye corner key points (2).

[0169] The expression of the first feature information (taking the left eye as an example) is as follows:

[0170] Next, according to the coordinates of the upper eyelid key points and the lower eyelid key points, the sum of the Euclidean distances of the two pairs of eye key points in the vertical direction is calculated, that is, the distance from the upper eyelid to the lower eyelid. At the same time, according to the coordinates of the eye key points at the eye corners, the horizontal distance between the left and right eye corners is calculated. The formula for calculating the eye aspect ratio EAR is:

[0171]

[0172] In the formula, p 1 to p 6 are the 6 key points of the eyes. The numerator ||p2 -p 6 ||, ||p 3 -p 5 || represents the Euclidean distance in the vertical direction between two pairs of corresponding eye key points on the upper and lower eyelids, ||p 1 -p 4 || represents the Euclidean distance in the horizontal direction between the eye key points at the left and right eye corners, ||p 1 -p 4 Multiply || by 2 to ensure that the numerator and denominator have the same weight in the calculation.

[0173] The formula for calculating the Euclidean distance between two pairs of corresponding eye key points on the upper and lower eyelids is:

[0174]

[0175] x 38 x 39 is the abscissa value of the coordinate point corresponding to the key point on the upper eyelid; x 42 x 41 is the abscissa value of the coordinate point corresponding to the key point on the lower eyelid; y 38 y 39 is the ordinate value of the coordinate point corresponding to the key point on the upper eyelid; y 42 y 41 is the ordinate value of the coordinate point corresponding to the key point on the lower eyelid;

[0176] Calculate the eye pixel width. Calculate the eye pixel width corresponding to both eyes based on the coordinates of the eye corner key points. The calculation formula is:

[0177]

[0178] where, ||p 1 -p 4 || is the pixel width value between the left and right eye corners of the eye; x 37 y 37 are respectively the abscissa and ordinate values of the coordinate point corresponding to the key point of the left eye corner; x 40 y 40 are respectively the abscissa and ordinate values of the coordinate point corresponding to the key point of the right eye corner;

[0179] Finally, according to the calculated ||p 2 -p 6 ||, ||p 3 -p 5 || and ||p 1 -p 4 ||, substitute them into the above-mentioned eye aspect ratio EAR formula for calculation, and finally obtain the EAR value.

[0180] Next, dynamic feature extraction is performed on the image data of the mouth opening and closing state to obtain the first dynamic feature information, the mouth key point coordinates. The mouth key point coordinates include the inner upper lip key point coordinates (3), the inner lower lip key point coordinates (3), and the corner of the mouth key point coordinates (2).

[0181] The expression of the first feature information is as follows:

[0182] According to the mouth key point coordinates, including the inner upper lip key point coordinates and the inner lower lip key point coordinates, the sum of the Euclidean distances of the upper and lower lip key points in the vertical direction is calculated through these coordinates. In addition, according to the coordinates of the inner corners of the lips, the Euclidean distance between the left and right corners of the mouth in the horizontal direction is calculated. The calculation formula for the mouth aspect ratio MAR is:

[0183]

[0184] where p 61 to p 65 are the 7 key points of the mouth. ||p 62 -p 68 ||, ||p 63 -p 67 ||, ||p 64 -p 66 || represent the Euclidean distances of the three pairs of key points inside the upper and lower lips; ||p 61 -p 65 || represents the Euclidean distance of the key points of the left and right corners of the mouth inside the lips; ||p 61 -p 65 || is multiplied by 3 to ensure that the numerator and denominator have the same weight in the calculation.

[0185] The formula for calculating the Euclidean distance of the three pairs of corresponding key points inside the upper and lower lips is:

[0186]

[0187] x 62 x 63 x 64 is the abscissa value of the corresponding coordinate point of the inner upper lip key point; x 66 x 67 x 68 is the abscissa value of the corresponding coordinate point of the inner lower lip key point; y 62 y 63 y 64 is the ordinate value of the corresponding coordinate point of the inner upper lip key point; y 66 y 67 y 68 is the ordinate value of the corresponding coordinate point of the inner lower lip key point;

[0188] Calculate the width of the corners of the mouth. Calculate the width corresponding to the corners of the mouth based on the coordinates of the key points at the corners of the mouth. The calculation formula is:

[0189]

[0190] where ||p 61 -p 65 || is the pixel width value between the left and right corners of the mouth; x 61 y 61 are respectively the abscissa value and the ordinate value of the coordinate point corresponding to the key point of the left corner of the mouth inside the lips; x 65 y 65 are respectively the abscissa value and the ordinate value of the coordinate point corresponding to the key point of the right corner of the mouth inside the lips.

[0191] Finally, according to the calculated ||p 62 -p 68 ||, ||p 63 -p 67 ||, ||p 64 -p 66 || and ||p 61 -p 65 ||, substitute them into the above-mentioned mouth aspect ratio MAR formula for calculation, and finally obtain the MAR value.

[0192] The specific implementation of Step 7 is as follows:

[0193] First, when the driver starts the vehicle and the system is initialized, start video frame extraction and extract the sequence of facial feature parameters, denoted as P = {p 1 , p 2 ,..., p n}, where p i is the facial feature value of the i-th frame. Obtain the change curves of the key parameters eye aspect ratio EAR and mouth aspect ratio MAR.

[0194] Next, calculate the average values of the EAR and MAR change curves, which are used to reflect the typical values of the parameters in the normal state, and use them as the reference benchmark values for the adaptive thresholds of the double-threshold filtering. The calculation formulas for the average values of the eye aspect ratio EAR and the mouth aspect ratio MAR are as follows:

[0195]

[0196] Then, set an action threshold value to determine whether an action has ended, and select a threshold setting method based on feature recovery. In blink detection, only when the Eye Aspect Ratio (EAR) parameter drops below this threshold value can a newly generated peak be used as the basis for counting blink actions. In yawn detection, only when the Mouth Aspect Ratio (MAR) parameter drops below this threshold value can a newly generated peak be used as the basis for counting yawn actions. This is used to prevent the influence of multiple adjacent peaks with unclear rising and falling trends caused by lens jitter due to bumpy road conditions on the counting of blink and yawn actions.

[0197] Finally, the Eye Aspect Ratio (EAR) threshold value (action threshold value) in Step 7 is improved on the basis of the traditional fixed empirical threshold value. A dynamic adjustment method based on the historical average of EAR is adopted, that is, the threshold value is dynamically adjusted by using the average value of EAR values over a period of time, and the historical record is updated according to the current EAR value in each frame. In addition, an adjustment coefficient α is introduced. By calculating the average EAR value of the last 20 frames and multiplying it by the coefficient α (usually between 1.05 and 1.20, which can be fine-tuned according to the actual situation), the dynamic adjustment of the action threshold value is realized. The adaptive threshold calculation formula is as follows:

[0198] dynamic_threshold = sum(historical_ears) / len(historical_ears) × α (16)

[0199] Among them, dynamic_threshold represents the adaptive action threshold value, historical_ears represents the historical Eye Aspect Ratio (EAR) values, sum(historical_ears) represents the sum of EAR values over a period of time, len(historical_ears) represents the total number of frames recorded over a period of time, and α represents the dynamic adjustment coefficient, with a value range of [1.05, 1.20].

[0200] When setting the Mouth Aspect Ratio (MAR) threshold value (action threshold value) and the value threshold value, first, select the MAR threshold value in Step 7. From the Mouth Aspect Ratio (MAR) calculation formula, it can be seen that the MAR value is positively correlated with the degree of mouth opening and closing: when the mouth is completely closed, the distance between the upper and lower lips is the smallest, corresponding to the minimum MAR value; as the opening degree increases, the MAR value increases accordingly. In the daily driving state, the MAR value during speaking is usually between the closed state and the yawn state. The MAR value during speaking is greater than the MAR value in the closed state and generally less than the MAR value during yawning.

[0201] Based on the reference value of the driver's personalized mouth aspect ratio MAR parameter obtained in Step Seven, a critical threshold M, i.e., the action threshold value, can be determined to distinguish between the yawning state and the normal speaking state. However, it should be noted that in practical applications, the MAR value during speaking may briefly exceed the threshold M at certain moments. This is due to the periodic opening and closing of the mouth during speaking. Even if it occasionally exceeds the threshold M, the duration is relatively short. Therefore, to improve the detection accuracy, this method adopts a dual judgment mechanism, which not only examines whether the MAR value exceeds the threshold M, but also takes the duration for which the MAR value exceeds the threshold M as an important reference index. Through this comprehensive evaluation method that combines time and space, the accuracy of yawning detection is effectively improved.

[0202] Next, according to the standard deviation of the statistical curve of facial feature parameters, which is used to reflect the fluctuation range of the parameters. The standard deviation calculation formulas for the eye aspect ratio EAR and the mouth aspect ratio MAR are as follows:

[0203]

[0204] Finally, set the value threshold. The value threshold is used to exclude invalid peaks caused by picture jitter or minor facial movements (such as slight eyelid movement), ensuring that only peaks exceeding a certain amplitude (reflecting clear blinking or yawning actions) will be counted as valid actions. The value threshold is set as:

[0205] T υ = μ - k·σ (18)

[0206] where T υ represents the value threshold, which is used to filter out invalid peaks below T υ in the feature curve. μ is the average value of the EAR curve, σ is the standard deviation of the EAR curve, and k is a hyperparameter, usually adjusted between [0.5, 1.5].

[0207] The specific implementation of Step Eight is as follows:

[0208] First, blink action determination and counting. According to the eye aspect ratio EAR obtained in Step Six, compare it with the EAR adaptive threshold (action threshold value) and the preset value threshold to determine whether the detected person has generated and completed a blink action. When the EAR values of both eyes are less than the preset EAR action threshold value, it is determined that the detected driver has started to generate a blink action. When the EAR values of consecutive image frames are detected to exceed the EAR value threshold again, it is determined that the detected driver has completed this blink action, and the blink count is incremented by one.

[0209] Next, perform yawn action determination and counting. According to the mouth aspect ratio MAR obtained in Step 6, compare it with the preset MAR action threshold value and the value threshold value to determine whether the detected person generates and completes a yawn action. When the mouth MAR value is less than the preset MAR action threshold value, it is determined that the detected person starts to generate a yawn action. When the MAR value of consecutive image frames is detected to exceed the MAR value threshold again, it is determined that the detected person has completed this yawn action, and the yawn action count is incremented by one.

[0210] The specific implementation of Step 9 is as follows:

[0211] First, the blink frequency calculation formula is as follows:

[0212]

[0213] where f wink is the blink frequency, N wink is the number of blinks within the unit time, and T is the unit time.

[0214] Then, the yawn frequency calculation formula is as follows:

[0215]

[0216] In the above formula, f yawn is the yawn frequency, n represents the number of yawns within the unit time T, and the yawn frequency can also be calculated by the number of frames:

[0217]

[0218] where, N represents the total number of frames within the unit time period, f n represents whether the mouth of the nth frame is in a yawning state. If f n is equal to 1, it is in a yawning state, and if f n is equal to 0, it is in a normal state.

[0219] The specific implementation of Step 10 is as follows:

[0220] First, set the blink frequency threshold for Step 10. Based on a large amount of driver behavior research data, this study found the regular changes in blink characteristics as the fatigue level evolves. In the normal waking state, the blink frequency of drivers usually remains at the level of 12 - 18 times per minute. As the fatigue level progresses, the blink characteristics show obvious phased changes. In the mild fatigue stage, drivers will instinctively increase blinking to combat fatigue, resulting in a significant increase in blink frequency and an extension of eye closure time. Research data shows that the blink frequency in the fatigue state will increase by 50% to 60% compared to the normal level, reaching 18 - 27 times per minute. In the moderate to severe fatigue stage, although drivers may keep their eyes open, they often experience phenomena such as distracted attention, dull looks, or persistent eye closure, resulting in a sharp drop in blink frequency. In addition, in terms of the blink cycle, the time required for a single blink (from fully open to fully closed) in the normal state is 200 - 400 milliseconds, while in the fatigue state, this cycle will be significantly extended to 600 - 700 milliseconds. Considering individual differences, especially that some groups may have a relatively high baseline blink frequency, after in-depth analysis, this study determined the fatigue determination threshold of the blink frequency to be 25 times per minute to balance the accuracy and universality of detection.

[0221] Then, conduct fatigue determination. When the blink frequency of the detected driver exceeds the preset upper threshold of the blink frequency or is lower than the preset lower threshold of the blink frequency, it is determined that the detected driver is in a fatigue driving state; or when the yawn frequency of the detected driver exceeds the preset yawn frequency threshold, it is determined that the detected driver is in a fatigue driving state.

[0222] Finally, based on the personalized facial feature parameters of the driver and combined with the adaptive threshold adjustment mechanism, the personalized dual-threshold filtering method can automatically adjust the action thresholds and value thresholds of the Eye Aspect Ratio (EAR) and Mouth Aspect Ratio (MAR) according to the driver's real-time state and the dynamic changes of the environment (such as lighting conditions, driving environment, and driver individual differences). Through this dynamic optimization mechanism, the system can maintain high precision in different scenarios. At the same time, the dual-threshold filtering effectively suppresses the lens jitter caused by road bumps, thus avoiding the interference of multiple adjacent invalid peaks without an obvious upward or downward trend on the counting of blinking and yawning actions, and greatly improving the overall accuracy, stability, and environmental adaptability of the vehicle fatigue driving detection system. In addition, by designing a new CNN structure as the facial feature parameter head, directly embedding the geometric relationship of key points into the heat map, and fusing the heat map information of specific key points in the eye and mouth regions to construct a complete feature representation, and finally learning and fitting the Eye Aspect Ratio (EAR) and Mouth Aspect Ratio (MAR) from the fused heat map, the efficient extraction of face key point features is realized. The design of the fatigue driving detection method based on the personalized dual-threshold filtering and adaptive threshold mechanism has opened up a new direction for the research and application in the field of fatigue driving detection, and promoted the practical and intelligent development of the system in complex scenarios.

[0223] It should be noted that the above content only illustrates the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. For those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and refinements can still be made, and these improvements and refinements all fall within the protection scope of the claims of the present invention.

Claims

1. A fatigue driving detection method based on personalized double threshold filtering and adaptive threshold mechanism, characterized by: The following steps are involved: Step S1: image acquisition, collecting real-time facial image data of the driver to be detected; Step S2: image preprocessing, standardizing the collected facial image data, including size unification, region cropping, color channel conversion and normalization; Step S3, input the preprocessed image into the trained FaRL model; Step S4: face region detection, using the face detection model MTCNN to extract the face region; Step S5, key point recognition, using the FaRL model visual module to perform face key point recognition, and inputting the face image data into a preset face detection model for calculation and processing; Step S6, based on the facial key point coordinate data, calculate the eye aspect ratio EAR and mouth aspect ratio MAR, two important facial feature parameters; Step S7, extracting the driver's personalized facial feature parameters and setting the eye aspect ratio EAR and mouth aspect ratio MAR action thresholds and value thresholds; sampling and analyzing the video frames of the driver's facial video when used for the first time, and extracting the driver's personalized facial feature parameters from the key frames as double threshold filtering reference values, and setting the eye aspect ratio EAR and mouth aspect ratio MAR action thresholds and value thresholds according to the reference values; Step S8, processing the eye aspect ratio EAR, comparing it with the adaptive eye aspect ratio EAR threshold, determining whether a blinking action occurs, and comparing it with the eye aspect ratio EAR value threshold, determining whether the blinking action is completed; processing the mouth aspect ratio MAR, comparing it with the adaptive mouth aspect ratio MAR threshold, determining whether a yawning action occurs, and comparing it with the mouth aspect ratio MAR value threshold, determining whether the yawning action is completed; Step S9, frequency statistics, counting the number of blinks per unit time according to the blink action determination method in S8, and then obtaining the blink frequency; counting the number of yawns per unit time according to the yawn action determination method in S8, and then obtaining the yawn frequency; Step S10, fatigue driving state determination, the blinking frequency per unit time in S9 and the yawning frequency per unit time are compared with the preset blinking frequency threshold and the yawning frequency threshold respectively to determine whether the driver is in a fatigue driving state; or whether the driver to be detected is in a fatigue driving state by judging whether the time when the driver to be detected is in a closed eye state exceeds the preset fatigue closed eye time threshold.

2. The fatigue driving detection method based on personalized double threshold filtering and adaptive threshold mechanism according to claim 1 is characterized by: The specific steps of step S1 are as follows: Step S11, image acquisition uses an independent image acquisition module, including a common image sensor and an infrared image sensor, and the infrared image sensor includes an infrared fill light and an infrared camera; when acquisition starts, the common image sensor is used for acquisition, and the image acquisition module determines whether the light is qualified. If it is unqualified, the infrared fill light is started, and then the infrared camera is switched to acquire the face image data of the driver to be detected.

3. The fatigue driving detection method based on personalized double threshold filtering and adaptive threshold mechanism according to claim 1 is characterized in that: The specific steps of step S2 are as follows: Step S21, perform image preprocessing, that is, perform size unification, cropping, color channel conversion and normalization extraction processing on the face image data. First, generate a unified image size of 447×447; Step S22: convert the color channel of the image. If the input image is in grayscale or BGR format, convert the image format to RGB format. Step S23, normalize the image and map the pixel values ​​from the range [0, 255] to [0, 1]; Step S24, perform standardization on the normalized image, further subtract the training set mean of the channel from each pixel value, and divide the result by the training set standard deviation of the channel; unify the numerical ranges of different features, and adjust the distribution of the image to the standardized state during model training.

4. The fatigue driving detection method based on personalized double threshold filtering and adaptive threshold mechanism according to claim 1 is characterized in that: The specific steps of step S4 are as follows: Step S41: Use the face detection model MTCNN to detect the face position in the image; output the bounding box of each face and return the coordinates (x box ,y box ,width,height); Step S42: To ensure that the entire face area is included, the width and height of the bounding box are expanded by 10%-20%. The calculation formula of the new face bounding box is as follows: Where new_width and new_height are the width and height of the face bounding box after expansion by 10%-20%, p is the expansion ratio, and the value range is [0.1, 0.2]. box_new and box_new They are the horizontal and vertical coordinates of the upper left corner of the new face bounding box respectively; Step S43: face cropping: cropping a sub-region containing the face from the original image according to the updated face bounding box.

5. The method for detecting fatigue driving based on personalized dual threshold filtering and adaptive threshold mechanism according to claim 1, characterized in that: The specific steps of step S5 are as follows: Step S51, using the trained FaRL visual encoder as backbone; Step S52, select the feature maps of the 4th, 6th, 7th and 12th layers for multi-level feature fusion, and use UperNet to integrate the multi-layer feature maps; Step S53, using 1×1 convolution to generate a heat map prediction of the key points as the output layer; Step S54, render the ground truth key points into a Gaussian heat map of size 127×127, use Gaussian distribution to model the key point positions and set its standard deviation to 1 pixel, and the heat map value range is [0,1]; Step S55, according to the relationship between the detected key points and the preset standard template, using affine transformation, applying translation, rotation, scaling geometric transformation to adjust the face to the position of the target template; Step S56, heat map fusion: using the heat map regression method soft-argmax, defined as follows: Where d is the given component x or y, and P is the weight matrix of W×H×2 corresponding to the coordinates (x, y); The matrix P is represented by its components P x and P y Indicates that both components are two-dimensional discrete normalized linear mappings, defined as follows: Here, Φ(h i,j ) represents the softmax result of a single heat map, defined as: This method is regarded as a convolution with a kernel size of H×W, where the content of the convolution kernel contains the position information of x or y and is calculated according to the pixel position Φ(h i,j ) arrangement; according to the change characteristics of the x-axis and y-axis, the weight heat map of the key point is convolved once, and after normalization and other processing, the x-coordinate and y-coordinate of the key point are obtained; Step S57, analyzing the geometric relationship between facial key points; skipping the regression and geometric calculation of each key point, and analyzing the overall facial features; Based on the definition of facial key points in the 300-W dataset, landmarks 36 to 47 are classified as eye feature key points, and inner lip landmarks 61 to 67 are classified as mouth feature key points; definition Overall facial feature weight Φ(h' i,j )for: Where n and m represent the serial numbers of the eye feature key points and mouth feature key points respectively; The overall heat map M is defined as: Step S58, coordinate embedding; When processing the overall heat map extracted from multiple images, we draw on the core idea of ​​the soft-argmax algorithm and design a dual coordinate embedding mechanism to enhance the feature expression in the heat map matrix; In the specific implementation process, the overall heat map matrix is ​​first preprocessed by zero padding. When H≤W, the (WH) columns of zeros are filled on the right side of M, and when H>W, the (HW) rows of zeros are filled at the bottom of M to convert it into a standard square matrix structure. Next, construct a diagonal matrix diag(0,1,...K-1); this diagonal matrix is ​​used to represent the coordinate information of each column or row in the heat map; finally, by using the diagonal matrix to multiply the square matrix M′ on the left and right, linear weighting by row and column is achieved, and the horizontal and vertical coordinate information is obtained respectively, and the position information is embedded in the feature map, and the coordinate embedding is achieved on the overall heat map: Where H and W represent the height and width of the heat map, respectively, and M x represents the result after the heat map is embedded in the x direction, M y represents the result after the heat map is embedded in the y direction; where K = max(W,H) ensures that the diagonal matrix dimension matches the square matrix M′; in this way, the geometric relationship between the relative coordinates of the key points in the entire heat map is emphasized; after the coordinates are embedded, the facial feature parameters required for driver fatigue detection are extracted; Step S59, extracting data containing facial key points from the fused heat map after coordinate embedding, wherein key points 36 to 47 are classified as eye key points, key points 61 to 67 are classified as mouth key points, and extracting corresponding facial image data.

6. The method for detecting fatigue driving based on personalized dual threshold filtering and adaptive threshold mechanism according to claim 5, characterized in that: The specific steps of step S6 are as follows: Step S61, extracting dynamic features from the image data of the eye opening and closing state to obtain first dynamic feature information binocular key point coordinates, wherein the binocular key point coordinates include two upper eyelid key point coordinates, two lower eyelid key point coordinates, and two eye corner key point coordinates; The expression of the left eye in the first feature information binocular key point coordinates is as follows: According to the coordinates of the upper eyelid key points and the lower eyelid key points, the sum of the vertical Euclidean distances of the two pairs of eye key points, that is, the distance from the upper eyelid to the lower eyelid, is calculated. At the same time, according to the coordinates of the eye key points at the corners of the eyes, the distance between the left and right eye corners in the horizontal direction is calculated. The formula for calculating the eye aspect ratio EAR is: Where p1 to p6 are the six key points of the eye; the numerators ||p2-p6|| and ||p3-p5|| represent the vertical Euclidean distances between the two pairs of eye key points on the upper and lower eyelids, ||p1-p4|| represents the horizontal Euclidean distance between the eye key points at the left and right corners of the eye, and ||p1-p4|| is multiplied by 2 to ensure that the numerator and denominator maintain the same weight in the calculation; The formula for calculating the Euclidean distance between the two pairs of eye key points corresponding to the upper and lower eyelids is: x 38 x 39 is the horizontal coordinate value of the coordinate point corresponding to the upper eyelid key point; 42 x 41 y is the horizontal coordinate value of the coordinate point corresponding to the key point of the lower eyelid; 38 y 39 y is the ordinate value of the coordinate point corresponding to the upper eyelid key point; 42 y 41 The ordinate value of the coordinate point corresponding to the key point of the lower eyelid; Calculate the eye pixel width, and calculate the eye pixel width corresponding to both eyes according to the eye corner key point coordinates. The calculation formula is: Among them, ||p1-p4|| is the pixel width value of the left and right corners of the eye; x 37 y 37 The x and y coordinates are the values ​​of the corresponding coordinates of the left eye corner key point respectively; 40 y 40 They are the horizontal and vertical coordinates of the coordinate points corresponding to the key points of the right eye corner; Step S62, extracting dynamic features from the image data of the mouth opening and closing state to obtain first dynamic feature information, namely, the coordinates of the key points of the mouth, wherein the coordinates of the key points of the mouth include three coordinates of the inner side of the upper lip, three coordinates of the inner side of the lower lip, and two coordinates of the key points of the corners of the mouth; The expression of the coordinates of the key points of the mouth of the first feature information is as follows: According to the coordinates of the key points of the mouth, including the coordinates of the key points on the inner side of the upper lip and the key points on the inner side of the lower lip, the vertical Euclidean distance of the key points of the upper and lower lips is calculated by these coordinates. In addition, according to the coordinates of the inner corners of the lips, the horizontal Euclidean distance of the left and right corners of the mouth is calculated. The calculation formula of the mouth aspect ratio MAR is: Among them, p 61 to p 65 There are 7 key points of the mouth; ||p 62 -p 68 ||、||p 63 -p 67 ||、||p 64 -p 66 || represents the Euclidean distance of three pairs of key points on the inner side of the upper and lower lips; ||p 61 -p 65 || represents the Euclidean distance between the key points of the left and right corners of the lips; || p 61 -p 65 ||Multiply by 3 to ensure that the numerator and denominator have the same weight in the calculation; The Euclidean distance formula for calculating the three pairs of mouth key points corresponding to the inner sides of the upper and lower lips is: x 62 x 63 x 64 is the horizontal coordinate value of the key point on the inner side of the upper lip; 66 x 67 x 68 is the horizontal coordinate value of the key point on the inner side of the lower lip; 62 y 63 y 64 y is the ordinate value of the key point on the inner side of the upper lip corresponding to the coordinate point; 66 y 67 y 68 The ordinate value of the key point on the inner side of the lower lip corresponding to the coordinate point; Calculate the mouth corner width, and calculate the width corresponding to the mouth corner according to the coordinates of the mouth corner key points. The calculation formula is: Among them, ||p 61 -p 65 || is the pixel width of the left and right corners of the mouth; x 61 y 61 are the horizontal and vertical coordinate values ​​of the key points of the left corner of the mouth on the inner side of the lip; 65 y 65 They are respectively the horizontal and vertical coordinate values ​​of the coordinate points corresponding to the key points of the right corner of the inner side of the lips.

7. The method for detecting fatigue driving based on personalized dual threshold filtering and adaptive threshold mechanism according to claim 1, characterized in that: The specific steps of step S7 are as follows: Step S71: When the driver just gets on the car and starts the system, start video frame cutting and extract facial feature parameter sequence, set as P = {p1, p2, ..., p n }, where p i is the facial feature value of the i-th frame; obtain the change curves of the key parameters eye aspect ratio EAR and mouth aspect ratio MAR; Step S72: Calculate the average values ​​of the eye aspect ratio EAR and mouth aspect ratio MAR change curves to reflect the typical values ​​of the parameters under normal conditions, and use them as reference values ​​for the adaptive thresholds of the dual-gate filtering. The formula for calculating the average values ​​of the eye aspect ratio EAR and mouth aspect ratio MAR is as follows: Where u is the average of the eye aspect ratio EAR and the mouth aspect ratio MAR, n is the total number of frames per unit time, and p i is the facial feature value of the i-th frame; Step S73, setting an action threshold value, used to determine whether an action is finished, and selecting a threshold setting method based on feature recovery; in the blink detection process, it is stipulated that only when the eye aspect ratio EAR parameter drops below the threshold, the newly generated peak value can be used as the basis for blink action counting; in the yawn detection process, it is stipulated that only when the mouth aspect ratio MAR parameter drops below the threshold, the newly generated peak value can be used as the basis for yawn action counting; Step S74: for the eye aspect ratio EAR threshold of step S7, a dynamic adjustment method based on the historical eye aspect ratio EAR average value is adopted, that is, the threshold is dynamically adjusted by using the average value of the eye aspect ratio EAR value over a period of time, and the historical record is updated in each frame according to the current eye aspect ratio EAR value; in addition, an adjustment coefficient α is introduced, and the average eye aspect ratio EAR value of the most recent 20 frames is calculated and multiplied by the coefficient α, so as to realize the dynamic adjustment of the action threshold value; The adaptive threshold calculation formula is as follows: dynamic_threshold=sum(historical_ears) / len(historical_ears)×α (16) Where dynamic_threshold represents the adaptive motion threshold, historical_ears represents the historical eye aspect ratio EAR value, sum(historical_ears) represents the sum of EAR values ​​in a period of time, len(historical_ears) represents the total number of frames recorded in a period of time, and α represents the dynamic adjustment coefficient, which ranges from [1.05, 1.20]. Step S75, selecting the mouth aspect ratio MAR threshold value in step S7. It can be known from the mouth aspect ratio MAR calculation formula in step S62 that the MAR value is positively correlated with the degree of mouth opening and closing: when the mouth is completely closed, the distance between the upper and lower lips is the smallest, and the corresponding mouth aspect ratio MAR value reaches the minimum; as the opening degree increases, the mouth aspect ratio MAR value increases accordingly; in daily driving conditions, the mouth aspect ratio MAR value when speaking is usually between the closed state and the yawning state; the mouth aspect ratio MAR value when speaking is greater than the MAR value when closed, and generally less than the MAR value when yawning; Based on the driver's personalized mouth aspect ratio MAR parameter reference value obtained in step S7, a critical threshold M, i.e., an action threshold value, is determined to distinguish between a yawning state and a normal speaking state; A dual judgment mechanism is adopted, which not only examines whether the mouth aspect ratio MAR value exceeds the threshold M, but also takes the duration of the MAR value exceeding the threshold M as an important reference indicator; Step S76: Calculate the standard deviation of the facial feature parameter variation curve according to step S71 to reflect the fluctuation range of the parameters; the standard deviation calculation formulas for the eye aspect ratio EAR and the mouth aspect ratio MAR are as follows: Set the value threshold. The value threshold is used to exclude invalid peaks caused by image jitter or slight facial movements, ensuring that only peaks exceeding a certain amplitude are counted as valid movements. Set the value threshold to: T υ =μ-k·σ (18) Where T υ Indicates the threshold value, used to filter the characteristic curve below T υ The invalid peak value of , μ is the mean value of the eye aspect ratio EAR curve, σ is the standard deviation of the eye aspect ratio EAR curve, and k is a hyperparameter adjusted between [0.5, 1.5].

8. The method for detecting fatigue driving based on personalized dual threshold filtering and adaptive threshold mechanism according to claim 1, characterized in that: The specific steps of step S8 are as follows: Step S81, blink action determination and counting, based on the eye aspect ratio EAR obtained in step S6, compare it with the EAR adaptive threshold and the preset threshold value to determine whether the detected person has produced and completed the blink action; when the EAR values ​​of both eyes are less than the preset EAR action threshold value, it is determined that the detected driver has started to produce a blink action, and when it is detected that the EAR value of the continuous image frame exceeds the EAR value threshold value again, it is determined that the detected driver has completed the blink action, and the blink count is increased by one; Step S82, yawning action determination and counting, based on the mouth aspect ratio MAR obtained in step S6, compare it with the preset MAR action threshold value and the value threshold value to determine whether the detected person has generated and completed the yawning action; when the mouth MAR value is less than the preset MAR action threshold value, it is determined that the detected person has started to yawn, and when it is detected that the MAR value of the continuous image frame exceeds the MAR value threshold value again, it is determined that the detected person has completed the yawning action, and the yawning action count is increased by one.

9. The method for detecting fatigue driving based on personalized dual threshold filtering and adaptive threshold mechanism according to claim 1, characterized in that: The specific steps of step S9 are as follows: Step S91: The blink frequency calculation formula is as follows: where f wink is the blink frequency, N wink is the number of blinks per unit time, T is the unit time; Step S92: The yawning frequency calculation formula is as follows: where f yawn is the yawning frequency, n represents the number of yawns in the unit time T, and the yawning frequency can also be calculated by the number of frames: Where N represents the total number of frames in a unit time period, f n Indicates whether the mouth of the nth frame is in a yawning state. If f n If it is equal to 1, it is a yawning state, f n A value of 0 indicates normal status.

10. The method for detecting fatigue driving based on personalized dual threshold filtering and adaptive threshold mechanism according to claim 1, characterized in that: The specific steps of step S10 are as follows: Step S101, setting the blink frequency threshold of step S10, and determining the blink frequency fatigue determination threshold to be 25 times / minute; Step S102, fatigue determination: when the blinking frequency of the detected driver exceeds the preset blinking frequency upper limit threshold or is lower than the preset blinking frequency lower limit threshold, it is determined that the detected driver is in a fatigue driving state; or when the yawning frequency of the detected driver exceeds the preset yawning frequency threshold, it is determined that the detected driver is in a fatigue driving state.

Citation Information

Patent Citations

  • Early fatigue detection method and system based on fine eye movement features

    CN112434611A

  • Fatigue detection method and system based on machine vision

    CN117392644A

  • Fatigue driving detection method and device, electronic equipment and storage medium

    CN117746400A

  • Method and apparatus for detecting driver state, and storage medium

    WO2023108364A1