Miner driver fatigue symptom detection method

Through the multimodal fusion fatigue sign detection method, combined with the SSD deep learning model and Faster R-CNN's Anchor mechanism for facial recognition, and combined with blinking, yawning and heart rate characteristics for fatigue determination, the problem of insufficient accuracy and stability of single-modal recognition technology in miner and driver fatigue detection is solved, and higher detection accuracy and environmental adaptability are achieved.

CN120241072APending Publication Date: 2025-07-04XI'AN UNIVERSITY OF ARCHITECTURE AND TECHNOLOGY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510304500.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing single-modal recognition technology is difficult to adapt to individual differences and complex environments in the fatigue detection of miners and drivers, resulting in insufficient accuracy and stability of identification results, especially in the working environment under the mine, which has a great impact on factors such as lighting changes, dust, and noise.

Method used

The fatigue sign detection method of multimodal fusion is used, and facial recognition is combined with the SSD deep learning model and Faster R-CNN's Anchor mechanism, and fatigue determination is performed by blinking, yawning and heart rate characteristics, so that detection accuracy is improved through feature fusion.

Benefits of technology

It improves the accuracy and stability of miner driver fatigue detection, can better adapt to individual differences and complex environmental changes, and enhances the robustness and adaptability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120241072A_ABST
    Figure CN120241072A_ABST
Patent Text Reader

Abstract

The invention discloses a miner driver fatigue symptom detection method. The method comprises the following steps: 1, recognizing the face of a miner driver; 2, detecting the driving fatigue based on the recognized facial features; step 3, detecting fatigue driving based on heart rate characteristics; and 4, performing feature fusion fatigue judgment according to the driving fatigue detection and the fatigue driving detection based on the heart rate features. According to the fatigue driving detection method, three modes of blink, yawn and heart rate detection are fused, recognized and analyzed; according to the multi-mode identification technology, multiple information sources can be comprehensively considered, the individual difference can be better adapted, and the identification accuracy and stability are improved. The method has the characteristics of multi-modal fusion and high adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of driver fatigue detection, and particularly to a method for detecting fatigue signs of miner drivers. Background Art

[0002] Physiological and behavioral responses of different miners in a fatigued state may vary significantly. Single-modal recognition technology often has difficulty adapting to such individual differences, resulting in the accuracy and stability of recognition results being affected.

[0003] Since single-modal recognition technology relies on specific information sources, its robustness and generalization ability are usually limited. When faced with new or unknown fatigue states, this technology may not be able to make accurate judgments.

[0004] In a complex underground mining working environment, factors such as light changes, dust, and noise can all affect the acquisition and recognition effect of a single information source.

[0005] Existing technologies for processing driver fatigue detection, such as physiological signal detection, behavior model analysis, heart rate variability analysis, etc., such as An, G. (2024) ‘Fatigue driving detection methods based on human-computer interaction’, Transactions on Computer Science and Intelligent Systems Research, 5, pp.621–625.doi: 10.62051 / 4rdcbr36;

[0006] However, these detection technologies have certain defects, such as poor individual differences and scene adaptability, single detection, and being prone to missing other fatigue characteristics by relying only on one indicator. Summary of the Invention

[0007] In order to overcome the above technical problems, the purpose of the present invention is to provide a method for detecting fatigue signs of miner drivers, which has the characteristics of multi-modal fusion and strong adaptability.

[0008] To achieve the above purpose, the technical solution adopted by the present invention is:

[0009] A method for detecting fatigue signs of miner drivers, comprising the following steps;

[0010] Step 1: Identify the face of the miner driver:

[0011] Step 2: Based on the identified facial features, perform driving fatigue detection;

[0012] Step 3: Fatigue driving detection based on heart rate features;

[0013] Step 4: Based on the driving fatigue detection and the fatigue driving detection based on heart rate characteristics, perform feature fusion fatigue determination.

[0014] The specific content of the said Step 1 is as follows:

[0015] Combine the SSD deep learning model with the Anchor mechanism of Faster R-CNN to perform face recognition on the collected video images; use the SSD deep learning model to identify the face region and the coordinate position information of the face in the video frame;

[0016] The video images are collected in real time from the camera, and the target category and position information are read from the annotation file during training. The annotation contains the category label and bounding box coordinates of each target.

[0017] The Anchor mechanism enables the RPN of Faster R-CNN to generate candidate boxes using multiple scales and aspect ratios to cover different target shapes.

[0018] The Adaptation of SSD targets the characteristic that the face is close to 1:1. Denser 1:1 ratio anchors are set in each feature layer of SSD, supplemented by a small number of other ratios to handle side faces or occlusions. Among them, the shallow feature map (such as 38x38) detects small-scale faces; the deep feature map (such as 1x1) detects large-scale faces;

[0019] The Anchor size of each layer of the feature map matches its receptive field; to avoid missing detection of too large or too small targets, the shallow layer is responsible for detecting faces of 16-48 pixels, the middle layer for 48-128 pixels, and the deep layer for processing larger sizes;

[0020] At the same time, adopt the bidirectional matching strategy of Faster R-CNN. Each image corresponds to a set of annotations, including the category and bounding box coordinates of each target. Match the ground truth box (GT) directly from the annotation file of the dataset to the Anchor with the highest IoU, and at the same time mark the Anchor with IoU>0.5 as a positive sample to improve the recall rate; the rest are marked as negative samples;

[0021] For negative samples, use hard sample mining, and preferentially select background Anchors with high classification confidence but low IoU to alleviate class imbalance.

[0022] In Faster R-CNN, dynamic weight adjustment is usually carried out during the model training phase. After determining the positive and negative samples, the model uses the Focal Loss function to calculate the loss value of each sample, and performs backpropagation and parameter update based on these loss values. By introducing loss functions such as Focal Loss, the model can more effectively handle class imbalance and hard example problems, thereby improving the overall performance.

[0023] Dynamic weight adjustment reduces the weight of easy-to-classify samples by introducing Focal Loss to replace cross-entropy, and focuses on training hard examples (such as occluded faces).

[0024] Using a two-stage heuristic design, a simplified version of RPN is embedded in the front end of SSD, and Anchor adjustment and feature fusion are transformed into a practical network structure. Potential defects in the previous steps are compensated by data augmentation and regularization to avoid overfitting simple samples. Generate high-quality face candidate boxes, and then balance speed and accuracy through SSD multi-feature layer classification and regression.

[0025] Add an attention mechanism to the regression branch, simulate the face detection scenario, increase random occlusion, blur, illumination changes and multi-angle rotation for targeted enhancement. Mix multiple images to generate complex backgrounds and improve the generalization of the model.

[0026] The loss function design selects Smooth L1 Loss to cope with challenges such as occlusion and illumination changes, and at the same time adjusts the weight parameters in Smooth L1 Loss to balance the contributions of positive and negative samples to the loss function, thereby avoiding the model from overfitting to negative or positive samples.

[0027] The loss function of the entire model is:

[0028]

[0029] where x is used to determine whether the detection box matches the true bounding box, represented by 1 and 0; N represents the number of matching boxes; l is the predicted bounding position parameter; g is the true bounding position parameter; c represents the category;

[0030] L conf L(x, c) represents the recognition result degree, which is the Softmax loss function for multiple classes; L loc L(x, l, g) uses the Smooth L1 Loss to predict the loss value between the position and the actual result;

[0031] After the above operations, the face is recognized to obtain the features of the face, eyes and mouth.

[0032] The specific steps of step 2 are as follows:

[0033] (1) Facial feature marking

[0034] Based on the facial images collected in real time by the camera, 68 facial feature points are extracted from the face according to the Dlib library. The extraction of the feature points depends on the principle of the HOG algorithm to locate the position and bounding box of the face in the image, and the bounding box of the facial area is output.

[0035] Then, a pre-trained model is loaded, and this model is trained by the ERT algorithm.

[0036] The position of the feature points is gradually optimized through a multi-level regression tree. From the initial estimate to the gradual adjustment, each regression tree learns the mapping relationship between the local texture features and the offset of the feature points, and finally the coordinates (x, y) of the feature points are output.

[0037] (2) Eye feature extraction

[0038] First, the coordinate data of the eye feature points are obtained through the dlib library, and the coordinate data are used as input values; by calculating the ratio of the distance between two points in the vertical direction to the distance between two points in the horizontal direction, the degree of eye closure is quantified.

[0039] Then, the EAR algorithm is used to calculate the six obtained feature points. For the left eye and the right eye respectively, the corresponding Euclidean distance is calculated, and its EAR value is obtained.

[0040] The formula for calculating the EAR value of the left eye is:

[0041] The formula for calculating the EAR value of the right eye is:

[0042] In the blinking action, the left eye and the right eye move synchronously. By calculating the average value of the EAR values of the left and right eyes, the EAR value of the entire eye is obtained, so as to more accurately reflect the state of eye movement, as shown in the formula

[0043]

[0044] (3) Mouth feature extraction

[0045] The Dlib library collects feature points on the mouth model, and the features cover the outer contour and inner contour of the lips.

[0046] First, the collected feature points are used as input data for the algorithm; subsequently, the MAR algorithm is used to calculate the Euclidean distance between these six feature points to obtain the MAR value, as shown in the formula:

[0047]

[0048] (4) Eye feature fatigue determination

[0049] Using the p80 judgment criterion, by extracting eye features, calculating the eye aspect ratio (EAR), and then based on the size of the EAR, determining whether the human eye is in a closed state;

[0050] (5) Mouth feature fatigue determination

[0051] According to the threshold set by the MAR algorithm, determine whether the degree of mouth opening has reached the yawn standard, and at the same time, further confirm whether it is a yawn action by analyzing the duration of the mouth opening state;

[0052] In the step (2), when the value of EAR is lower than the preset threshold, the eye completes a blink. Set the threshold of EAR to 0.2. When EAR is less than 0.2, it is recorded as a blink.

[0053] In the step (3), set the threshold of MAR to 0.6. When MAR is greater than 0.6, it is determined as an abnormal mouth opening action.

[0054] In the step (4), set the threshold of EAR to 0.2. When EAR is less than 0.2, it is recorded as a blink. PERCLOS is defined as the degree of eye closure within a certain period, and its calculation formula is as shown.

[0055]

[0056] Statistical value of PERCLOS for the driver's left and right eye closures within 1 minute is K eye ; When K eye value is greater than 0.2, it is determined that the tram driver is in a fatigued state.

[0057] In the step (5), when the MAR value exceeds 0.65 and this state lasts for more than 3 seconds, it is determined as a yawn. The number of yawns of a normal adult within 1 minute is 1. Statistical number of yawns of the driver within 1 minute is recorded as K mouth ;

[0058] When K mouth value is greater than 1, it is determined that the tram driver is in a fatigued state.

[0059] The specific content of the step 3 is as follows:

[0060] (1) Heart rate feature signal extraction

[0061] The user's face is facing the camera. Based on the real-time captured face image by the camera, 68 feature points of the face can be successfully extracted. Select the area with rich blood flow as the ROI (region of interest). Stabilize the position of the ROI through feature point tracking (such as Lucas-Kanade) to compensate for the small head movement.

[0062] Extract the time-series signals of the red (R), green (G), and blue (B) channels from the video frames. Perform color space conversion to separate luminance (Y) and chrominance (Cb, Cr), reduce the influence of illumination, and simultaneously extract hue (H) and saturation (S) information. Then perform signal normalization operations to eliminate baseline drift, and the formula is

[0063]

[0064] where μ c (t) is the mean value of the color channels within the ROI.

[0065] Separate the heart rate signal from other interference sources through independent component analysis (ICA), and use a band-pass filter to retain the frequency range of 0.5 - 4 Hz (corresponding to 30 - 240 BPM).

[0066] Perform Fourier transform (FFT) to convert the time-domain signal to the frequency domain, find the energy peak, and then locate the frequency f corresponding to the maximum power through the power spectral density (PSD) peak .

[0067] (2) Heart rate calculation

[0068] Select the frequency f corresponding to the maximum power peak peak , then the corresponding heart rate

[0069] v hr is:

[0070] v hr = 60 * f peak

[0071] (3) Heart rate feature fatigue determination

[0072] When the value of v hr stabilizes at 100 or higher, the tram driver is in a fatigued state.

[0073] The specific steps of step 4 are as follows:

[0074] The degree of fatigue has reached a certain stage, and the KSS values of blinking, yawning, and heart rate fatigue states are set to 7;

[0075] When the value of K eye is greater than 0.2 and less than 0.4, assign it a KSS value of 7; when the value of K eye is greater than 0.4 and less than 0.5, assign it a KSS value of 8; when the value of K eye is greater than 0.5, assign it a KSS value of 9;

[0076] When the value of K mouth is 2, assign it a KSS value of 7; when the value of K mouthWhen the value is 3, assign its KSS value as 8; when K mouth is greater than 3, assign its KSS value as 9;

[0077] When v hr is greater than 100 and less than 110, assign its KSS value as 7; when v hr is greater than 110 and less than 130, assign its KSS value as 8; when v hr is greater than 130, assign its KSS value as 9;

[0078] When KSS e +KSS m +KSS h is greater than or equal to 21 and less than 24, the miner driver belongs to mild fatigue; when KSS e +KSS m +KSS h is greater than or equal to 24 and less than 27, the miner driver belongs to moderate fatigue; when KSS e +KSS m +KSS h is equal to 27, the miner driver belongs to severe fatigue.

[0079] Advantages of the present invention:

[0080] The present invention is a fatigue driving detection method that fuses and analyzes three modalities of blink, yawn, and heart rate detection; the multi-modal recognition technology of the present invention can comprehensively consider multiple information sources, better adapt to this individual difference, and improve the accuracy and stability of recognition.

[0081] The multi-modal recognition technology of the present invention: By fusing information from multiple modalities, the multi-modal recognition technology can better handle data noise and input changes. For example, in the field of face recognition, the multi-modal model can simultaneously utilize image and voice inputs, thus better coping with problems such as face occlusion and light changes. This multi-modal fusion improves the robustness and stability of the system.

[0082] The multi-modal recognition technology of the present invention can reduce the dependence on a single information source by integrating multiple information sources, thereby improving the adaptability to environmental changes.

[0083] The multi-modal recognition technology of the present invention: Since it can process multiple types of data inputs, the multi-modal recognition technology has a wider range of application scenarios. It can be applied to multiple fields such as natural language processing, computer vision, speech recognition, and intelligent interaction, meeting the recognition requirements in various complex scenarios. Description of the Drawings

[0084] Figure 1 is a schematic flow diagram of the present invention.

[0085] Figure 2 It is a schematic diagram of 68 facial feature points of the Dlib library.

[0086] Figure 3 It is a schematic diagram of the ROI area. Specific implementation manner

[0087] The present invention will be further described in detail below with reference to the accompanying drawings.

[0088] A method for detecting fatigue signs of miner drivers includes the following steps;

[0089] Step 1: Perform face recognition:

[0090] In order to identify the facial area and the coordinate position information of the face in the video frame, and considering the accuracy and timeliness of the deep convolutional neural network framework at the same time, the SSD model proposed by Wei Liu is adopted. This model has the following advantages:

[0091] 1) An end-to-end detection and recognition method, similar to YOLO, but faster;

[0092] 2) Combining the Anchor mechanism in Faster R-CNN on the basis of YOLO, and using the network feature maps of different scales to predict the targets at each position, achieving an effect approximate to the object detection and recognition based on the region proposal method;

[0093] Although the detection effect of the SSD model on small sizes is not ideal, for this application, the distance between the tester's head area and the camera is kept at about 1.5 meters, which can better avoid the occurrence of this problem.

[0094] The input parameters of the SSD deep learning model include images, object classes, and the position information of the objects. The overall structure includes the extraction of image depth information and the discrimination of the deep network.

[0095] The images are collected in real time from the camera, and the object categories and position information are read from the annotation file during training. The annotation contains the class labels and bounding box coordinates of each object.

[0096] Combining the SSD deep learning model with the Anchor mechanism of Faster R-CNN for face recognition, mainly by optimizing the Anchor design, feature layer selection, matching strategy, and training method, to improve the detection accuracy and robustness. The core improvement of the Anchor mechanism is that the RPN of Faster R-CNN uses multiple scales and aspect ratios to generate candidate boxes, covering different object shapes.

[0097] The shallow feature map (such as 38x38) detects small-scale faces, and a smaller reference size (such as 16x16) is set;

[0098] Deep feature maps (such as 1x1) are used to detect large-scale faces, and the reference size gradually increases (such as 256x256).

[0099] Feature layer optimization is achieved by adding shallow high-resolution feature maps on the basis of the ResNet backbone network of SSD to enhance the small object detection ability. The deep feature maps are fused with the shallow details through deconvolution or skip connections to improve the sensitivity to faces of different scales.

[0100] The Anchor size of each layer of feature maps matches its receptive field to avoid missing detection of too large or too small objects. The shallow layer is responsible for detecting faces of 16 - 48 pixels, the middle layer for 48 - 128 pixels, and the deep layer for larger sizes.

[0101] At the same time, the bidirectional matching strategy of Faster R-CNN is adopted. Each ground truth box (GT) is matched to the Anchor with the highest IoU, and at the same time, the Anchors with IoU > 0.5 are marked as positive samples to improve the recall rate. Hard negative mining is used for negative samples, and the background Anchors with high classification confidence but low IoU are preferentially selected to alleviate class imbalance.

[0102] Dynamic weight adjustment is achieved by introducing Focal Loss to replace cross-entropy, reducing the weights of easy-to-classify samples and concentrating on training hard examples (such as occluded faces).

[0103] A two-stage heuristic design is used. A simplified version of RPN is embedded at the front end of SSD to generate high-quality face candidate boxes, and then classification and regression are performed through multiple feature layers of SSD to balance speed and accuracy. An attention mechanism is added to the regression branch to improve the context reasoning ability for occluded faces. The face detection scenario is simulated, and random occlusion, blur, illumination change, and multi-angle rotation are added for targeted enhancement. Multiple images are mixed to generate complex backgrounds to improve the generalization of the model. The loss function design uses SmoothL1 Loss to optimize the bounding box regression, paying more attention to the center point distance and aspect ratio.

[0104] The loss function of the entire model is as follows:

[0105]

[0106] Among them, x is used to judge whether the detection box matches the real bounding box, represented by 1 and 0; N represents the number of matching boxes; l is the predicted boundary position parameter; g is the real boundary position parameter; c represents the category;

[0107] L conf (x, c) represents the recognition result degree, which is the Softmax loss function for multiple classes; L loc(x, l, g) uses the Smooth L1 Loss to predict the loss value between the position and the actual result.

[0108] Step 2: Fatigue driving detection based on facial features; both fatigue driving detection based on facial features and facial recognition rely on facial feature analysis technologies (such as face detection and feature point localization), but their objectives are different: fatigue detection evaluates the driving state through dynamic features (such as blinking and yawning), while facial recognition verifies identity through static features (such as facial features).

[0109] (1) Facial feature marking

[0110] Based on the face images captured in real time by the camera, 68 feature points of the face can be successfully extracted. The extraction of these feature points mainly relies on the principle of the HOG algorithm to locate the position and bounding box of the face in the image, and outputs the bounding box (Bounding Box) of the face area.

[0111] Then load the pre-trained model (such as shape_predictor_68_face_landmarks.dat), which is trained by the ERT algorithm. The position of the feature points is gradually optimized through a multi-level regression tree, from the initial estimate to the gradual adjustment, with the aim of minimizing the error between the predicted points and the real points. Each regression tree learns the mapping relationship between the local texture features and the offset of the feature points, and finally outputs the coordinates (x, y) of 68 feature points. It includes 17 points on the chin contour, 5 points on each of the left and right eyebrows, 9 points on the nose bridge and tip, 6 points on each of the left and right eyes, 12 points on the outer contour of the lips and 8 points on the inner contour.

[0112] The dlib's shape_predictor_68_face model is designed specifically for face feature point detection after training. Through this model library, these feature points can be accurately marked on the face image, as Figure 2 shown.

[0113] The 68-point annotation is derived from the widely used iBug 68-point dataset in the academic community (ICCV 2013) and has become a benchmark for face alignment. This standard balances accuracy and computational efficiency. 68 points are sufficient to cover the key areas of the face while avoiding redundancy. The 68-point distribution can accurately describe the key structures of the face, including the dynamic changes of the eyebrows, eyes, and lips; the geometric features of the nose bridge and chin, and the head pose can be calculated through symmetric points.

[0114] (2) Eye feature extraction

[0115] The EAR algorithm, also known as the eye aspect ratio, the core of this algorithm is to identify whether fatigue blinking occurs by calculating the aspect ratio of the eyes to evaluate whether an individual is fatigued.

[0116] When the human eye closes, the EAR value drops sharply. When the EAR value is lower than the preset threshold, it can be inferred that the eye has completed a blink. By detecting the number and frequency of blinks, the fatigue state can be further determined. In the fatigue detection algorithm of this application, the threshold of EAR is set to 0.2. When EAR is less than 0.2, it is recorded as one blink.

[0117] First, the coordinate data of six feature points of the eyes are obtained through the dlib library, and these six points are used as input values. For each eye, 6 key points on the upper and lower eyelids (1 on the left and right, 2 on the upper and lower) are taken. By calculating the ratio of the distance between two points in the vertical direction to the distance in the horizontal direction, the degree of eye closure is quantified. The 6 points can simplify the calculation, only 2 sets of vertical distances and 1 set of horizontal distances need to be measured, and the real-time performance is high.

[0118] Then, the EAR algorithm is used to calculate the six feature points obtained. For the left eye and the right eye respectively, the corresponding Euclidean distance is calculated, and its EAR value is obtained;

[0119] The formula for calculating the EAR value of the left eye is:

[0120] The formula for calculating the EAR value of the right eye is:

[0121] In the blink action, the left eye and the right eye move synchronously. By calculating the average value of the EAR values of the left and right eyes, the EAR value of the entire eye is obtained, so as to more accurately reflect the state of eye movement, as shown in the formula

[0122]

[0123] Through the above steps, the EAR algorithm can be used to effectively detect the fatigue state of the eyes;

[0124] (3) Mouth feature extraction

[0125] The action of yawning of the mouth is similar to the performance when the eyes are fatigued, and both can be used as the basis for judging the fatigue state. With the help of the feature point detection function of the Dlib library, the outline of the mouth can be accurately depicted. Subsequently, by applying the MAR algorithm, the degree of mouth opening can be calculated, and a suitable threshold can be set. In this way, yawning can be effectively distinguished from normal mouth speaking, providing an important basis for subsequent fatigue driving detection.

[0126] The MAR algorithm, the mouth aspect ratio algorithm, as an algorithm that relies on the mouth feature points of the Dlib library to evaluate the fatigue state, has significant application value. The Dlib library has collected 20 feature points on the mouth model, and these points cover the outer and inner contours of the lips. However, considering that the inner contour is more vulnerable to skin color noise interference during movement, resulting in inaccurate data, it is decided to use the outer contour feature points for calculation. When the mouth is in motion, each feature point will change synchronously. Therefore, 6 typical feature points are selected for calculating MAR, which not only ensures the accuracy of the calculation but also significantly improves the calculation speed. Through this method, the feature point detection function of the Dlib library can be utilized more effectively, providing a more reliable basis for fatigue driving detection.

[0127] Considering that the mouth movements can be effectively reflected by the six collected feature points, the calculation process of the MAR algorithm has certain similarities with the EAR algorithm. The main purpose of selecting six feature points in the MAR algorithm is to accurately and stably quantify the mouth opening and closing state. Through the collaborative calculation of multiple key points, the accuracy and reliability of mouth opening and closing detection are enhanced, supporting application scenarios such as fatigue monitoring.

[0128] The specific calculation steps are as follows: First, the six collected feature points are used as the input data of the algorithm; subsequently, using the MAR algorithm, the Euclidean distances between these six feature points are calculated to further analyze the degree of mouth opening, thereby judging the fatigue state of the driver. Thus, the value of MAR is obtained. This value will be used as an important basis for judging the fatigue state, as shown in the formula:

[0129]

[0130] In the fatigue detection algorithm of this paper, the threshold of MAR is set to 0.6. When MAR is greater than 0.6, it is determined as an abnormal mouth opening action.

[0131] (4) Eye feature fatigue determination

[0132] There are usually three measurement methods for the PRECLOS standard: p70, p80, and EM, which respectively indicate that when the pupil area is covered by the eyelid by more than 70%, 80%, and 50%, it is judged as the closed-eye state, and the proportion of the duration of the closed-eye state within the unit time is statistically calculated. This paper adopts the p80 evaluation standard. Compared with P70, P80 allows the system to process longer context or more candidate results, reducing the risk of missing key content due to premature truncation; compared with EM, P80 has stronger fault tolerance for the latter 20%, avoiding sacrificing practicality due to excessive pursuit of complete accuracy, especially suitable for scenarios that tolerate a small amount of error but require efficient coverage. In a question-and-answer or dialogue system, P80 allows some non-critical errors but can ensure the correctness of the main content, being closer to the real application requirements.

[0133] By extracting eye features, calculating the eye aspect ratio (EAR), and then judging whether the human eye is in a closed state according to the size of the EAR.

[0134] The fatigue detection algorithm of the present invention sets the threshold of EAR to 0.2. When EAR is less than 0.2, it is recorded as a blink, and then the PERCLOS criterion is combined to judge whether it is in a fatigued state. PERCLOS is defined as the degree of eye closure within a certain period, and its calculation formula is shown.

[0135]

[0136] Because the actual fatigue driving detection has real-time requirements, in order to avoid risks and errors during the detection process as much as possible, the PERCLOS values of the driver's left and right eye closures within 1 minute are statistically recorded as K eye 。

[0137] When K eye is greater than 0.2, it is determined that the trolley truck driver is in a fatigued state.

[0138] (5) Fatigue determination based on mouth features

[0139] When the body is fatigued, the mouth often shows a yawning motion. To accurately determine this motion, the degree of mouth opening can be judged according to the threshold set by the MAR algorithm to see if it reaches the standard of yawning. At the same time, it can also be further confirmed whether it is a yawning motion by analyzing the duration of the mouth opening state, which can be used as another important indicator for judging the fatigued state.

[0140] When the MAR value exceeds 0.65 and this state lasts for more than 3 seconds, it can be determined as a yawn. The frame rate of the video in this dataset is 30 frames per second. In consecutive video frames, if the number of frames with the mouth open reaches 90 frames, it is recorded as one yawn. The number of yawns of a normal adult within 1 minute is 1, and the number of yawns of the driver within 1 minute is statistically recorded as K mouth 。

[0141] When K mouth is greater than 1, it is determined that the trolley truck driver is in a fatigued state.

[0142] Step 3: Fatigue driving detection based on heart rate features;

[0143] (1) Heart rate feature signal extraction

[0144] The user's face is facing the camera, and 68 feature points of the human face can be successfully extracted based on the real-time captured face image by the camera. Select areas with rich blood flow such as Figure 3Shown as the Region of Interest (ROI), the position of the ROI is stabilized through feature point tracking (such as Lucas-Kanade) to compensate for small head movements.

[0145] Extract the time-series signals of the red (R), green (G), and blue (B) channels from the video frames. Perform color space conversion to separate luminance (Y) and chrominance (Cb, Cr), reduce the influence of illumination, and at the same time extract hue (H) and saturation (S) information. Then perform signal normalization operation to eliminate baseline drift, and its formula is

[0146]

[0147] where μ c (t) is the mean value of the color channels within the ROI.

[0148] Separate the heart rate signal from other interference sources through independent component analysis (ICA), and use a band-pass filter to retain the frequency range of 0.5 - 4 Hz (corresponding to 30 - 240 BPM).

[0149] Perform Fourier transform (FFT) to convert the time-domain signal into the frequency domain, find the energy peak, and then locate the frequency f corresponding to the maximum power through the power spectral density (PSD) peak .

[0150] (2) Heart rate calculation

[0151] The Fourier transform can convert the time-domain signal into the frequency domain space, so as to analyze the frequency-domain characteristics of the signal. After using the fast Fourier transform FFT on the extracted BVP signal, the spectrogram in the shown flowchart is obtained. Since the heart rate of adults at rest is between 50 and 150 beats per minute (corresponding to the frequency between 0.8 Hz - 2.5 Hz), within this range, select the frequency f corresponding to the maximum power peak peak , then the corresponding heart rate

[0152] v hr is:

[0153] v hr = 60 * f peak

[0154] In addition, when using the Fourier transform to perform spectrum analysis, high-frequency noise can be removed, realizing the denoising process of the signal.

[0155] (3) Heart rate feature fatigue determination

[0156] When the value of v hr stabilizes at 100 or higher, the tram driver is in a fatigued state.

[0157] Step 4: Feature fusion fatigue determination

[0158] Considering various fatigue behavior characteristics comprehensively, the research on driver fatigue state detection based on facial features is carried out, and the KSS value is used for comprehensive judgment to improve the accuracy and reliability of the fatigue detection algorithm.

[0159] In the definition of fatigue level, the KSS scale based on drowsiness feeling and using numerical scoring is usually used. The specific scoring can be seen in the following table

[0160]

[0161] According to the "Terminal Technical Specification for Intelligent Video Surveillance and Alarm System of Road Transport Vehicles", because the fatigue level has reached a certain stage, the KSS values of blinking, yawning and heart rate fatigue states are set to 7.

[0162] When the K eye value is greater than 0.2 and less than 0.4, its KSS value is given as 7; when the K eye value is greater than 0.4 and less than 0.5, its KSS value is given as 8; when the K eye value is greater than 0.5, its KSS value is given as 9.

[0163] When the K mouth value is 2, its KSS value is given as 7; when the K mouth value is 3, its KSS value is given as 8; when the K mouth value is greater than 3, its KSS value is given as 9.

[0164] When the v hr value is greater than 100 and less than 110, its KSS value is given as 7; when the v hr value is greater than 110 and less than 130, its KSS value is given as 8; when the v hr value is greater than 130, its KSS value is given as 9.

[0165] When the value of KSS e +KSS m +KSS h is greater than or equal to 21 and less than 24, the miner driver belongs to mild fatigue; when the value of KSS e +KSS m +KSS h is greater than or equal to 24 and less than 27, the miner driver belongs to moderate fatigue; when the value of KSS e +KSS m +KSS h is equal to 27, the miner driver belongs to severe fatigue.

[0166] The SSD model of the deep learning framework is selected in the present invention for facial expression recognition. This model has powerful feature extraction capabilities and can extract effective facial expression features from facial images in complex backgrounds, overcoming the interference of the mine environment to a certain extent and improving the recognition accuracy.

[0167] The present invention adopts the KSS scale based on the feeling of sleepiness and using numerical scoring, assigns different KSS values to different fatigue levels, and finally conducts comprehensive fatigue determination through the total KSS value.

[0168] The present invention combines the EAR algorithm and the MAR algorithm to first judge yawning and blinking behaviors, and then judges the fatigue of miners' drivers through the KSS scale based on the feeling of sleepiness and using numerical scoring, providing a clear standard for accurately judging the stress state of miners.

Claims

1. A method for detecting fatigue signs of miner drivers, characterized in that, Including the following steps; Step 1: Identify the face of the miner driver; Step 2: Based on the identified facial features, conduct driving fatigue detection; Step 3: Fatigue driving detection based on heart rate features; Step 4: Based on the driving fatigue detection and the fatigue driving detection based on heart rate features, conduct feature fusion fatigue determination.

2. The method for detecting fatigue signs of miner drivers according to claim 1, wherein The specific content of Step 1 is as follows: Combine the SSD deep learning model with the Anchor mechanism of Faster R-CNN to perform face recognition on the collected video images; Use the SSD deep learning model to identify the facial area and the coordinate position information of the face in the video frame.

3. The method for detecting fatigue signs of miner drivers according to claim 2, characterized in that, The video images are collected in real time from the camera. The target category and position information are read from the annotation file during training. The annotation contains the category label and bounding box coordinates of each target; The Anchor mechanism enables the RPN of Faster R-CNN to generate candidate boxes using multiple scales and aspect ratios, covering different target shapes; Set anchors on each feature layer of SSD. The shallow feature map detects small-scale faces; the deep feature map detects large-scale faces; the Anchor size of each layer of the feature map matches its receptive field; At the same time, adopt the bidirectional matching strategy of Faster R-CNN. Each image corresponds to a set of annotations, including the category and bounding box coordinates of each target. Match the ground truth box (GT) of each annotation file directly from the dataset to the Anchor with the highest IoU, and at the same time mark the Anchor with IoU>0.5 as a positive sample; the rest are marked as negative samples; Adopt hard sample mining for negative samples; Dynamic weight adjustment is achieved by introducing FocalLoss to replace cross-entropy; Use a two-stage heuristic design. Embed a simplified version of RPN at the front end of SSD, and transform the Anchor adjustment and feature fusion into a practical network structure; generate high-quality face candidate boxes, and then through the classification and regression of multiple feature layers of SSD, balance speed and accuracy; Add an attention mechanism to the regression branch to simulate the face detection scenario.

4. The miner driver fatigue symptom detection method according to claim 3, characterized in that, The loss function design selects Smooth L1 Loss, The loss function of the entire model is: Where x is used to judge whether the detection box matches the real bounding box, represented by 1 and 0; N represents the number of matching boxes; l is the predicted boundary position parameter; g is the real boundary position parameter; c represents the category; L conf (x, c) represents the recognition result degree and is the Softmax loss function for multiple classes; L loc (x, l, g) uses the SmoothL1 Loss to predict the loss value between the position and the actual result; After the above operations, identify the face and obtain the features of the face, eyes, and mouth.

5. A method for detecting fatigue signs of miner drivers according to claim 1, characterized in that, The specific content of Step 2 is as follows: (1) Facial feature marking Based on the face map collected in real time by the camera, extract the feature points of the face according to the 68 facial feature points of the Dlib library. The extraction of the feature points depends on the principle of the HOG algorithm to locate the position and bounding box of the face in the image, and output the bounding box of the face area; Then load the pre-trained model, which is trained by the ERT algorithm; Gradually optimize the position of the feature points through a multi-level regression tree, from the initial estimate to the gradual adjustment. Each regression tree learns the mapping relationship between the local texture features and the offset of the feature points, and finally outputs the coordinates (x, y) of the feature points; (2) Eye feature extraction First, the coordinate data of eye feature points was obtained through the dlib library, and the coordinate data was used as the input value; By calculating the ratio of the distance between two points in the vertical direction to the distance in the horizontal direction, the degree of eye closure was quantified; Then, the EAR algorithm was used to calculate the six obtained feature points. For the left and right eyes respectively, the corresponding Euclidean distances were calculated, and their EAR values were obtained; The calculation formula for the EAR value of the left eye is: The calculation formula for the EAR value of the right eye is: During the blinking action, by calculating the average value of the EAR values of the left and right eyes, the EAR value of the entire eye was obtained, as shown in Equation (3) Mouth feature extraction The Dlib library collected feature points on the mouth model, and the features covered the outer and inner contours of the lips; First, the collected feature points were used as the input data of the algorithm; subsequently, using the MAR algorithm, the Euclidean distances between these six feature points were calculated to obtain the MAR value, as shown in the equation: (4) Eye feature fatigue determination Adopting the p80 judgment criterion, by extracting eye features, calculating the aspect ratio of the eye width to height, and then according to the size of the EAR, it was judged whether the human eye was in a closed state; (5) Mouth feature fatigue determination According to the threshold set by the MAR algorithm, it was judged whether the degree of mouth opening reached the yawn standard, and at the same time, by analyzing the duration of the mouth opening state, it was further confirmed whether it was a yawn action.

6. A method for detecting fatigue signs of miner drivers according to claim 5, characterized in that, In the step (2), when the EAR value was lower than the preset threshold, the eye completed a blink. The threshold of EAR was set to 0.

2. When EAR was less than 0.2, it was recorded as a blink; In the step (3), the threshold of MAR was set to 0.

6. When MAR was greater than 0.6, it was determined as an abnormal mouth-opening action.

7. A method for detecting fatigue signs of miner drivers according to claim 4, characterized in that, In the step (4), the threshold of EAR was set to 0.

2. When EAR was less than 0.2, it was recorded as a blink. PERCLOS was defined as the degree of eye closure within a certain period, and its calculation formula was shown. Statistically calculate the PERCLOS value K of the driver's left and right eye closures within 1 minute eye ; When K eye is greater than 0.2, it is determined that the tramcar driver is in a fatigued state.

8. A method for detecting fatigue signs of miner drivers according to claim 4, characterized in that, In the step (5), when the MAR value exceeds 0.65 and this state lasts for more than 3 seconds, it is determined as yawning. The number of yawns of a normal adult within 1 minute is 1, and the number of yawns of the driver within 1 minute is recorded as K mouth ; When K mouth is greater than 1, it is determined that the tram driver is in a fatigued state.

9. A method for detecting fatigue signs of miner drivers according to claim 4, characterized in that, The specific content of the step 3 is as follows: (1) Heart rate feature signal extraction The user's face was facing the camera. Based on the real-time collection of the face image by the camera, the feature points of the face could be successfully extracted. The area with rich blood flow was selected as the ROI (Region of Interest), and the position of the ROI was stabilized through feature point tracking to compensate for the small head movement; The time series signals of the red (R), green (G), and blue (B) channels were extracted from the video frame, and color space conversion was performed to separate the luminance (Y) and chrominance (Cb, Cr), and then signal normalization operation was performed to eliminate the baseline drift. The formula was where μ c (t) is the mean value of the color channels within the ROI; The heart rate signal and other interference sources were separated through independent component analysis, and a band-pass filter was used to retain the frequency range of 0.5 - 4 Hz; Perform Fourier transform to convert the time-domain signal to the frequency domain, find the energy peak, and then locate the frequency f corresponding to the maximum power through the power spectral density peak ; (2) Heart rate calculation Select the frequency f corresponding to the maximum power peak peak , and the corresponding heart rate v hr is: v hr = 60 * f peak (3) Heart rate feature fatigue determination When v hr is stable at 100 or higher, the tram driver is in a fatigued state.

10. A method for detecting fatigue signs of miner drivers according to claim 9, characterized in that, The specific content of the step 4 is as follows: When the degree of fatigue has reached a certain stage, the KSS values of the fatigue states of blinking, yawning, and heart rate are set to 7; When K eye value is greater than 0.2 and less than 0.4, assign its KSS value as 7; when K eye value is greater than 0.4 and less than 0.5, assign its KSS value as 8; when K eye value is greater than 0.5, assign its KSS value as 9; When K mouth has a value of 2, assign it a KSS value of 7; when K mouth has a value of 3, assign it a KSS value of 8; when K mouth has a value greater than 3, assign it a KSS value of 9; When v hr value is greater than 100 and less than 110, assign its KSS value to 7; when v hr value is greater than 110 and less than 130, assign its KSS value to 8; when v hr value is greater than 130, assign its KSS value to 9; When KSS e + KSS m + KSS h is greater than or equal to 21 and less than 24, the miner driver is mildly fatigued; when KSS e + KSS m + KSS h is greater than or equal to 24 and less than 27, the miner driver is moderately fatigued; when KSS e + KSS m + KSS h is equal to 27, the miner driver is severely fatigued.

Citation Information

Cited By

  • Petroleum miner fatigue monitoring method and system based on multi-modal information fusion

    CN122291049A