Facial feature-based fatigue driving recognition method and related device
By combining a multi-task neural network for head pose estimation and facial landmark detection with a lightweight multi-classification neural network, and comprehensively analyzing the features of the driver's head, mouth, and eyes, this method solves the problems of high invasiveness, low accuracy, and poor robustness of existing fatigue driving recognition methods, and achieves fatigue driving recognition with high accuracy and low resource consumption.
Patent Information
- Application Number
- CN202411647775.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-18
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2044-11-18
AI Technical Summary
Existing methods for identifying fatigued driving suffer from problems such as high intrusiveness, low accuracy, and poor robustness. In particular, methods based on facial features suffer from poor recognition performance in complex environments due to their reliance on a single feature.
A multi-task neural network for head pose estimation and facial landmark detection is adopted, combined with a lightweight multi-classification neural network. Driver images are collected through vehicle-mounted camera equipment, and fatigue status is identified by comprehensively analyzing head, mouth and eye features.
It achieves highly accurate fatigue driving recognition without requiring drivers to wear equipment, and has high robustness and adaptability to low-complexity environments, with low hardware resource consumption and recognition time.
Smart Images

Figure CN119888694B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of fatigue detection technology, specifically to a fatigue driving recognition method and related device based on facial features. Background Technology
[0002] With the rapid growth of my country's economy and the gradual improvement of its highway system, automobiles have become the most frequently used means of transportation for Chinese citizens. Real-time and accurate identification of driver fatigue characteristics is of great significance in reducing traffic accidents.
[0003] Current methods for identifying fatigued driving can be categorized into three types based on the characteristics used: those based on driver physiological signals, those based on vehicle features, and those based on driver facial features. Methods based on driver physiological signals acquire real-time physiological signals such as electrocardiogram (ECG), electroencephalogram (EEG), and pulse using medical devices to analyze and identify the driver's fatigue state. While these methods offer high accuracy, they require drivers to wear specialized equipment, making them highly invasive. Methods based on vehicle status features analyze vehicle braking, lane departure, and speed changes to assess driver fatigue. However, these features are affected by external environment, road conditions, and the driver's skill level and habits, resulting in low reliability. Methods based on facial features use computer vision to identify fatigue characteristics in the driver's mouth, eyes, and head. These methods offer advantages such as good real-time performance, high accuracy, low invasiveness, and low cost. However, existing methods suffer from limitations such as using only a single feature and poor robustness. Summary of the Invention
[0004] This application provides a fatigue driving recognition method and related device based on facial features, which comprehensively analyzes the three major facial features of the driver's head, mouth and eyes to achieve high accuracy and robustness in real-time recognition of the driver's driving status.
[0005] The first aspect of this application provides a fatigue driving recognition method based on facial features, the fatigue driving recognition method comprising:
[0006] The driver's images are collected by an in-vehicle camera device while the driver is driving. The driver's image set is a collection of driver images collected within a preset time interval.
[0007] A multi-task neural network for head pose estimation and facial landmark detection is used to process driver images in a driver image set to obtain the head deflection angle and facial landmark coordinates corresponding to each driver image.
[0008] The facial key point coordinates are used to crop the corresponding driver image to obtain the left eye image, right eye image and mouth image corresponding to each driver image;
[0009] The left eye image, right eye image, and mouth image corresponding to each driver image are input into a lightweight multi-classification neural network to perform driver state recognition, thereby obtaining the mouth state and eye state corresponding to each driver image.
[0010] The driver's fatigue state is obtained by using the head deflection angle, mouth state, and eye state corresponding to each driver image.
[0011] Compared with existing technologies, this invention has the following advantages and beneficial effects: 1. This invention identifies driver fatigue by acquiring driver images from a frontal view using a camera device, without requiring the driver to wear any equipment, thus avoiding strong invasiveness. 2. This invention utilizes the high correlation between head pose estimation and facial landmark detection tasks, introducing a multi-task learning mechanism, setting head pose estimation as the primary task, and improving the accuracy of head pose estimation with the help of facial landmark detection. 3. This invention uses multiple indicators to comprehensively analyze driver fatigue status, and different indicators are not identified by a single network. Even under complex environmental conditions, this invention still has high generalization and robustness. 4. This invention uses a shared backbone network with only added branch networks in the multi-task detection network, and a lightweight network framework in the multi-classification neural network based on the task, ensuring recognition performance without significantly increasing hardware resource consumption and recognition time.
[0012] A second aspect of this application provides a fatigue driving recognition device based on facial features, the device comprising:
[0013] The acquisition unit is used to acquire a set of driver images while driving using an in-vehicle camera device. The set of driver images is a collection of driver images acquired within a preset time interval.
[0014] The processing unit is used to process the driver images in the driver image set using a multi-task neural network for head pose estimation and facial key point detection, so as to obtain the head deflection angle and facial key point coordinates corresponding to each driver image.
[0015] The cropping unit is used to crop the corresponding driver image using the facial key point coordinates to obtain the left eye image, right eye image and mouth image corresponding to each driver image;
[0016] The recognition unit is used to input the left eye image, right eye image and mouth image corresponding to each driver image into a lightweight multi-classification neural network to perform driver state recognition, and obtain the mouth state and eye state corresponding to each driver image;
[0017] The detection unit is used to detect driver fatigue by using the head deflection angle, mouth state and eye state corresponding to each driver image, and to obtain the driver's fatigue state.
[0018] A third aspect of this application provides a terminal including a processor, an input device, an output device, and a memory, wherein the processor, input device, output device, and memory are interconnected, wherein the memory is used to store a computer program, the computer program including program instructions, and the processor is configured to invoke the program instructions to execute the step instructions as described in the first aspect of this application.
[0019] A fourth aspect of this application provides a computer-readable storage medium storing a computer program for electronic data interchange, wherein the computer program causes a computer to perform some or all of the steps described in the first aspect of this application.
[0020] A fifth aspect of this application provides a computer program product, wherein the computer program product includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps described in the first aspect of this application. The computer program product may be a software installation package. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This application provides a flowchart illustrating a fatigue driving recognition method based on facial features.
[0023] Figure 2 This application provides a schematic diagram of the location of facial key points in an embodiment;
[0024] Figure 3 A schematic diagram of a head deflection angle is provided for an embodiment of this application;
[0025] Figure 4This is a schematic diagram of the structure of a terminal provided in an embodiment of this application;
[0026] Figure 5 This application provides a schematic diagram of the structure of a fatigue driving recognition device based on facial features. Detailed Implementation
[0027] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0028] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.
[0029] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application can be combined with other embodiments.
[0030] Please see Figure 1 , Figure 1 This application provides a flowchart illustrating a fatigue driving recognition method based on facial features. Figure 1 As shown, the method includes:
[0031] 101. A set of driver images collected by an in-vehicle camera device while the driver is driving, wherein the set of driver images is a collection of driver images collected within a preset time interval.
[0032] The vehicle-mounted camera can be installed on the dashboard. It can capture real-time images from the driver's frontal view, creating a set of driver images. The preset time interval can be a continuous time interval.
[0033] Specifically, a 1080P resolution camera with a frame rate of 30FPS is typically used and installed above the dashboard to capture a complete image from the driver's frontal view.
[0034] 102. A multi-task neural network for head pose estimation and facial landmark detection is used to process driver images in the driver image set to obtain the head yaw angle and facial landmark coordinates corresponding to each driver image.
[0035] The head pose estimation and facial landmark detection multi-task neural network (HFMNet) is described in detail below:
[0036] HFMNet selected 8,000 images from the open-source dataset 300-LP and 2,000 images from the BIWI dataset as the training and testing sets for the network, respectively. The annotation information for each image in the dataset includes 68 facial key points and the angles of the head in three directions.
[0037] HFMNet consists of three parts: a backbone network, a face landmark detection branch, and a head pose estimation branch. HFMN selects the first four segments of GhostNet as the backbone network to extract shared features. GhostNet can generate redundant feature maps through linear operations, thus simplifying the feature extraction process. The extracted feature maps are then fed into the face landmark detection branch and the head pose estimation branch, respectively. Both the face landmark detection branch and the head pose estimation branch have the last segment of GhostNet used to extract specific features for the branch task. Finally, the shared features and specific features are passed through pooling layers and fully connected layers to finally regress and obtain the results of each branch task.
[0038] HFMNet divides the entire head pose estimation time into 12 non-overlapping sub-intervals [-90°, 90°], each representing a 15-degree range; each sub-interval is further divided into 3 equally divided intervals. Finally, the classification result is transformed into a regression problem for head pose estimation by calculating the final expected value, specifically calculated using the following formula:
[0039]
[0040] Where p k It represents the probability that the angle falls within the k-th large interval. It is the probability that the angle is in the i-th sub-interval of the k-th large interval, α(i k ) is the representative value of the i-th cell in the k-th large interval, and θ is the head deflection angle;
[0041] The sum of the losses from facial landmark detection and head pose estimation is the total loss of HFMN, and the loss function is:
[0042] Loss total =Loss HPE +φLoss FKD ,
[0043] Among them, Loss total Here is the total loss function for HFMN;
[0044] Loss HPE and Loss FKD The loss functions for the head pose estimation task and the facial landmark task are as follows:
[0045]
[0046] in These are the predicted values for the three head deflection angles. These are the actual values of the three head deflection angles. These are the predicted and actual values of the i-th key point on the face, respectively.
[0047] Specifically, such as Figure 2 As shown, Figure 2 A schematic diagram showing the locations of facial key points is provided. Multiple key point labels are shown in the diagram; the facial key point labels used subsequently correspond to those in the diagram. Figure 3 A schematic diagram of a head tilt angle is shown.
[0048] 103. Using the facial key point coordinates, crop the corresponding driver image to obtain the left eye image, right eye image, and mouth image corresponding to each driver image.
[0049] Among them, six key points were selected for the left and right eyes, and four key points for the mouth were numbered 49, 52, 55 and 58.
[0050] For the left eye region, select the maximum x-coordinate among the 6 key points. max and minimum value x min And the maximum value y in the y-coordinate max and minimum value y min The center coordinates (x, y) of the candidate region are calculated using the method shown in the following formula. center y center ):
[0051]
[0052] Around the central coordinates, and The left eye image was cropped from the length and width of the candidate regions, and the right eye and mouth images were obtained using the same principle.
[0053] 104. Input the left eye image, right eye image and mouth image corresponding to each driver image into a lightweight multi-classification neural network to perform driver state recognition, and obtain the mouth state and eye state corresponding to each driver image.
[0054] A lightweight multi-class neural network is an improved ShuffleNetV2 network:
[0055] The dataset for the improved ShuffleNetV2 network consists of 2500 images each of the mouth and eyes obtained from the Kaggle platform. Based on the aspect ratio of the eyes (EAR) and the aspect ratio of the mouth (MAR), the images are divided into four categories: closed eyes, open eyes, yawning, and no yawning. There are 1316 images with closed eyes, 1184 with open eyes, 1236 with yawning, and 1264 without yawning. The images are divided into training and testing sets in a 7:1 ratio.
[0056] The eye aspect ratio (EAR) and mouth aspect ratio (MAR) can be calculated using the following formula:
[0057]
[0058] Among them, M i M is the coordinate information of the i-th key point of the face. 62 M is the coordinate information of the 62nd key point of the face. 38 M is the coordinate information of the 38th key point of the face. 42 M is the coordinate information of the 42nd key point of the face. 39 M is the coordinate information of the 39th key point of the face. 41 M is the coordinate information of the 41st key point of the face. 68 M is the coordinate information of the 68th key point of the face. 64 M is the coordinate information of the 64th key point of the face. 65 M is the coordinate information of the 65th key point of the face. 61 The coordinates of the 61st key point on the face;
[0059] The ShuffleNetV2 network is mainly composed of three modular layers: ShuffleNetV2 Unit2 and ShuffleNetV2 Unit1.
[0060] The αH-swish activation function replaces the ReLU activation function used in the original ShuffleNetV2 network. α is a very small number. The mathematical definition of the αH-swish activation function is as follows:
[0061]
[0062] 105. The driver's fatigue status is obtained by using the head deflection angle, mouth state and eye state corresponding to each driver image.
[0063] The fatigue state includes blinking frequency, longest eye-closing time, yawning frequency, nodding frequency, and longest head-down time, with the specific judgment criteria as follows:
[0064] A driver typically blinks for 0.2-0.3 seconds each time. Therefore, a blink is considered to occur when the eyes change from open to closed to open again within 6 consecutive frames.
[0065] If the eyes are identified as closed for more than 60 frames, it is considered as a prolonged period of eye closure.
[0066] A driver's mouth is usually open in a yawning state for no more than 1 second, so 30 consecutive frames of the mouth being in a yawning state is considered as one yawn.
[0067] The normal range for the pitch angle of the driver's head is 0-15°. Within 30 consecutive frames, a pitch angle that changes from the normal range to greater than 15° and then back to the normal range is considered a nod.
[0068] If the head tilt angle exceeds 90° or is greater than 15°, it is identified as a prolonged period of head tilting.
[0069] Furthermore, based on the fatigue level measurement table in Table 1, when a driver is detected to meet the standard of a certain fatigue level, it is identified as driving while fatigued to the corresponding degree.
[0070] Table 1 Fatigue Level Measurement Table
[0071]
[0072] Therefore, after extracting the blinking frequency, the longest time the eyes are closed, whether there is yawning, the nodding frequency, and the longest time the head is lowered, the driver's fatigue state can be determined by looking up Table 1.
[0073] In one specific implementation, this application also provides another fatigue driving recognition method based on facial features, as follows:
[0074] Step 1: Install the vehicle-mounted camera above the dashboard to capture real-time images of the driver from a frontal view.
[0075] Specifically, a 1080P resolution camera with a frame rate of 30FPS is typically used and installed above the dashboard to capture a complete image from the driver's frontal view.
[0076] Step 2: Use a multi-task neural network for head pose estimation and facial landmark detection to detect the data collected in Step 1 and obtain the driver's head yaw angle and facial landmark coordinates;
[0077] Specifically, the Head Pose Estimation and Facial Landmark Detection Multi-Task Neural Network (HFMNet) selected 8000 images from the open-source dataset 300-LP and 2000 images from the BIWI dataset as the training and testing sets, respectively. Each image in the dataset contains annotations for 68 facial landmarks and the angles of head rotation in three directions. The diagrams illustrating the locations of the facial landmarks and the head rotation angles are shown below. Figure 2 and Figure 3 As shown.
[0078] Furthermore, HFMNet consists of three parts: a backbone network, a face landmark detection branch, and a head pose estimation branch. HFMN selects the first four segments of GhostNet as the backbone network to extract shared features. GhostNet can generate redundant feature maps through linear operations, thus simplifying the feature extraction process. The extracted feature maps are then fed into the face landmark detection branch and the head pose estimation branch, respectively. Both branches have specific features extracted from the last segment of GhostNet for their respective tasks. Finally, each branch combines the shared features with its specific features through pooling layers and fully connected layers to obtain the results for its task.
[0079] Furthermore, in the head pose estimation problem, the head pose angle has a continuous distribution characteristic. If direct regression is used, there is often a large difference between the predicted value and the true value. HFMNet transforms the entire head pose estimation from a regression problem into a classification problem, dividing the entire interval [90°, 90°] into 12 non-overlapping sub-intervals, each sub-interval representing a range of 15 degrees. To further improve accuracy, each interval is further divided into 3 equally divided intervals. Finally, the classification result is transformed into a regression problem for head pose estimation by calculating the final expected value, as shown in Equation (1):
[0080]
[0081] Where p k It represents the probability that the angle falls within the k-th large interval. It is the probability that the angle is in the i-th sub-interval of the k-th large interval, α(i k ) is the representative value of the i-th cell in the k-th large interval, and θ is the head deflection angle.
[0082] Furthermore, the sum of the losses from the two tasks of face landmark detection and head pose estimation is the total loss of HFMN. To ensure that head pose estimation is the primary task, a weight parameter φ is introduced into the loss function, as shown in equation (2):
[0083] Loss total =Loss HPE +φLoss FKD (2)
[0084] Among them, Loss total Here is the total loss function for HFMN;
[0085] LOss HPE and Loss FKD The loss functions for the head pose estimation task and the face landmark task are respectively, and the calculation formulas are shown in equations (3) and (4):
[0086]
[0087] in These are the predicted values for the three head deflection angles. These are the actual values of the three head deflection angles. These are the predicted and actual values of the i-th key point on the face, respectively.
[0088] Furthermore, the initial learning rate was set to 0.001, the batch size to 4, and training was performed for a total of 100 iterations, with the learning rate decreasing by 1 / 10 every 10 iterations. The trained network was tested on the test set, and the absolute errors for Yaw, Pitch, and Roll were 4.39, 3.97, and 4.91, respectively. The mean absolute error of facial landmarks was 3.74, demonstrating the network's effectiveness. The HFMNet recognition results are as follows... Figure 4 and Figure 5 .
[0089] Step 3: Based on the facial landmark coordinate data obtained by the multi-task neural network for head pose estimation and facial landmark detection, locate and crop the driver's left eye, right eye and mouth images;
[0090] Specifically, six key points were selected for each of the left and right eyes, as well as four key points for the mouth numbered 49, 52, 55, and 58. For the left eye region, the maximum x-coordinate among the six key points was selected. max and minimum value x min And the maximum value y in the y-coordinate max and minimum value y min The center coordinates (x, y) of the candidate region are calculated using equations (5) and (6). center y center ):
[0091]
[0092] Furthermore, around the central coordinates, and The left eye image is cropped from the length and width of the candidate regions, and the right eye and mouth images are obtained using the same principle.
[0093] Step 4: Input the eye and mouth images obtained in Step 3 into a lightweight multi-classification neural network, and classify the input into four categories: yawning, not yawning, closed eyes, and open eyes, to identify the state of the driver's mouth and eyes.
[0094] Specifically, 2500 images each of the mouth and eyes were obtained from the Kaggle platform. Based on the eye aspect ratio (EAR) and mouth aspect ratio (MAR), the images were categorized into four types: closed eyes, open eyes, yawning, and not yawning. Specifically, 1316 images were closed eyes, 1184 were open eyes, 1236 were yawning, and 1264 were not yawning. These were then divided into training and testing sets at a 7:1 ratio. The formulas for calculating the eye aspect ratio (EAR) and mouth aspect ratio (MAR) can be expressed by equations (7) and (8):
[0095]
[0096]
[0097] Among them, M i The coordinate information of the i-th key point of the face;
[0098] Furthermore, the lightweight multi-class neural network is based on the ShuffleNetV2 network. The ShuffleNetV2 network is mainly composed of three module layers combining ShuffleNetV2 Unit2 and ShuffleNetV2 Unit1. This invention connects an efficient channel attention module (ECA) to each ShuffleNetV2 Unit1 and Unit2, enabling the network to perform global average pooling without dimensionality reduction. This avoids the adverse effects of dimensionality reduction on channel attention learning. Moreover, this module is simple and has a low computational burden.
[0099] Furthermore, the original ShuffleNetV2 network uses the ReLU activation function. While simple and efficient, this function has a value of 0 for inputs less than 0, preventing some neurons from being activated. This invention uses the αH-swish activation function instead of ReLU, maintaining the simplicity and efficiency of the activation function while also addressing the issue of gradient vanishing when the input is less than -3, which prevents neurons from being activated. α For a very small number, 0.05 is used in this invention. The mathematical definition of the H-swish activation function is Equation (9):
[0100]
[0101] Furthermore, the initial learning rate was set to 0.01, the batch size to 8, and training was performed for a total of 50 iterations, with the learning rate decreasing by 1 / 10 every 10 iterations. The trained network was tested on the test set, achieving an accuracy of 0.973, demonstrating the effectiveness of the network in eye and mouth pose classification.
[0102] Step 5: Based on the driver's facial features obtained in Step 2 and Step 4, statistically analyze the driver's blinking frequency, longest eye-closing time, whether yawning, nodding frequency, and longest head-down time over a period of time. Finally, identify the driver's fatigue state based on the initial thresholds set by multiple indicators.
[0103] Specifically, the results obtained in steps 2 and 4 are single-frame states. Drivers also exhibit characteristics such as closing their eyes and opening their mouths during normal driving; therefore, these results cannot be used as a basis for judgment. Therefore, statistical analysis of single-frame states is needed to generate features that can be used to identify fatigued driving, such as blinking frequency, yawning frequency, nodding frequency, and prolonged nodding or eye closure. The judgment criteria for each feature are as follows (video frame rate 30FPS):
[0104] (1) It usually takes 0.2-0.3 seconds for a driver to blink. Therefore, if the eyes change from open to closed to open again in 6 consecutive frames, it is considered as one blink.
[0105] (2) A driver's mouth is usually open in a yawning state for no more than 1 second. Therefore, if the mouth is in a yawning state for 30 consecutive frames, it is considered as one yawn.
[0106] (3) The normal range of the pitch angle of the driver's head is 0-15°. If the pitch angle of the head changes from the normal range to greater than 15° and then back to the normal range within 30 consecutive frames, it is considered as one nod.
[0107] (4) If the eyes are identified as closed for more than 60 frames, it is considered as a prolonged period of eye closure.
[0108] (5) If the head tilt angle exceeds 90° or is greater than 15°, it is identified as a prolonged head tilt.
[0109] Furthermore, based on the fatigue level measurement table in Table 1, when a driver is detected to meet the standard of a certain fatigue level, it is identified as driving while fatigued to the corresponding degree.
[0110] For examples consistent with the above embodiments, please refer to... Figure 4 , Figure 4A schematic diagram of a terminal structure provided in an embodiment of this application is shown in the figure. It includes a processor, an input device, an output device, and a memory. The processor, input device, output device, and memory are interconnected. The memory is used to store a computer program, which includes program instructions. The processor is configured to call the program instructions. The program includes instructions for performing the following steps.
[0111] The driver's images are collected by an in-vehicle camera device while the driver is driving. The driver's image set is a collection of driver images collected within a preset time interval.
[0112] A multi-task neural network for head pose estimation and facial landmark detection is used to process driver images in a driver image set to obtain the head deflection angle and facial landmark coordinates corresponding to each driver image.
[0113] The facial key point coordinates are used to crop the corresponding driver image to obtain the left eye image, right eye image and mouth image corresponding to each driver image;
[0114] The left eye image, right eye image, and mouth image corresponding to each driver image are input into a lightweight multi-classification neural network to perform driver state recognition, thereby obtaining the mouth state and eye state corresponding to each driver image.
[0115] The driver's fatigue state is obtained by using the head deflection angle, mouth state, and eye state corresponding to each driver image.
[0116] The above mainly describes the solutions of the embodiments of this application from the perspective of the method execution process. It is understood that, in order to achieve the above functions, the terminal includes the corresponding hardware structure and / or software modules for executing each function. Those skilled in the art should readily recognize that, in conjunction with the units and algorithm steps of the various examples described in the embodiments provided herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0117] This application embodiment can divide the terminal into functional units according to the above method example. For example, each function can be divided into a separate functional unit, or two or more functions can be integrated into one processing unit. The integrated unit can be implemented in hardware or as a software functional unit. It should be noted that the unit division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.
[0118] For those consistent with the above, please refer to Figure 5 , Figure 5 This application provides a schematic diagram of a fatigue driving recognition device based on facial features. Figure 5 As shown, the device includes:
[0119] The acquisition unit 501 is used to acquire a set of driver images while driving through an in-vehicle camera device. The set of driver images is a collection of driver images acquired within a preset time interval.
[0120] The processing unit 502 is used to process the driver images in the driver image set using a multi-task neural network for head pose estimation and facial key point detection, so as to obtain the head deflection angle and facial key point coordinates corresponding to each driver image.
[0121] The cropping unit 503 is used to crop the corresponding driver image using the facial key point coordinates to obtain the left eye image, right eye image and mouth image corresponding to each driver image;
[0122] The recognition unit 504 is used to input the left eye image, right eye image and mouth image corresponding to each driver image into a lightweight multi-classification neural network to perform driver state recognition, and obtain the mouth state and eye state corresponding to each driver image.
[0123] The detection unit 505 is used to detect driver fatigue by using the head deflection angle, mouth state and eye state corresponding to each driver image, and to obtain the driver's fatigue state.
[0124] In one possible implementation, the head pose estimation and facial landmark detection multi-task neural network (HFMNet) is specifically as follows:
[0125] HFMNet selected 8,000 images from the open-source dataset 300-LP and 2,000 images from the BIWI dataset as the training and testing sets for the network, respectively. The annotation information for each image in the dataset includes 68 facial key points and the angles of the head in three directions.
[0126] HFMNet consists of three parts: a backbone network, a face landmark detection branch, and a head pose estimation branch. HFMN selects the first four segments of GhostNet as the backbone network to extract shared features. GhostNet can generate redundant feature maps through linear operations, thus simplifying the feature extraction process. The extracted feature maps are then fed into the face landmark detection branch and the head pose estimation branch, respectively. Both the face landmark detection branch and the head pose estimation branch have the last segment of GhostNet used to extract specific features for the branch task. Finally, the shared features and specific features are passed through pooling layers and fully connected layers to finally regress and obtain the results of each branch task.
[0127] HFMNet divides the entire head pose estimation time into 12 non-overlapping sub-intervals [-90°, 90°], each representing a 15-degree range; each sub-interval is further divided into 3 equally divided intervals. Finally, the classification result is transformed into a regression problem for head pose estimation by calculating the final expected value, specifically calculated using the following formula:
[0128]
[0129] Where p k It represents the probability that the angle falls within the k-th large interval. It is the probability that the angle is in the i-th sub-interval of the k-th large interval, α(i k ) is the representative value of the i-th cell in the k-th large interval, and θ is the head deflection angle;
[0130] The sum of the losses from facial landmark detection and head pose estimation is the total loss of HFMN, and the loss function is:
[0131] Loss total =Loss HPE +φLoss FKD
[0132] Among them, Loss total Here is the total loss function for HFMN;
[0133] Loss HPE and Loss FKD The loss functions for the head pose estimation task and the facial landmark task are as follows:
[0134]
[0135] in These are the predicted values for the three head deflection angles. These are the actual values of the three head deflection angles. These are the predicted and actual values of the i-th key point on the face, respectively.
[0136] In one possible implementation, the clipping unit is specifically used for:
[0137] Select 6 key points for each of the left and right eyes, and 4 key points for the mouth numbered 49, 52, 55 and 58;
[0138] For the left eye region, select the maximum x-coordinate among the 6 key points. max and minimum value x min And the maximum value y in the y-coordinate max and minimum value y min The center coordinates (x, y) of the candidate region are calculated using the method shown in the following formula. center y center ):
[0139]
[0140] Around the central coordinates, and The left eye image was cropped from the length and width of the candidate regions, and the right eye and mouth images were obtained using the same principle.
[0141] In one possible implementation, a lightweight multi-class neural network is an improved ShuffleNetV2 network:
[0142] The dataset for the improved ShuffleNetV2 network consists of 2500 mouth and 2500 eye images obtained from the Kaggle platform. Based on the aspect ratio of the eyes (EAR) and the aspect ratio of the mouth (MAR), the images are divided into four categories: closed eyes, open eyes, yawning, and no yawning. There are 1316 closed eyes, 1184 open eyes, 1236 yawning, and 1264 no yawning images. They are divided into training and testing sets in a 7:1 ratio.
[0143] The eye aspect ratio (EAR) and mouth aspect ratio (MAR) can be calculated using the following formula:
[0144]
[0145] Among them, M i The coordinate information of the i-th key point of the face;
[0146] The ShuffleNetV2 network is mainly composed of three modular layers: ShuffleNetV2 Unit2 and ShuffleNetV2 Unit1.
[0147] The αH-swish activation function replaces the ReLU activation function used in the original ShuffleNetV2 network. α is a very small number. The mathematical definition of the αH-swish activation function is as follows:
[0148]
[0149] In one possible implementation, the fatigue state includes blinking frequency, longest eye-closing time, yawning frequency, nodding frequency, and longest head-down time, with the specific judgment criteria as follows:
[0150] A driver typically blinks for 0.2-0.3 seconds each time. Therefore, a blink is considered to occur when the eyes change from open to closed to open again within 6 consecutive frames.
[0151] If the eyes are identified as closed for more than 60 frames, it is considered as a prolonged period of eye closure.
[0152] A driver's mouth is usually open in a yawning state for no more than 1 second, so 30 consecutive frames of the mouth being in a yawning state is considered as one yawn.
[0153] The normal range for the pitch angle of the driver's head is 0-15°. Within 30 consecutive frames, a pitch angle that changes from the normal range to greater than 15° and then back to the normal range is considered a nod.
[0154] A head tilt angle exceeding 90° (greater than 15°) is identified as a prolonged period of head tilting down. This application also provides a computer storage medium storing a computer program for electronic data interchange, which causes a computer to perform some or all of the steps of any of the facial feature-based fatigue driving recognition methods described in the above method embodiments.
[0155] This application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program that causes a computer to perform some or all of the steps of any of the facial feature-based fatigue driving recognition methods described in the above method embodiments.
[0156] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0157] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0158] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical or other forms.
[0159] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0160] Furthermore, the functional units in the various embodiments of the application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software program module.
[0161] If the integrated unit is implemented as a software program module and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0162] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage device, which may include: a flash drive, a read-only memory, a random access memory, a magnetic disk, or an optical disk, etc.
[0163] The embodiments of this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A face feature-based fatigue driving recognition method, characterized in that, The fatigue driving recognition method comprises: Collecting a driver image set of a driver during driving through a vehicle-mounted camera device, the driver image set being a set of driver images collected within a preset time interval; Processing the driver images in the driver image set using a head pose estimation and facial key point detection multi-task neural network to obtain a head deflection angle and facial key point coordinates corresponding to each driver image; Image cropping of the corresponding driver images using the facial key point coordinates to obtain a left eye image, a right eye image and a mouth image corresponding to each driver image; Inputting the left eye image, the right eye image and the mouth image corresponding to each driver image into a lightweight multi-classification neural network to recognize the state of the driver to obtain a mouth state and an eye state corresponding to each driver image; Fatigue detection of the driver using the head deflection angle, the mouth state and the eye state corresponding to each driver image to obtain a fatigue state of the driver; The HFMNet selects 8000 pictures of an open source dataset 300-LP and 2000 pictures of a BIWI dataset as a training set and a test set of the network respectively, and the annotation information of each picture of the dataset includes 68 facial key points and the angles of three head direction deflection angles; The HFMNet is composed of a backbone network, a facial key point detection branch and a head pose estimation branch, the HFMNet selects the first four segments of a GhostNET as the backbone network to extract features for sharing, the GhostNet can generate redundant feature maps through linear operation, thereby simplifying the feature extraction process, and the extracted feature maps are sent into the facial key point detection branch and the head pose estimation branch respectively, the facial key point detection branch and the head pose estimation branch both have the last segment of the GhostNET for extracting specific features of the branch task, and finally each shared feature and specific feature are subjected to a pooling layer and a fully connected layer to finally regress the results of the tasks of the branches; HFMNet divides the angle interval into 12 disjoint sub-intervals, each of which represents a 15-degree range; each interval is further divided into 3 equal intervals, and the classification result is converted into a regression problem by calculating the final expected value to estimate the head pose. The head pose is calculated by the following formula: , wherein is the probability that the angle lies in the kth macro interval, is the probability that the angle lies in the ith sub-interval of the kth macro interval, is the representative value of the ith sub-interval of the kth macro interval, and θ is the head deflection angle. The sum of the losses of the two tasks of facial key point detection and head pose estimation is the total loss of the HFMNet, and the loss function is: , where Loss total is the total loss function of HFMNet; and are loss functions for the head pose estimation task and the facial landmark task, respectively, and are as follows: , , wherein is a predicted value of the three head yaw angles, is an actual value of the three head yaw angles, are a predicted value and an actual value of the i-th key point of the face, respectively. The ShuffleNetV2 network is mainly composed of a module layer composed of three ShuffleNetV2 Unit2 and ShuffleNetV2 Unit1 combination layers; using a ReLU activation function instead of the aH-swish activation function, The aH-swish activation function is mathematically defined as follows: 。 2.The face feature based driver fatigue recognition method according to claim 1, characterized in that, Image cropping of the corresponding driver images using the facial key point coordinates to obtain a left eye image, a right eye image and a mouth image corresponding to each driver image, comprising: Selecting six key points of the left eye and the right eye and four key points marked as 49, 52, 55 and 58 of the mouth; For the left eye region, the maximum value of the 6 key point x coordinates is selected and the minimum value and the maximum value of the y coordinates and the minimum value are selected The center coordinates of the candidate region are calculated by the following formula : , , Around the center coordinate, cut out the left eye image, right eye and mouth image with the same principle. and respectively as the long and width of the candidate region. 3.The face feature based driver fatigue recognition method according to claim 2, characterized in that, The lightweight multi-classification neural network is an improved ShuffleNetV2 network: The improved ShuffleNetV2 network dataset obtains 2500 mouth and eye pictures from the Kaggle platform, and the pictures are divided into four categories of closed eyes, open eyes, yawning and non-yawning according to the eye aspect ratio EAR and the mouth aspect ratio MAR, wherein there are 1316 closed eyes, 1184 open eyes, 1236 yawning and 1264 non-yawning, and the pictures are divided into a training set and a test set according to a ratio of 7:1; The eye aspect ratio EAR and the mouth aspect ratio MAR can be calculated by the following formula: , , wherein, is coordinate information of the i-th key point of the face. is coordinate information of the i-th key point of the face. 4.The face feature based driver fatigue recognition method according to claim 3, characterized in that, The fatigue state includes the blink frequency, the longest closed eye time, whether yawning, the nodding frequency and the longest head-down time, and the specific judgment basis is as follows: Generally, the driver needs 0.2-0.3s for each blink, so in the continuous 6 frames, the eye appears the state change of open eye, closed eye to open eye, which is regarded as a blink; The eye is more than 60 frames and is recognized as a long time closed eye; Generally, the driver's mouth opening to yawn state will not exceed 1s, so the continuous 30 frames of mouth are all yawning, which is regarded as a yawn; The normal range of the driver's head pitch angle is 0-15°, and in the continuous 30 frames, the head pitch angle changes from the normal range to more than 15° and then to the normal range, which is regarded as a nod. The head pitch angle is more than 90 frames and is recognized as a long time head-down.
5. A face feature-based drowsy driving recognition apparatus, characterized by comprising: The device comprises: The collection unit is configured to collect a driver image set of a driver while driving by using a vehicle-mounted camera device, wherein the driver image set is a set of driver images collected within a preset time interval; The processing unit is configured to process the driver images in the driver image set by using a head pose estimation and facial key point detection multi-task neural network to obtain a head deflection angle and facial key point coordinates corresponding to each driver image; The cropping unit is configured to crop the corresponding driver image by using the facial key point coordinates to obtain a left eye image, a right eye image and a mouth image corresponding to each driver image; The recognition unit is configured to input the left eye image, the right eye image and the mouth image corresponding to each driver image into a lightweight multi-classification neural network to recognize the state of the driver and obtain a mouth state and an eye state corresponding to each driver image; The detection unit is configured to detect the fatigue of the driver by using the head deflection angle, the mouth state and the eye state corresponding to each driver image to obtain the fatigue state of the driver. The collection unit is specifically configured to: The HFMNet selects 8000 pictures of the open source dataset 300-LP and 2000 pictures of the BIWI dataset as the training set and the test set of the network respectively, and the annotation information of each picture of the dataset includes 68 facial key points and the angles of three head direction deflection angles. The HFMNet is composed of a backbone network, a face key point detection branch and a head pose estimation branch, the HFMNet selects the first four segments of the GhostNET as the backbone network to extract features for sharing, the GhostNET can generate redundant feature maps through linear operation, thereby simplifying the feature extraction process, the extracted feature maps are respectively sent to the face key point detection branch and the head pose estimation branch, the face key point detection branch and the head pose estimation branch both have the last segment of the GhostNET for extracting specific features of the branch task, finally, each shared feature and the specific feature are subjected to a pooling layer and a full connection layer to finally regress the results of the tasks of each branch; HFMNet divides the angle interval into 12 disjoint sub-intervals, each of which represents a 15-degree range; each interval is further divided into 3 equal intervals, and the classification result is converted into a regression problem by calculating the final expected value to estimate the head pose. The head pose is calculated by the following formula: , wherein is the probability that the angle lies in the kth large interval, is the probability that the angle lies in the ith sub-interval of the kth large interval, is the representative value of the ith sub-interval of the kth large interval, and θ is the head deflection angle. The sum of the losses of the two tasks of face key point detection and head pose estimation is the total loss of the HFMN, and the loss function is: wherein Loss total is the HFMN total loss function; and are the loss functions of the head pose estimation task and the facial landmark task, respectively, and are as follows: , , wherein is a predicted value of the three head yaw angles, is an actual value of the three head yaw angles, are a predicted value and an actual value of the i-th key point of the face, respectively. The ShuffleNetV2 network is mainly composed of a module layer composed of three ShuffleNetV2 Unit2 and ShuffleNetV2 Unit1; using a ReLU activation function instead of the aH-swish activation function, The aH-swish activation function is mathematically defined as follows: 。 6. A terminal, characterized by comprising: The computer readable storage medium stores a computer program, the computer program includes program instructions, and the processor is configured to invoke the program instructions to execute the method in any one of claims 1-4.
7. A computer readable storage medium characterized in that, The computer readable storage medium stores a computer program, the computer program includes program instructions, and the processor is configured to invoke the program instructions to execute the method in any one of claims 1-4.
Citation Information
Patent Citations
Deep learning-based ship driver fatigue detection method and system
CN113158850A
Fatigue state detection method and system based on key point detection and head posture
CN114360041A